Image generation method and device, electronic equipment and storage medium
By combining style description information and object recognition information, and using a low-rank adapter model and a control network model, the image generation method solves the problem of fixed image content features, generates stylized images with personalized object features and style features, and improves image diversity.
Patent Information
- Application Number
- CN202410431240.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-17
AI Technical Summary
The image generation method in the prior art cannot control the image content, resulting in fixed content features of the generated images, high content similarity, and inability to achieve personalized portrait features.
By acquiring style description information and object recognition information, an image generation model containing a first low-rank adapter model and a second low-rank adapter model is invoked to generate a stylized image. The first low-rank adapter model is used to generate the main object, and the second low-rank adapter model is used to set the image style. Fine control is achieved using a control network model.
It enables the generation of stylized images with personalized object features and specified style features, thereby improving the diversity and personalization of image content.
Smart Images

Figure CN120807269A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of artificial intelligence, and particularly relate to an image generation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Currently, in various types of applications (Applicant, APP) with image editing and generation functions, beautification special effect tools for portrait pictures are provided. Through such tools, the image style of a portrait picture can be changed to improve the visual performance of the portrait picture.
[0003] However, the existing solution can only adjust the image style on the basis of a fixed portrait picture, and cannot control the image content, resulting in the generated image having fixed image content features and high content similarity. SUMMARY
[0004] Embodiments of the present disclosure provide an image generation method and device, an electronic device, and a storage medium to overcome the problem of fixed image content features and high content similarity of the generated image.
[0005] In a first aspect, embodiments of the present disclosure provide an image generation method, comprising:
[0006] obtaining style description information, the style description information being used to represent a target image style; calling an image generation model based on the style description information and object recognition information to generate a stylized image, wherein a subject object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style, the image generation model comprising a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model being used to generate the subject object in the stylized image, and the second low-rank adapter model being used to set the image style in the stylized image.
[0007] In a second aspect, embodiments of the present disclosure provide an image generation device, comprising:
[0008] an interaction module configured to obtain style description information, the style description information being used to represent a target image style;
[0009] The generating module is configured to invoke an image generation model based on the style description information and the object recognition information to generate a stylized image, wherein a subject object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style, and the image generation model comprises a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model is configured to generate the subject object in the stylized image, and the second low-rank adapter model is configured to set an image style in the stylized image.
[0010] In a third aspect, an electronic device is provided, and the electronic device comprises a processor and a memory.
[0011] The memory stores computer-executable instructions.
[0012] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image generation method according to the first aspect and various possible designs of the first aspect.
[0013] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the image generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0014] In a fifth aspect, a computer program product is provided, and the computer program product comprises a computer program. When a processor executes the computer program, the image generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0015] The image generation method, device, electronic device, and storage medium provided in this embodiment obtain style description information, which is used to characterize the style of a target image; based on the style description information and object recognition information, call an image generation model to generate a stylized image, wherein the main object in the stylized image has object features corresponding to the object recognition information, and the stylized image has style features corresponding to the style of the target image; the image generation model includes a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model is used to generate the main object in the stylized image, and the second low-rank adapter model is used to set the image style in the stylized image. By determining the target image style of the generated image based on the style description information, and then calling the image generation model to generate a stylized image based on the style description information and object recognition information, the capabilities of the first low-rank adapter model and the second low-rank adapter model in the image generation model are utilized to generate a stylized image with both personalized object features and specified style features, thereby improving the diversity and personalization of image content. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] Figure 1 A diagram of an application scenario of the image generation method provided in an embodiment of the present disclosure;
[0018] Figure 2 Schematic diagram of the process of the image generation method provided in the embodiment of the present disclosure Figure 1 ;
[0019] Figure 3 A schematic diagram of a process for generating a stylized image provided by an embodiment of the present disclosure;
[0020] Figure 4 for Figure 2 A flowchart of a specific implementation method of step S102 in the embodiment shown;
[0021] Figure 5 A schematic diagram of the structure of an image generation model provided in an embodiment of the present disclosure;
[0022] Figure 6 Schematic diagram of the process of the image generation method provided in the embodiment of the present disclosure Figure 1 ;
[0023] Figure 7 As Figure 6 a flowchart illustrating a specific implementation manner of step S203 in the embodiment shown in the figure;
[0024] Figure 8 a flowchart of a specific implementation manner of step S2030;
[0025] Figure 9 a structural block diagram of an image generation apparatus provided by an embodiment of the present disclosure;
[0026] Figure 10 a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0027] Figure 11 a hardware structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0030] The application scenarios of the embodiments of the present disclosure are explained as follows:
[0031] The image generation method provided by the embodiments of the present disclosure can be applied in an application program (APP, Application) with image generation and image editing functions, such as a camera application program, a video editing application program, an AI assistant application program capable of generating images, etc., or in a character image generation module of other application programs, such as a role creation module in a game application program. More specifically, it can be applied in various application scenarios of generating character images based on user requirements. The execution subject of the present embodiment can be a terminal device running the above-mentioned application program with image generation and image editing functions, or a server deploying a server corresponding to the above-mentioned application program, or other electronic devices with similar functions.
[0032] In some embodiments, the terminal device or the server can implement the image generation method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; it can be a local application program, i.e., a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, i.e., a program running based on a browser environment. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins, and the specific implementation form can be configured as needed. Further, in some embodiments, the server can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud function, network service, middleware service, domain name service, security service, content delivery network (CDN), and big data and artificial intelligence platform, etc. basic cloud computing services, wherein the cloud service can be an interactive processing service for calling by the terminal device.
[0033] Figure 1 An application scenario diagram of the image generation method provided by the embodiments of the present disclosure is shown in Figure 1 For example, in the case of taking the server as the execution subject, there is a target APP (for example, a game application program) running in the terminal device, and in the image generation interface of the target APP, the user can specify a specific image style by operation, and the specific operation mode includes inputting descriptive text or selecting from the optional items provided by the APP; then, the terminal device sends the corresponding style description information to the server according to the image style specified by the user, for example, the "Cyberpunk" style shown in the figure, the style description information can refer to the information corresponding to the above-mentioned descriptive text, or the identifier representing the image style, then the server generates a user's portrait according to the style information and returns it to the terminal device side for display. Then, the user can further create a virtual role, set an avatar, and perform subsequent steps based on the portrait. In another implementation manner, the execution subject can also be the terminal device, i.e., the terminal device alone completes the above-mentioned process of generating a portrait, which can be set as needed.
[0034] In the prior art, in an application scenario such as generating a character image, an application usually provides a setting option for an image style, so that a user can generate an image with different styles. However, such an image is usually generated on the basis of a fixed portrait picture, and since the content of the portrait picture is fixed, although the generated image has different image style features, the image content features are similar, the image content cannot be controlled, personalized portrait features cannot be realized, and the generated image has the problems of fixed image content features and high content similarity.
[0035] Embodiments of the present disclosure provide an image generation method to solve the above problems.
[0036] Reference Figure 2 , Figure 2 The flowchart of the image generation method provided by the embodiments of the present disclosure is shown in FIG. 1. Figure 1 The method of the present embodiment can be applied in a server, and the image generation method comprises the following steps.
[0037] Step S101: Obtain style description information, which is used to represent a target image style.
[0038] Step S102: Based on the style description information and the object recognition information, call an image generation model to generate a stylized image, wherein the main object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style. The image generation model comprises a first low-rank adapter model and a second low-rank adapter model. The first low-rank adapter model is used to generate the main object in the stylized image, and the second low-rank adapter model is used to set the image style in the stylized image.
[0039] Reference Figure 1The application scenario diagram shows that the server receives an image generation request sent by the terminal device, and obtains information for representing a specific image style, i.e., style description information, based on user operation in response to the image generation request. The style description information can include an identifier for representing an image style, for example, the style description information includes a style field. When the field value of the style field is 01, the target image style represented is "science fiction style"; when the field value of the style field is 02, the target image style represented is "cartoon style"; and when the field value of the style field is 00, the target image style represented is "no style" or "original style". In another possible implementation, the style description information can include multiple keywords or corresponding keyword identifiers, for example, the style description information includes keywords such as "grass", "snow", and "sunset". The target image style is represented by the keywords. Of course, when the style description information includes keyword identifiers corresponding to the keywords, the style description information can be represented in the form of a matrix or a vector, and no further examples are given.
[0040] The style description information can be generated on the terminal device side and sent to the server by the terminal device, for example, the style description information is included in the image generation request and sent to the server. The style description information can also be sent to other third-party devices by the terminal device, and then obtained by the server from the third-party devices in response to the image generation request. The specific implementation of the style description information and the specific implementation of the server obtaining the style description information can be set as needed, and are not limited herein.
[0041] Further, after obtaining the style description information through the image generation request, the server can further determine the corresponding sender information, i.e., object recognition information, through the image generation request. Specifically, for example, the application program running on the terminal device has a target user's account logged in. Before or at the same time as the terminal device sends the image generation request to the server through the application program, the terminal device sends the target user's account information to the server. Therefore, the server can determine the corresponding object recognition information according to the account information before or after receiving the image generation request.
[0042] Afterwards, the server calls the image generation model based on the style description information and object recognition information to generate a stylized image that matches the style description information and object recognition information, wherein the main object in the stylized image has object features corresponding to the object recognition information, and the stylized image has style features corresponding to the style of the target image. For example, the stylized image P1 generated by the image generation model is a single-person portrait image, in which the main object is the main person in the single-person portrait image. More specifically, the main person in the stylized image P1 matches the object recognition information, that is, the stylized image P1 is an image with significant user-personalized features, such as a self-portrait or portrait of the user User_1 corresponding to the object recognition information. At the same time, the stylized image P1 also has the target image style described by the above-mentioned style description information, specifically referring to the style corresponding to the background and clothing of the person in the stylized image P1. In other possible implementations, the stylized image P2 generated by the image generation model may also be an image containing a pet (such as a cat or dog). Similar to the stylized image P1, the pet in the stylized image P2 is the main object. The stylized image P2 also has the target image style described by the above-mentioned style description information, specifically referring to the style corresponding to the environmental background and the pet's clothing in the stylized image P2.
[0043] Figure 3 A schematic diagram of a process for generating a stylized image provided by an embodiment of the present disclosure is shown as follows: Figure 3 As shown, exemplarily, user User_1 inputs style description information representing the style of the target image on the terminal device side, and the content of the style description information is the text "cheerleaders". After that, the terminal device sends an image generation request containing the above style description information to the server. The server obtains the style description information and object recognition information based on the image generation request. After that, the image generation model M1 is called using the style description information and the object recognition information. After being processed by the image generation model M1, an image P1 is generated. The image P1 contains a "character" with a face similar to that of the user User_1, that is, it has object features corresponding to the object recognition information. At the same time, the "character" in the image P1 is wearing "cheerleaders" clothes, that is, the image P1 has style features corresponding to the style of the target image.
[0044] Further, the image generation model in the embodiment can be directly inputted with the style description information and the object recognition information, that is, the image generation model has a calling function or an interface for inputting the style description information and the object recognition information, so as to directly call the image generation model to generate the stylized image. Alternatively, the image generation model can be constructed through at least one of the style description information and the object recognition information, and then the image generation model is called. Alternatively, through at least one of the style description information and the object recognition information, a target image generation model is determined from a plurality of image generation models, and then the target image generation model is called. Alternatively, a combination of the above two methods is used.
[0045] In a possible implementation, the image generation model includes a control network model, a first low-rank adapter model, and a second low-rank adapter model, as shown in Figure 4 The specific implementation of step S102 includes the following steps.
[0046] Step S1021: Obtain a target image generation model corresponding to the object recognition information, and the target image generation model includes the first low-rank adapter model, the second low-rank adapter model, and the control network model.
[0047] Step S1022: Generate a first image through the first low-rank adapter model and the control network model, and the subject object in the first image has the object feature corresponding to the object recognition information.
[0048] Step S1023: Process the first image through the second low-rank adapter model and the style description information to generate a stylized image.
[0049] Exemplarily, first, a plurality of image generation models can be pre-installed in the server, after obtaining the object recognition information, first, according to the object recognition information, determine the target image generation model matched therewith. Wherein, the image generation model is an integrated model, which contains a control network (ControlNet) model, a first low-rank adapter model and a second low-rank adapter model. Specifically, the control network model and the low-rank adapter model are both plug-ins based on Stable Diffusion, and Stable Diffusion is a commonly used text-to-image model, wherein the control network model is a plug-in for controlling image generation, which is realized based on Conditional Generative Adversarial Networks (CGAN) technology, through which users can control the generated images more finely, thus obtaining images that meet user needs more. The low-rank adapter model is also called a large language model low-rank adapter (Lora) model, which allows model training with a small amount of training data, so as to quickly adjust the style of the output content or add new content without changing the underlying large language model. The specific principles of the above control network model and low-rank adapter model are prior art, which will not be described here.
[0050] In this embodiment, by integrating two low-rank adapter models (first low-rank adapter model, second low-rank adapter model) and a control network model, a stylized image is realized. As shown above, first, a first image based on the object recognition information is generated through the first low-rank adapter model and the control network model; wherein the first low-rank adapter model is generated by pre-training with a small amount of sample images, and the sample images are images of the target user corresponding to the object recognition information. Therefore, the first low-rank adapter model learned the appearance characteristics of the target user after training, so that the subject (main character) in the first image generated based on the control network model and the first low-rank adapter model has similar appearance characteristics to the target user. In short, the control network model and the first low-rank adapter model can output a portrait of the target user corresponding to the object recognition information. Then, the control network model and the second low-rank adapter model are used to process the first image with the style description information as input, change the image style of the first image based on the first image, and obtain a stylized image, so that the character object in the stylized image has the object characteristics corresponding to the object recognition information, and also has the style characteristics corresponding to the target image style.
[0051] Figure 5 A structural diagram of an image generation model provided by the embodiment of the present disclosure is shown in Figure 5As shown, the image generation model includes a pre-trained ID_Lora model (i.e., a first low-rank adapter model), a pre-trained Style_Lora model (i.e., a second low-rank adapter model), and a basic generation model (i.e., a control network model). The ID_Lora model and the Style_Lora model are integrated into the basic generation model to generate the image generation model. Then, based on the capabilities of the three models integrated in the image generation model, a stylized image with personalized object features and style features set by a user can be obtained according to input information (style description information and object recognition information).
[0052] Further, the first low-rank adapter model and the second low-rank adapter model in the embodiment are both multi-layer models. Different layers affect different content dimensions of the generated image. For example, the i-th layer of the low-rank adapter model has a greater impact on the facial expression of the person in the image, and the j-th layer of the low-rank adapter model has a greater impact on the clothing expression of the person in the image. Based on the above reasons, before generating a stylized image using the target image generation model, the embodiment further includes:
[0053] Step S100: obtaining at least one first model layer corresponding to the first low-rank adapter model and at least one second model layer corresponding to the second low-rank adapter model.
[0054] Correspondingly, the specific implementation of step S1022 includes generating a first image by the first model layer of the first low-rank adapter model and the control network model.
[0055] Correspondingly, the specific implementation of step S1023 includes processing the first image by the second model layer of the second low-rank adapter model and the style description information to generate a stylized image.
[0056] More specifically, due to the functional differences between different model layers, some model layers have a stronger correlation with object features, and some model layers have a stronger correlation with style features. Therefore, when using the first low-rank adapter model and the second low-rank adapter model, by selecting the first model layer in the first low-rank adapter model and the second model layer in the second low-rank adapter model to perform corresponding image processing steps, the purpose of better expressing object features and style features can be achieved, thereby improving the image quality of the finally generated stylized image. The first model layer and the second model layer can be pre-set, and the specific implementation is not described here.
[0057] In the step of the embodiment, the style description information is obtained, and the style description information is used to represent the target image style; based on the style description information and the object recognition information, an image generation model is called to generate a stylized image, wherein the subject object in the stylized image has the object characteristics corresponding to the object recognition information, and the stylized image has the style characteristics corresponding to the target image style. By determining the target image style of the generated image according to the style description information, and then calling the image generation model to generate the stylized image based on the style description information and the object recognition information, the image generation model is used to generate the stylized image with personalized object characteristics and specified style characteristics, thereby improving the diversity and personalization degree of image content.
[0058] Reference Figure 6 , Figure 6 The flowchart of the image generation method provided by the embodiment of the present disclosure is shown in Figure 1 The embodiment is based on the embodiment shown in Figure 2 The embodiment is based on the embodiment shown in
[0059] Step S201: Obtain at least two subject object images corresponding to the object recognition information.
[0060] Step S202: Train a first low-rank adapter model according to the at least two subject object images, and the first low-rank adapter model is used to generate an image with the object characteristics of the subject object in the subject object image.
[0061] Exemplarily, the embodiment is based on the embodiment shown in Figure 2 In order to enable the image generation model to have the ability to output the stylized image, the model needs to be constructed for the object recognition information, that is, each user has an image generation model corresponding to the user, which is used to generate the stylized image of the user. The image generation model includes the first low-rank adapter model, which is first trained. Specifically, first, training samples are obtained, and the training samples include at least two subject object images. In a possible implementation, the subject object image is a portrait corresponding to the object recognition information, or a picture containing pets such as cats and dogs. The subject object image can be uploaded by the user through the terminal device, that is, for example, the user User_1 uploads N personal photos to the server through the terminal device, and then the server trains the corresponding first low-rank adapter model Lora_01 based on the N personal photos uploaded by the user User_1, and further constructs the image generation model M_001 belonging to the user User_1, and in the subsequent step, the image generation model M_001 is used to generate the stylized image with the portrait characteristics of the user User_1.
[0062] The general implementation process of training the low-rank adapter model based on the training samples will not be described in detail here.
[0063] Optionally, before step S202, there is also:
[0064] Step S201A: Set the subject object image to a first image size.
[0065] Step S201B: Based on a random cutting strategy, cut the subject object image into a target number of sample blocks, the target number being determined based on the first image size.
[0066] For example, in a possible implementation, before training the first low-rank adapter model, the training samples can be preprocessed. Specifically, the at least two subject object images in the training samples and the sample labels representing the object recognition information corresponding to the subject object images are preprocessed. Then, before using the training samples, the subject object images are first set to a first image size, for example, the first image size is set to 768*576. Then, based on a random cutting (randomcrop) strategy, the subject object images are cut into a target number of sample blocks. Specifically, the above process can be implemented by first setting the target number and then cutting the subject object images into sample blocks of the target number. Alternatively, the above process can be implemented by first setting a target size and then cutting the subject object images into sample blocks of the target size. For example, the sexualized image is cut into sample blocks of 256*256.
[0067] Correspondingly, in another possible implementation, the specific implementation of step S202 includes training the initial low-rank adapter model based on the target number of sample blocks until the initial low-rank adapter model converges to the first low-rank adapter model.
[0068] In the steps of this embodiment, by randomly cutting the sexualized image of the first image size, sample blocks of a target size are obtained, and then the first low-rank adapter model is trained based on the sample blocks, thereby reducing the degree of overfitting and alleviating the conflict between the first low-rank adapter model and the second low-rank adapter model. At the same time, the training time is shortened.
[0069] Step S203: Construct an image generation model according to the first low-rank adapter model.
[0070] For example, after obtaining the first low-rank adapter model by training, the first low-rank adapter model and the pre-trained second low-rank adapter model are integrated into the control network model as the base model, and the construction of the image generation model is realized.
[0071] For example, as shown in Figure 7 The specific implementation of step S203 includes:
[0072] Step S2031: obtaining a pre-trained second low-rank adapter model, the second low-rank adapter model having a first model input parameter, the second low-rank adapter model being configured to generate an image of a corresponding image style according to the first model input parameter.
[0073] Step S2032: integrating the first low-rank adapter model and the second low-rank adapter model based on the control network model to construct an image generation model.
[0074] Exemplarily, similar to the first low-rank adapter model, the second low-rank adapter model needs to be trained before use. The training step for the second low-rank adapter model can be offline training, i.e., the second low-rank adapter model is generated after training on other devices. The second low-rank adapter model has a first model input parameter, which is configured to receive style description information, and then control the model to perform image style transfer on the input image based on the target image style represented by the style description information, thereby generating an image of the target image style. Then, the first low-rank adapter model and the pre-trained second low-rank adapter model are integrated into the control network model as the base model to obtain the construction of the image generation model. The specific integration process is, for example, connecting the first low-rank adapter model and the first low-rank adapter model to the control network model respectively to form the image generation model; or connecting the control network model, the first low-rank adapter model, and the second low-rank adapter model in series to form the image generation model. The specific implementation manner can be set as needed, which will not be described here.
[0075] Further, optionally, before step S2031, it further includes:
[0076] Step S2030: training the second low-rank adapter model.
[0077] Exemplarily, as shown in Figure 8 , the specific implementation manner of step S2030 includes:
[0078] Step S2030-1: obtaining high-quality style data, the high-quality style data including high-quality style images and corresponding style description texts.
[0079] Step S2030-2: parsing the style description text into at least one style description word, the style description word being configured to represent a style feature constituting the image style.
[0080] Exemplarily, for the second low-rank adapter model, high-quality style data needs to be used for training, so that it learns the style features in the image. In the high-quality style data, high-quality style images and corresponding style description texts are included, which are used to describe the image style of the high-quality style images. By using high-quality style data for training, the second low-rank adapter model can learn the mapping relationship between image style and image features.
[0081] In a possible implementation, before the server trains the second low-rank adapter model based on the high-quality style data, the server further analyzes the active style description text to generate one or more style description words. For example, the style description text is “outdoor style”, and the server analyzes it into multiple style description words “grass”, “green trees”, “blue sky”, “stream”, etc. Then, the model is trained by using the above style description words, which is equivalent to concretizing and liking the style features. This can further improve the output control of the second low-rank adapter model, and thus make the generated stylized images have better quality and accuracy.
[0082] Further, optionally, the method further comprises:
[0083] The high-quality style image is detected, at least one supplementary style description word is added to the high-quality style image, and / or at least one abnormal style description word is deleted.
[0084] Specifically, referring to the implementation process of step S2030, after the server generates one or more style description words by analyzing the style description text, it can further add additional style description words, i.e., supplementary style description words, to the sample information to improve the sample quality. For example, the style description text corresponding to the high-quality style image is “outdoor style”, and the server analyzes it into multiple style description words “grass”, “green trees”, and “blue sky”. Then, by checking the content in the high-quality style image (this process may need to call other image recognition models), additional style description words are determined, wherein exemplarily, the supplementary style description words are used to represent at least one of the following style features: color of an object in the image; posture of an object in the image; spatial position relationship between at least two objects in the image. Specifically, for example, “blue-green stream”, “flying sparrow”, etc., i.e., at least one supplementary style description word, so as to expand the style description words, make the label information corresponding to the high-quality style image more abundant, and thus provide the quality of the high-quality style image as a training sample.
[0085] On the other hand, the server can eliminate one or more generated style description words, for example, eliminate style description words with ambiguous meanings, that is, abnormal style description words, such as "traditional", "day", and the like. For the determination of the above-mentioned abnormal style description words, the determination can be based on a preset part-of-speech rule, or can be realized based on a pre-trained identification model, which is not limited here.
[0086] In the embodiment, by further supplementing or eliminating the abnormal style prompt words after analyzing the style description text into at least one style description word, the accuracy and richness of the label information corresponding to the high-quality style image are improved, the sample quality of the training sample is improved, and the training effect of the second low-rank adapter model is improved.
[0087] Step S204: Obtain style description information, the style description information being used to represent a target image style.
[0088] Step S205: Based on the style description information and the object recognition information, call an image generation model to generate a stylized image, wherein the subject object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style.
[0089] In the embodiment, the implementation manners of steps S204-S205 are the same as those of steps S101-S102 in the embodiment of the disclosure Figure 2 The implementation manners of steps S204-S205 in the embodiment are the same as those of steps S101-S102 in the embodiment of the disclosure
[0090] Corresponding to the image generation method of the above embodiment, Figure 9 A structural block diagram of an image generation apparatus provided by the embodiment of the disclosure is shown in FIG. 3. The method introduced in the above embodiment can be executed by the image generation apparatus. The apparatus can be realized in the form of software and / or hardware. The apparatus can be integrated in an electronic device with certain data processing function. The electronic device can include but is not limited to a mobile terminal with large data processing capacity, and a fixed terminal with large data processing capacity such as a desktop computer and a supercomputer.
[0091] For the convenience of description, only parts related to the embodiments of the disclosure are shown. For details, refer to Figure 9 The image generation apparatus 3 includes:
[0092] An interaction module 31 is configured to obtain style description information, the style description information being used to represent a target image style.
[0093] The generation module 32 is used to call the image generation model based on the style description information and the object recognition information to generate a stylized image, wherein the main object in the stylized image has object features corresponding to the object recognition information, and the stylized image has style features corresponding to the target image style. The image generation model includes a first low-rank adapter model and a second low-rank adapter model. The first low-rank adapter model is used to generate the main object in the stylized image, and the second low-rank adapter model is used to set the image style in the stylized image.
[0094] According to one or more embodiments of the present disclosure, the generation module 32 is specifically used to: obtain a target image generation model corresponding to the object recognition information, the target image generation model including a first low-rank adapter model, a second low-rank adapter model and a control network model; generate a first image through the first low-rank adapter model and the control network model, and the main object in the first image has object features corresponding to the object recognition information; process the first image through the second low-rank adapter model and the style description information to generate a stylized image.
[0095] According to one or more embodiments of the present disclosure, the generation module 32 is further configured to: obtain at least one first model layer corresponding to the first low-rank adapter model, and at least one second model layer corresponding to the second low-rank adapter model; when the generation module 32 generates the first image through the first low-rank adapter model, the generation module 32 is specifically configured to: generate the first image through the first model layer of the first low-rank adapter model; when the generation module 32 processes the first image through the second low-rank adapter model and the style description information to generate a stylized image, the generation module 32 is specifically configured to: generate the stylized image through the second model layer of the second low-rank adapter model;
[0096] According to one or more embodiments of the present disclosure, the generation module 32 is further used to: obtain at least two subject object images corresponding to object recognition information; train a first low-rank adapter model based on the at least two subject object images, and the first low-rank adapter model is used to generate an image having object features of the subject object in the subject object image; and construct an image generation model based on the first low-rank adapter model.
[0097] According to one or more embodiments of the present disclosure, when constructing an image generation model based on the first low-rank adapter model, the generation module 32 is specifically used to: obtain a pre-trained second low-rank adapter model, the second low-rank adapter model having a first model input parameter, and the second low-rank adapter model is used to generate an image of a corresponding image style based on the first model input parameter; based on the control network model, integrate the first low-rank adapter model and the second low-rank adapter model to construct an image generation model.
[0098] According to one or more embodiments of the present disclosure, before obtaining the pre-trained second low-rank adapter model, the generation module 32 is further configured to: obtain high-quality style data, the high-quality style data including high-quality style images and corresponding style description texts; parse the style description texts into at least one style description word, the style description word being used to represent one style feature constituting an image style; and train the initialized low-rank adapter model by using the high-quality style images in the high-quality style data and the corresponding at least one style description word, to generate the second low-rank adapter model.
[0099] According to one or more embodiments of the present disclosure, the generation module 32 is further configured to: add at least one supplementary style description word to the high-quality style image by detecting the high-quality style image, and / or delete at least one abnormal style description word; wherein the supplementary style description word is used to represent at least one of the following style features: color of an object in the image; posture of the object in the image; and spatial position relationship between at least two objects in the image.
[0100] According to one or more embodiments of the present disclosure, before training the first low-rank adapter model, the generation module 32 is further configured to: set the subject object image to a first image size; and split the subject object image into a target number of sample blocks based on a random cutting strategy, the target number being determined based on the first image size; and when training the first low-rank adapter model based on the at least two subject object images, the generation module 32 is specifically configured to: train the initial low-rank adapter model based on the target number of sample blocks, until the initial low-rank adapter model converges to the first low-rank adapter model.
[0101] The interaction module 31 and the generation module 32 are connected. The image generation apparatus 3 provided in this embodiment can execute the technical solutions of the above-mentioned method embodiments, and has similar implementation principles and technical effects, which will not be described here again.
[0102] Figure 10 A structural schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown in FIG. 4, which includes: Figure 10
[0103] a processor 41 and a memory 42 connected with the processor 41;
[0104] The memory 42 stores computer execution instructions.
[0105] The processor 41 executes the computer execution instructions stored in the memory 42, to implement the image generation method in the embodiment shown in FIG. 3. Figures 2-8
[0106] Optionally, the processor 41 and the memory 42 are connected through a bus 43.
[0107] The relevant description can be referred to Figures 2-8 The relevant description and effects of the steps in the corresponding embodiments are understood, and will not be described in detail here.
[0108] The embodiment of the present disclosure provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are used for realizing the present disclosure Figures 2-8 The image generation method provided by any of the embodiments of the corresponding embodiments.
[0109] The embodiment of the present disclosure provides a computer program product, including a computer program, and the computer program is executed by a processor to realize the present disclosure Figures 2-8 The image generation method provided by any of the embodiments of the corresponding embodiments.
[0110] In order to realize the above-mentioned embodiments, the embodiment of the present disclosure further provides an electronic device.
[0111] Reference Figure 11 , which shows a structural schematic diagram of an electronic device 900 suitable for realizing the embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers (PAD), portable media players (PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 11 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0112] As Figure 11 shown, the electronic device 900 can include a processing device (such as a central processor, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 902 or programs loaded from a storage device 908 to a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0113] In general, the following devices can be connected to the I / O interface 905: input devices 906, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 907, including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, and the like; storage devices 908, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate wirelessly or via a wire with other devices to exchange data. Although Figure 11 The electronic device 900 is shown with various devices, but it is understood that all of the illustrated devices are not required to implement or be present. More or fewer devices can alternatively be implemented or present.
[0114] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication devices 909, or installed from the storage devices 908, or installed from the ROM 902. When the computer program is executed by the processing devices 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0115] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0116] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.
[0117] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0118] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0119] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or by combinations of dedicated hardware and computer instructions.
[0120] The units or modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself.
[0121] The functions described in this specification can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0123] In a first aspect, according to one or more embodiments of the present disclosure, an image generation method is provided, comprising:
[0124] obtaining style description information, the style description information being used to represent a target image style; calling an image generation model based on the style description information and object recognition information to generate a stylized image, wherein a subject object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style, the image generation model comprising a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model being used to generate the subject object in the stylized image, and the second low-rank adapter model being used to set an image style in the stylized image.
[0125] According to one or more embodiments of the present disclosure, the calling of the image generation model based on the style description information and the object recognition information to generate the stylized image comprises: obtaining a target image generation model corresponding to the object recognition information, the target image generation model comprising a first low-rank adapter model, a second low-rank adapter model, and a control network model; generating a first image through the first low-rank adapter model and the control network model, a subject object in the first image having an object feature corresponding to the object recognition information; and processing the first image through the second low-rank adapter model and the style description information to generate the stylized image.
[0126] According to one or more embodiments of the present disclosure, the method further comprises: obtaining at least one first model layer corresponding to the first low-rank adapter model, and at least one second model layer corresponding to the second low-rank adapter model; the generating, by the first low-rank adapter model, the first image comprises: generating, by the first model layer of the first low-rank adapter model, the first image; and the processing, by the second low-rank adapter model and the style description information, the first image to generate the stylized image comprises: generating, by the second model layer of the second low-rank adapter model, the stylized image.
[0127] According to one or more embodiments of the present disclosure, the method further comprises: obtaining at least two subject object images corresponding to the object recognition information; training a first low-rank adapter model according to the at least two subject object images, the first low-rank adapter model being used to generate an image with an object feature of a subject object in the subject object image; and constructing the image generation model according to the first low-rank adapter model.
[0128] According to one or more embodiments of the present disclosure, the constructing the image generation model according to the first low-rank adapter model comprises: obtaining a pre-trained second low-rank adapter model, the second low-rank adapter model having a first model input parameter, the second low-rank adapter model being used to generate an image with a corresponding image style according to the first model input parameter; and integrating the first low-rank adapter model and the second low-rank adapter model to construct the image generation model based on a control network model.
[0129] According to one or more embodiments of the present disclosure, before the obtaining the pre-trained second low-rank adapter model, the method further comprises: obtaining high-quality style data, the high-quality style data including high-quality style images and corresponding style description texts; parsing the style description texts into at least one style description word, the style description word being used to represent a style feature constituting the image style; and training an initialized low-rank adapter model by the high-quality style images in the high-quality style data and the corresponding at least one style description word to generate the second low-rank adapter model.
[0130] According to one or more embodiments of the present disclosure, the method further comprises: adding at least one supplementary style description word to the high-quality style image and / or deleting at least one abnormal style description word by detecting the high-quality style image; wherein the supplementary style description word is used to represent at least one of the following style features: a color of an object in the image; a pose of an object in the image; and a spatial position relationship between at least two objects in the image.
[0131] According to one or more embodiments of the present disclosure, before the training of the first low-rank adapter model, further comprising: setting the subject object image to a first image size; based on a random cutting strategy, cutting the subject object image into a target number of sample blocks, the target number being determined based on the first image size; and the training of the first low-rank adapter model based on the at least two subject object images comprises: training an initial low-rank adapter model based on the target number of sample blocks until the initial low-rank adapter model converges to the first low-rank adapter model.
[0132] In a second aspect, according to one or more embodiments of the present disclosure, an image generation apparatus is provided, comprising:
[0133] An interaction module configured to obtain style description information, the style description information being used to represent a target image style;
[0134] A generation module configured to invoke an image generation model based on the style description information and object recognition information to generate a stylized image, wherein a subject object in the stylized image has an object feature corresponding to the object recognition information, and the stylized image has a style feature corresponding to the target image style, the image generation model comprising a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model being used to generate the subject object in the stylized image, and the second low-rank adapter model being used to set an image style in the stylized image.
[0135] According to one or more embodiments of the present disclosure, the generation module is specifically configured to: obtain a target image generation model corresponding to the object recognition information, the target image generation model comprising a first low-rank adapter model, a second low-rank adapter model, and a control network model; generate a first image through the first low-rank adapter model and the control network model, a subject object in the first image having an object feature corresponding to the object recognition information; and process the first image through the second low-rank adapter model and the style description information to generate the stylized image.
[0136] According to one or more embodiments of the present disclosure, the generation module is further configured to: obtain at least one first model layer corresponding to the first low-rank adapter model and at least one second model layer corresponding to the second low-rank adapter model; when generating the first image by using the first low-rank adapter model, the generation module is specifically configured to generate the first image by using the first model layer of the first low-rank adapter model; and when processing the first image by using the second low-rank adapter model and the style description information to generate the stylized image, the generation module is specifically configured to generate the stylized image by using the second model layer of the second low-rank adapter model.
[0137] According to one or more embodiments of the present disclosure, the generation module is further configured to: obtain at least two subject object images corresponding to the object recognition information; train a first low-rank adapter model according to the at least two subject object images, the first low-rank adapter model being used to generate an image with an object feature of a subject object in the subject object image; and construct the image generation model according to the first low-rank adapter model.
[0138] According to one or more embodiments of the present disclosure, when constructing the image generation model according to the first low-rank adapter model, the generation module is specifically configured to: obtain a pre-trained second low-rank adapter model, the second low-rank adapter model having a first model input parameter, the second low-rank adapter model being used to generate an image with a corresponding image style according to the first model input parameter; and integrate the first low-rank adapter model and the second low-rank adapter model to construct the image generation model based on a control network model.
[0139] According to one or more embodiments of the present disclosure, before the generation module obtains the pre-trained second low-rank adapter model, the generation module is further configured to: obtain high-quality style data, the high-quality style data including high-quality style images and corresponding style description texts; parse the style description texts into at least one style description word, the style description word being used to represent a style feature constituting the image style; and train an initialized low-rank adapter model by using the high-quality style images in the high-quality style data and the corresponding at least one style description word to generate the second low-rank adapter model.
[0140] According to one or more embodiments of the present disclosure, the generation module is further configured to: add at least one supplementary style description word to the high-quality style image and / or delete at least one abnormal style description word by detecting the high-quality style image; wherein the supplementary style description word is used to represent at least one of the following style features: a color of an object in an image; a pose of an object in an image; and a spatial position relationship between at least two objects in an image.
[0141] According to one or more embodiments of the present disclosure, before the training of the first low-rank adapter model, the generation module is further configured to: set the subject object image to a first image size; and split the subject object image into a target number of sample blocks based on a random cutting strategy, the target number being determined based on the first image size; and when training the first low-rank adapter model based on the at least two subject object images, the generation module is specifically configured to: train an initial low-rank adapter model based on the target number of sample blocks until the initial low-rank adapter model converges to the first low-rank adapter model.
[0142] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;
[0143] The memory stores computer-executable instructions.
[0144] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the image generation method according to the first aspect and various possible designs of the first aspect.
[0145] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions, when a processor executes the computer-executable instructions, the image generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0146] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program, when a processor executes the computer program, the image generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0147] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0148] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0149] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An image generation method, characterized in that: include: Acquire style description information, where the style description information is used to characterize the style of the target image; Based on the style description information and object recognition information, an image generation model is called to generate a stylized image, wherein the main object in the stylized image has object features corresponding to the object recognition information, and the stylized image has style features corresponding to the target image style, and the image generation model includes a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model is used to generate the main object in the stylized image, and the second low-rank adapter model is used to set the image style in the stylized image.
2. The method according to claim 1, characterized in that Based on the style description information and the object recognition information, calling an image generation model to generate a stylized image includes: Acquire a target image generation model corresponding to the object recognition information, wherein the target image generation model further includes a control network model; generating a first image by using the first low-rank adapter model and the control network model, wherein a subject object in the first image has an object feature corresponding to the object recognition information; The first image is processed using the second low-rank adapter model and the style description information to generate the stylized image.
3. The method according to claim 2, characterized in that The method further comprises: Obtain at least one first model layer corresponding to the first low-rank adapter model, and at least one second model layer corresponding to the second low-rank adapter model; Generating a first image by using the first low-rank adapter model includes: generating the first image through a first model layer of the first low-rank adapter model; The step of processing the first image using the second low-rank adapter model and the style description information to generate the stylized image includes: The stylized image is generated by a second model layer of the second low-rank adapter model.
4. The method according to claim 1, wherein The method further comprises: Obtaining at least two subject object images corresponding to the object recognition information; training a first low-rank adapter model based on the at least two subject object images, wherein the first low-rank adapter model is used to generate an image having object features of the subject object in the subject object images; The image generation model is constructed based on the first low-rank adapter model.
5. The method according to claim 4, characterized in that The constructing the image generation model according to the first low-rank adapter model includes: Obtaining a pre-trained second low-rank adapter model, the second low-rank adapter model having a first model input parameter, and the second low-rank adapter model is used to generate an image corresponding to the image style according to the first model input parameter; Based on the control network model, the first low-rank adapter model and the second low-rank adapter model are integrated to construct the image generation model.
6. The method according to claim 5, characterized in that Before obtaining the pre-trained second low-rank adapter model, the method further includes: Acquire high-quality style data, wherein the high-quality style data includes a high-quality style image and corresponding style description text; parsing the style description text into at least one style description word, where the style description word is used to represent a style feature constituting the style of the image; The initialized low-rank adapter model is trained using a high-quality style image and at least one corresponding style descriptor in the high-quality style data to generate the second low-rank adapter model.
7. The method according to claim 6, characterized in that The method further comprises: By detecting the high-quality style image, adding at least one supplementary style descriptor to the high-quality style image, and / or deleting at least one abnormal style descriptor; The supplementary style descriptor is used to represent at least one of the following style features: The color of the object in the image; the posture of the object in the image; the spatial position relationship between at least two objects in the image.
8. The method according to claim 4, characterized in that Before training the first low-rank adapter model, the method further includes: Setting the subject object image to a first image size; Based on a random cutting strategy, the subject object image is divided into a target number of sample blocks, where the target number is determined based on the first image size; The step of training a first low-rank adapter model based on the at least two subject object images comprises: Based on the target number of sample blocks, the initial low-rank adapter model is trained until the initial low-rank adapter model converges to the first low-rank adapter model.
9. An image generating device, characterized in that: include: An interaction module, configured to obtain style description information, wherein the style description information is used to characterize the style of a target image; A generation module is used to call an image generation model based on the style description information and object recognition information to generate a stylized image, wherein the main object in the stylized image has object features corresponding to the object recognition information, and the stylized image has style features corresponding to the target image style, and the image generation model includes a first low-rank adapter model and a second low-rank adapter model, the first low-rank adapter model is used to generate the main object in the stylized image, and the second low-rank adapter model is used to set the image style in the stylized image.
10. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image generation method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the image generation method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image generating method according to any one of claims 1 to 8 is implemented.