Image generation method and device, equipment and medium
By integrating image functional requirements to determine the image processing method, the shortcomings of image generation models in the prior art in high-resolution image generation and detail restoration are solved, the image generation quality and model generalization capabilities are improved, and more accurate image presentation is achieved.
Patent Information
- Application Number
- CN202510131630.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-27
AI Technical Summary
When generating high-resolution images, existing image generation models are difficult to maintain the clarity of details, resulting in blurring or distortion of images after enlargement, and lack of details restoration, making it difficult to accurately present subtle features and complex scenes in text descriptions.
By acquiring the image and text data to be processed and corresponding image generation requirements, the image generation method of the corresponding image generation model is determined, and the image and text data to be processed are input to the image generation model, and the target image is generated according to the image generation method corresponding to the image generation model. Integrating image functional requirements to determine the corresponding image processing methods, improve the quality of image generation and improve the generalization ability and adaptability of the model.
It improves the quality of image generation, improves the generalization ability and adaptability of the model, and can present subtle features and complex scenes in text descriptions more accurately, reducing the blur or distortion of images after enlargement.
Smart Images

Figure CN120047564A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, device and medium. Background Art
[0002] Image generation models are playing an increasingly important role in the field of artificial intelligence, and they play a key role in multiple online services and creative industries, such as digital art creation, content automatic generation platforms and social media.
[0003] In film and television production software, an image generation model can automatically generate corresponding scene images according to the script description or plot summary, greatly improving the creation efficiency and the expressiveness of the works. With the continuous growth of content creation needs, the creative challenges faced by creators are also increasing day by day. Image generation models are particularly important in helping creators quickly visualize their ideas.
[0004] Although image generation models show strong potential and application value in multiple fields, there are still some significant defects and challenges. First of all, the quality and details of the generated images are insufficient. Especially when generating high-resolution images, it is difficult to maintain the clarity of the details, resulting in blurring or distortion of the images after enlargement. At the same time, the model still has deficiencies in detail restoration, and it is difficult to accurately present the subtle features and complex scenes in the text description. Summary of the Invention
[0005] The present invention provides an image generation method, apparatus, device and medium. By integrating the image function requirements to determine the corresponding image processing method for processing, the quality of image generation is improved, and the generalization ability and adaptability of the model are enhanced.
[0006] According to one aspect of the present invention, there is provided an image generation method, including:
[0007] Obtaining the text-image data to be processed and the corresponding image generation requirements;
[0008] Determining the image generation method of the corresponding image generation model according to the image generation requirements;
[0009] Inputting the text-image data to be processed into the image generation model, and generating a target image according to the image generation method corresponding to the image generation model.
[0010] According to another aspect of the present invention, there is provided an image generation apparatus, including:
[0011] A data acquisition module, configured to obtain the text-image data to be processed and the corresponding image generation requirements;
[0012] A generation method determination module, configured to determine an image generation method of a corresponding image generation model according to the image generation requirement;
[0013] A target image generation module, configured to input the to-be-processed text-image data into the image generation model, and generate a target image according to the image generation method corresponding to the image generation model.
[0014] According to another aspect of the present invention, there is provided an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the image generation method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the image generation method according to any embodiment of the present invention when executed.
[0019] In the technical solution of the embodiment of the present invention, by obtaining the to-be-processed text-image data and the corresponding image generation requirement; determining the image generation method of the corresponding image generation model according to the image generation requirement; inputting the to-be-processed text-image data into the image generation model, and generating a target image according to the image generation method corresponding to the image generation model. With this technical solution, by integrating the image function requirements to determine the corresponding image processing method for processing, the quality of image generation is improved, and the generalization ability and adaptability of the model are enhanced.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0022] Figure 1 is a flowchart of an image generation method according to Embodiment 1 of the present invention;
[0023] Figure 2 It is an example diagram of an image generation model framework provided according to Embodiment 1 of the present invention;
[0024] Figure 3 It is a flowchart of an image generation method provided according to Embodiment 2 of the present invention;
[0025] Figure 4 It is a schematic structural diagram of an image generation device provided according to Embodiment 3 of the present invention;
[0026] Figure 5 It is a schematic structural diagram of an electronic device provided according to Embodiment 4 of the present invention. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] Figure 1 It is a flowchart of an image generation method provided according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of image generation for graphic and text data. This method can be executed by an image generation device, and the image generation device can be implemented in the form of hardware and / or software. The image generation device can be configured in an electronic device with data processing capabilities. As Figure 1 shown, the method includes:
[0031] S110. Obtain the text-image data to be processed and the corresponding image generation requirements.
[0032] Among them, the text-image data to be processed can be considered as the data that needs to be processed by the image generation model. In this embodiment, the text-image data to be processed usually includes the original image or the original text related thereto, and may also include the reference image. In this embodiment, the text-image data to be processed may include the original image data, text data, or text-image combined data, and may also include data such as the reference image. In this embodiment, there is no limitation on the specific input text-image data to be processed, and it can be one or more. It can be understood that the more detailed the input text-image data to be processed, the higher the quality of the generated image. The image generation requirements can be considered as the functional requirements for the image to be generated. In this embodiment, the image generation requirements may include the image style, image enhancement, image category, or other requirements, etc. In this embodiment, the corresponding image generation requirements can be configured according to the actual business needs of the user. In this embodiment, the text-image data to be processed input by the user and its corresponding image generation functional requirements can be obtained.
[0033] In this embodiment, optionally, the text-image data to be processed includes the text data to be processed, picture data, or reference image.
[0034] Among them, the text data can be considered as the descriptive text, label, or annotation of the image to be generated, etc. In this embodiment, the text content included in the specific text data may be related to the image content, and may include the name, attributes, and relationships of the object, etc.; it may also include the text format, etc. The picture data in this embodiment can be considered as the original image to be processed, and may include the image type, image content, or image quality, etc. The image type in this embodiment may be types such as photos, illustrations, charts, etc., and may include the specific image content, such as describing specific objects, scenes, or events. The image quality may refer to the resolution, color depth, or clarity, etc. The reference image can be considered as the reference example image corresponding to the image to be generated. In this embodiment, the corresponding reference image can be determined according to the actual needs. The text-image data to be processed in this embodiment may include at least one of the text data to be processed, picture data, or reference image. Through such a setting in this embodiment, the corresponding generated image can be obtained by combining multiple different data to be processed with the image generation model, improving the diversity of the input object.
[0035] S120. Determine the image generation method of the corresponding image generation model according to the image generation requirements.
[0036] Among them, the image generation model can be a pre-trained image generation module, which can be used to generate images that meet the requirements. The image generation model in this embodiment can be an improved text-to-image model. The image generation model architecture in this embodiment can include a basic image generation method and an advanced image generation method, and can execute different image generation methods according to different image generation requirements.
[0037] Specifically, the image generation model in this embodiment can include parts for processing basic image generation and high-order image generation workflows, improving the accuracy and customization ability of image generation. Figure 2 It is an example diagram of an image generation model framework provided in this embodiment. As Figure 2As shown in the figure, in the framework of this image generation model, the image generation part is divided into a basic image generation module and an advanced image generation module. The basic image generation module is responsible for basic image generation tasks and can use base models such as Stable Diffusion XL and Flux for image generation. The processing method of the advanced image generation module is more complex, and it can include the generation capabilities of various media forms such as images, videos, and animated gifs. Specifically, various functions such as super-resolution technology and ComfyUI workflows are integrated in the advanced image generation module, and these modules together improve the quality and diversity of image generation. The image training module is also an indispensable part. It uses widely used image generation base models such as Stable Diffusion XL and Flux for image training to further optimize the effect of the generated images. The processes and working methods of the entire system are uniformly managed and optimized by the framework to ensure the coordinated and efficient operation of each module. Regarding the image training module, in this embodiment, a method of integrating multiple pre-trained large models and using the diffuser library and accelerate library for LoRA training is adopted. First, the training image set to be trained is processed. The training image set can meet the training requirements with a relatively small amount of data, such as 5-20 pictures. Then, the image reverse-push algorithm model is used for cropping and labeling, and different training parameters are configured by specifying different training base models on the algorithm logic side, so as to achieve the training effect of the corresponding base model and provide users with a customized training path. The algorithm process of the image training module in this embodiment specifically includes: The first step is cropping and labeling. The training image set input by the user is pre-processed. After cropping and scaling to the specified size (generally the size for base model training), the image reverse-push algorithm model is used for labeling processing. The second step is to configure the training base model and training parameters. According to the different configured base models, the corresponding configurable training parameters are different, and the number of training times is determined according to the requirements. The third step is to perform LoRA training using the cropped and labeled training set and the configured parameters. Finally, according to the trained model, the trained model is notified to the basic image generation module and the advanced image generation module in the image generation model through pipeline operations. In this embodiment, by adopting a solution of integrating multiple pre-trained large models and using the diffuser library and accelerate library for LoRA training, the training logic is integrated and simplified, and more attention can be paid to the training effect.
[0038] In this embodiment, optionally, the image generation requirements include the first image generation requirement and the second image generation requirement; correspondingly, determining the image generation method of the corresponding image generation model according to the image generation requirements includes: determining the first image generation method of the corresponding image generation model according to the first image generation requirement; determining the second image generation method of the corresponding image generation model according to the second image generation requirement.
[0039] Among them, the first image generation requirement can be regarded as the basic image generation requirement. The second image generation requirement can be regarded as the advanced image generation requirement. It can be understood that the first image generation requirement is only to generate a single image requirement, and the second image generation requirement can be an image requirement that needs to generate multiple media forms. The corresponding generation requirement can be selected according to actual needs. The first image generation method can be regarded as the generation method corresponding to the basic image. The second image generation method can be regarded as the generation method corresponding to the advanced image. In this embodiment, the corresponding image generation method can be determined according to different requirements. In this embodiment, the workflows executed by the first image generation method and the second image generation method can be different; the first image generation method can be an image generation method that executes a single function according to the basic image generation function; the second image generation method can be an image production method corresponding to multiple selected functions, or other configured settings, which can be determined according to actual needs.
[0040] In this embodiment, the basic image generation method of the corresponding image generation model can be determined according to the first image generation requirement in the image generation requirement, so as to execute the generation of the corresponding target image according to the basic image generation method; the advanced image generation method of the corresponding image generation model can also be determined according to the second image generation requirement in the image generation requirement, so that the image generation model executes the generation of the corresponding target image according to the advanced image generation method. Through such a setting in this embodiment, different image generation methods can be selected according to the actual image generation requirements, improving the convenience and diversity of image generation.
[0041] S130. Input the text and image data to be processed into the image generation model, and generate a target image according to the image generation method corresponding to the image generation model.
[0042] Among them, the target image can be regarded as the image finally generated based on the corresponding image generation method determined according to different requirements. It can be understood that the target images generated by different image generation methods in this embodiment can be different. In this embodiment, the text and image data to be processed can be input into the image generation model, and the corresponding workflow can be executed according to the pre-determined image generation method, and then the corresponding target image can be generated. The basic image generation method corresponding to the image generation model in this embodiment can generate a single target image, while the advanced image generation method corresponding to the image generation model can generate target images in multiple media forms, which can be determined according to actual needs; among them, multiple media forms can include multiple media forms such as images, videos, and gifs.
[0043] In this embodiment, optionally, input the text and image data to be processed into an image generation model, and generate a target image according to the first image generation method corresponding to the image generation model, including: input the text and image data to be processed into the image generation model, and determine the corresponding image generation base parameters according to the model parameters included in the image generation model; determine whether the image generation base parameters include style adaptation parameters; if the image generation base parameters include style adaptation parameters, then determine the corresponding target style adaptation parameters according to the image generation base parameters; generate a target image according to the image generation base parameters, style adaptation parameters, and the text and image data to be processed.
[0044] Among them, the model parameters can be understood as various pre-set operation specified parameters. The base models corresponding to different operation specified parameters can be different. The model parameters in this embodiment can be corresponding to the first image generation method. The model parameters in this embodiment can include determining the loaded base model and the corresponding VAE and CLIP models used according to the model operation specified parameters. The image generation base parameters can be considered as the operation parameters corresponding to the base model. In this embodiment, the required loaded base model and the corresponding style adaptation model can be distinguished according to the image generation base parameters. The style adaptation parameters can be considered as the respective operation parameters of the fine-tuning or stylized model corresponding to the image generation base parameters. Among them, the target style adaptation parameters can be considered as the style adaptation parameters corresponding to the base model. The style adaptation parameters in this embodiment can be the operation parameters corresponding to the LoRA model.
[0045] In this embodiment, the corresponding model parameters can be determined according to the determined first image generation method, and the corresponding base model can be determined according to the image generation base parameters in the model parameters; then determine whether the image generation base parameters include the corresponding style adaptation parameters, that is, by confirming whether there is a style adaptation model associated with the base model. The LoRA model in this embodiment is equivalent to a small model superimposed on the base model and can be trained separately; the small model can be considered as an adaptation model, and its function can be to perform customized training on people and objects, or to adjust the generated image style. It can be selected to be superimposed on the base model or not. Moreover, the LoRA model in this embodiment must correspond to the base model. The LoRA model trained according to a specific base model can only be applied to that base model. For example, the LoRA model trained based on the Stable Diffusion 1.5 base model can only be used on StableDiffusion 1.5.
[0046] In this embodiment, different processing methods can be performed according to whether the style adaptation parameters are included in the raw image generation base parameters. In this embodiment, if the style adaptation parameters are included in the raw image generation base parameters, the corresponding target style adaptation parameters are determined according to the raw image generation base parameters. Then, according to the raw image generation base model, the style adaptation model, and the text and image data to be processed, the image size information is generated in combination with the configured sampling steps, and the corresponding VAE and CLIP models supporting the base model are selected according to the specified parameters of the model parameters and other processing operations, so as to obtain the generated target image.
[0047] In this embodiment, for the processing flow of the basic image generation module, a solution of integrating multiple pre-trained large models for inference is adopted and distinguished by the parameters specified during runtime. On the algorithm logic side, the base model and the corresponding VAE and CLIP models used are determined according to the specified parameters. At the same time, the synchronous and asynchronous processing logics can also be compatible according to the user's needs, and the asynchronous design unifies the callback logic to ensure the normal return of results. In order to communicate information and share models with the image training module, a set of LoRA model loading logics is designed, and the models trained in the image training module are notified to the image generation module through pipeline operations for loading on the image generation side.
[0048] With such a setting in this embodiment, when integrating and inferring using the pre-trained image generation method, a unified loading logic and a compatible synchronous and asynchronous processing logic are designed, making the efficiency of the image generation part on the inference side relatively high. And the LoRA model loading and intercommunication logics are designed, and there is no need to add other modules to process the loading and training part, improving the accuracy of image generation.
[0049] In this embodiment, optionally, determining the corresponding target style adaptation parameters according to the raw image generation base parameters includes: obtaining a style adaptation parameter list according to the raw image generation base parameters; wherein, the style adaptation parameter list includes multiple style adaptation parameters; and determining the style adaptation parameter corresponding to the raw image generation base parameters from the style adaptation parameter list as the target style adaptation parameter.
[0050] Among them, the specific style adaptation parameters included in the style adaptation parameter list can be determined according to the actual situation. The multiple style adaptation parameters can be considered as the parameters corresponding to multiple trained LoRA models. In this embodiment, it can be considered that a style adaptation parameter list is obtained according to the raw image generation base parameters, and a list of all trained LoRA models is obtained according to the base model. Then, the style adaptation parameter corresponding to the raw image generation base model is determined from the LoRA model list as the target style adaptation parameter, and thus the LoRA model corresponding to the target style adaptation parameter is selected according to the base model. With such a setting in this embodiment, the corresponding LoRA model can be obtained according to the base model, which can improve various adaptabilities of the generated image processing and facilitate the efficiency of image generation.
[0051] In the technical solution of the embodiment of the present invention, by obtaining the text and image data to be processed and the corresponding image generation requirements; determining the image generation method of the corresponding image generation model according to the image generation requirements; inputting the text and image data to be processed into the image generation model, and generating the target image according to the image generation method corresponding to the image generation model. In this technical solution, by integrating the image function requirements to determine the corresponding image processing method for processing, the quality of image generation is improved, and the generalization ability and adaptability of the model are enhanced.
[0052] Embodiment 2
[0053] Figure 3 FIG. is a flowchart of an image generation method provided according to Embodiment 2 of the present invention. This embodiment is optimized based on the above embodiment. The specific optimization is as follows: inputting the text and image data to be processed into the image generation model, and generating the target image according to the second image generation method corresponding to the image generation model, including: inputting the text and image data to be processed into the image generation model, and converting the text and image data to be processed into the text and image data in a set format corresponding to the second image generation method; processing the text and image data in the set format through the second image generation method to obtain the corresponding target image. As Figure 3 shown, the method includes:
[0054] S310. Obtain the text and image data to be processed and the corresponding image generation requirements.
[0055] S320. Determine the image generation method of the corresponding image generation model according to the image generation requirements.
[0056] In this embodiment, the image generation requirements include the first image generation requirement and the second image generation requirement; correspondingly, determining the image generation method of the corresponding image generation model according to the image generation requirements includes: determining the first image generation method of the corresponding image generation model according to the first image generation requirement; determining the second image generation method of the corresponding image generation model according to the second image generation requirement.
[0057] S330. Input the text and image data to be processed into the image generation model, and convert the text and image data to be processed into the text and image data in a set format corresponding to the second image generation method.
[0058] The second image generation method in this embodiment may be an advanced image generation method. The setting format may be a format corresponding to the second image generation method, and the setting format in this embodiment may be a json format. The second image generation method in this embodiment may be processed by the unified input and output processing method of ComfyUI, so the image and text data to be processed may be input into the image generation model, and the image and text data to be processed may be converted into image and text data in json format according to the input and output specifications of ComfyUI.
[0059] S340: Process the graphic data in the set format through a second image generation method to obtain a corresponding target image.
[0060] In this embodiment, the graphic data in json format can be input into the unified input and output processing mode of ComfyUI for processing, and the ComfyUI workflow is executed accordingly to obtain the corresponding target image. The ComfyUI workflow in this embodiment can be pre-configured with a variety of image processing functions for processing the graphic data to generate the corresponding target image.
[0061] In the second image generation method of the advanced image generation module in this embodiment, a processing method combining multiple function combinations and the ComfyUI workflow is adopted. Since the traditional ComfyUI workflow processing method is based on visualization, users can directly perform operations such as dragging, adding, and connecting different modules through a graphical interface to complete the construction of the desired workflow. Although it is intuitive to display the work modules through the graphical interface, it may appear cumbersome for complex processes. Users need to directly operate the modules, which requires a relatively high professional level from the users. They need to understand in detail the functions of each module and their connection rules. On the other hand, when a large number of workflow modules need to be built or a large number of modules are required for assistance, this method will be time-consuming. Since different modules have different functions, it makes the definition of the process more complex and difficult, and at the same time increases the difficulty of creating and maintaining the workflow. Therefore, in this embodiment, the input text and image data to be processed are simplified, and only the uploaded pictures or texts need to be concerned. At the same time, a ComfyUI unified input and output processing module is designed to process the given json workflow into an API format for invocation. The specific processing method can be as follows: in the first step, obtain the input text and image data to be processed, such as information including text, images, reference images, etc.; in the second step, process the user input, enter the ComfyUI unified input and output processing module, process the user input into a json format input, and capture the output of the node of ComfyUI after processing; in the third step, execute the ComfyUI workflow, and obtain the result file according to the output of the node of ComfyUI after capture processing. This result file may be a picture, a video, or an animated picture, that is, the target image. In this embodiment, a callback notification service can also be performed according to the final result file to notify the user that the generation has been completed. In addition, in this embodiment, other function module combinations can also be used to replace the ComfyUI workflow, and it can be dynamically adjusted according to the usage effect and hardware usage. Only the ComfyUI workflow processing part needs to be replaced in the processing flow, which will not be elaborated here.
[0062] It can be understood that the ComfyUI workflow in this embodiment can be replaced by other function module combinations for the ComfyUI workflow. It can be freely combined and is not a fixed workflow processing method. It can be pre-configured in real time or adjusted in real time, and the corresponding configuration can be performed according to the selected functions.
[0063] Through such settings in this embodiment, a processing method combining multiple function combinations and the ComfyUI workflow is adopted, the input part is simplified, and users only need to focus on uploading pictures or texts. The designed ComfyUI unified input and output processing module processes and executes the workflow, improving the stability of the system, and at the same time ensuring the timeliness and stability of the processing.
[0064] In this embodiment, optionally, the second image generation method includes at least two workflow processing methods for setting image processing functions; processing the graphic and text data in a set format through the second image generation method to obtain a corresponding target image, including: processing the graphic and text data in a set format through at least two workflow processing methods for setting image processing functions to obtain a target image.
[0065] Among them, at least two set image processing functions may include super-resolution processing function, redrawing processing function, image enhancement, image repair, and image transformation function, etc., and different processing functions can also be set according to actual needs. In this embodiment, the graphic and text data in json format can be processed through the processing methods of at least two workflow of the set image processing functions included in the second image generation method, so as to obtain the corresponding target image. It can be understood that the specific image processing functions in this embodiment can be determined according to actual image requirements, and this embodiment does not limit this.
[0066] Through such a setting in this embodiment, the accuracy, efficiency, and customizable ability of the target image for image generation are improved. It also combines the unified processing ability of the ComfyUI workflow to increase efficiency and system robustness, and has the advantage of a wider adaptation range.
[0067] Exemplarily, the graphic and text data to be processed in this embodiment may include text input or reference image input; for example, if the input is a product reference image, if only basic image generation is used, that is, using the first image generation method to process will only generate a similar image by generating a label for the reference product image. If there is a corresponding reference image, there may also be slight changes in the corresponding image main body. If processed through the image advanced module, that is, through the second image generation method, the required background can be customized on the basis of maintaining the image main body; such as a product introduction image, specifically including the refined generation of the background. There may be an image of a product in the middle, but the background can be generated based on the text prompt, which is more in line with the text description of the product background. Then, the product image and the background image version are fused to obtain the target image, which improves the accuracy of the target image generation and also enhances the fit between the generated image and the image generation requirements.
[0068] The technical solution of the embodiment of the present invention is as follows: obtain the text and image data to be processed and the corresponding image generation requirements; determine the image generation method of the corresponding image generation model according to the image generation requirements; input the text and image data to be processed into the image generation model, and convert the text and image data to be processed into text and image data in a set format corresponding to the second image generation method; process the text and image data in the set format through the second image generation method to obtain the corresponding target image. Through this technical solution, by integrating the image function requirements to determine the corresponding image processing method for processing, the quality of image generation is improved, and the generalization ability and adaptability of the model are enhanced.
[0069] Embodiment III
[0070] Figure 4 It is a schematic structural diagram of an image generation device provided according to Embodiment III of the present invention.
[0071] As Figure 4 shown, the device includes:
[0072] A data acquisition module 410, configured to acquire the text and image data to be processed and the corresponding image generation requirements;
[0073] A generation method determination module 420, configured to determine the image generation method of the corresponding image generation model according to the image generation requirements;
[0074] A target image generation module 430, configured to input the text and image data to be processed into the image generation model, and generate a target image according to the image generation method corresponding to the image generation model.
[0075] Optionally, the image generation requirements include a first image generation requirement and a second image generation requirement;
[0076] Correspondingly, the generation method determination module 420 is specifically configured to:
[0077] Determine the first image generation method of the corresponding image generation model according to the first image generation requirement;
[0078] Determine the second image generation method of the corresponding image generation model according to the second image generation requirement.
[0079] Optionally, the target image generation module 430 includes:
[0080] A base parameter determination unit for generating an image, configured to input the text and image data to be processed into the image generation model, and determine the corresponding base parameter for generating an image according to the model parameters included in the image generation model;
[0081] A judgment unit, configured to judge whether the base parameter for generating an image includes a style adaptation parameter;
[0082] a target style adaptation parameter determination unit, configured to determine corresponding target style adaptation parameters according to the raw image base parameters if the raw image base parameters include style adaptation parameters;
[0083] The image generation unit is used to generate a target image according to raw image base parameters, style adaptation parameters and image and text data to be processed.
[0084] Optionally, the target style adaptation parameter determination unit is specifically used to:
[0085] Obtaining a style adaptation parameter list according to the raw image base parameter; wherein the style adaptation parameter list includes a plurality of style adaptation parameters;
[0086] The style adaptation parameters corresponding to the raw image base parameters are determined from the style adaptation parameter list as target style adaptation parameters.
[0087] Optionally, the target image generation module 430 includes:
[0088] A data format conversion unit, used for inputting the image and text data to be processed into the image generation model, and converting the image and text data to be processed into image and text data of a set format corresponding to the second image generation mode;
[0089] The image processing unit is used to process the graphic data in a set format through a second image generation method to obtain a corresponding target image.
[0090] Optionally, the second image generation method includes at least two workflow processing methods for setting image processing functions;
[0091] The image processing unit is specifically used to process the graphic data of a set format through at least two workflow processing methods with set image processing functions to obtain a target image.
[0092] Optionally, the graphic data to be processed includes text data to be processed, picture data or reference images.
[0093] An image generating device provided by an embodiment of the present invention can execute an image generating method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0094] Embodiment 4
[0095] Figure 5It is a schematic structural diagram of an electronic device provided according to Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0096] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0097] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0098] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the image generation method.
[0099] In some embodiments, the image generation method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image generation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the image generation method by any other suitable means (e.g., by means of firmware).
[0100] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0102] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0105] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0106] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0107] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image generation method, characterized in that: include: Obtain the graphic data to be processed and the corresponding image generation requirements; Determine an image generation method of a corresponding image generation model according to the image generation requirement; The image and text data to be processed are input into the image generation model, and a target image is generated according to an image generation method corresponding to the image generation model.
2. The method according to claim 1, characterized in that The image generation requirement includes a first image generation requirement and a second image generation requirement; Correspondingly, determining the image generation method of the corresponding image generation model according to the image generation requirement includes: Determining a first image generation method of a corresponding image generation model according to the first image generation requirement; Determine a second image generation method of the corresponding image generation model according to the second image generation requirement.
3. The method according to claim 2, characterized in that Inputting the image and text data to be processed into the image generation model, and generating a target image according to the first image generation method corresponding to the image generation model, including: Inputting the image data to be processed into the image generation model, and determining corresponding raw image base parameters according to the model parameters included in the image generation model; Determining whether the raw image base parameters include style adaptation parameters; If the raw image base parameters include the style adaptation parameters, determining corresponding target style adaptation parameters according to the raw image base parameters; A target image is generated according to the raw image base parameters, the style adaptation parameters and the image and text data to be processed.
4. The method according to claim 3, characterized in that Determining corresponding target style adaptation parameters according to the raw image base parameters includes: Acquire a style adaptation parameter list according to the raw image base parameter; wherein the style adaptation parameter list includes a plurality of style adaptation parameters; A style adaptation parameter corresponding to the raw image base parameter is determined from the style adaptation parameter list as a target style adaptation parameter.
5. The method according to claim 2, characterized in that: Inputting the to-be-processed graphic data into the image generation model, and generating a target image according to the second image generation mode corresponding to the image generation model, comprises: Inputting the to-be-processed graphic data into the image generation model, and converting the to-be-processed graphic data into graphic data of a set format corresponding to the second image generation method; The graphic data in the set format is processed by a second image generation method to obtain a corresponding target image.
6. The method according to claim 5, characterized in that The second image generation method includes at least two workflow processing methods for setting image processing functions; Processing the graphic data in the set format by a second image generation method to obtain a corresponding target image includes: The graphic data in the set format is processed by at least two workflow processing methods with set image processing functions to obtain a target image.
7. The method according to claim 1, characterized in that The graphic data to be processed includes text data to be processed, picture data or reference images.
8. An image generating device, characterized in that: include: A data acquisition module is used to acquire the graphic data to be processed and the corresponding image generation requirements; A generation method determination module, used to determine the image generation method of the corresponding image generation model according to the image generation requirement; The target image generation module is used to input the image and text data to be processed into the image generation model, and generate a target image according to the image generation method corresponding to the image generation model.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the image generating method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image generation method according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Intelligent welding system based on casting machining
CN121598048A