High-resolution image generation method and system, terminal and storage medium

Through the high-resolution image generation method of multi-model collaboration and hardware acceleration optimization, the problems of slow generation speed and unstable image quality in the prior art are solved, and high-quality high-resolution images are efficiently generated on ordinary hardware.

CN120510031APending Publication Date: 2025-08-19SHENZHEN KUKAI SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510527909.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art is difficult to find a balance between taking into account computing resource consumption, generation speed and image quality, which leads to difficulty in running on ordinary hardware, slow generation speed and unstable image quality.

Method used

Using multi-model collaborative workflow and combining hardware acceleration technology, through modules such as DualCLIPLoader, FluxGuidance, SamplerCustomAdvanced, VAEDecode and UpscalerTensorrt, we optimize text encoding, potential space generation and image decoding processes, integrate integrated workflow design, and use TorchCompileModel and TorchCompileVAE for model compilation and optimization, and achieve end-to-end efficient image generation.

Benefits of technology

It significantly improves the image generation speed, the generated image resolution and quality, reduces the system's demand for hardware resources, and meets the real-time application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510031A_ABST
    Figure CN120510031A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution image generation method and system, a terminal and a storage medium, and the method comprises the steps: obtaining a text prompt, carrying out the word segmentation of the text prompt, obtaining a segmented text, carrying out the coding of the segmented text, and obtaining a coded text; representing the encoded text by using a low-dimensional vector to generate a potential spatial representation, the low-dimensional vector being used for representing potential features and semantic information of data; performing decoding processing on the potential space representation to obtain a decoded image, and performing super-resolution processing on the decoded image to obtain an amplified image; and carrying out optimization processing on the amplified image to generate a high-resolution image. Through multi-model collaboration and hardware acceleration, the image generation speed is remarkably improved, the requirement of the system for hardware resources is reduced, the generated image has higher resolution and better visual effect, the high-quality image can be generated in a short time, and the real-time application requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a high-resolution image generation method, system, terminal and computer-readable storage medium. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, text-to-image generation technology has become an important research direction in the field of computer vision. Driven by the wave of digitalization, the demand for high-quality, high-resolution images in various industries has exploded. In the field of film and television production, in order to create an immersive visual experience, the resolution requirements for special effects scenes and virtual characters are becoming increasingly stringent; the gaming industry requires high-resolution texture maps and character models to enhance the player's sense of immersion; and in the field of advertising design, the demand for high-resolution posters and promotional images continues to rise to attract consumers' attention. However, traditional image generation systems based on a single GAN (Generative Adversarial Network) or diffusion model have difficulty balancing computing resource consumption, generation speed, and image quality, and are unable to meet the urgent needs of the industry.

[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0004] The main purpose of the present invention is to provide a high-resolution image generation method, system, terminal and computer-readable storage medium, aiming to solve the problem in the existing technology that text-to-image generation technology is difficult to balance computing resource consumption, generation speed and image quality and cannot meet user needs.

[0005] To achieve the above object, the present invention provides a method for generating a high-resolution image, the method comprising the following steps:

[0006] Obtaining a text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, and encoding the segmented text to obtain an encoded text;

[0007] Representing the encoded text using a low-dimensional vector to generate a latent space representation, wherein the low-dimensional vector is used to represent the latent features and semantic information of the data;

[0008] Decoding the latent space representation to obtain a decoded image, and performing super-resolution processing on the decoded image to obtain an enlarged image;

[0009] The amplified image is optimized to generate a high-resolution image.

[0010] Optionally, the high-resolution image generation method, wherein the obtaining of text prompts, performing word segmentation processing on the text prompts to obtain a segmented text, and encoding the segmented text to obtain an encoded text, specifically includes:

[0011] Obtain a text prompt, input the text prompt into DualCLIPLoader, load the corresponding model and word segmenter through DualCLIPLoader to perform word segmentation on the text prompt, and obtain the text after word segmentation;

[0012] The segmented text is input into CLIPTextEncode, and the segmented text is encoded by CLIPTextEncode to output the encoded text.

[0013] Optionally, the high-resolution image generation method, wherein the step of representing the encoded text using a low-dimensional vector to generate a latent space representation, specifically includes:

[0014] Input the encoded text into FluxGuidance, extract and process text features through FluxGuidance, transform the text encoding vector through a multi-layer neural network to capture semantic features and obtain a feature vector;

[0015] The feature vector is input to SamplerCustomAdvanced, and SamplerCustomAdvanced performs sampling in the latent space according to the feature vector to generate a latent space representation that conforms to a specific distribution.

[0016] Optionally, the high-resolution image generation method, wherein the generating of the latent space representation further comprises:

[0017] Introducing random noise in the process of generating latent space representation;

[0018] The scheduler adjusts the intensity of the noise according to the preset strategy;

[0019] The process of generating the latent space representation is guided by a guide according to the encoded text.

[0020] Optionally, the high-resolution image generation method, wherein the decoding process on the latent space representation to obtain a decoded image and the super-resolution process on the decoded image to obtain an enlarged image, specifically comprises:

[0021] Inputting the latent space representation into VAEDecode and loading the VAE model through VAEDecode;

[0022] Decoding the latent space representation through a decoder component in a VAE model to obtain a decoded image, and inputting the decoded image into an UpscalerTensorrt;

[0023] The decoded image is super-resolution processed by UpscalerTensorRT using a deep learning model to obtain an enlarged image.

[0024] Optionally, in the high-resolution image generation method, the step of optimizing the magnified image specifically includes:

[0025] Performing image quality enhancement and detail reconstruction on the enlarged image through ImageUpscaleWithModel;

[0026] The amplified image is subjected to resolution adaptation and data format conversion through ImageScale4KBase64.

[0027] Optionally, in the high-resolution image generation method, the image quality enhancement includes sharpening processing, noise reduction processing and color correction;

[0028] The sharpening process is used to enhance the edges and details of the image through a specific convolution kernel or algorithm;

[0029] The noise reduction process is used to remove the introduced noise through Gaussian filtering, median filtering or a noise reduction model based on deep learning;

[0030] The color correction is used to adjust the color balance, saturation and contrast of the image through color space conversion and histogram equalization;

[0031] The detail reconstruction is used to reconstruct and optimize the details of the image using a pre-trained deep learning model;

[0032] The resolution adaptation is used to adjust the resolution of the image to the 4K standard through an interpolation algorithm;

[0033] The data format conversion converts the optimized image into a Base64 encoding format.

[0034] In addition, to achieve the above-mentioned object, the present invention further provides a high-resolution image generation system, wherein the high-resolution image generation system comprises:

[0035] A text encoding module is used to obtain text prompts, perform word segmentation processing on the text prompts to obtain a segmented text, and perform encoding processing on the segmented text to obtain an encoded text;

[0036] A latent space generation module is used to represent the encoded text using a low-dimensional vector to generate a latent space representation, wherein the low-dimensional vector is used to represent the potential features and semantic information of the data;

[0037] An image decoding and magnification module, configured to decode the latent space representation to obtain a decoded image, and perform super-resolution processing on the decoded image to obtain an enlarged image;

[0038] The image optimization module is used to optimize the amplified image to generate a high-resolution image.

[0039] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a high-resolution image generation program stored in the memory and runnable on the processor, and when the high-resolution image generation program is executed by the processor, the steps of the high-resolution image generation method described above are implemented.

[0040] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a high-resolution image generation program, and when the high-resolution image generation program is executed by a processor, the steps of the high-resolution image generation method as described above are implemented.

[0041] In the present invention, a text prompt is obtained, the text prompt is segmented to obtain a segmented text, the segmented text is encoded to obtain an encoded text; the encoded text is represented using a low-dimensional vector to generate a latent space representation, the low-dimensional vector is used to represent the potential features and semantic information of the data; the latent space representation is decoded to obtain a decoded image, the decoded image is super-resolution processed to obtain an enlarged image; the enlarged image is optimized to generate a high-resolution image. Through multi-model collaboration and hardware acceleration, the present invention significantly improves the image generation speed and reduces the system's demand for hardware resources. The generated images have higher resolution and better visual effects, and can generate high-quality images in a relatively short time to meet the needs of real-time applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flow chart of a preferred embodiment of the high-resolution image generation method of the present invention;

[0043] Figure 2 is a schematic diagram of an efficient image generation process in a preferred embodiment of the high-resolution image generation method of the present invention;

[0044] Figure 3 is a structural diagram of a preferred embodiment of the high-resolution image generation system of the present invention;

[0045] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0047] With the rapid development of artificial intelligence and deep learning technologies, text-to-image generation technology has become a key research area in computer vision. Driven by the wave of digitalization, the demand for high-quality, high-resolution images has skyrocketed across various industries. In film and television production, the resolution requirements for special effects scenes and virtual characters are becoming increasingly stringent to create an immersive visual experience. The gaming industry requires high-resolution texture maps and character models to enhance player engagement. In advertising design, the demand for high-resolution posters and promotional images continues to rise to attract consumers. However, traditional image generation systems based on single GANs or diffusion models struggle to balance computational resource consumption, generation speed, and image quality, making them unable to meet the urgent needs of the industry. In recent years, researchers have actively explored and developed numerous cutting-edge technologies, such as Stable Diffusion, DALL-E2, and Midjourney.

[0048] Stable Diffusion has received a major update, version 3.5, offering more options for users. Stable Diffusion 3.5Large boasts 8 billion parameters and excels in cue-following and image quality, generating high-quality images at 1 megapixel resolution, suitable for professional-level creative scenarios. A streamlined version, Stable Diffusion 3.5LargeTurbo, generates high-quality images in just four steps, significantly improving speed. Stable Diffusion 3.5Medium, with 2.6 billion parameters, is optimized for edge computing deployments and works out-of-the-box on consumer hardware, generating images at resolutions from 0.25 to 2 megapixels. In terms of technological innovation, Stable Diffusion 3.5 integrates Query-Key Normalization into the transformer block, making model fine-tuning and development easier while also enhancing the stability of the training and fine-tuning process. Furthermore, optimizations to the Multimodal Diffusion Transformer (MMDiT-X) architecture further enhance image quality and multi-resolution generation capabilities. In the future, Stable Diffusion 3.5 also plans to launch ControlNets capabilities to provide more control options for professional use cases.

[0049] DALL-E2 was developed by OpenAI and is based on the unCLIP model, which draws on some of the concepts of the GLIDE system. Text is encoded using BPE tokens, and images are encoded using special image tokens generated by a discrete variational autoencoder (dVAE). The model is trained on a dataset of 250M image-text pairs, and the generated results are ranked using the CLIP model. Although the images generated by DALL-E 2 are of high quality, their computational overhead makes real-time generation difficult. The CLIP model consists of two encoders, one for text and one for image, which can calculate similarity scores between images and text and can be used for tasks such as zero-shot classification.

[0050] Midjourney has released version 6.1, which boasts significant improvements in image generation quality and functionality. In terms of character generation, human skin rendering is more natural, with finer textures, and the depiction of body parts such as arms and legs has been optimized. Text rendering is now clearer and easier to read, effectively resolving the challenge of adding text to images. Additionally, Midjourney v6.1 boasts a 25% increase in image generation speed, improved magnification capabilities, and a new –q2 mode that adds more texture to images for enhanced realism. Regarding text accuracy, enclosing words in quotation marks allows the model to more accurately represent relevant content in images.

[0051] However, existing technologies consume a lot of computing resources. Existing systems require a large amount of GPU resources and are difficult to run on ordinary hardware. The generation speed is slow, and high-resolution image generation takes a long time, which cannot meet real-time requirements. The image quality is unstable, and the generated images may have artifacts, distortion and other problems. There is a lack of hardware acceleration, and existing systems fail to fully utilize hardware acceleration technologies such as TensorRT. The model integration is low, and the collaboration between different models is not optimized, resulting in low system efficiency.

[0052] That is to say, in the current field of image generation, especially image generation models based on deep learning (such as Stable Diffusion), there are the following technical problems:

[0053] (1) Slow generation speed: The traditional image generation process usually requires a long inference time, especially in the high-resolution image generation scenario, which consumes a lot of computing resources and has a slow generation speed.

[0054] (2) Low resource utilization: The existing image generation process fails to fully utilize the advantages of hardware acceleration (such as TensorRT, Torch Compile, etc.) during model inference, resulting in low computing resource utilization.

[0055] (3) Image quality and efficiency cannot be balanced: While pursuing generation speed, it is often difficult to ensure high image quality, especially in terms of detail processing and resolution improvement.

[0056] (4) Complex workflow: The existing image generation process usually requires multiple independent steps (such as text encoding, latent space generation, decoding, amplification, etc.), and lacks an integrated and efficient solution.

[0057] The present invention improves the stability and quality of image generation results through resource management and optimization.

[0058] The high-resolution image generation method described in the preferred embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the high-resolution image generation method includes the following steps:

[0059] Step S10: Acquire a text prompt, perform word segmentation processing on the text prompt to obtain a segmented text, and perform encoding processing on the segmented text to obtain an encoded text.

[0060] Specifically, a text prompt is obtained, and the text prompt is input into DualCLIPLoader. The corresponding CLIP model (such as t5xxl_fp8_e4m3fn.safetensors and clip_l.safetensors) and the word segmentor are loaded through DualCLIPLoader to perform word segmentation on the text prompt to obtain a segmented text; the segmented text is input into CLIPTextEncode (CLIPTextEncode is a commonly used component in Stable Diffusion (SD) or other generative AI systems based on the CLIP model, mainly used to encode text prompts (Prompt) into CLIP text embeddings (Text Embeddings), thereby guiding the image generation process), and the segmented text is encoded through CLIPTextEncode to output the encoded text.

[0061] Text is a natural language composed of characters and words that computers cannot directly understand and process. Therefore, encoding converts text into numerical vectors that computers can process. This allows text data to interact and be calculated with other types of data (such as image feature vectors) within the same model framework. Neural network models typically require fixed-length numerical vectors as input. Text encoding converts text of varying lengths into vectors of fixed dimensions, facilitating model processing and training.

[0062] Step S20: Represent the encoded text using a low-dimensional vector to generate a latent space representation, where the low-dimensional vector is used to represent the latent features and semantic information of the data.

[0063] Specifically, the encoded text is input into FluxGuidance, and text features are extracted and processed by FluxGuidance. The text encoding vector is transformed by a multi-layer neural network to capture semantic features and obtain a feature vector. The feature vector is input into SamplerCustomAdvanced, and SamplerCustomAdvanced samples the feature vector in the latent space to generate a latent space representation that conforms to a specific distribution.

[0064] Latent space is a low-dimensional vector space obtained by encoding high-dimensional data (such as images, text, etc.). In this space, the features of the data are compressed and represented as a low-dimensional vector. These vectors can be used to represent the latent features and semantic information of the data. In other words, the latent space representation is to represent the encoded text information with a low-dimensional vector. This vector contains the semantics, sentiment and other features of the text, which can be used by the model to generate images or other tasks. The process of generating the latent space representation is as follows:

[0065] (1) Feature extraction: The encoded text is input into the FluxGuidance model (FluxGuidance is a guidance technology used in diffusion models to improve the generation process, usually combined with Classifier-Free Guidance (CFG) or Energy-Based Models (EBMs) to enhance the quality, controllability, or alignment of the generated images). FluxGuidance usually further extracts and processes the text features, and may transform the text encoding vector through a multi-layer neural network to capture more advanced semantic features. For example, it may generate a more representative feature vector based on the different words and phrases in the text and the relationship between them.

[0066] (2) Latent space mapping: The feature vector processed by FluxGuidance will be input into SamplerCustomAdvanced (SamplerCustomAdvanced is an advanced sampler (Sampler) used to control the image generation process in diffusion models (DiffusionModels). It is usually used to replace the default sampling strategy (such as DDIM, Euler or DPM++) to achieve more refined generation control, faster convergence or higher quality output). The role of SamplerCustomAdvanced is to sample in the latent space based on the input feature vector and generate a latent space vector that conforms to a specific distribution. This process may involve some random sampling operations, and will also map the text features to the latent space according to the rules learned by the model. For example, it may find a corresponding area in the latent space based on the semantic information of the text and sample a vector from this area as the latent space representation of the text.

[0067] (3) Output latent space representation: SamplerCustomAdvanced finally outputs the latent space representation of the encoded text. This representation is a low-dimensional vector that contains the semantic and structural information of the original text and can be used by subsequent models (such as the VAE decoder) to generate high-quality images and other tasks.

[0068] The present invention can also load the UNet model (such as flux1-dev.safetensors) through UNETLoader and combine it with ModelSamplingFlux to generate the latent space.

[0069] When using the UNet model in combination with ModelSamplingFlux for potential space generation, the following steps are usually performed:

[0070] (1) Model loading: Use UNETLoader to load a pre-trained UNet model, such as the flux1-dev.safetensors model. This model has been trained on a large amount of data and has learned some of the inherent patterns and features of the data.

[0071] (2) Input encoding: Input data (such as images or text representations) is fed into the UNet model. The encoder part of the model processes the input, maps it into the latent space, and generates a low-dimensional latent vector. This process can be seen as feature extraction and compression of the input data, condensing the complex information of the original data into the latent vector.

[0072] (3) Sampling Generation: Combined with ModelSamplingFlux, it may operate on the latent space according to a certain sampling strategy (such as sampling based on probability distribution) to generate new latent vectors from the latent space. These new latent vectors can be regarded as a sampling of the original data distribution, and they represent different data points that may exist in the latent space.

[0073] (4) Decoding generation: Through the decoder part of the UNet model, the generated latent vector is decoded into corresponding output data, such as generating new images or text.

[0074] Textual prompts can be used as a conditional information to guide the generation of latent space. Specifically, the relationship between latent space generation and textual prompts is as follows:

[0075] (1) Guiding the generation direction: The text prompt can describe the characteristics, themes, or attributes of the content you want to generate. For example, given a text prompt such as "generate a beautiful landscape image", the model will search and generate latent vectors related to the landscape image in the latent space based on this prompt, thereby guiding the generation process in the direction that conforms to the text description.

[0076] (2) Controlling generated content: Different text prompts will cause the model to explore different regions in the latent space and generate outputs with different content. For example, two different text prompts, "Generate a cute kitten" and "Generate a red car", will cause the model to generate images related to kittens and cars, respectively. This is because the text prompts provide the model with specific semantic information, helping the model determine the target content to be generated in the latent space.

[0077] (3) Incorporating semantic information: In some generative models, textual prompts are encoded as a vector and then fused or interacted with the vectors in the latent space. This allows the model to incorporate the semantic information of the text into the generation process of the latent space, making the generated result more consistent with the meaning expressed by the textual prompt.

[0078] In addition, RandomNoise, BasicScheduler, and BasicGuider are combined to optimize the generation process, specifically as follows:

[0079] RandomNoise: Introduces random noise into the generation process. Random noise can increase the diversity of generated results and prevent the model from generating overly similar latent space representations. It can enable the model to explore a wider range of regions in the latent space, thereby generating richer and more creative results.

[0080] BasicScheduler: The scheduler is typically used to control the noise attenuation process. During the generation process, as iterations progress, the noise gradually decreases, stabilizing the generated latent space representation. The scheduler can adjust the noise intensity according to a preset strategy to ensure the stability and convergence of the generation process.

[0081] BasicGuider: The guide guides the generation process based on the encoded text. It uses the semantic information of the text to guide the model in the latent space toward a direction consistent with the text description. The guide strengthens the relevance of the generated results to the text input, improving generation quality.

[0082] Step S30: Decode the latent space representation to obtain a decoded image, and perform super-resolution processing on the decoded image to obtain an enlarged image.

[0083] Specifically, the latent space representation is input into VAEDecode (VAEDecode is the decoder part of the variational autoencoder (VAE), which is mainly used to convert the low-dimensional representation of the latent space (Latent Space) back to the pixel space (Pixel Space) image in diffusion models such as Stable Diffusion (SD), and the VAE model is loaded through VAEDecode; the latent space representation is decoded by the decoder component in the VAE model to obtain a decoded image, and the decoded image is input into UpscalerTensorRT (UpscalerTensorRT is a super-resolution processing tool implemented based on the TensorRT framework, which usually uses a deep learning model to improve the resolution of the image); UpscalerTensorRT uses a deep learning model to perform super-resolution processing on the decoded image to obtain an enlarged image, that is, UpscalerTensorRT is used to load the TensorRT engine (such as 4x-UltraSharp.engine), and super-resolution processing is performed on the decoded image to significantly improve the image quality.

[0084] Hardware Acceleration: TorchCompileModel and TorchCompileVAE optimize the inference speed of UNet and VAE models, respectively. These tools leverage PyTorch's inductor backend to compile and optimize models, improving inference speed. PyTorch's compilation mechanism analyzes and optimizes the model's computational graph, reducing unnecessary computation steps by fusing operators, reducing memory accesses, and utilizing more efficient kernels, thereby improving model inference speed. TorchCompileModel and TorchCompileVAE can apply these optimization strategies to UNet and VAE models, respectively.

[0085] Efficient Sampling and Scheduling: Using SamplerCustomAdvanced and BasicScheduler, combined with the Euler sampler and simple scheduling strategy, we reduce sampling steps while ensuring image quality. The FluxGuidance module optimizes conditional control during the generation process, improving generation efficiency.

[0086] Integrated workflow design: Integrates text encoding, latent space generation, decoding, and amplification steps into a single JSON configuration file, enabling efficient end-to-end image generation. Supports customizable parameters (such as resolution, sampling steps, and noise seed) to meet the needs of different scenarios.

[0087] Step S40: Optimize the amplified image to generate a high-resolution image.

[0088] Specifically, the image quality of the enlarged image is enhanced and details are reconstructed by using ImageUpscaleWithModel; the image quality enhancement includes sharpening, noise reduction and color correction; the details are as follows:

[0089] Sharpening: Using a specific convolution kernel or algorithm, the edges and details of an image are enhanced to make it appear clearer. For example, using a sharpening filter like the Laplacian operator can highlight high-frequency information in an image, making the outlines of objects more distinct.

[0090] Noise reduction: Some noise may be introduced during the super-resolution process. ImageUpscaleWithModel can remove this noise using methods such as Gaussian filtering and median filtering, or use more advanced deep learning-based noise reduction models to effectively reduce noise while preserving image details.

[0091] Color correction: Adjusts the color balance, saturation, and contrast of an image to make the colors more vivid and realistic. This can be achieved through techniques such as color space conversion and histogram equalization.

[0092] Model-driven detail reconstruction: Utilize pre-trained deep learning models to further reconstruct and optimize image details. These models are typically trained on large-scale image datasets and can learn various features and patterns of images.

[0093] The amplified image is subjected to resolution adaptation and data format conversion through ImageScale4KBase64.

[0094] Resolution adaptation: Adjusting the image resolution to the 4K standard (3840×2160 pixels) to meet the requirements of high-definition display devices. This may involve using interpolation algorithms, such as bilinear interpolation and bicubic interpolation, to adjust the image size. Furthermore, to avoid image distortion, intelligent scaling may be performed based on the image's content and structure.

[0095] Data format conversion: Convert the optimized image to Base64 encoding. Base64 encoding converts binary data into ASCII characters and is commonly used to transmit or store binary data in text environments, such as embedding images in web pages. Converting to Base64 format makes it easier to integrate images into various applications and reduces network request overhead.

[0096] Resource Management and Optimization: Generate high-quality noise using the RandomNoise module to reduce the impact of randomness during the generation process. Use ImageUpscaleWithModel and ImageScale4KBase64 to further optimize and save the generated images.

[0097] The innovative features of the present invention are as follows:

[0098] (1) Multi-model collaborative workflow:

[0099] Addressing flaws: Traditional methods typically use a single model, resulting in low generation efficiency and unstable quality. This invention significantly improves generation efficiency and quality by collaborating across multiple models.

[0100] Innovation: Dual CLIP models (such as t5xxl_fp8_e4m3fn.safetensors and clip_l.safetensors) are loaded through DualCLIPLoader, combined with UNETLoader and VAELoader to achieve efficient collaboration of text encoding, latent space generation, and image decoding.

[0101] Effect or purpose: Achieve end-to-end efficient image generation and reduce redundant calculations in intermediate links.

[0102] (2) Hardware acceleration optimization:

[0103] Addressing the drawback: Existing technologies fail to fully utilize hardware acceleration capabilities, resulting in slow generation speeds. This invention significantly improves inference speed through hardware acceleration optimization.

[0104] Innovation: Introducing TorchCompileModel and TorchCompileVAE, using PyTorch's inductor backend to compile and optimize the model, and using UpscalerTensorRT to load the TensorRT engine (such as 4x-UltraSharp.engine) for super-resolution processing.

[0105] Effect or purpose: Greatly shorten the generation time while ensuring image quality.

[0106] (3) Efficient sampling and scheduling:

[0107] Addressing drawbacks: Traditional sampling methods have multiple steps, are time-consuming, and lack precise control over conditions. This invention reduces the number of sampling steps and improves generation efficiency through efficient sampling and scheduling.

[0108] Innovation: SamplerCustomAdvanced and BasicScheduler are used, combined with the Euler sampler and simple scheduling strategy, to optimize the sampling process. The FluxGuidance module improves the efficiency of conditional control during the generation process.

[0109] Effect or purpose: Reduce sampling steps and increase generation speed while ensuring image quality.

[0110] (4) Integrated workflow design:

[0111] Addressing drawbacks: Existing technologies typically require multiple independent steps, making operations complex and prone to errors. This invention simplifies the operational process through an integrated design.

[0112] Innovation: Integrates text encoding, latent space generation, decoding, amplification and other steps into a JSON configuration file, supporting custom parameters (such as resolution, sampling step, noise seed, etc.).

[0113] Effect or purpose: To provide a simple, easy-to-use, efficient and reliable image generation solution.

[0114] (5) Resource management and optimization:

[0115] Resolved drawbacks: Traditional methods have deficiencies in noise generation and image optimization, resulting in unstable results. This invention improves the stability and quality of the generated results through resource management and optimization.

[0116] Innovation: Generates high-quality noise using the RandomNoise module, reducing the impact of randomness during the generation process. Uses ImageUpscaleWithModel and ImageScale4KBase64 to further optimize and save the generated image.

[0117] Effect or purpose: Provide high-quality and highly stable image generation results.

[0118] The present invention realizes the efficient collaboration of multiple models and improves the overall efficiency of image generation; through hardware acceleration technology, it significantly increases the generation speed of high-resolution images; optimizes the image generation quality and reduces artifacts and distortion; and reduces the system's requirements for hardware resources, enabling it to run on ordinary hardware.

[0119] The technical effects that the present invention can bring are as follows:

[0120] (1) Efficiency: Through multi-model collaboration and hardware acceleration, the image generation speed is significantly improved.

[0121] (2) High quality: The generated images have higher resolution and better visual effects.

[0122] (3) Scalability: Modular design makes the system easy to expand and upgrade.

[0123] (4) Resource optimization: Through model optimization and hardware acceleration, the system's demand for hardware resources is reduced.

[0124] (5) Real-time: It can generate high-quality images in a short time to meet the needs of real-time applications.

[0125] Furthermore, if Figure 3 As shown, based on the above-mentioned high-resolution image generation method, the present invention also provides a high-resolution image generation system, wherein the high-resolution image generation system includes:

[0126] A text encoding module 51 is used to obtain a text prompt, perform word segmentation processing on the text prompt to obtain a segmented text, and perform encoding processing on the segmented text to obtain an encoded text;

[0127] a latent space generation module 52 for representing the encoded text using low-dimensional vectors to generate a latent space representation, wherein the low-dimensional vectors are used to represent the latent features and semantic information of the data;

[0128] An image decoding and magnification module 53 is configured to decode the latent space representation to obtain a decoded image, and perform super-resolution processing on the decoded image to obtain an enlarged image;

[0129] The image optimization module 54 is used to optimize the amplified image to generate a high-resolution image.

[0130] Furthermore, if Figure 4 As shown, based on the above-mentioned high-resolution image generation method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0131] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code of the installation terminal, etc. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a high-resolution image generation program 40 is stored on the memory 20, and the high-resolution image generation program 40 can be executed by the processor 10, thereby realizing the high-resolution image generation method in the present application.

[0132] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 20, such as executing the high-resolution image generation method.

[0133] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, memory 20, and display 30 of the terminal communicate with each other via a system bus.

[0134] In one embodiment, when the processor 10 executes the high-resolution image generation program 40 in the memory 20, the following steps are implemented:

[0135] Obtaining a text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, and encoding the segmented text to obtain an encoded text;

[0136] Representing the encoded text using a low-dimensional vector to generate a latent space representation, wherein the low-dimensional vector is used to represent the latent features and semantic information of the data;

[0137] Decoding the latent space representation to obtain a decoded image, and performing super-resolution processing on the decoded image to obtain an enlarged image;

[0138] The amplified image is optimized to generate a high-resolution image.

[0139] The obtaining of the text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, and encoding the segmented text to obtain an encoded text, specifically includes:

[0140] Obtain a text prompt, input the text prompt into DualCLIPLoader, load the corresponding model and word segmenter through DualCLIPLoader to perform word segmentation on the text prompt, and obtain the text after word segmentation;

[0141] The segmented text is input into CLIPTextEncode, and the segmented text is encoded by CLIPTextEncode to output the encoded text.

[0142] The step of representing the encoded text using a low-dimensional vector to generate a latent space representation specifically includes:

[0143] Input the encoded text into FluxGuidance, extract and process text features through FluxGuidance, transform the text encoding vector through a multi-layer neural network to capture semantic features and obtain a feature vector;

[0144] The feature vector is input to SamplerCustomAdvanced, and SamplerCustomAdvanced performs sampling in the latent space according to the feature vector to generate a latent space representation that conforms to a specific distribution.

[0145] The generating of the latent space representation further includes:

[0146] Introducing random noise in the process of generating latent space representation;

[0147] The scheduler adjusts the intensity of the noise according to the preset strategy;

[0148] The process of generating the latent space representation is guided by a guide according to the encoded text.

[0149] The decoding of the latent space representation to obtain a decoded image, and the super-resolution processing of the decoded image to obtain an enlarged image specifically include:

[0150] Inputting the latent space representation into VAEDecode and loading the VAE model through VAEDecode;

[0151] Decoding the latent space representation through a decoder component in a VAE model to obtain a decoded image, and inputting the decoded image into an UpscalerTensorrt;

[0152] The decoded image is super-resolution processed by UpscalerTensorRT using a deep learning model to obtain an enlarged image.

[0153] The optimizing process on the amplified image specifically includes:

[0154] Performing image quality enhancement and detail reconstruction on the enlarged image through ImageUpscaleWithModel;

[0155] The amplified image is subjected to resolution adaptation and data format conversion through ImageScale4KBase64.

[0156] Wherein, the image quality enhancement includes sharpening processing, noise reduction processing and color correction;

[0157] The sharpening process is used to enhance the edges and details of the image through a specific convolution kernel or algorithm;

[0158] The noise reduction process is used to remove the introduced noise through Gaussian filtering, median filtering or a noise reduction model based on deep learning;

[0159] The color correction is used to adjust the color balance, saturation and contrast of the image through color space conversion and histogram equalization;

[0160] The detail reconstruction is used to reconstruct and optimize the details of the image using a pre-trained deep learning model;

[0161] The resolution adaptation is used to adjust the resolution of the image to the 4K standard through an interpolation algorithm;

[0162] The data format conversion converts the optimized image into a Base64 encoding format.

[0163] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a high-resolution image generation program, and when the high-resolution image generation program is executed by a processor, the steps of the high-resolution image generation method described above are implemented.

[0164] In summary, the present invention provides a method, system, terminal, and computer-readable storage medium for generating high-resolution images, the method comprising: obtaining a text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, encoding the segmented text to obtain an encoded text; representing the encoded text using a low-dimensional vector to generate a latent space representation, the low-dimensional vector being used to represent the potential features and semantic information of the data; decoding the latent space representation to obtain a decoded image, performing super-resolution processing on the decoded image to obtain an enlarged image; and optimizing the enlarged image to generate a high-resolution image. The present invention significantly improves the image generation speed and reduces the system's demand for hardware resources through multi-model collaboration and hardware acceleration. The generated image has higher resolution and better visual effects, and can generate high-quality images in a relatively short time, meeting the needs of real-time applications.

[0165] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0166] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0167] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for generating a high-resolution image, characterized in that: The high-resolution image generation method comprises: Obtaining a text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, and encoding the segmented text to obtain an encoded text; Representing the encoded text using a low-dimensional vector to generate a latent space representation, wherein the low-dimensional vector is used to represent the latent features and semantic information of the data; Decoding the latent space representation to obtain a decoded image, and performing super-resolution processing on the decoded image to obtain an enlarged image; The amplified image is optimized to generate a high-resolution image.

2. The high-resolution image generation method according to claim 1, characterized in that: The obtaining of the text prompt, performing word segmentation processing on the text prompt to obtain a segmented text, and encoding the segmented text to obtain an encoded text specifically includes: Obtain a text prompt, input the text prompt into DualCLIPLoader, load the corresponding model and word segmenter through DualCLIPLoader to perform word segmentation on the text prompt, and obtain the text after word segmentation; The segmented text is input into CLIPTextEncode, and the segmented text is encoded by CLIPTextEncode to output the encoded text.

3. The high-resolution image generation method according to claim 1, wherein: The encoded text is represented by a low-dimensional vector to generate a latent space representation, specifically comprising: Input the encoded text into FluxGuidance, extract and process text features through FluxGuidance, transform the text encoding vector through a multi-layer neural network to capture semantic features and obtain a feature vector; The feature vector is input to SamplerCustomAdvanced, and SamplerCustomAdvanced performs sampling in the latent space according to the feature vector to generate a latent space representation that conforms to a specific distribution.

4. The high-resolution image generation method according to claim 1 or 3, characterized in that: The generation of the latent space representation also includes: Introducing random noise in the process of generating latent space representation; The scheduler adjusts the intensity of the noise according to the preset strategy; The process of generating the latent space representation is guided by a guide according to the encoded text.

5. The high-resolution image generation method according to claim 1, wherein: The decoding process of the latent space representation to obtain a decoded image, and the super-resolution process of the decoded image to obtain an enlarged image specifically includes: Inputting the latent space representation into VAEDecode and loading the VAE model through VAEDecode; Decoding the latent space representation through a decoder component in a VAE model to obtain a decoded image, and inputting the decoded image into an UpscalerTensorrt; The decoded image is super-resolution processed by UpscalerTensorRT using a deep learning model to obtain an enlarged image.

6. The high-resolution image generation method according to claim 1 or 5, characterized in that: The optimizing process on the enlarged image specifically includes: Performing image quality enhancement and detail reconstruction on the enlarged image through ImageUpscaleWithModel; The amplified image is subjected to resolution adaptation and data format conversion through ImageScale4KBase64.

7. The high-resolution image generation method according to claim 6, characterized in that: The image quality enhancement includes sharpening processing, noise reduction processing and color correction; The sharpening process is used to enhance the edges and details of the image through a specific convolution kernel or algorithm; The noise reduction process is used to remove the introduced noise through Gaussian filtering, median filtering or a noise reduction model based on deep learning; The color correction is used to adjust the color balance, saturation and contrast of the image through color space conversion and histogram equalization; The detail reconstruction is used to reconstruct and optimize the details of the image using a pre-trained deep learning model; The resolution adaptation is used to adjust the resolution of the image to the 4K standard through an interpolation algorithm; The data format conversion converts the optimized image into a Base64 encoding format.

8. A high-resolution image generation system, characterized in that: The high-resolution image generation system comprises: A text encoding module is used to obtain text prompts, perform word segmentation processing on the text prompts to obtain a segmented text, and perform encoding processing on the segmented text to obtain an encoded text; A latent space generation module is used to represent the encoded text using a low-dimensional vector to generate a latent space representation, wherein the low-dimensional vector is used to represent the potential features and semantic information of the data; An image decoding and magnification module, configured to decode the latent space representation to obtain a decoded image, and perform super-resolution processing on the decoded image to obtain an enlarged image; The image optimization module is used to optimize the amplified image to generate a high-resolution image.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a high-resolution image generation program stored in the memory and executable on the processor. When the high-resolution image generation program is executed by the processor, the steps of the high-resolution image generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a high-resolution image generation program, and when the high-resolution image generation program is executed by a processor, the steps of the high-resolution image generation method according to any one of claims 1 to 7 are implemented.