Customized image generation method based on large model fusion, system, medium, terminal, and program product
By fusing image and text data and utilizing various large-scale modeling techniques to generate customized images, the problem of insufficient capture of user needs in existing technologies has been solved, achieving high-quality and creative image generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI QICHUANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-07-18
- Publication Date
- 2026-05-07
AI Technical Summary
Existing image generation technologies struggle to fully capture users' comprehensive needs, particularly in handling complex scenes, maintaining consistency in detail, and generating creative content.
By fusing sample image data and text data, and utilizing the Stable Diffusion model, CycleGAN model, and Generative Adversarial Network model, combined with graph neural network and Transformer model, customized images are generated.
It enables the generation of high-quality, creative, and stylized images based on user needs, satisfying users' personalized and deeply customized requirements.
Smart Images

Figure CN2025109280_07052026_PF_FP_ABST
Abstract
Description
Customized image generation methods, systems, media, terminals, and application products based on large model fusion Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to customized image generation methods, systems, media, terminals and program products based on large model fusion. Background Technology
[0002] While significant progress has been made in the field of image generation technology, most techniques remain limited to single-modal input, such as accepting only images or only text as input. This single-modal input approach struggles to fully capture the comprehensive needs of users, resulting in generated images that do not meet user requirements. Furthermore, existing image generation technologies have limitations in handling complex scenes, maintaining detail consistency, and generating creative content.
[0003] Therefore, it is necessary to provide a customized image generation method, system, medium, terminal, and program product based on large model fusion to solve the above-mentioned problems in the prior art. Summary of the Invention
[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a customized image generation method, system, medium, terminal and program product based on large model fusion, so as to solve the technical problem that the prior art is unable to capture the full range of user needs.
[0005] To achieve the above and other related objectives, a first aspect of this application provides a customized image generation method based on large model fusion, comprising:
[0006] The acquired sample image data and text data are fused to obtain a cross-modal fusion vector;
[0007] The cross-modal fusion vectors are input into the Stable Diffusion model and the CycleGAN model for encoding to obtain quality vectors and style vectors, respectively. The quality vectors and style vectors are then fused based on the attention mechanism and input into the generative adversarial network model for training to obtain a customized image generation model.
[0008] The customized image generation model is deployed to generate customized images; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs the customized image.
[0009] In some embodiments of the first aspect of this application, the process of fusing the acquired sample image data and text data to obtain a cross-modal fusion vector includes: acquiring sample image data and text data respectively; using a graph neural network model to extract features from the sample image data to obtain image features; using a Transformer model to extract features from the text data to obtain text semantic features; and fusing the image features and the text semantic features based on the cross-modal attention mechanism of the Transformer model to obtain a cross-modal fusion vector.
[0010] In some embodiments of the first aspect of this application, the method further includes: fine-tuning the generated customized image based on real-time acquired user interaction behavior features.
[0011] In some embodiments of the first aspect of this application, the step of fine-tuning the generated customized image based on real-time acquired user interaction behavior features includes: preprocessing the real-time acquired user interaction behavior features to convert them into a data format suitable for the customized image generation model; inputting the format-converted user interaction behavior data into the customized image generation model to adjust the customized image generated by the customized image generation model.
[0012] In some embodiments of the first aspect of this application, the preprocessing methods include: data deduplication, missing value processing, outlier processing, and standardization processing.
[0013] In some embodiments of the first aspect of this application, the method further includes: collecting user feedback data on the generated customized image and inputting it into the generative adversarial network model to optimize the customized image generation model.
[0014] To achieve the above and other related objectives, a second aspect of this application provides a customized image generation system based on large model fusion, comprising:
[0015] The cross-modal fusion module is used to fuse the acquired sample image data and text data to obtain a cross-modal fusion vector;
[0016] The multi-model collaboration module is used to input the cross-modal fusion vector into the Stable Diffusion model and the CycleGAN model for encoding to obtain quality vector and style vector respectively; after fusing the quality vector and style vector using an attention mechanism, the vector is input into the generative adversarial network model for training to obtain a customized image generation model.
[0017] A customized image generation module is used to deploy the customized image generation model; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs a customized image.
[0018] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the customized image generation method based on large model fusion.
[0019] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, enables the computer to implement the customized image generation method based on large model fusion.
[0020] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the customized image generation method based on large model fusion.
[0021] As described above, the customized image generation method, system, medium, terminal, and program product based on large model fusion of this application have the following beneficial effects:
[0022] This application can not only receive sample images provided by users as style references, but also receive detailed text descriptions as content guidance. By integrating multiple advanced large-scale model technologies such as Stable Diffusion model, CycleGAN model and Generative Adversarial Network model, it achieves unprecedented image generation effects, thereby meeting users' personalized requirements and deep customization needs for image creation, and providing users with a more intelligent, flexible and innovative image generation solution. Attached Figure Description
[0023] Figure 1 shows a flowchart of a customized image generation method based on large model fusion in one embodiment of this application.
[0024] Figure 2 shows a schematic diagram of the framework of a customized image generation method based on large model fusion in one embodiment of this application.
[0025] Figure 3 shows a flowchart of obtaining cross-modal fusion vectors in one embodiment of this application.
[0026] Figure 4 shows a flowchart illustrating the process of fine-tuning a customized image in one embodiment of this application.
[0027] Figure 5 shows a schematic diagram of the structure of a customized image generation system based on large model fusion in one embodiment of this application.
[0028] Figure 6 shows a schematic diagram of the structure of an electronic terminal in one embodiment of this application. Detailed Implementation
[0029] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0030] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first XX" and "second XX" are merely used to distinguish different XXs and do not limit their order. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that terms such as "first" and "second" do not necessarily imply that they are different.
[0031] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0032] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0033] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:
[0034] <1> Graph Neural Networks (GNNs) are deep learning models specifically designed for processing graph data. Graph data consists of nodes and edges, and the graph structure is typically represented using an adjacency matrix to capture the relationships between nodes and edges. The core of GNNs is to learn and optimize attribute vectors in the graph, including node information, edge information, and overall graph information, which are usually represented by vectors.
[0035] <2> Stable Diffusion (SDD) is an advanced text-to-image generation model that uses deep learning and a diffusion process to produce high-quality images. Its core is the Latent Diffusion Model (LDM), which works through a Variational Autoencoder (VAE), a U-Net (neural network), and a text encoder. It first transforms the image into a low-dimensional latent space, then adds and removes Gaussian noise within that space, and finally uses a VAE decoder to transform the latent representation back into pixel space to generate the output image.
[0036] <3> CycleGAN is a deep learning model for image style transfer that can transform images from one style to another without paired training data. The model consists of two generators and two discriminators, learning style transfer through adversarial training and a cycle consistency loss.
[0037] <4> Generative Adversarial Networks (GANs) are deep learning-based generative models that learn the distribution of data and generate new data samples through an adversarial process between two neural networks—a generator and a discriminator.
[0038] To facilitate understanding of the embodiments of this application, a detailed description will first be provided with reference to Figure 1. Figure 1 shows a flowchart illustrating a customized image generation method based on large model fusion according to an embodiment of the present invention. The customized image generation method based on large model fusion in this embodiment mainly includes the following steps:
[0039] Step S11: Fuse the acquired sample image data and text data to obtain a cross-modal fusion vector.
[0040] In some embodiments of this application, as shown in Figures 2-3, the process of fusing the acquired sample image data and text data to obtain a cross-modal fusion vector includes:
[0041] S31: Acquire sample image data and text data respectively; S32: Use a graph neural network model to extract features from the sample image data to obtain image features; use a Transformer model to extract features from the text data to obtain text semantic features; S33: Based on the cross-modal attention mechanism of the Transformer model, fuse the image features and the text semantic features to obtain a cross-modal fusion vector.
[0042] In step S31, a sample image is obtained as a style reference. For example, a user can upload a painting in the style of Van Gogh as a similar artistic style to the image they wish to generate. Detailed text descriptions input by the user are also obtained to guide the content of the generated image. For example, a user can describe "a scene of ginkgo leaves falling on both sides of the road on a sunny autumn day," thus forming the input sample image data and text data. Based on the user-input sample image and related text descriptions, the final generated image meets the user's personalized needs, achieving the effect of a customized image for the user.
[0043] In step S32, using the examples listed above, a graph neural network model (see the glossary above for graph neural networks) is used to analyze the image data and extract features, such as features representing the style of Van Gogh's paintings. Simultaneously, a Transformer model is used to analyze the text data and extract content elements, such as roads and ginkgo leaves. In step S33, a cross-modal attention mechanism based on the Transformer model is used to achieve fusion between different modalities, that is, to fuse image features and text semantic features to obtain a cross-modal fusion vector. It should be noted that the cross-modal attention mechanism is implemented through a self-attention mechanism and an encoder-decoder structure, thereby fusing image features and text semantic features.
[0044] In this embodiment, the Transformer model is a deep learning architecture based on a self-attention mechanism, which has achieved great success in the field of Natural Language Processing (NLP). The core of the Transformer model is the self-attention mechanism, which allows the model to dynamically calculate the correlation between each position in the input sequence and other positions when processing sequence data, thereby better capturing long-distance dependencies between sequences. This mechanism enables the Transformer to process sequence data in parallel, greatly improving the efficiency of training and inference. Specifically, the Transformer model is used to extract text data features to form text semantic features. The Transformer model consists of two parts: an encoder and a decoder. The encoder is responsible for encoding the input sequence into a hidden representation, while the decoder generates the target sequence based on the encoder's output and the generated partial sequences. Each encoder and decoder consists of multiple stacked Transformer blocks, each of which includes a multi-head self-attention layer and a fully connected feedforward network layer. Furthermore, a cross-modal attention mechanism is designed to fuse features from different modalities, ultimately outputting a cross-modal fused vector.
[0045] Specifically, using the examples listed above, the self-attention layer of the Transformer model's self-attention mechanism calculates the similarity between image features of "Van Gogh style" and text features such as "road, ginkgo leaves," generating attention weights to identify which image features are associated with keywords (road, ginkgo leaves) in the text data. The encoder processes image features, extracting key visual information, such as sharpness, while the decoder processes text features and transforms them into semantic information that can guide image generation, such as road, ginkgo leaves. The cross-attention mechanism enables information interaction and integration, ensuring that the final customized image not only contains the content of the text description but also incorporates the style of Van Gogh's paintings.
[0046] Step S12: Input the cross-modal fusion vector into the Stable Diffusion model (see the glossary above) and the CycleGAN model (see the glossary above) for encoding to obtain the quality vector and style vector respectively; after fusing the quality vector and style vector based on the attention mechanism, input them into the generative adversarial network model for training to obtain a customized image generation model.
[0047] It should be understood that in this embodiment, the cross-modal fusion vector is input into the Stable Diffusion model and encoded into a latent space representation through an encoder network. This means the input cross-modal fusion vector is mapped to a low-dimensional latent space. It should be noted that the latent space is a compressed representation space of the data learned by the model, in which the latent features of the data are extracted and can exist in a compact form. "Latent" means capturing the latent visual content and stylistic features of the image. These latent representations are quality vectors that contain important visual features of the image, such as sharpness.
[0048] It should also be understood that, in this embodiment, the cross-modal fusion vector is input into the CycleGAN model, and the generator learns a mapping from one style to another, which can be analogous to a "style vector." The style vector captures the stylistic features of the image, such as color usage.
[0049] The attention mechanism in this embodiment dynamically adjusts the weights of the quality and style vectors based on the user's description, and then performs a weighted fusion of the quality and style vectors according to the adjusted weights to generate a comprehensive feature vector. This feature vector contains both the quality and style features of the image and is used as input to the generative adversarial network (GAN) model. Through continuous training of the GAN model, a customized image generation model is obtained.
[0050] Step S13: Deploy the customized image generation model; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs the customized image.
[0051] Specifically, customized image generation models are deployed in real-world application scenarios, such as advertising design and game development, to generate customized images. For example, in the scenario of advertising design, for a specific product, the user inputs a textual description of the product and their desired image style, based on the image content they want to generate. The customized image generation model then generates an image specific to that product, thus meeting the user's customized image requirements.
[0052] In some embodiments of this application, the customized image generation method based on large model fusion further includes: fine-tuning the generated customized image according to real-time acquired user interaction behavior features.
[0053] Furthermore, the step of fine-tuning the generated customized image based on real-time acquired user interaction behavior features includes: S41: preprocessing the real-time acquired user interaction behavior features to convert them into a data format suitable for the customized image generation model; S42: inputting the format-converted user interaction behavior data into the customized image generation model to adjust the customized image generated by the customized image generation model.
[0054] It should be understood that user interaction behavior includes positive behavior, negative behavior, and neutral behavior. For example, a user describing the generated customized image using positive language such as "good," "great," or "really good"; or giving the generated customized image four stars or higher out of five, is considered positive behavior. Negative behavior includes, but is not limited to: a user being dissatisfied with the generated customized image and choosing to regenerate it; or giving it a score of two stars or lower. Neutral behavior includes, but is not limited to: a user giving it three stars; or a user not commenting on the generated customized image.
[0055] After the customized image generation model generates a customized image, users can edit text descriptions and input them into the model according to their further needs, thereby adjusting parameters such as the style and content of the customized image. Specifically, using the example above, when a user inputs the text "The overall style of the presented image is quite good, but I hope the falling ginkgo leaves have an atmospheric feel," the customized image generation model analyzes and interprets the user's input text description, adjusts the initially generated customized image, and modifies parameters such as color saturation. This allows for continuous adjustments to the customized image until the desired image is generated.
[0056] In some embodiments of this application, the preprocessing methods include, but are not limited to: data deduplication, missing value handling, outlier handling, and standardization. Data deduplication refers to identifying and deleting duplicate records in the dataset; missing value handling refers to processing missing or lost data in the dataset, with common methods including deleting records containing missing values and filling in missing values using the mean or median; outlier handling refers to processing points in the dataset that significantly deviate from other observations, with common methods including deletion and replacement with the median or mean; standardization refers to scaling the data proportionally to fit it into a small, specific interval, with common standardization methods including min-max standardization, scaling all data to the [0, 1] interval. It should be noted that the preprocessing method should be selected according to the specific scenario and is not limited to the methods described above.
[0057] In some embodiments of this application, the customized image generation method based on large model fusion further includes: collecting user feedback data on the generated customized image and inputting it into the generative adversarial network model to optimize the customized image generation model.
[0058] Specifically, users can rate and provide feedback on the generated images, which can take the form of a rating system or a text comment box. The processed feedback data is then combined with the original image feature vectors to form a new training set. This feedback data serves as additional supervisory signals to guide the training process of the generative adversarial network model, thereby continuously optimizing the customized image generation model.
[0059] Figure 5 is a schematic block diagram of a customized image generation system based on large model fusion provided in an embodiment of this application. As shown in Figure 5, the customized image generation system 500 based on large model fusion provided in an embodiment of this application includes:
[0060] The cross-modal fusion module 501 is used to fuse the acquired sample image data and text data to obtain a cross-modal fusion vector;
[0061] The multi-model collaboration module 502 is used to input the cross-modal fusion vector into the Stable Diffusion model and the CycleGAN model for encoding, respectively, to obtain quality vector and style vector; after fusing the quality vector and style vector using an attention mechanism, it is input into the generative adversarial network model for training to obtain a customized image generation model;
[0062] A customized image generation module 503 is used to deploy the customized image generation model; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs a customized image.
[0063] The cross-modal fusion module 501 of this application embodiment can seamlessly integrate two input modalities: sample images and text descriptions. Utilizing techniques such as attention mechanisms and graph neural network (GNN) models, the cross-modal fusion module 501 achieves deep fusion of image features and text semantics, ensuring that the generated image not only conforms to the user-specified style but also accurately conveys the intent of the text description. This satisfies the user's customization needs for images, generating the desired image style and content.
[0064] The multi-model collaboration module 502 in this embodiment integrates multiple advanced large-scale model technologies, such as the Stable Diffusion model, CycleGAN model, and generative adversarial network model, and overcomes the challenges of integrating these large-scale models. Through continuous model training, a customized image generation model is obtained. The trained customized image generation model is deployed in specific application scenarios to generate customized images, meeting users' personalized needs for creating images.
[0065] The customized image generation module 503 of this application embodiment generates high-quality, highly realistic images based on the content of the text data according to the image style of the provided sample images through adversarial training between the generator and the discriminator.
[0066] In some embodiments of this application, the customized image generation system 500 based on large model fusion also includes an interactive interface. The interactive interface allows users to adjust parameters such as the style and content of the image in real time during the image generation process. For example, users can adjust parameters by inputting text descriptions according to their further needs and view the effects of the adjustments promptly. Simultaneously, the interactive interface also integrates a user feedback mechanism to collect user satisfaction and suggestions regarding the generated images, which are used to continuously optimize the model and improve the user experience.
[0067] It should be understood that the specific process of each module performing the above-mentioned steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.
[0068] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0069] Figure 6 is a schematic block diagram of an electronic terminal provided in an embodiment of this application. As shown in Figure 6, the electronic terminal 600 includes at least one processor 601, a memory 602, at least one network interface 603, and a user interface 605. The various components in the electronic terminal 600 are coupled together via a bus system 604. It is understood that the bus system 604 is used to implement communication between these components. In addition to a data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as a bus system in Figure 6.
[0070] The user interface 605 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0071] It is understood that memory 602 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0072] In this embodiment of the invention, the memory 602 is used to store various types of data to support the operation of the electronic terminal 600. Examples of this data include: any executable program for operation on the electronic terminal 600, such as the operating system 6021 and application programs 6022; the operating system 6021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 6022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The customized image generation method based on large model fusion provided in this embodiment of the invention can be included in the application program 6022.
[0073] The methods disclosed in the above embodiments of the present invention can be applied to processor 601, or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 601 or by instructions in the form of software. The processor 601 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 601 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 601 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0074] In an exemplary embodiment, the electronic terminal 600 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.
[0075] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to perform the method of any of the embodiments shown in FIG1 to FIG4.
[0076] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code that, when run on a computer, causes the computer to perform the method of any of the embodiments shown in FIG1 to FIG4.
[0077] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0078] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0079] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0080] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0082] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0083] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs, etc.).
[0084] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0086] In summary, this invention provides a customized image generation method, system, medium, terminal, and program product based on large-scale model fusion. It can not only receive user-provided image samples as style references but also receive detailed text descriptions as content guidance. By integrating multiple advanced large-scale model technologies such as the Stable Diffusion model, CycleGAN model, and Generative Adversarial Network model, it achieves unprecedented image generation effects, thereby meeting users' personalized requirements and deep customization needs for image creation. It provides users with a more intelligent, flexible, and innovative image generation solution. Therefore, this application effectively overcomes the various shortcomings of existing technologies and has high industrial application value.
[0087] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A customized image generation method based on large model fusion, characterized in that, include: The acquired sample image data and text data are fused to obtain a cross-modal fusion vector; The cross-modal fusion vectors are input into the Stable Diffusion model and the CycleGAN model for encoding to obtain quality vectors and style vectors, respectively. The quality vectors and style vectors are then fused based on the attention mechanism and input into the generative adversarial network model for training to obtain a customized image generation model. The customized image generation model is deployed; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs the customized image.
2. The customized image generation method based on large model fusion according to claim 1, characterized in that, The process of fusing the acquired sample image data and text data to obtain a cross-modal fusion vector includes: Obtain the sample image data and text data respectively; The sample image data is used to extract features using a graph neural network model to obtain image features; the text data is used to extract features using a Transformer model to obtain text semantic features. The cross-modal attention mechanism based on the Transformer model fuses the image features and the text semantic features to obtain a cross-modal fusion vector.
3. The customized image generation method based on large model fusion according to claim 1, characterized in that, The method further includes: fine-tuning the generated customized image based on real-time acquired user interaction behavior features.
4. The customized image generation method based on large model fusion according to claim 3, characterized in that, The step of fine-tuning the generated customized image based on real-time acquired user interaction behavior features includes: The real-time acquired user interaction behavior features are preprocessed to convert them into a data format suitable for the customized image generation model; The user interaction behavior data after format conversion is input into the customized image generation model, and the customized image generated by the customized image generation model is adjusted.
5. The customized image generation method based on large model fusion according to claim 4, characterized in that, The preprocessing methods include: data deduplication, missing value handling, outlier handling, and standardization.
6. The customized image generation method based on large model fusion according to claim 1, characterized in that, The method further includes: collecting user feedback data on the generated customized image and inputting it into the generative adversarial network model to optimize the customized image generation model.
7. A customized image generation system based on large model fusion, characterized in that, include: The cross-modal fusion module is used to fuse the acquired sample image data and text data to obtain a cross-modal fusion vector; The multi-model collaboration module is used to input the cross-modal fusion vector into the Stable Diffusion model and the CycleGAN model for encoding to obtain quality vector and style vector respectively; after fusing the quality vector and style vector using an attention mechanism, the vector is input into the generative adversarial network model for training to obtain a customized image generation model. A customized image generation module is used to deploy the customized image generation model; the customized image generation model converts the text data into corresponding image data according to the image style of the sample image and then outputs a customized image.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
9. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to implement the method as described in any one of claims 1 to 6.
10. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1 to 6.