Text-to-image generation model training method based on artificial intelligence, and image generation method

By directly using image feature extraction and alignment to text feature space, the problem of image-text pair training dependent on image-text pairs is solved, and efficient image generation and semantic understanding ability are improved.

WO2025176092A1PCT designated stage Publication Date: 2025-08-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/077612
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-17
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The training of existing text-general graphics models relies on high-quality image-text pairs, resulting in poor training results, and the automatically generated text annotation quality is average, making it difficult to effectively generate high-quality images.

Method used

By directly using image features to extract and align them to text feature space, the dependence on image-text pairs is broken, and the rich semantic information in the image is used for model training to generate target images.

Benefits of technology

It improves the semantic understanding and image generation ability of literary graphics models, reduces dependence on high-quality text annotations, and improves training efficiency and generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025077612_28082025_PF_FP_ABST
    Figure CN2025077612_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A text-to-image generation model training method based on artificial intelligence. The method comprises: acquiring a sample image (101); performing image feature extraction on the sample image, so as to obtain a first image feature of a first sample image (102); aligning the first image feature of the sample image with a text feature space, so as to obtain a second image feature of the first sample image (103); performing image generation on the basis of the second image feature of the first sample image, so as to obtain a prediction result (104); and performing model training on the basis of the prediction result, so as to obtain a text-to-image generation model (105).
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence-based cultural graph model training method and image generation method

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2024101890930, filed on February 20, 2024, entitled “Artificial Intelligence-Based Text-Graph Model Training Method and Image Generation Method,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to artificial intelligence technology, and in particular to an artificial intelligence-based text graph model training method, image generation method, device, electronic device, computer-readable storage medium and computer program product. Background Art

[0004] Text-to-image generation (TIG) is a major area of ​​AI-generated content (AIGC) and one of the most promising research areas in the field of AI. By converting text descriptions into images, TIG models can assist humans in content creation and are widely used in various application scenarios, such as original art design.

[0005] In the solutions provided by related technologies, the training of the text-graph model is usually achieved through a large number of image-text pairs. The effect of model training is highly dependent on the quality of the image-text pairs. However, in actual situations, the quality of text annotations in image-text pairs is relatively average, resulting in poor training effect of the text-graph model, and further resulting in poor image generation effect based on the trained text-graph model. Summary of the Invention

[0006] According to various embodiments provided in the present application, a method for training a cultural graph model based on artificial intelligence, an image generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product are provided.

[0007] This application provides an artificial intelligence-based text graph model training method, which is executed by a computer device, and the method includes:

[0008] Get a sample image;

[0009] performing image feature extraction on the sample image to obtain a first image feature of the sample image;

[0010] Aligning the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image;

[0011] Performing image generation based on the second image feature of the sample image to obtain a prediction result; and

[0012] Model training is performed based on the prediction results to obtain a cultural graph model.

[0013] In another aspect, the present application provides an artificial intelligence-based image generation method, performed by a computer device, the method comprising:

[0014] Get prompt information;

[0015] Extracting prompt features from the prompt information to obtain prompt features of the prompt information aligned with a text feature space;

[0016] Generate target condition features according to the prompt features of the prompt information; and

[0017] The target image is generated by guiding the cultural graph model according to the target condition characteristics; wherein the cultural graph model is trained according to the cultural graph model training method based on artificial intelligence.

[0018] On the other hand, the present application provides an artificial intelligence-based text graph model training device, comprising:

[0019] A first acquisition module is used to acquire a sample image;

[0020] a first feature extraction module, configured to extract image features from the sample image to obtain a first image feature of the sample image;

[0021] A first alignment module is used to align the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image;

[0022] A first prediction module is configured to generate an image based on a second image feature of the sample image to obtain a prediction result; and

[0023] The training module is used to perform model training according to the prediction results to obtain a cultural graph model.

[0024] On the other hand, the present application provides an image generation device based on artificial intelligence, comprising:

[0025] The second acquisition module is used to obtain prompt information;

[0026] A second feature extraction module is used to extract prompt features from the prompt information to obtain prompt features of the prompt information aligned to a text feature space;

[0027] A second generating module is configured to generate target condition features according to the prompt features of the prompt information; and

[0028] The second prediction module is used to guide the Wensheng graph model to generate an image according to the target condition characteristics to obtain a target image; wherein the Wensheng graph model is trained according to the Wensheng graph model training method based on artificial intelligence.

[0029] On the other hand, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the steps of the method embodiments of the present application when executing the computer program.

[0030] In another aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of each method embodiment of the present application.

[0031] On the other hand, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, performs the steps of each method embodiment of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0033] FIG1 is a schematic diagram of an architecture of an artificial intelligence-based image generation system provided in an embodiment of the present application;

[0034] FIG2 is a schematic diagram of the structure of a server provided in an embodiment of the present application;

[0035] FIG3 is a schematic structural diagram of a terminal device provided in an embodiment of the present application;

[0036] FIG4 is a schematic diagram of a first flow chart of a method for training a text graph model based on artificial intelligence according to an embodiment of the present application;

[0037] FIG5 is a schematic diagram of a training multimodal model provided in an embodiment of the present application;

[0038] FIG6 is a flow chart of an artificial intelligence-based image generation method provided in an embodiment of the present application;

[0039] FIG7A is a comparative diagram of style binding provided in an embodiment of the present application;

[0040] FIG7B is a comparative schematic diagram of object binding provided in an embodiment of the present application;

[0041] FIG8 is a schematic diagram of a process of training a LoRA model provided in an embodiment of the present application;

[0042] FIG9 is a schematic diagram of a training alignment network provided in an embodiment of the present application;

[0043] FIG10 is a schematic diagram of training embedded image features provided by an embodiment of the present application;

[0044] FIG11 is a schematic diagram of a training LoRA model provided in an embodiment of the present application;

[0045] FIG12 is a schematic diagram of outputting a target image for a prompt image according to an embodiment of the present application;

[0046] FIG13 is a schematic diagram of outputting a target image for a prompt text according to an embodiment of the present application;

[0047] FIG14 is a schematic diagram of outputting a target image for a prompt image and prompt text provided by an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0049] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0050] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. In the following description, the term "plurality" refers to at least two.

[0051] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0053] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0054] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0055] 1) Artificial Intelligence (AI): It is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0056] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0057] In the embodiments of the present application, the cultural graph model involved is a model constructed based on the principles of artificial intelligence.

[0058] 2) Computer Vision (CV): Computer vision is the science of enabling machines to "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying, tracking, and measuring objects, and then performing image processing to transform the images into images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0059] 3) Natural Language Processing (NLP): This is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life, and is closely related to linguistics research; it also involves computer science and mathematics. Similarly, large model technology can be used in the field of NLP. For example, pre-trained models in the field of NLP, such as large language models (LLM), can be quickly and widely applied to specific downstream tasks after fine-tuning. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0060] In the embodiments of this application, computer vision technology and natural language processing technology are used to study how to convert text descriptions into suitable images.

[0061] 4) Machine Learning (ML): This is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0062] In an embodiment of the present application, the training of the document graph model can be implemented based on the principle of machine learning.

[0063] 5) Text-to-Graph Model: This model is used to convert text descriptions into images. It guides image generation based on conditional features in a conditional feature space. In traditional text-to-graph models, the conditional feature space is the text feature space.

[0064] The embodiment of the present application does not limit the type of the Wensheng graph model, for example, it can be a stable diffusion (SD) model, an SDXL model, etc.

[0065] 6) Multimodal model: A model that can accept multiple different input methods (such as text, images, voice, and video). Multimodal models can process and analyze multiple types of data, thereby more comprehensively understanding and utilizing various information.

[0066] In an embodiment of the present application, the multimodal model includes at least an image feature extraction network for implementing image feature extraction processing and a text feature extraction network for implementing text feature extraction processing. The image feature extraction network is also called an image encoder (Image Encoder), and the text feature extraction network is also called a text encoder (Text Encoder). The embodiment of the present application does not limit the type of multimodal model. For example, it can be a contrastive language-image pre-training (CLIP) model, a large language and vision assistant (LLAVA) model, a bootstrapping language-image pre-training (BLIP) model, etc.

[0067] 7) Pre-training Model (PTM): Also known as a cornerstone model or large model, this refers to a large-parameter deep neural network (DNN). It is trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, the PTM extracts common features from the data. Through techniques such as fine tuning, efficient parameter fine tuning (PEFT), and prompt-tuning, it is then adapted for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. Based on the data modality processed, PTMs can be categorized into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Pre-trained models are a crucial tool for AIGC and can also serve as a universal interface for connecting multiple task-specific models.

[0068] In the embodiment of the present application, the text graph model and the multimodal model used can both be pre-trained models, which can improve the training efficiency, that is, better results can be achieved after a small amount of training.

[0069] 8) Feature space: This is a key concept in machine learning. Sample features can be represented as vectors, and the space composed of these vectors is called a feature space. By mapping samples into the feature space, various machine learning algorithms can be executed, including classification, clustering, and regression.

[0070] 9) Binding ID: This refers to training a cultural graph model based on sample images with a specific ID, so that the trained cultural graph model is bound to the ID, that is, the trained cultural graph model can generate images that conform to the ID during the model inference stage. ID refers to a certain characteristic of an image, such as style, objects contained (such as people), quality, etc. For example, a cultural graph model can be trained based on multiple sample images with a specific style, so that the images generated by the trained cultural graph model during the model inference stage also have the style; the cultural graph model can be trained based on multiple sample images including a specific object, so that the images generated by the trained cultural graph model during the model inference stage also include the object.

[0071] In the solutions provided by related technologies, text-based graph models are usually trained by using a large number of image-text pairs. However, this solution has at least the following problems:

[0072] The effectiveness of text-to-graph models is highly dependent on the quality of image-text pairs. In practice, achieving effective text-to-graph models requires significant time and effort to generate high-quality text annotations (text annotations refer to the text within the image-text pair). However, text annotations automatically generated by multimodal models (such as BLIP) are often of mediocre quality and difficult to avoid inaccurate annotations caused by model "hallucinations."

[0073] Since text cannot fully summarize the rich semantic information in images, and multiple sample images with severe homogeneity are usually used in the model training stage (multiple sample images all have a certain ID), the model parameters will inevitably converge to a narrow space during the model training stage, which will have a negative impact on the semantic understanding and generation capabilities of the text-to-image model.

[0074] The embodiments of the present application provide an artificial intelligence-based text graph model training method, image generation method, device, electronic device, computer-readable storage medium, and computer program product, which can use only images to drive the training of text graph models, breaking the dependence on image-text pairs; at the same time, directly using the rich semantic information contained in the image to guide model training can greatly retain the capabilities of the original model. The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as various types of terminal devices or as a server.

[0075] Refer to Figure 1, which is an architectural diagram of an artificial intelligence-based image generation system 100 provided in an embodiment of the present application. The terminal device 400 is connected to the server 200 through the network 300, and the server 200 is connected to the database 500, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.

[0076] In some embodiments, taking the electronic device as a server as an example, the artificial intelligence-based text graph model training method and image generation method provided in the embodiments of the present application can be implemented by the server. For example, the server 200 can obtain a sample image from the database 500, perform image feature extraction on the sample image, and obtain the first image feature of the sample image. The server 200 aligns the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image, performs image generation processing based on the second image feature of the first sample image, obtains a prediction result, performs model training based on the prediction result, and obtains a text graph model. Among them, the storage location of the first sample image is not limited, and is not limited to the database 500. For example, it can also be stored in the distributed file system, blockchain, and other locations of the server 200.

[0077] After completing the training of the text-based graph model, the server 200 can obtain the prompt information sent by the terminal device 400, extract the prompt features of the prompt information, obtain the prompt features of the prompt information aligned to the text feature space, generate the target condition features based on the prompt features of the prompt information, guide the trained text-based graph model to perform image generation processing based on the target condition features, obtain the target image, and send the target image to the terminal device 400 as a response to the prompt information.

[0078] In some embodiments, taking the electronic device as a terminal device as an example, the artificial intelligence-based Wensheng graph model training method and image generation method provided in the embodiments of the present application can be implemented by the terminal device. For example, the terminal device 400 can obtain a first sample image from a local or other storage location, enter the model training phase based on the first sample image, and obtain a trained Wensheng graph model. Then, the terminal device 400 can receive prompt information input by the user, enter the model inference phase based on the prompt information, and obtain the target image output by the trained Wensheng graph model. The terminal device 400 can display the target image as a response to the prompt information input by the user.

[0079] In some embodiments, the artificial intelligence-based Wensheng graph model training method provided in the embodiments of the present application can be implemented by a server, and the artificial intelligence-based image generation method provided in the embodiments of the present application can be implemented by a terminal device. For example, the server 200 can obtain a first sample image from the database 500, enter the model training phase according to the first sample image, and obtain a trained Wensheng graph model. Then, the server 200 sends the trained Wensheng graph model (which can be the model parameters of the trained Wensheng graph model) to the terminal device 400, so that the terminal device 400 deploys the trained Wensheng graph model locally. In this way, the computing power of the server 200 can be used to improve the training efficiency; after the terminal device 400 deploys the trained Wensheng graph model locally, it can provide model reasoning capabilities locally.

[0080] In some embodiments, the terminal device 400 or the server 200 can implement the artificial intelligence-based text graph model training method and the artificial intelligence-based image generation method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP, and the small program can be controlled by the user to run or close. In short, the above-mentioned computer program can be an application, module or plug-in in any form.

[0081] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0082] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. For example, the embodiments of the present application can be applied to the design of game concept art. Designers only need to input prompt information, and the trained text graph model can automatically generate the target image. In this way, the automatic design of game concept art can be realized, or it is convenient for designers to modify the generated target image to obtain the game concept art, which can effectively reduce the workload and improve design efficiency.

[0083] In some embodiments, various data involved in the embodiments of the present application (such as various samples, model parameters, etc.) can be stored in the blockchain, and the data credibility can be guaranteed based on the tamper-proof characteristics of the blockchain.

[0084] Taking the electronic device provided in the embodiment of the present application as a server as an example, refer to Figure 2, which is a structural diagram of the server 200 provided in the embodiment of the present application. The server 200 shown in Figure 2 includes: at least one processor 210, a memory 250 and at least one network interface 220. The various components in the server 200 are coupled together through a bus system 240. It can be understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 240 in Figure 2.

[0085] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0086] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0087] The memory 250 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0088] In some embodiments, the memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0089] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0090] A network communication module 252 for reaching other computing devices via one or more (wired or wireless) network interfaces 220 , exemplary network interfaces 220 including Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB);

[0091] In some embodiments, the AI-based Wensheng graph model training device provided in the embodiments of the present application can be implemented in software. FIG2 shows an AI-based Wensheng graph model training device 255 stored in memory 250. The device 255 can be software in the form of a program or plug-in, and includes the following software modules: a first acquisition module 2551, a first feature extraction module 2552, a first alignment module 2553, a first prediction module 2554, and a training module 2555. These modules are logical and can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0092] Taking the electronic device provided in the embodiment of the present application as a terminal device as an example, it is understandable that, for the case where the electronic device is a server, parts of the structure shown in Figure 3 (such as a user interface, a presentation module, and an input processing module) can be omitted. Referring to Figure 3, Figure 3 is a structural diagram of a terminal device 400 provided in an embodiment of the present application. The terminal device 400 shown in Figure 3 includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understandable that the bus system 440 is used to realize connection communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 440 in Figure 3.

[0093] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0094] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0095] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0096] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0097] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0098] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0099] A network communication module 452 for reaching other computing devices via one or more (wired or wireless) network interfaces 420 , exemplary network interfaces 420 including Bluetooth, WiFi, and USB;

[0100] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0101] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0102] In some embodiments, the AI-based image generation device provided in the embodiments of the present application can be implemented in software. FIG3 shows an AI-based image generation device 455 stored in memory 450. The AI-based image generation device 455 can be software in the form of a program or plug-in, and includes the following software modules: a second acquisition module 4551, a second feature extraction module 4552, a second generation module 4553, and a second prediction module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions implemented. The functions of each module will be described below.

[0103] The artificial intelligence-based text graph model training method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the electronic device provided in the embodiment of the present application.

[0104] [Corrected 12.03.2025 according to Rule 91] See Figure 4, which is a flow chart of the artificial intelligence-based text graph model training method provided in an embodiment of the present application, which will be explained in conjunction with the steps shown in Figure 4.

[0105] In step 101, a sample image is acquired.

[0106] Here, a training data set of the Wensheng graph model is obtained, wherein the training data set of the Wensheng graph model includes a plurality of first sample images.

[0107] In some embodiments, in order to improve the training effect of the text graph model, the number of sample images can be multiple.

[0108] In some embodiments, multiple sample images in the training dataset of the culture graph model all have the same ID, so that the trained culture graph model is bound to the ID. The ID refers to a certain characteristic of the image, such as style, objects (such as people) contained, quality, etc. For example, the training dataset of the culture graph model includes multiple sample images with the same style, so that the target image generated by the trained culture graph model during the model inference stage also has the style; the training dataset of the culture graph model includes multiple sample images containing the same object, so that the target image generated by the trained culture graph model during the model inference stage also contains the object.

[0109] In step 102, image feature extraction processing is performed on the sample image to obtain a first image feature of the sample image.

[0110] Here, image feature extraction processing is performed on the sample image to extract semantic information in the sample image and obtain a first image feature of the sample image.

[0111] In some embodiments, an image feature extraction network may be used to extract image features from the sample image to obtain a first image feature of the sample image. The image feature extraction network may be pre-trained, such as an image feature extraction network in a multimodal model.

[0112] In step 103 , the first image feature of the sample image is aligned to the text feature space to obtain the second image feature of the sample image.

[0113] A common use of text-based graph models is to generate images based on text. That is, text-based graph models typically guide image generation based on text features in a text feature space. Therefore, in this embodiment of the present application, the first image features of a sample image are aligned (linearly aligned) to the text feature space to obtain the second image features of the sample image. In this way, the second image features of the sample image can be treated as text features in the text feature space to guide image generation.

[0114] Among them, aligning features to the text feature space means transforming (or mapping) the features so that the dimensions of the transformed features conform to the form of the dimensions in the text feature space. The form of the dimensions in the text feature space is related to the type of text graph model and can be pre-set. For example, the form of the dimensions in the text feature space can be set to n*768, where 1*768 represents a token. Based on the form of the dimensions in the text feature space, the dimensions of the second image feature can also be pre-set, for example, to 3*768.

[0115] For example, the dimension in the text feature space is n*768, and the dimension of the first image feature is 1*1024. The first image feature with a dimension of 1*1024 can be converted into a second image feature with a dimension of 3*768, so that the second image feature with a dimension of 3*768 is located in the text feature space.

[0116] In some embodiments, the conversion from the first image feature to the second image feature can be achieved by using an alignment network, wherein the alignment network can be preset or trained.

[0117] It is worth noting that the alignment network, image feature extraction network, and text feature extraction network involved in the embodiments of the present application all refer to artificial neural networks (ANN).

[0118] In step 104, an image is generated based on the second image feature of the sample image to obtain a prediction result.

[0119] In this embodiment, the sample condition feature is determined according to the second image feature of the sample image, and the image is generated based on the sample condition feature to obtain a prediction result.

[0120] The sample conditional features serve as guiding conditions for the image generation of the text-based graph model to be trained.

[0121] Since the second image features of the sample image are already in the text feature space, sample conditional features can be generated based on the second image features of the sample image. For example, the second image features of the sample image can be directly determined as the sample conditional features, or the second image features of the sample image can be further processed to obtain the sample conditional features. Therefore, directly determining the second image features of the sample image as the sample conditional features makes the training process more lightweight, which can reduce the amount of data processing during the training process and improve processing efficiency. Further processing of the second image features of the sample image to obtain the sample conditional features has richer detailed features, making the text graph model more effective in generating images from the model.

[0122] In step 105, model training is performed based on the prediction results to obtain a text graph model.

[0123] In this embodiment, a loss value is determined according to the prediction result, and a model is trained based on the loss value to obtain a text graph model.

[0124] Here, the to-be-trained Wensheng graph model is guided to perform image generation processing according to the sample conditional features, that is, the to-be-trained Wensheng graph model refers to the sample conditional features to perform image generation processing. A loss value is calculated based on the prediction results generated during the image generation process to obtain a loss value, and the Wensheng graph model is trained based on the loss value until the Wensheng graph model training is completed. Among them, the loss value can be a mean squared error (MSE) loss value, and of course it can also be other types of loss values. According to the prediction results, the loss generated by the Wensheng graph model during the training process can be determined, and the model can be adjusted based on the generated loss and then continue training, so that the model gradually enhances the model's feature extraction ability, feature alignment ability and image generation ability during the training process, thereby gradually reducing the generated loss.

[0125] During the training phase of the text-based graph model, only sample images are required, eliminating the time and effort required for text annotation. This eliminates the reliance on image-text pairs and enables an automated training process that requires no human intervention. Furthermore, by aligning the primary image features of the sample images to the text feature space, the resulting secondary image features contain rich semantic information from the sample images, effectively improving the semantic understanding and image generation capabilities of the trained text-based graph model.

[0126] The conditions for completing the training of the raw graph model in the embodiment of the present application are not limited. For example, it can be reaching a preset number of training times, or the performance index reaching an index threshold.

[0127] In one embodiment, the sample image includes at least one of a first sample image, a second sample image, or a third sample image, the first sample image is used to train a text-image model, the second sample image is used to train an alignment network, the alignment network is used to align a first image feature of the first sample image to a text feature space, and the third sample image is used to train a multimodal model, the multimodal model is used to extract the first image feature of the first sample image;

[0128] The sample condition feature includes at least one of the first sample condition feature, the second sample condition feature or the third sample condition feature, the prediction result includes at least one of the first prediction result, the second prediction result or the third prediction result, and the loss value includes at least one of the first loss value, the second loss value or the third loss value.

[0129] At least one of the first sample image, the second sample image, or the third sample image can be used to train the Vincent graph model. The second sample image can be used to train the alignment network, and the third sample image can be used to train the multimodal model. The third sample image is also used to train and update the embedded image features.

[0130] In this embodiment, when the sample image includes a first sample image, the sample conditional feature includes a first sample conditional feature, the prediction result includes a first prediction result, and the loss value includes a first loss value. Image feature extraction is performed on the first sample image to obtain a first image feature of the first sample image. The first image feature of the first sample image is aligned to the text feature space to obtain a second image feature of the first sample image. A first sample conditional feature is determined based on the second image feature of the first sample image, and an image is generated based on the first sample conditional feature to obtain a first prediction result. A first loss value is determined based on the first prediction result, and a model is trained based on the first loss value to obtain a text-based image model.

[0131] When the sample image includes a second sample image, the sample conditional feature includes a second sample conditional feature, the prediction result includes a second prediction result, and the loss value includes a second loss value. Image feature extraction is performed on the second sample image to obtain a first image feature of the second sample image. The first image feature of the second sample image is aligned to the text feature space to obtain a second image feature of the second sample image. A second sample conditional feature is determined based on the second image feature of the second sample image. An image is generated based on the second sample conditional feature to obtain a second prediction result. A second loss value is determined based on the second prediction result. A model is trained based on the second loss value to obtain a text-based image model.

[0132] Furthermore, when the sample image includes a second sample image, a second sample text corresponding to the second sample image is obtained, and text features of the second sample text are extracted. Feature fusion processing is performed on the second image features of the second sample image and the text features of the second sample text to obtain a second sample conditional feature. A second sample conditional feature is generated based on the second image features of the second sample image, and image generation is performed based on the second sample conditional feature to obtain a second prediction result.

[0133] When the sample image includes a third sample image, the sample conditional feature includes the third sample conditional feature, the prediction result includes a third prediction result, and the loss value includes a third loss value. Image feature extraction is performed on the third sample image to obtain a first image feature of the third sample image. The first image feature of the third sample image is aligned to the text feature space to obtain a second image feature of the third sample image. A third sample conditional feature is determined based on the second image feature of the third sample image. An image is generated based on the third sample conditional feature to obtain a third prediction result. A third loss value is determined based on the third prediction result. Model training is performed based on the third loss value to obtain a text-based graph model.

[0134] In this embodiment, model training can be performed in combination with at least one sample image, and the alignment network, multimodal model, and text graph model can all be trained separately or jointly.

[0135] The multimodal model is trained separately, for example, by obtaining a plurality of third sample images and third sample texts corresponding to the plurality of third sample images. The first image features of each third sample image are extracted using an image feature extraction network in the multimodal model. The text features of each third sample text are extracted using a text feature extraction network in the multimodal model. A feature extraction loss value is determined based on the first image features of each third sample image and the text features of each third sample text, and the multimodal model is trained based on the feature extraction loss value to obtain a trained multimodal model.

[0136] After the multimodal model training is completed, the alignment network is trained separately. For example, a second sample image and a corresponding second sample text are obtained, and the first image features of the second sample image are extracted through the image feature extraction network in the trained multimodal model. The text features of the second sample text are extracted through the text feature extraction network in the trained multimodal model. The first image features of the second sample image are aligned to the text feature space through the alignment network to be trained to obtain the second image features of the second sample image. The second image features of the second sample image and the text features of the second sample text are subjected to feature fusion processing to obtain second sample conditional features, and an image is generated based on the second sample conditional features to obtain a second prediction result. A second loss value is determined based on the second prediction result, and the alignment network is trained based on the second loss value to obtain a trained alignment network.

[0137] After the multimodal model and alignment network training are complete, the text graph model is trained separately. For example, a first sample image is obtained, and the first image features of the first sample image are extracted using the image feature extraction network in the trained multimodal model. The first image features of the first sample image are aligned to the text feature space using the trained alignment network to obtain the second image features of the first sample image. A first sample conditional feature is generated based on the second image feature of the first sample image, and image generation is performed based on the first sample conditional feature to obtain a prediction result. A first loss value is determined based on the prediction result, and the text graph model is trained based on the first loss value.

[0138] In this embodiment, the alignment network, multimodal model, and text-based graph model are all trained separately, which reduces the complexity and difficulty of training and helps improve training efficiency. The multimodal model is used to extract image and text features, the alignment network is used to align image features to the text feature space to obtain sample-conditioned features, and the text-based graph model is used to generate images using the sample-conditioned features. Joint training can better learn the potential connections between image feature extraction, text feature extraction, feature alignment, and image generation, thereby improving the accuracy of image feature extraction, text feature extraction, and the accuracy of image feature alignment to the text feature space, thereby improving the accuracy of image generation.

[0139] It is worth noting that in the above steps, only one sample image is used as an example to illustrate the model training stage of the text-generated graph model, but multiple sample images may also participate in the model training stage of the text-generated graph model.

[0140] It is worth noting that training the Vincent graph model based on the loss value may refer to backpropagating the loss value through the Vincent graph model and updating the model parameters of the Vincent graph model along the gradient descent direction during the backpropagation process. The model parameters involved in the embodiments of the present application are used to define the basic structure and feature representation capabilities of the artificial neural network. The model parameters may include weight parameters and bias parameters. The essence of training the model is to update the model parameters.

[0141] In some embodiments, the above-mentioned training of the Vincent graph model based on the loss value can be achieved by performing any of the following processing: learning the first sub-incremental parameter and the second sub-incremental parameter based on the loss value, performing parameter fusion processing on the first sub-incremental parameter and the second sub-incremental parameter to obtain the incremental parameter, and updating the model parameters of the Vincent graph model based on the incremental parameter; learning the incremental parameter based on the loss value, and updating the model parameters of the Vincent graph model based on the incremental parameter; wherein the incremental parameter and the model parameter are the same in dimension.

[0142] Here, two methods are provided for training the Wensheng graph model based on the loss value:

[0143] 1) learning a first sub-incremental parameter and a second sub-incremental parameter based on the loss value, performing parameter fusion processing on the first sub-incremental parameter and the second sub-incremental parameter to obtain an incremental parameter, and updating the model parameters of the Wensheng graph model based on the incremental parameter, wherein the incremental parameter and the model parameter have the same dimension, and the parameter fusion processing can be a product processing.

[0144] For example, the model parameters of the Wensheng graph model are represented by W, which is an m×n matrix. The incremental parameters ΔW that need to be learned are also m×n matrices. Based on this, the incremental parameters ΔW to be learned can be decomposed into a first sub-incremental parameter ΔW1 and a second sub-incremental parameter ΔW2 according to the rank r, where ΔW1 is an m×r matrix and ΔW2 is an r×n matrix. The rank r can be set according to the actual application scenario. After learning the first sub-incremental parameter ΔW1 and the second sub-incremental parameter ΔW2 based on the loss value, the first sub-incremental parameter ΔW1 and the second sub-incremental parameter ΔW2 are fused to obtain the incremental parameter, i.e., ΔW=ΔW1×ΔW2. Then, the model parameters W of the Wensheng graph model are updated according to the incremental parameters ΔW, i.e., W=W+ΔW is executed.

[0145] In essence, this method is to regard the Wensheng graph model as the base model and train the low-rank adaptation (LoRA) model according to the loss value. The model parameters of the trained LoRA model include the first sub-incremental parameters and the second sub-incremental parameters learned according to the loss value. Then, the trained LoRA model is inserted into the base model to obtain the trained Wensheng graph model, that is, the trained Wensheng graph model = base model + trained LoRA model.

[0146] Method 1) requires fewer parameters to learn, which can improve model training efficiency and reduce computing resource consumption during model training. At the same time, the LoRA model uses matrix multiplication to store model parameters, so it occupies less storage space, is easy to deploy, can be plug-and-play, and has good portability on different base models (different base models have the same network structure but different model parameters).

[0147] 2) Learn incremental parameters based on the loss value and update the model parameters of the Wensheng graph model based on the incremental parameters. For example, you can directly learn the incremental parameter ΔW and perform the operation W = W + ΔW. Method 2) requires a large number of parameters to learn, but the trained Wensheng graph model has better image generation performance.

[0148] In actual application scenarios, you can choose method 1) or method 2) for training based on your focus. For example, if you are more concerned about training efficiency and training consumption, choose method 1); if you are more concerned about accuracy and do not consider training efficiency and training consumption, choose method 2).

[0149] In this embodiment, the loss value in method 1) or method 2) may be at least one of the first loss value, the second loss value, or the third loss value.

[0150] [Corrected 12.03.2025 according to Rule 91] As shown in FIG4 , the embodiment of the present application obtains a sample image, performs image feature extraction on the sample image, and obtains the first image feature of the sample image. In this way, the second image feature of the sample image can be regarded as a text feature, that is, image generation processing is performed based on the second image feature of the sample image to obtain a prediction result. Model training is performed based on the prediction result to obtain a text graph model. In the model training stage of the text graph model in the embodiment of the present application, only the first sample image is required to drive it, and there is no need to spend time and manpower to perform text annotation, which breaks the dependence on image-text pairs and can establish an automatic training process without human intervention; at the same time, the second image feature of the sample image contains rich semantic information in the first sample image, which can effectively improve the semantic understanding ability and image generation ability of the trained text graph model.

[0151] In some embodiments, aligning the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image includes:

[0152] Obtain a second sample text corresponding to the second sample image, and extract text features of the second sample text; align the first image features of the first sample image to the text feature space through an alignment network to obtain a second image feature of the first sample image; align the first image features of the second sample image to the text feature space through an alignment network to obtain a second image feature of the second sample image;

[0153] Determining a sample condition feature based on a second image feature of the sample image includes: determining a first sample condition feature based on the second image feature of the first sample image, and performing feature fusion processing on the second image feature of the second sample image and a text feature of the second sample text to obtain a second sample condition feature;

[0154] Model training is performed based on the loss value to obtain a text-graph model, including: training the text-graph model to be trained based on a first loss value, and training an alignment network based on a second loss value; wherein the trained alignment network is used to align a first image feature of a first sample image to a text feature space to obtain a second image feature of the first sample image.

[0155] Here, the first image feature of the first sample image can be aligned to the text feature space through the alignment network to obtain the second image feature of the first sample image.

[0156] To help the alignment network better align the first image features to the text feature space, the alignment network can be pre-trained. The training dataset for the alignment network includes several image-text pairs to learn the potential correlation between image and text features. To facilitate differentiation, the image in each image-text pair is named the second sample image, and the text in each image-text pair is named the second sample text. It is worth noting that the second sample image and the second sample text correspond to each other, meaning that the second sample text is used to describe the second sample image.

[0157] The embodiment of the present application does not limit the method for obtaining the training dataset of the alignment network. For example, an open source image-text pair dataset can be obtained as the training dataset of the alignment network.

[0158] The embodiment of the present application does not limit the network structure of the alignment network. For example, the alignment network may include a linear layer and a normalization layer (LN).

[0159] Similarly, here, image feature extraction is performed on the second sample image to obtain the first image feature of the second sample image, and text feature extraction is performed on the second sample text to obtain the text feature of the second sample text.

[0160] In some embodiments, an image feature extraction network is used to extract image features from a second sample image to obtain first image features of the second sample image; and a text feature extraction network is used to extract text features from a second sample text to obtain text features of the second sample text. The image feature extraction network and the text feature extraction network can be trained separately or jointly. The latter can improve the performance of image and text feature extraction processing and better learn the potential correlation between image features and text features.

[0161] In some embodiments, it also includes: obtaining third sample texts corresponding to multiple third sample images respectively; extracting the first image features of each third sample image through the image feature extraction network in the multimodal model; extracting the text features of each third sample text through the text feature extraction network in the multimodal model; determining the feature extraction loss value based on the first image features of each third sample image and the text features of each third sample text, and training the multimodal model based on the feature extraction loss value; wherein, the image feature extraction network in the trained multimodal model is used to perform image feature extraction on the second sample image and the first sample image; and the text feature extraction network in the trained multimodal model is used to perform text feature extraction on the second sample text.

[0162] Here, a training dataset for the multimodal model can be obtained, and the image feature extraction network and the text feature extraction network in the multimodal model can be jointly trained based on the training dataset of the multimodal model. The training dataset of the multimodal model includes several image-text pairs. For ease of distinction, the images in the image-text pairs are named third sample images, and the texts in the image-text pairs are named third sample texts. The training dataset of the multimodal model and the training dataset of the alignment network can be the same or different, and this is not limited.

[0163] In the model training stage of the multimodal model, the first image features of each third sample image and the text features of each third sample text are extracted through the image feature extraction network in the multimodal model. The feature extraction loss value is determined based on the first image features of each third sample image and the text features of each third sample text, and the multimodal model is trained based on the feature extraction loss value, that is, the image feature extraction network and the text feature extraction network in the multimodal model are trained (jointly trained).

[0164] In which, a loss function can be constructed based on contrastive learning, and the first image feature of each third sample image and the text feature of each third sample text are substituted into the loss function to obtain a feature extraction loss value, wherein the type of loss function is not limited, for example, it can be a cross-entropy loss function. For example, the third sample image and the corresponding third sample text (referring to the third sample image and the third sample text being in the same image-sample pair) can constitute a positive sample, and the third sample image and the non-corresponding third sample text can constitute a negative sample. Based on this, a loss function can be constructed, the purpose of which is to maximize the similarity corresponding to the positive sample and minimize the similarity corresponding to the negative sample, wherein the similarity corresponding to the positive sample refers to the similarity between the first image feature of the third sample image in the positive sample and the text feature of the third sample text in the positive sample, and the similarity corresponding to the negative sample refers to the similarity between the first image feature of the third sample image in the negative sample and the text feature of the third sample text in the negative sample.

[0165] For ease of understanding, the embodiment of the present application also provides a flow chart of training a multimodal model as shown in Figure 5. In Figure 5, the third sample image includes N, namely, third sample image 1, third sample image 2...third sample image N, the first image feature of the third sample image 1 is represented as I1, and so on; the third sample text also includes N, namely, third sample text 1, third sample text 2...third sample text N, the third sample text 1 corresponds to the third sample image 1, the text feature of the third sample text 1 is represented as T1, the similarity between the first image feature I1 of the third sample image 1 and the text feature T1 of the third sample text 1 is represented as I1.T1, and so on. Based on this, an N×N similarity matrix can be obtained, and the purpose of the loss function is to maximize the similarity on the diagonal in the similarity matrix and minimize the similarity not on the diagonal in the similarity matrix. The diagonal here refers to the diagonal from I1.T1 in the upper left corner to IN.TN in the lower right corner.

[0166] In one embodiment, the training of the multimodal model and the training of the document graph model can be performed simultaneously, or the training of the multimodal model can be completed first, the parameters of the multimodal model can be fixed, and then the document graph model can be trained.

[0167] The training of the multimodal model and the training of the text graph model can be performed simultaneously, that is, the text graph model to be trained is trained based on the first loss value, and the alignment network is trained according to the second loss value.

[0168] In one embodiment, after completing the training of the multimodal model, the parameters of the multimodal model are fixed, and then the text graph model is trained. That is, after completing the training of the multimodal model, the first image features of the second sample image are extracted by the image feature extraction network in the trained multimodal model, and the text features of the second sample text are extracted by the text feature extraction network in the trained multimodal model. The image feature extraction network in the trained multimodal model performs image feature extraction processing on the first sample image to obtain the first image features of the first sample image. The above method can improve the effects of image feature extraction processing and text feature extraction processing by training the multimodal model, thereby better training the alignment network and accurately extracting the first image features of the first sample image and the first image features of the second sample image.

[0169] The first image features of the second sample image are aligned to the text feature space through the alignment network to obtain the second image features of the second sample image. The second image features of the second sample image are fused with the text features of the second sample text to obtain the second sample conditional features.

[0170] Here, the second image features of the second sample image have been aligned to the text feature space. Therefore, the second image features of the second sample image can be fused with the text features of the second sample text to obtain a second sample conditional feature. The method of feature fusion processing is not limited; for example, a splicing process can be used. The text-based graph model to be trained is guided to generate an image based on the second sample conditional feature to obtain a second prediction result. A second loss value is determined based on the second prediction result, and the alignment network is trained based on the second loss value.

[0171] Here, the image generation process is performed on the to-be-trained text graph model based on the second sample conditional feature, and a second loss value is determined based on the second prediction result generated during the image generation process. Unlike step 105, during the model training phase of the alignment network, all networks except the alignment network are frozen (i.e., all model parameters except the model parameters of the alignment network are kept unchanged). That is, the second loss value is only used to train the alignment network.

[0172] It is worth noting that the loss function used to determine the second loss value may be the same as the loss function used to determine the first loss value.

[0173] In the embodiment of the present application, an alignment network is trained based on image-text pairs. During the training of the alignment network, all networks except the alignment network are frozen, so that during the model training phase of the alignment network, it is possible to learn how to efficiently and accurately align the first image features to the text feature space. In this way, the trained alignment network can be used to align the first image features of the first sample image to the text feature space to obtain the second image features of the first sample image. This can improve the richness of the semantic information contained in the second image features of the first sample image and reduce the loss of semantic information during the alignment process.

[0174] In some embodiments, the first sample condition feature includes at least one of a second image feature of the first sample image or an image fusion feature, where the image fusion feature is obtained by fusing the second image feature of the first sample image with an embedded image feature in the text feature space.

[0175] In this embodiment, the second image feature of the first sample image is fused with the embedded image feature in the text feature space to obtain the first sample conditional feature.

[0176] Here, embedded image features can be introduced into the text feature space to achieve binding with the ID of the first sample image. Based on this, the second image features of the first sample image can be fused with the embedded image features in the text feature space to obtain the first sample conditional features, making the first sample conditional features more comprehensive and accurate.

[0177] It is worth noting that based on the introduction of embedded image features, the embedded image features are not only used in the model training stage of the cultural graph model, but also in the model inference stage of the cultural graph model, thereby making the target image output by the cultural graph model in the model inference stage more effective.

[0178] In some embodiments, fusing the second image feature of the first sample image with the embedded image feature in the text feature space includes: initializing the embedded image feature in the text feature space, and fusing the second image feature of the first sample image with the embedded image feature to obtain a third sample conditional feature;

[0179] Determine the loss value based on the prediction results, perform model training based on the loss value, and obtain the Wensheng graph model, including:

[0180] A third loss value is determined according to the third prediction result; the embedded image feature is updated according to the third loss value, and the to-be-trained document image model is trained based on the first loss value.

[0181] In this embodiment, the to-be-trained Vincent graph model is guided to generate an image based on the first sample conditional feature to obtain a first prediction result. The to-be-trained Vincent graph model is guided to generate an image based on the third sample conditional feature to obtain a third prediction result. A first loss value is determined based on the first prediction result, and a third loss value is determined based on the third prediction result. The to-be-trained Vincent graph model is trained based on the first loss value, and the embedded image features are updated based on the third loss value.

[0182] Here, the embedded image features can be trained and updated based on the training data set of the embedded image features, so that the effect of binding the embedded image features to the ID is better. The training data set of the embedded image features can be the same as the training data set of the cultural graph model, or it can be a subset of the training data set of the cultural graph model.

[0183] During the training and updating phase of the embedded image features, the embedded image features are first initialized (e.g., randomly initialized) in the text feature space, wherein the dimensions of the embedded image features conform to the dimensions of the text feature space, and the embedded image features and the second image features may be the same or different in dimension. Next, the second image features of the first sample image are fused with the embedded image features to obtain a third sample conditional feature. The third sample conditional feature is used to guide the text-based graph model to be trained to generate an image, thereby obtaining a third prediction result. A third loss value is determined based on the third prediction result, and the embedded image features are trained and updated based on the third loss value.

[0184] It is worth noting that the training and updating stage of the embedded image features can be carried out simultaneously with the training of the text-based graph model. Alternatively, the training and updating of the embedded image features can be performed first, and then the training of the text-based graph model can be performed after completion.

[0185] For example, during the training and updating phase of the embedded image features, all model parameters except for the embedded image features remain unchanged. After completing the training and updating of the embedded image features, the second image features of the first sample image are fused with the embedded image features in the text feature space to obtain a first sample conditional feature of the first sample image. The first sample conditional feature is used to guide the text graph model to be trained to generate an image, obtaining a first prediction result. A first loss value is determined based on the first prediction result, and the text graph model to be trained is trained based on the first loss value to obtain a text graph model.

[0186] It is worth noting that the loss function used to determine the third loss value may be the same as the loss function used to determine the first loss value.

[0187] In some embodiments, there are multiple first sample images, and the embedded image feature includes the second image feature of any one of the first sample images.

[0188] Here, if the training dataset of the Wensheng graph model includes multiple first sample images, the second image feature of any of the first sample images can be determined as the embedded image feature. This method can quickly determine the embedded image feature, and the binding ID of the embedded image feature determined by training (updating) is more effective.

[0189] In this embodiment, the second image feature of the first sample image can be fused with the embedded image feature in the text feature space to obtain the first sample conditional feature. The second image feature of the first sample image can also be directly determined as the first sample conditional feature. Compared with obtaining the first sample conditional feature through feature fusion, determining the second image feature as the first sample conditional feature is more lightweight, but obtaining the first sample conditional feature through feature fusion can bind the ID according to the embedded image feature, so that the target image output by the text graph model in the model inference stage is better, and can be selected according to the needs of the actual application scenario.

[0190] The embodiment of the present application provides two methods for determining the first sample condition feature based on the second image feature of the first sample image. Either method can be selected according to the needs of the actual application scenario, which is highly flexible.

[0191] In some embodiments, the first prediction result includes prediction noise; and determining the loss value according to the prediction result includes:

[0192] The first sample image is mapped from the pixel space to the latent space to obtain the third image feature of the first sample image; noise is added to the third image feature of the first sample image to obtain a noisy image feature; noise is extracted from the noisy image feature according to the first sample conditional feature to obtain predicted noise; and the difference between the added noise and the predicted noise is used as the first loss value.

[0193] During the model training phase of the Wensheng graph model, noise can be first added to the image through a forward process (also called forward process, forward diffusion process), and then the added noise can be predicted through a reverse process (also called reverse process, reverse reconstruction process), thereby training a powerful image generation capability.

[0194] First, the first sample image is mapped from pixel space to latent space to obtain the third image feature of the first sample image. Latent space refers to a feature space smaller than pixel space. The forward and reverse processes described above can be implemented in latent space, thereby achieving image compression and improving training efficiency.

[0195] In some embodiments, a matching target encoder and target decoder may be pre-trained, and the target encoder may be used to map the first sample image from the pixel space to the latent space to obtain the third image feature of the first sample image. That is, the target encoder may be used to perform image feature extraction processing on the first sample image to map the first sample image from the pixel space to the latent space to obtain the third image feature of the first sample image. For example, the target encoder may refer to the encoder in an autoencoder (AE), and the target decoder may refer to the decoder in the autoencoder.

[0196] In the forward process, adding noise refers to applying noise to the third image feature of the first sample image to obtain a noisy image feature. For example, the third image feature can be fused with the image feature of the noisy image to obtain the noisy image feature. The noise added by the noise addition process is known. The noisy image feature can be used to restore the noisy image.

[0197] Noise extraction involves predicting and extracting the characteristic noise of a noisy image. The extracted noise is called the predicted noise. Denoising produces denoised image features. Using these features, the denoised image can be restored to its original state.

[0198] In the reverse process, the noisy image features are denoised according to the first sample conditional features through the Wensheng graph model, that is, the noise in the noisy image features is predicted and removed to obtain the denoised image features.

[0199] In some embodiments, the above-mentioned noise addition processing of the third image feature of the first sample image to obtain the noisy image feature can be achieved in the following manner: T rounds of noise addition iterations are performed, and the following processing is performed during the tth round of noise addition iteration: noise is added to the image feature input in the tth round of noise addition iteration to obtain the image feature input in the t+1th round of noise addition iteration; wherein the third image feature of the first sample image is used as the image feature input in the 1st round of noise addition iteration; and the image feature obtained in the Tth round of noise addition iteration is the noisy image feature.

[0200] The above-mentioned denoising process of the noisy image features according to the first sample conditional features by the Wensheng graph model to obtain the denoised image features can be achieved in the following manner: T rounds of denoising iterations are performed, and the following process is performed during the tth round of denoising iteration: the noise in the image features input to the tth round of denoising iteration is predicted according to the first sample conditional features by the Wensheng graph model, and the predicted noise is removed from the image features input to the tth round of denoising iteration to obtain the image features input to the t+1th round of denoising iteration; wherein the noisy image features are used as the image features input to the 1st round of denoising iteration; and the image features obtained by the Tth round of denoising iteration are the denoised image features; wherein T is an integer greater than 1, and t is an integer greater than 0 and not exceeding T.

[0201] Here, the noise addition process may include T rounds of noise addition iterations to gradually add noise. For ease of explanation, the process of the tth round of noise addition iteration is taken as an example, where T is an integer greater than 1 and can be set according to the actual application scenario, such as being set to 30; and t is an integer greater than 0 but not exceeding T. In the tth round of noise addition iteration, noise is added to the image features input in the tth round of noise addition iteration to obtain the image features input in the t+1th round of noise addition iteration. The third image feature of the first sample image is used as the image feature input in the first round of noise addition iteration; the image feature obtained in the Tth round of noise addition iteration (i.e., the image feature input in the T+1th round of noise addition iteration) is the noisy image feature.

[0202] Correspondingly, the denoising process may include T rounds of denoising iterations to gradually predict and remove noise. For ease of explanation, the process of the tth round of denoising iteration is taken as an example. In the tth round of denoising iteration, the noise in the image features input to the tth round of denoising iteration is predicted by the Vincent graph model based on the first sample conditional features, and the predicted noise is removed from the image features input to the tth round of denoising iteration to obtain the image features input to the t+1th round of denoising iteration. Among them, the noisy image features are used as the image features input to the first round of denoising iteration; the image features obtained from the Tth round of denoising iteration (i.e., the image features input to the T+1th round of denoising iteration) are the denoised image features. The forward process and the reverse process are realized through the above-mentioned Markov architecture, which makes the model training phase of the Vincent graph model more stable and interpretable.

[0203] Here, the noise added by the noise addition process is regarded as the expected result, the predicted noise extracted by the noise is regarded as the predicted result, and the difference between the expected result and the predicted result is calculated as the first loss value. The embodiment of the present application does not limit the type of loss function used to calculate the first loss value.

[0204] It is worth noting that, since the noise predicted by the denoising process is obtained in the model training stage of the text graph model, it is not necessary to restore the denoised image based on the denoised image features.

[0205] In some embodiments, the Vincent graph model includes a feature cross network; the Vincent graph model can be used to perform noise extraction on the noisy image features based on the first sample condition features to obtain denoised image features: the first sample condition features and the noisy image features are subjected to feature cross processing through the feature cross network to predict and remove noise in the noisy image features to obtain denoised image features; the above-mentioned training of the Vincent graph model based on the first loss value can be achieved in the following manner: the feature cross network is trained based on the first loss value.

[0206] Here, the Wensheng graph model includes a feature cross network. In the reverse process, the first sample conditional features and the noisy image features can be subjected to feature cross processing through the feature cross network to predict and remove the noise in the noisy image features to obtain denoised image features. Based on this, the model training phase of the Wensheng graph model is to train the feature cross network. The embodiment of the present application does not limit the network structure of the feature cross network. For example, it can include a multi-head attention layer, and of course it can also include other network layers.

[0207] If the denoising process includes T rounds of denoising iterations, then during the tth round of denoising iteration, the first sample conditional features are subjected to feature cross-processing with the image features inputted at the tth round of denoising iteration through a feature cross-network to predict and remove noise from the image features inputted at the tth round of denoising iteration, thereby obtaining the image features inputted at the t+1th round of denoising iteration.

[0208] In an embodiment of the present application, noise is added to an image in a forward process, the added noise is predicted by a Vincent graph model in a reverse process, a first loss value is calculated based on the added noise and the predicted noise, and the Vincent graph model is trained based on the first loss value, so that the noise predicted by the trained Vincent graph model is more accurate.

[0209] The artificial intelligence-based image generation method provided in the embodiments of the present application will be explained in combination with the exemplary application and implementation of the electronic device provided in the embodiments of the present application.

[0210] See Figure 6, which is a flow chart of the artificial intelligence-based image generation method provided in an embodiment of the present application, which will be explained in conjunction with the steps shown in Figure 6.

[0211] In step 501, prompt information is obtained.

[0212] After the text graph model is trained, we can enter the model inference stage.

[0213] First, prompt information is obtained. The prompt information is used to indicate the requirement for image generation. The prompt information may include at least one of an image (named as prompt image for easy distinction) and text (named as prompt text for easy distinction).

[0214] In step 502, prompt features are extracted from the prompt information to obtain prompt features of the prompt information aligned with the text feature space.

[0215] Here, according to the modality of the prompt information, prompt features of the prompt information are extracted to obtain prompt features of the prompt information. The prompt features of the prompt information have been aligned to the text feature space.

[0216] In some embodiments, the above-mentioned prompt feature extraction processing of the prompt information can be implemented in the following manner to obtain the prompt feature of the prompt information aligned to the text feature space: when the prompt information is a prompt image, the prompt image is subjected to image feature extraction processing to obtain the first image feature of the prompt image, the first image feature of the prompt image is aligned to the text feature space to obtain the second image feature of the prompt image, and the second image feature of the prompt image is determined as the prompt feature; when the prompt information is a prompt text, the prompt text is subjected to text feature extraction processing to obtain the text feature of the prompt text, and the text feature of the prompt text is determined as the prompt feature; when the prompt information includes a prompt image and a prompt text, the second image feature of the prompt image is subjected to feature fusion processing with the text feature of the prompt text to obtain the prompt feature.

[0217] Here, when the prompt information only includes a prompt image, image feature extraction processing is performed on the prompt image to obtain a first image feature of the prompt image, the first image feature of the prompt image is aligned to the text feature space to obtain a second image feature of the prompt image, and the second image feature of the prompt image is determined as the prompt feature. Specifically, image feature extraction processing can be performed on the prompt image using an image feature extraction network in a trained multimodal model to obtain the first image feature of the prompt image; and the first image feature of the prompt image can be aligned to the text feature space using a trained alignment network to obtain the second image feature of the prompt image.

[0218] When the prompt information only includes the prompt text, a text feature extraction process is performed on the prompt text to obtain text features of the prompt text, and the text features of the prompt text are determined as the prompt features. The text features of the prompt text can be obtained by performing text feature extraction on the prompt text using a text feature extraction network in a trained multimodal model.

[0219] When the prompt information includes both a prompt image and prompt text, image feature extraction is performed on the prompt image to obtain a first image feature of the prompt image. The first image feature of the prompt image is then aligned to the text feature space to obtain a second image feature of the prompt image. Simultaneously, text feature extraction is performed on the prompt text to obtain a text feature of the prompt text. The second image feature of the prompt image is then fused with the text feature of the prompt text to obtain the prompt feature.

[0220] By adopting the above method, the prompt feature is aligned with the text feature space while being able to reflect the information of all modes in the prompt information.

[0221] In step 503, target condition features are determined according to the prompt features of the prompt information.

[0222] Here, the prompt feature of the prompt information is located in the text feature space, so the target condition feature can be determined according to the prompt feature of the prompt information.

[0223] In some embodiments, the target condition feature includes at least one of a prompt feature of the prompt information or a target fusion feature, where the target fusion feature is obtained by fusing the prompt feature of the prompt information with an embedded image feature in a text feature space.

[0224] For example, the prompt feature of the prompt information is fused with the embedded image feature in the text feature space to obtain the target condition feature, or the prompt feature of the prompt information is determined as the target condition feature.

[0225] If, in the model training phase of the text-graph model, the second image features of the first sample image are fused with the embedded image features in the text feature space to obtain the first sample conditional features, then in the model inference phase of the text-graph model, the prompt features of the prompt information are similarly fused with the embedded image features in the text feature space to obtain the target conditional features.

[0226] If the second image feature of the first sample image is determined as the first sample conditional feature in the model training phase of the culture graph model, then the prompt feature of the prompt information is also determined as the target conditional feature in the model inference phase of the culture graph model.

[0227] In step 504, the text graph model is guided to generate an image according to the target condition characteristics to obtain a target image; wherein the text graph model is trained according to an artificial intelligence-based text graph model training method.

[0228] Here, the trained text graph model is guided to generate an image according to the target condition features. The generated image is called a target image, and the target image can be used as a response to the prompt information.

[0229] In some embodiments, the above-mentioned image generation processing guided by the Wensheng graph model according to the target condition characteristics can be achieved in the following manner to obtain the target image: image features of a random noise image are generated in the latent space; image features of the random noise image are denoised according to the target condition characteristics by the Wensheng graph model to obtain image features of the target image; and image features of the target image are mapped from the latent space to the pixel space to obtain the target image.

[0230] The model inference phase of the Vincent graph model only involves the reverse process. First, image features of a random noise image (e.g., a random Gaussian noise image) are generated in the latent space. The image features of the random noise image are then denoised using the trained Vincent graph model based on the target conditional features to obtain the image features of the target image. The image features of the target image are then mapped from the latent space to the pixel space to obtain the target image. The image features of the target image can be mapped from the latent space to the pixel space using a target decoder, such as the decoder in an autoencoder.

[0231] In some embodiments, the above-mentioned denoising processing of the image features of the random noise image according to the target condition features by the Wensheng graph model to obtain the image features of the target image can be achieved in the following manner: T rounds of denoising iterations are performed, and the following processing is performed during the tth round of denoising iteration: the noise in the image features input to the tth round of denoising iteration is predicted according to the target condition features by the Wensheng graph model, and the predicted noise is removed from the image features input to the tth round of denoising iteration to obtain the image features input to the t+1th round of denoising iteration; wherein the image features of the random noise image are used as the image features input to the 1st round of denoising iteration; the image features obtained from the Tth round of denoising iteration are the image features of the target image; wherein T is an integer greater than 1, and t is an integer greater than 0 and not exceeding T.

[0232] The denoising process may include T rounds of denoising iterations. During the tth round of denoising iteration, the trained text graph model is used to predict the noise in the image features input to the tth round of denoising iteration based on the target condition features, and the predicted noise is removed from the image features input to the tth round of denoising iteration to obtain the image features input to the t+1th round of denoising iteration.

[0233] In some embodiments, the Wensheng graph model includes a feature cross network; the above-mentioned denoising processing of the image features of the random noise image according to the target condition features by the Wensheng graph model to obtain the image features of the target image can be achieved in the following manner: the target condition features and the image features of the random noise image are subjected to feature cross processing by the feature cross network to predict and remove the noise in the image features of the random noise image to obtain the image features of the target image.

[0234] When the Wensheng graph model includes a feature cross network, the model training phase of the Wensheng graph model is to train the feature cross network. During the model inference phase of the Wensheng graph model, the trained feature cross network performs feature cross processing on the target conditional features and the image features of the random noise image to predict and remove noise from the image features of the random noise image, thereby obtaining the image features of the target image.

[0235] If the denoising process includes T rounds of denoising iterations, then during the tth round of denoising iteration, the trained feature cross network is used to perform feature cross processing on the target conditional features and the image features inputted at the tth round of denoising iteration, so as to predict and remove the noise in the image features inputted at the tth round of denoising iteration, and obtain the image features inputted at the t+1th round of denoising iteration.

[0236] As shown in FIG6 , after the training of the text-based graph model is completed, the embodiment of the present application only requires the user to input prompt information to realize automatic image generation processing through the trained text-based graph model; at the same time, it can support the input of image modality and text modality, has strong flexibility, and can meet various image generation needs.

[0237] The following describes an exemplary application of an embodiment of the present application in a practical application scenario. The embodiment of the present application can implement single-modal training of the LoRA model driven solely by images, breaking the reliance on image-text pairs and enabling the LoRA model to be trained using only image data. Furthermore, images themselves often contain more details than text. The embodiment of the present application uses the rich semantic information contained in images to guide LoRA model training, which can greatly preserve the original capabilities of the text-based graph model and reduce the adverse effects that may be caused to the original model parameters of the text-based graph model when training based on image-text pairs.

[0238] Taking the method of training the LoRA model as an example, an embodiment of the present application provides a comparative schematic diagram of style binding as shown in Figure 7A. Figure 7A shows a target image 712 generated by the Wensheng graph model (i.e., base model + trained LoRA model) trained according to the solution provided by the relevant technology, and also shows a target image 713 generated by the Wensheng graph model (i.e., base model + trained LoRA model) trained according to the solution provided by the embodiment of the present application, wherein the prompt information used to generate the target image 712 is the same as the prompt information used to generate the target image 713. In addition, a sample image 711 in the training data set of the LoRA model (or the training data set of the Wensheng graph model) is exemplarily shown. All sample images in the training data set have the same style, so the process of training the LoRA model is also a process of style binding. It is worth noting that the process of training the LoRA model according to the solution provided by the embodiment of the present application and the process of training the LoRA model according to the solution provided by the relevant technology use the same sample image, but in the solution provided by the relevant technology, the sample text corresponding to the sample image is also used.

[0239] The embodiment of the present application also provides a comparative schematic diagram of object binding as shown in FIG7B , in which FIG7B shows a target image 722 generated by the Wensheng graph model (i.e., base model + trained LoRA model) trained according to the solution provided by the relevant technology, and also shows a target image 723 generated by the Wensheng graph model (i.e., base model + trained LoRA model) trained according to the solution provided by the embodiment of the present application, wherein the prompt information used to generate the target image 722 is the same as the prompt information used to generate the target image 723. In addition, a sample image 721 in the training data set of the LoRA model is exemplarily shown, and all sample images in the training data set contain the same object (i.e., the game character in the sample image 721), so the process of training the LoRA model is also a process of object binding.

[0240] According to Figures 7A and 7B, compared to the solutions provided by related technologies, the embodiments of the present application can shorten the LoRA model training link while ensuring good image generation effects, avoid manual intervention in the training process, and establish a standard LoRA automated training process for different IDs. As shown in Figure 8, the embodiments of the present application avoid the process of implementing text annotation through automatic labeling, manual modification, and other operations in the solutions provided by related technologies.

[0241] Next, the solution provided by the embodiments of the present application will be described in detail.

[0242] The embodiments of the present application use a pre-trained text-based graph model and a multimodal model. There is no limitation on the type of the text-based graph model. Here, the SD model (Unet network in the SD model) is taken as an example; there is no limitation on the type of the multimodal model. The multimodal model needs to include an image feature extraction network and a text feature extraction network. Here, the CLIP model is taken as an example.

[0243] Among them, the SD model is a conditional diffusion model based on the latent space. The SD model uses the encoder in the Auto Encoder to encode the image into a latent variable (Latent), and then uses the diffusion model to introduce conditions (usually referring to text features in the text feature space), converting the latent variable encoded by the encoder into another latent variable, and finally restoring the converted latent variable to an image through the decoder of the Auto Encoder. The conditional features of the SD model are text features, where the text features are extracted through the text feature extraction network of CLIP. For example, the SD model uses the text feature extraction network in the CLIP ViT-L model to encode text into 77*768-dimensional text features, and introduces text features into Latent through the cross attention mechanism (corresponding to the feature cross network mentioned above).

[0244] The CLIP model is a multimodal model pre-trained through contrastive learning on image-text pairs. It mainly consists of an image feature extraction network and a text feature extraction network. The image feature extraction network is also called an image encoder, and the text feature extraction network is also called a text encoder.

[0245] Based on the above premise, the training process of the embodiment of the present application can be divided into three stages, which will be explained below respectively.

[0246] 1) Train the alignment network.

[0247] In an embodiment of the present application, the SD model is used as the base model. Since the SD model generates images based on the text features in the text feature space, the features in the text feature space are also required as conditions for guidance when training the LoRA model. Therefore, in order to achieve image driving, an alignment network is designed to align the image features to the text feature space.

[0248] The alignment network can be a lightweight network including a Linear layer and a LayerNorm layer. The alignment network can be trained based on large-scale open source image-text pairs. The training process is shown in Figure 9. For example, the Text Encoder in the CLIP ViT-L model is used to extract the text features of the second sample text. The dimension of the text features is 77*768 (which can be regarded as 77 tokens). At the same time, the Image Encoder in the CLIP ViT-L model is used to extract the first image features of the second sample image. The dimension of the first image features of the second sample image is 1*768. Subsequently, the alignment network is used to convert the first image features of the second sample image into the text feature space, and the dimension is expanded to 3*768 so that the model can perceive finer-grained features. In this way, the second image features of the second sample image can be obtained. Finally, the second image features of the second sample image and the text features of the second sample text are spliced ​​into 80*768 second sample conditional features to guide the SD model for image generation.

[0249] During the alignment network training phase, all networks except the alignment network are frozen, meaning the model parameters of all networks except the alignment network remain unchanged. Adam can be used as the training optimizer, with a large learning rate (e.g., 1e-4) to enable the randomly initialized alignment network to quickly learn how to accurately align the first image features to the text feature space. The alignment network can be trained on large-scale open-source image-text pair data, enabling the trained alignment network to robustly transform the first image features of any image.

[0250] 2) Training embedded image features (Image Inversion).

[0251] After completing the training of the alignment network, Image Inversion can be introduced into the text feature space to store the commonality of the data, thereby achieving ID binding. The dimension of Image Inversion can be 3*768, of course, there is no limitation to this. Here, all or part of the training dataset of the LoRA model can be used as the training dataset of Image Inversion, wherein the training dataset of the LoRA model includes multiple first sample images, and the multiple first sample images have the same ID.

[0252] First, Image Inversion can be randomly initialized, as shown in Figure 10. During the training phase of Image Inversion, the first image feature of the first sample image is aligned to the text feature space through the trained alignment network to obtain the second image feature of the first sample image. The second image feature of the first sample image and Image Inversion are then concatenated into a 6*768 third sample conditional feature to guide the SD model for image generation.

[0253] Similarly, during the training phase of Image Inversion, all model parameters except Image Inversion are kept unchanged.

[0254] 3) Train the LoRA model.

[0255] After completing the training of the alignment network and Image Inversion, we can enter the model training phase of the LoRA model. As shown in Figure 11, the trained alignment network aligns the first image features of the first sample image to the text feature space, obtaining the second image features of the first sample image. The second image features of the first sample image are then concatenated with the trained Image Inversion to form a 6*768 first sample conditional feature, which guides the text-to-graph model (SD model + LoRA model) for image generation.

[0256] Likewise, during the training phase of the LoRA model, all networks except the LoRA model are frozen.

[0257] After completing the above three training stages, the model inference stage can be entered based on the trained alignment network, the trained Image Inversion, and the trained text graph model (SD model + trained LoRA model). The model inference stage can include the following steps.

[0258] 1) Get the prompt information entered by the user.

[0259] 2) Perform prompt feature extraction processing on the prompt information to obtain prompt features of the prompt information aligned to the text feature space, and generate target condition features based on the prompt features of the prompt information.

[0260] Wherein, depending on the mode of the prompt information, step 2) may include the following situations:

[0261] ① As shown in Figure 12, when the prompt information is a prompt image, the prompt image is subjected to image feature extraction processing through Image Encoder to obtain the first image feature of the prompt image (dimension is 1*768). Then, the first image feature of the prompt image is aligned to the text feature space through the trained alignment network to obtain the second image feature of the prompt image (dimension is 3*768). The second image feature of the prompt image is spliced ​​with the trained Image Inversion to form a 6*768 target condition feature.

[0262] ② As shown in Figure 13, when the prompt information is a prompt text, the prompt text is extracted through the Text Encoder to obtain the text features of the prompt text (dimension is 77*768). Then, the text features of the prompt text are spliced ​​with the trained Image Inversion to form the target condition features of 80*768.

[0263] ③ As shown in Figure 14, when the prompt information includes both a prompt image and prompt text, the prompt image is extracted using the Image Encoder to obtain the first image feature of the prompt image (dimension 1*768). The trained alignment network then aligns the first image feature of the prompt image to the text feature space to obtain the second image feature of the prompt image (dimension 3*768). Simultaneously, the Text Encoder extracts text features from the prompt text to obtain the text feature of the prompt text (dimension 77*768). The second image feature of the prompt image, the text feature of the prompt text, and the trained Image Inversion are then concatenated into the 83*768 target conditional feature.

[0264] 3) The trained text-based graph model is guided to perform image generation processing according to the target condition characteristics to obtain the target image.

[0265] The following continues to describe an exemplary structure of the artificial intelligence-based text graph model training device 255 provided in an embodiment of the present application implemented as a software module. In some embodiments, as shown in Figure 2, the software modules in the artificial intelligence-based text graph model training device 255 stored in the memory 250 may include: a first acquisition module 2551, used to acquire a sample image; a first feature extraction module 2552, used to extract image features of the sample image to obtain a first image feature of the sample image; a first alignment module 2553, used to align the first image feature of the sample image to the text feature space to obtain a second image feature of the sample image; a first prediction module 2554, used to generate an image based on the second image feature of the sample image to obtain a prediction result; and a training module 2555, used to perform model training according to the prediction result to obtain a text graph model.

[0266] In some embodiments, the Wensheng graph model training device 255 further includes a first generation module, the first generation module being configured to determine the sample condition feature based on the second image feature of the sample image;

[0267] The first prediction module 2554 is used to generate an image based on the sample conditional features to obtain a prediction result.

[0268] In some embodiments, the training module 2555 is further configured to determine a loss value according to the prediction result, perform model training based on the loss value, and obtain a text graph model.

[0269] In some embodiments, the sample image includes at least one of a first sample image, a second sample image, or a third sample image, the first sample image is used to train a text image model, the second sample image is used to train an alignment network, the alignment network is used to align a first image feature of the first sample image to a text feature space, and the third sample image is used to train a multimodal model, the multimodal model is used to extract the first image feature of the first sample image;

[0270] The sample condition feature includes at least one of the first sample condition feature, the second sample condition feature or the third sample condition feature, the prediction result includes at least one of the first prediction result, the second prediction result or the third prediction result, and the loss value includes at least one of the first loss value, the second loss value or the third loss value.

[0271] In some embodiments, the training module 2555 includes a text-image model training module and an alignment network training module; the first alignment module 2553 is further configured to: obtain a second sample text corresponding to the second sample image, extract text features of the second sample text; align the first image features of the first sample image to the text feature space through the alignment network to obtain the second image features of the first sample image; align the first image features of the second sample image to the text feature space through the alignment network to obtain the second image features of the second sample image;

[0272] The first generating module is further configured to determine a first sample conditional feature based on the second image feature of the first sample image, and to perform feature fusion processing on the second image feature of the second sample image and the text feature of the second sample text to obtain a second sample conditional feature;

[0273] a Wensheng graph model training module, configured to train the Wensheng graph model to be trained based on the first loss value;

[0274] An alignment network training module is used to train the alignment network according to the second loss value; wherein the trained alignment network is used to align the first image feature of the first sample image to the text feature space to obtain the second image feature of the first sample image.

[0275] In some embodiments, the artificial intelligence-based text-image model training device 255 also includes a multimodal model training module, which is used to: obtain multiple third sample images and third sample texts corresponding to the multiple third sample images; extract the first image features of each third sample image through the image feature extraction network in the multimodal model; extract the text features of each third sample text through the text feature extraction network in the multimodal model; determine the feature extraction loss value based on the first image features of each third sample image and the text features of each third sample text, and train the multimodal model based on the feature extraction loss value; wherein, the image feature extraction network in the trained multimodal model is used to extract image features of the second sample image and the first sample image; and the text feature extraction network in the trained multimodal model is used to extract text features of the second sample text.

[0276] In some embodiments, the first sample conditional feature of the first prediction module 2554 includes at least one of the second image feature or image fusion feature of the first sample image, and the image fusion feature is obtained by fusing the second image feature of the first sample image with the embedded image feature in the text feature space.

[0277] In some embodiments, the first prediction module 2554 is also used to perform any of the following processing: performing feature fusion processing on the second image feature of the first sample image and the embedded image feature in the text feature space to obtain the first sample conditional feature; determining the second image feature of the first sample image as the first sample conditional feature.

[0278] In some embodiments, the artificial intelligence-based text graph model training device 255 also includes an embedded image feature training module, which is used to: initialize the embedded image feature in the text feature space; perform feature fusion processing on the second image feature of the first sample image and the embedded image feature to obtain a third sample condition feature; guide the text graph model to perform image generation processing based on the third sample condition feature, and determine a third loss value based on the third prediction result generated in the image generation processing process, and update the embedded image feature based on the third loss value.

[0279] In some embodiments, there are multiple first sample images; the artificial intelligence-based text graph model training device 255 also includes an embedded image feature determination module, which is used to: determine the second image feature of any first sample image as the embedded image feature.

[0280] In some embodiments, the training module 2555 is further used to learn incremental parameters based on the loss value, and update the model parameters of the to-be-trained document graph model based on the incremental parameters to obtain the document graph model; wherein the incremental parameters and the model parameters are the same in dimension.

[0281] In some embodiments, the training module 2555 is also used to perform any of the following processing: learning the first sub-incremental parameter and the second sub-incremental parameter based on the first loss value, performing parameter fusion processing on the first sub-incremental parameter and the second sub-incremental parameter to obtain the incremental parameter, and updating the model parameters of the Wensheng graph model based on the incremental parameter; learning the incremental parameter based on the first loss value, and updating the model parameters of the Wensheng graph model based on the incremental parameter; wherein the incremental parameter and the model parameter are the same in dimension.

[0282] In some embodiments, the first prediction result includes predicted noise; the training module 2555 is also used to: map the first sample image from the pixel space to the latent space to obtain the third image feature of the first sample image; add noise to the third image feature of the first sample image to obtain a noisy image feature; extract noise from the noisy image feature according to the first sample conditional feature to obtain predicted noise; and use the difference between the added noise and the predicted noise as the first loss value.

[0283] The following continues to describe an exemplary structure of the artificial intelligence-based image generation device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, as shown in Figure 3, the software modules in the artificial intelligence-based image generation device 455 stored in the memory 450 may include: a second acquisition module 4551, used to obtain prompt information; a second feature extraction module 4552, used to extract prompt features of the prompt information to obtain prompt features of the prompt information aligned to the text feature space; a second generation module 4553, used to determine target condition features based on the prompt features of the prompt information; a second prediction module 4554, used to guide the text graph model to generate an image based on the target condition features to obtain a target image; wherein the text graph model is trained according to the artificial intelligence-based text graph model training method.

[0284] In some embodiments, the second feature extraction module 4552 is also used to: when the prompt information is a prompt image, extract the first image feature of the prompt image, and align the first image feature of the prompt image to the second image feature obtained in the text feature space as the prompt feature; when the prompt information is a prompt text, extract the text feature of the prompt text as the prompt feature; when the prompt information includes a prompt image and prompt text, perform feature fusion processing on the second image feature of the prompt image and the text feature of the prompt text to obtain the prompt feature.

[0285] In some embodiments, the target condition feature includes at least one of a prompt feature of the prompt information or a target fusion feature, where the target fusion feature is obtained by fusing the prompt feature of the prompt information with an embedded image feature in a text feature space.

[0286] In some embodiments, the second generation module 4553 is also used to perform any of the following processing: performing feature fusion processing on the prompt feature of the prompt information and the embedded image feature in the text feature space to obtain the target condition feature; determining the prompt feature of the prompt information as the target condition feature.

[0287] In some embodiments, the second prediction module 4554 is also used to: generate image features of a random noise image in a latent space; denoise the image features of the random noise image according to the target condition features through a text graph model to obtain image features of a target image; and map the image features of the target image from the latent space to the pixel space to obtain a target image.

[0288] The present invention provides a computer program product or computer program, which includes executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, causing the electronic device to perform the artificial intelligence-based text graph model training method and artificial intelligence-based image generation method described in the present invention.

[0289] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the artificial intelligence-based text graph model training method and the artificial intelligence-based image generation method provided in the embodiment of the present application.

[0290] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0291] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0292] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0293] By way of example, executable instructions may be deployed to be executed on one computing device or on multiple computing devices at one site or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0294] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0295] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for training a cultural graph model based on artificial intelligence, characterized in that: Executed by a computer device, the method includes: Get a sample image; performing image feature extraction on the sample image to obtain a first image feature of the sample image; Aligning the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image; Performing image generation based on the second image feature of the sample image to obtain a prediction result; and Model training is performed based on the prediction results to obtain a cultural graph model.

2. The method according to claim 1, characterized in that The generating an image based on the second image feature of the sample image to obtain a prediction result includes: determining a sample condition feature according to a second image feature of the sample image; and An image is generated based on the sample conditional features to obtain a prediction result.

3. The method according to claim 2, characterized in that The performing model training according to the prediction results to obtain a cultural graph model includes: A loss value is determined according to the prediction result, and a model is trained based on the loss value to obtain a Wensheng graph model.

4. The method according to claim 3, characterized in that The sample images include at least one of a first sample image, a second sample image, or a third sample image, the first sample image is used to train a text image model, the second sample image is used to train an alignment network, the alignment network is used to align a first image feature of the first sample image to a text feature space, and the third sample image is used to train a multimodal model, the multimodal model is used to extract the first image feature of the first sample image; The sample condition feature includes at least one of a first sample condition feature, a second sample condition feature, or a third sample condition feature, the prediction result includes at least one of a first prediction result, a second prediction result, or a third prediction result, and the loss value includes at least one of a first loss value, a second loss value, or a third loss value.

5. The method according to claim 4, characterized in that The aligning the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image includes: Obtaining a second sample text corresponding to the second sample image, and extracting text features of the second sample text; Aligning the first image feature of the first sample image to the text feature space through an alignment network to obtain a second image feature of the first sample image; Aligning the first image feature of the second sample image to the text feature space through an alignment network to obtain a second image feature of the second sample image; The determining of the sample condition feature according to the second image feature of the sample image includes: determining a first sample conditional feature based on the second image feature of the first sample image, and performing feature fusion processing on the second image feature of the second sample image and the text feature of the second sample text to obtain a second sample conditional feature; and The performing model training based on the loss value to obtain a cultural graph model includes: Training the to-be-trained document graph model based on the first loss value, and training the alignment network based on the second loss value; The trained alignment network is used to align the first image feature of the first sample image to the text feature space to obtain the second image feature of the first sample image.

6. The method according to claim 5, characterized in that The method further comprises: Obtaining third sample texts corresponding to the plurality of third sample images respectively; extracting the first image feature of each third sample image through an image feature extraction network in the multimodal model; Extracting text features of each third sample text through a text feature extraction network in the multimodal model; and determining a feature extraction loss value based on the first image feature of each third sample image and the text feature of each third sample text, and training the multimodal model based on the feature extraction loss value; Among them, the image feature extraction network in the trained multimodal model is used to extract image features of the second sample image and the first sample image; the text feature extraction network in the trained multimodal model is used to extract text features of the second sample text.

7. The method according to claim 4, characterized in that The first sample condition feature includes at least one of a second image feature or an image fusion feature of the first sample image, and the image fusion feature is obtained by fusing the second image feature of the first sample image with an embedded image feature in a text feature space.

8. The method according to claim 7, characterized in that The fusing the second image feature of the first sample image with the embedded image feature in the text feature space includes: Initializing the embedded image feature in the text feature space, and performing feature fusion on the second image feature of the first sample image and the embedded image feature to obtain a third sample conditional feature; The step of determining a loss value according to the prediction result, and performing model training based on the loss value to obtain a Wensheng graph model includes: determining a third loss value according to the third prediction result; and The embedded image features are updated according to the third loss value, and the to-be-trained document graph model is trained based on the first loss value.

9. The method according to claim 7, characterized in that There are multiple first sample images, and the embedded image features include the second image features of any one of the first sample images.

10. The method according to claim 3, characterized in that The performing model training based on the loss value to obtain a cultural graph model includes: learning incremental parameters according to the loss value, and updating model parameters of the to-be-trained Wensheng graph model according to the incremental parameters to obtain the Wensheng graph model; The incremental parameters and the model parameters have the same dimensions.

11. The method according to claim 4, characterized in that The first prediction result includes prediction noise; and determining the loss value according to the prediction result includes: Mapping the first sample image from the pixel space to the latent space to obtain a third image feature of the first sample image; adding noise to the third image feature of the first sample image to obtain a noisy image feature; Performing noise extraction on the noisy image feature according to the first sample condition feature to obtain predicted noise; and The difference between the added noise and the predicted noise is taken as a first loss value.

12. An image generation method based on artificial intelligence, characterized in that: Executed by a computer device, the method includes: Get prompt information; Extracting prompt features from the prompt information to obtain prompt features of the prompt information aligned with a text feature space; Determining target condition features according to the prompt features of the prompt information; and The target image is generated by guiding the Wensheng graph model according to the target condition characteristics; wherein the Wensheng graph model is trained according to the method according to any one of claims 1 to 11.

13. The method according to claim 12, characterized in that The extracting prompt features of the prompt information to obtain prompt features of the prompt information aligned to a text feature space includes: When the prompt information is a prompt image, extracting a first image feature of the prompt image, and aligning the first image feature of the prompt image with a second image feature obtained in a text feature space as a prompt feature; When the prompt information is a prompt text, extracting text features of the prompt text and determining the text features of the prompt text as prompt features; and When the prompt information includes the prompt image and the prompt text, the second image feature of the prompt image is fused with the text feature of the prompt text to obtain a prompt feature.

14. The method according to claim 12, characterized in that The target condition feature includes at least one of a prompt feature of the prompt information or a target fusion feature, and the target fusion feature is obtained by fusing the prompt feature of the prompt information with an embedded image feature in a text feature space.

15. The method according to any one of claims 12 to 14, characterized in that The step of guiding the text graph model to generate an image according to the target condition feature to obtain a target image includes: Generate image features of random noise images in latent space; Denoising the image features of the random noise image according to the target condition features using a cultural graph model to obtain image features of the target image; and The image features of the target image are mapped from the latent space to the pixel space to obtain the target image.

16. A device for training a cultural graph model based on artificial intelligence, characterized in that: include: A first acquisition module is used to acquire a sample image; a first feature extraction module, configured to extract image features from the sample image to obtain a first image feature of the sample image; A first alignment module is used to align the first image feature of the sample image to the text feature space to obtain the second image feature of the sample image; A first prediction module, configured to generate an image based on a second image feature of the sample image to obtain a prediction result; and The training module is used to perform model training according to the prediction results to obtain a cultural graph model.

17. An image generation device based on artificial intelligence, characterized in that: include: The second acquisition module is used to obtain prompt information; A second feature extraction module is used to extract prompt features from the prompt information to obtain prompt features of the prompt information aligned to a text feature space; A second generating module is configured to determine target condition features according to the prompt features of the prompt information; and The second prediction module is used to guide the Wensheng graph model to generate an image according to the target condition characteristics to obtain a target image; wherein the Wensheng graph model is trained according to the method according to any one of claims 1 to 11.

18. An electronic device, characterized in that: include: a memory for storing executable instructions; A processor, configured to implement the method according to any one of claims 1 to 15 when executing the executable instructions stored in the memory.

19. A computer-readable storage medium, characterized in that Executable instructions are stored, and when executed by a processor, they are used to implement the method described in any one of claims 1 to 15.

20. A computer program product, characterized in that The method comprises executable instructions for implementing the method according to any one of claims 1 to 15 when executed by a processor.

Citation Information

Patent Citations

  • Image generation method and device based on contrast learning, equipment and medium

    CN116050481A

  • Text-based image diffusion model training method and text-based image generation method

    CN116051668A

  • Image generation method and device, electronic equipment and storage medium

    CN117173497A

  • Image generation method and device

    CN117392260A

  • Training visual language grounding models using separation loss

    US20230061647A1

Cited By

  • Vector model fine tuning method and device, equipment, storage medium and program product

    CN121502271A

  • Customer reply generation method based on three-route mixed expert big language model

    CN122221922A