Image generation method and device, equipment and medium
The graphic and text coding model and reverse optimization technology generate text prompt words similar to the target image features, and the diffusion model generates semantic style images, which solves the problem of low semantic alignment between text and pictures in the prior art, and improves the accuracy and efficiency of image generation.
Patent Information
- Application Number
- CN202510278859.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The low semantic alignment between existing text and pictures results in deviations in accuracy of generated images, especially in the financial and medical fields.
The target image is encoded using a preset graphic encoding model, an image feature vector is generated, and text prompt words similar to the target image features are generated through reverse optimization. Then, the text prompt word is vectorized and the diffusion model is input to generate a style image similar to the text prompt word semantics.
Through this method, the semantic correlation between text and image and the accuracy of generating images are improved, the demand for high accuracy in the financial and medical fields is met, and the consumption of computing resources is reduced.
Smart Images

Figure CN120198526A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an image generation method, device, equipment and medium. Background Art
[0002] In the field of multimodal generation, the task of generating pictures from text, as a key link in realizing cross-modal content creation, has received extensive attention and research in recent years. Especially in the two key fields of financial business and medical business, its application prospects are particularly broad, but at the same time, it also faces unique challenges.
[0003] In the financial field, the technology of generating pictures from text has great potential. For example, in the automatic generation of financial reports, the system can automatically generate corresponding charts or images according to the input financial data or descriptive text to intuitively display the change trends, comparison situations, etc. of the financial data. However, most current multimodal generation methods follow a one-way generation process, that is, either generate pictures only through text prompts, or reverse-infer text only through pictures. Due to the complexity and professionalism of financial data, this traditional one-way generation process often fails to accurately capture the semantic association between text and pictures, resulting in deviations in the accuracy of the generated pictures or images. In addition, training or loading independent encoders separately not only leads to a huge consumption of computing resources, but also reduces the generation efficiency and is difficult to meet the high requirements of the financial business for timeliness.
[0004] In the medical field, the technology of generating pictures from text also has broad application prospects. For example, in the automatic generation of medical imaging reports, the system can automatically generate corresponding medical images according to the description information input by doctors. However, the complexity and diversity of medical images pose greater challenges to multimodal tasks. When traditional one-way generation methods process medical texts, they often fail to accurately capture the fine semantic relationship between text and pictures, resulting in inaccuracies in the details of the generated medical images.
[0005] More critically, in these two fields of finance and medicine, the precision requirements for semantic alignment between text and pictures are extremely high. The accurate display of financial data and the precise presentation of medical images are both key links related to decision-making and diagnosis, and any minor deviation may bring serious consequences. Therefore, traditional multimodal generation methods seem inadequate when dealing with such high-demand tasks. Although there have been some attempts to optimize this process through joint training or parameter sharing, the effects are still not satisfactory, and there is still a large room for improvement in the precision of semantic alignment and computational efficiency. The semantic alignment degree between text and pictures is relatively low, and the accuracy of the generated images is not high. Summary of the Invention
[0006] The purpose of the embodiments of the present application is to propose an image generation method, device, equipment and medium to solve the problem that the semantic alignment between existing text and pictures is relatively low and the accuracy of the generated images is not high.
[0007] To solve the above technical problems, the embodiments of the present application provide an image generation method, which adopts the following technical solutions:
[0008] Obtain a target image; use a preset text-image encoding model to encode the target image to obtain an image feature vector of the target image; generate a text prompt similar to the target image feature through reverse optimization according to the image feature vector; use the text-image encoding model to vectorize the text prompt to obtain a text feature vector of the text prompt; input the text feature vector into a preset diffusion model to generate a style image similar to the semantics of the text prompt.
[0009] Further, the step of using a preset text-image encoding model to encode the target image to obtain an image feature vector of the target image specifically includes:
[0010] Use the image encoder of the preset text-image encoding model to extract features from the target image to obtain multiple feature maps of the target image;
[0011] Perform a pooling operation on the multiple feature maps to reduce the size of the multiple feature maps and extract the key features of the multiple feature maps to obtain multiple pooled feature information;
[0012] Map the multiple pooled feature information to a preset dimension to obtain an image feature vector of the target image.
[0013] Further, the step of generating a text prompt similar to the target image feature through reverse optimization according to the image feature vector specifically includes:
[0014] Obtain a pre-trained general text embedding vector, input the general text embedding vector into the text encoder of the text-image encoding model to obtain an initial text feature vector; calculate the first similarity between the image feature vector and the initial text feature vector, and construct a loss function based on the first similarity; optimize the loss function through the gradient descent algorithm, update the value of the general text embedding vector until the loss function reaches a preset convergence condition to obtain an optimized text embedding vector; decode the optimized text embedding vector to generate a text prompt with feature similarity to the target image.
[0015] Further, the step of optimizing the loss function through the gradient descent algorithm, updating the value of the text embedding vector until the loss function reaches a preset convergence condition to obtain an optimized text embedding vector specifically includes:
[0016] Obtain the initial value of the general text embedding vector as the starting point of the gradient descent algorithm; take the general text embedding vector as the target text embedding vector, and calculate the loss value of the target text embedding vector according to the loss function; based on the loss value and the initial value, calculate the gradient of the loss function with respect to the target text embedding vector, and determine the direction and step size of the gradient descent of the target text embedding vector in the space of the loss function; according to the direction and step size, update the value of the target text embedding vector to obtain a new text embedding vector; determine whether the loss value corresponding to the new text embedding vector satisfies the preset convergence condition; if the loss value corresponding to the new text embedding vector satisfies the convergence condition, it is determined that the loss function corresponding to the new text embedding vector satisfies the convergence condition, and the new text embedding vector is used as the optimized text embedding vector; if the loss value corresponding to the new text embedding vector does not satisfy the convergence condition, the new text embedding vector is used as the target text embedding vector, and return to execute the step of calculating the loss value until the loss value corresponding to the obtained text embedding vector satisfies the convergence condition, and the text embedding vector that satisfies the convergence condition is used as the optimized text embedding vector.
[0017] Further, the step of vectorizing the text prompt using a text-image encoding model to obtain the text feature vector of the text prompt specifically includes:
[0018] Perform word segmentation on the text prompt to obtain a sequence of prompt words; use a pre-trained word vector model to convert the sequence of prompt words into a sequence of word vectors; extract the text features of the sequence of word vectors through the text encoder of the text-image encoding model to obtain the text feature vector of the text prompt.
[0019] Further, the step of inputting the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt specifically includes:
[0020] Input the text feature vector into a preset diffusion model, and generate an initial image through the multi-layer neural network of the diffusion model;
[0021] According to the semantic features of the text prompt, obtain a reference image with a matching degree to the semantic features from a preset style image library;
[0022] Based on the reference image, optimize the style features of the initial image to obtain a style image semantically similar to the text prompt.
[0023] Further, after the step of inputting the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt, it further includes:
[0024] Use a text-image encoding model to encode the style image to obtain a style feature vector;
[0025] Calculate the second similarity between the text feature vector and the style feature vector;
[0026] Based on the second similarity, construct an alignment loss function, and based on the alignment loss function, optimize the image-text encoding model.
[0027] To solve the above technical problems, an embodiment of the present application further provides an image generation device, which adopts the following technical solutions:
[0028] An acquisition module, configured to acquire a target image;
[0029] A first encoding module, configured to encode the target image by using a preset image-text encoding model to obtain an image feature vector of the target image;
[0030] A word generation module, configured to generate a text prompt similar to the target image feature according to the image feature vector through reverse optimization;
[0031] A vectorization module, configured to vectorize the text prompt by using the image-text encoding model to obtain a text feature vector of the text prompt;
[0032] An image generation module, configured to input the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt.
[0033] To solve the above technical problems, an embodiment of the present application further provides a computer device, including a memory and a processor, where the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the above image generation method are implemented.
[0034] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable at least one processor to execute the steps of the above image generation method.
[0035] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: Through a series of carefully designed steps, reverse optimization from an image to a text prompt and generation of a style image based on the prompt are achieved. Specifically, first, a preset image-text encoding model is used to encode the target image into an image feature vector, which ensures accurate capture of image information. Subsequently, through a reverse optimization strategy, a text prompt similar to the image features is generated, which enhances the semantic relevance between the text and the image. Further, the generated text prompt is vectorized again using the image-text encoding model to obtain a text feature vector. This model design with shared parameters not only saves computing and storage resources but also improves the execution efficiency of the model. Finally, the text feature vector is input into a preset diffusion model to generate a style image semantically similar to the text prompt, which enhances the semantic accuracy of the generated content and improves the accuracy of image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;
[0038] Figure 2 is a flowchart of a method for generating an image provided by the present application;
[0039] Figure 3 is a schematic structural diagram of an image generation device provided by the present application;
[0040] Figure 4 is a schematic structural diagram of a computer device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of the embodiments of this application in this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the description and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the description and claims of this application or the above drawings are used to distinguish different objects and are not used to describe a specific order.
[0042] References to "embodiments" in this specification mean that particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0043] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0044] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0045] A user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0046] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, and a desktop computer, etc.
[0047] The server 103 may be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.
[0048] It should be noted that the image generation method provided in the embodiments of the present application is generally executed by the server. Correspondingly, the image generation device is generally arranged in the server.
[0049] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers therein are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0050] Continuing to refer to Figure 2 , a flowchart of an embodiment of the image generation method according to the present application is shown. The image generation method includes the following steps:
[0051] Step S201, obtain a target image.
[0052] In this embodiment, the electronic device (such as Figure 1 the server shown) on which the image generation method runs can obtain the target image through a wired connection or a wireless connection. It should be noted that the above wireless connection methods can include but are not limited to 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultrawideband) connections, and other currently known or future-developed wireless connection methods.
[0053] Among them, the target image refers to the image data to be processed. The target image can be sourced from various ways such as digital camera shooting, network acquisition, user upload, etc., and is the basis for subsequent steps such as encoding, feature extraction, and image generation. For example, a landscape photo is a target image, which will be used to extract image features and generate related text prompts and style images.
[0054] Step S202, encode the target image using a preset image-text encoding model to obtain an image feature vector of the target image.
[0055] Among them, the image-text encoding model is a model that can map image and text data to the same feature space. For example, a shared CLIP encoder (Contrastive Language-Image Pre-training) can be used as the image-text encoding model. The image-text encoding model includes two parts: a text encoder and a picture encoder, which can encode image and text data into high-dimensional vectors respectively, so that similar images and texts are close in the feature space. For example, through the picture encoder, a photo of a cat can be encoded into an image feature vector, and the image feature vector is close to the text feature vector describing the cat in the feature space.
[0056] Among them, the image feature vector is obtained after the target image is processed by a text-image encoding model (such as the image encoder of shared CLIP), and it is a high-dimensional vector used to represent the image content. It contains an abstract representation of the key information in the target image, such as color, shape, texture, etc., and is an important basis for subsequent generation of text prompts and style images. For example, the image feature vector of an image containing a sunset can reflect features such as the color and shape of the sunset and the gradient of the sky.
[0057] Step S203: According to the image feature vector, generate a text prompt similar to the target image feature through reverse optimization.
[0058] Among them, reverse optimization is a method of reverse derivation and generation of a text prompt similar to the image feature vector according to the image feature vector. For example, through the gradient descent algorithm, according to the information in the image feature vector, the generated text prompt can be continuously iteratively adjusted to make it more semantically matched with the image feature vector. For example, starting from the image feature vector corresponding to a damaged vehicle picture used in the auto insurance claim of property insurance, text prompts such as "the left front bumper of the vehicle is damaged, the headlights are broken, and there are scratches on the body" can be generated through reverse optimization.
[0059] Among them, the text prompt is a text description that is generated by the reverse optimization method and is semantically similar to the image feature vector. It is used to further guide the diffusion model to generate a style image related to the target image content. The text prompt can be a phrase, a sentence, or a paragraph, and the specific form depends on the application requirements and model design. For example, "a woman in a red dress is smiling in the garden" is a text prompt describing the image content.
[0060] Step S204: Use the text-image encoding model to vectorize the text prompt to obtain the text feature vector of the text prompt.
[0061] Among them, the text feature vector is obtained after the text prompt is processed by a text-image encoding model (such as the text encoder of shared CLIP), and it is a high-dimensional vector used to represent the text content. It is the key input for subsequent generation of style images. For example, inputting the text prompt "beautiful sunset scenery" into the text encoder can obtain the corresponding text feature vector.
[0062] Step S205: Input the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt.
[0063] Among them, the diffusion model is a deep learning model based on a probabilistic generative framework that can gradually generate high-quality images from noisy data. The diffusion model accepts text feature vectors as input and, through a series of iterative steps, gradually removes noise and generates style images that are semantically similar to the text prompt words. The diffusion model can capture the semantic information in the text feature vectors and transform it into visual elements in the images. For example, starting from a text feature vector describing "rolling mountains and lingering clouds", the diffusion model can generate a landscape painting with the same artistic conception.
[0064] Among them, the style image is an image generated by the diffusion model based on the text feature vector and is semantically similar to the input text prompt word. It combines the content information of the image feature vector and the semantic information of the text prompt word, presenting a unique artistic style or visual effect. Style images can be used in various application scenarios such as art creation, image editing, and virtual reality. For example, starting from the text prompt word "classical-style palace building", the generated style image can be a palace building drawing with classical aesthetic features.
[0065] The embodiments of this application can achieve the reverse optimization from the image to the text prompt word and the generation of style images based on the prompt word through a series of carefully designed steps. Specifically, first, use a preset image-text encoding model to encode the target image into an image feature vector, which ensures the accurate capture of image information. Subsequently, through a reverse optimization strategy, generate text prompt words similar to the image features, which enhances the semantic relevance between the text and the image. Further, use the image-text encoding model again to vectorize the generated text prompt words to obtain text feature vectors. This model design with shared parameters not only saves computing and storage resources but also improves the execution efficiency of the model. Finally, input the text feature vectors into a preset diffusion model to generate style images that are semantically similar to the text prompt words, which enhances the semantic accuracy of the generated content and improves the accuracy of image generation.
[0066] In one embodiment, the image generation method of this application is illustrated by taking financial images as an example. In the financial field, especially in the process of automatic generation of financial reports, traditional text-to-picture generation methods often have difficulty accurately capturing the complex semantic relationships between financial data and charts, resulting in deviations in the generated charts in terms of showing data trends, comparison situations, etc. This not only affects the intuitiveness and readability of the report but may also mislead the judgment of investors or decision-makers. Therefore, this embodiment proposes an improved text-to-picture generation method aimed at improving the accuracy and efficiency of financial chart generation.
[0067] Specifically, a historical chart containing financial data can be extracted from a financial database as the target image, which shows the profit change trend of a company in the past five years. The preset image-text encoding model is used to encode the target image to obtain its image feature vector. The model can make full use of the features of text and image to achieve cross-modal semantic understanding. Next, according to the image feature vector, a reverse optimization algorithm (such as gradient descent method) is used to generate text prompt words similar to the target image features. For example, the generated text prompt words may include "profits have continued to grow in the past five years, reaching a peak in 2022". The image-text encoding model is used again to vectorize the generated text prompt words to obtain their text feature vectors. Finally, the text feature vector is input into the preset diffusion model to generate a style image with semantics similar to the text prompt word. While maintaining data accuracy, the image also has a certain artistic style, which enhances the visual effect of the report.
[0068] Through the method of this embodiment, the generated financial charts not only accurately reflect the changing trends of historical financial data, but also reflect the company's brand image and reporting style in details (such as color matching, line thickness, etc.). Compared with traditional methods, the charts generated by the method of this embodiment have been significantly improved in accuracy and aesthetics, effectively improving the quality and readability of financial reports. Through the graphic encoding model and reverse optimization algorithm, accurate semantic alignment between text and image is achieved, and the accuracy of the chart is improved. It avoids separate training or loading of independent encoders, reduces the consumption of computing resources, and improves generation efficiency. The style image generated by the diffusion model enhances the visual effect and attractiveness of the financial report.
[0069] In one embodiment, the image generation method of the present application is applied to the generation of text to pictures in a medical business scenario. In the medical field, the automatic generation of medical imaging reports is one of the important applications of text to picture generation technology. However, when processing medical text, traditional one-way generation methods often find it difficult to accurately capture the fine semantic relationship between text and medical images, resulting in the generated medical images being inaccurate in details. This not only affects the accuracy of the doctor's diagnosis, but may also delay the patient's treatment. Therefore, this embodiment proposes an improved text to picture generation method, which aims to improve the accuracy and efficiency of medical image generation.
[0070] Specifically, a CT scan image containing a specific lesion is extracted from a medical image database as the target image. The same pre-set text-image encoding model is used to encode the target image to obtain its image feature vector. According to the image feature vector, through a reverse optimization algorithm, a text prompt similar to the target image feature is generated. For example, the generated text prompt may include "There is a nodular high-density shadow visible in the lower right lobe of the lung, with clear edges", etc. The text prompt generated is vectorized again using the text-image encoding model. The text feature vector is input into a pre-set generative adversarial network to generate a medical image semantically similar to the text prompt. While maintaining the lesion characteristics, this image also has a certain degree of realism and clarity.
[0071] Through the method of this embodiment, the generated medical image not only accurately reflects the characteristics and location of the lesion, but also embodies the characteristics of real medical images in details (such as tissue texture, contrast, etc.). Compared with traditional methods, the medical image generated by the method provided in this embodiment has significant improvements in both accuracy and realism, effectively assisting doctors in diagnosis and decision-making. Through the text-image encoding model and the reverse optimization algorithm, an accurate semantic alignment between text and medical images is achieved, improving the accuracy of medical images. The medical image generated by the generative adversarial network has higher realism and clarity, which helps doctors make more accurate diagnoses. The method of this embodiment avoids complex preprocessing steps and independent encoder training, improves the generation efficiency, and can flexibly adjust the parameters of the generative adversarial network according to different requirements to adapt to different types of medical image generation tasks.
[0072] In some alternative implementation manners of this embodiment, in step 202, using a pre-set text-image encoding model to encode the target image to obtain the image feature vector of the target image specifically includes the following steps:
[0073] Using the image encoder of the pre-set text-image encoding model to extract features from the target image to obtain multiple feature maps of the target image; performing a pooling operation on the multiple feature maps to reduce the size of the multiple feature maps and extract the key features of the multiple feature maps to obtain multiple pooled feature information; mapping the multiple pooled feature information to a pre-set dimension to obtain the image feature vector of the target image.
[0074] Among them, the image encoder is a technical component that extracts features from the target image using a pre-set text-image encoding model (such as the image encoding part in the shared CLIP encoder). It converts the image data into a series of abstract feature representations by analyzing visual elements such as pixel information, color distribution, and texture in the image. These feature representations can reflect the key information in the image. For example, the image encoder can extract the features of elements such as trees, mountains, and sky from a landscape photo to form multiple feature maps.
[0075] Among them, multiple feature maps refer to multiple two-dimensional matrices or tensors generated by an image encoder when processing a target image according to different attributes of the image (such as color, texture, shape, etc.). Each feature map corresponds to the spatial distribution of a specific attribute in the image and is an abstract representation of the image features. For example, one feature map may represent the distribution of edge information in the image, and another feature map may represent the distribution of colors in the image.
[0076] Among them, the pooling operation is a technique for downsampling feature maps, which reduces the computational amount by reducing the size of the feature maps while retaining the key features of the image. For example, through the max-pooling operation, the maximum value of each local region can be extracted from a feature map to form a new, smaller-sized feature map.
[0077] Among them, key features refer to the features in the feature maps that are of great significance for subsequent processing tasks. These features can reflect the main content and attributes of the feature maps and are the core components of the image information.
[0078] Among them, multiple pooling feature information refers to multiple feature representations obtained by performing pooling operations on multiple feature maps. These feature representations retain the key features of the image while reducing the size and computational amount of the feature maps.
[0079] In one example, first, the target images in the financial field (such as checks, ID cards, etc.) can be subjected to convolution operations through the image encoder of the image-text encoding model to obtain multiple feature maps. These feature maps contain key information such as the color, texture, and shape in the image. Next, pooling operations are performed on the multiple feature maps to reduce the size of the feature maps and extract key features. The pooling operation can select strategies such as max-pooling or average-pooling to retain the most important information in the image while reducing the computational amount. After the pooling operation, multiple pooling feature information is obtained, and these feature information are more compact and contain the key features of the image. Finally, the multiple pooling feature information is mapped to a preset dimension to obtain the image feature vector of the target image. For example, taking check processing as an example, first, the check image is preprocessed, such as denoising and grayscaling. Then, the image encoder of the image-text encoding model is used to extract features from the check image to obtain multiple feature maps. These feature maps contain key information such as the amount, date, and payee on the check. Next, max-pooling operations are performed on the feature maps to extract the maximum value in each feature map as the key feature. After the pooling operation, multiple pooling feature information is obtained, and these feature information are more compact and contain the key information of the check. Finally, the pooling feature information is mapped to a preset dimension (such as 256 dimensions) to obtain the image feature vector of the check image.
[0080] In the embodiment of the present application, through the image encoder of the graphic-text encoding model, multi-dimensional features in the target image are deeply mined to generate a rich set of feature maps, which depict the details and structure of the image in detail. Further, a pooling operation is performed on multiple feature maps, which not only effectively reduces the size of the feature maps and the computational complexity, but also accurately extracts the key features in the image, enhancing the robustness and representativeness of the feature information. This step ensures that even in the case of image size changes or noise interference, the key features can still be stably captured. Finally, multiple pooled feature information is mapped to a preset dimension, and the obtained image feature vector is highly condensed and informative.
[0081] In some alternative implementation manners of this embodiment, in step S203, according to the image feature vector, a text prompt similar to the target image feature is generated through reverse optimization, which specifically includes the following steps:
[0082] Obtain a pre-trained general text embedding vector, input the general text embedding vector into the text encoder of the graphic-text encoding model to obtain an initial text feature vector; calculate the first similarity between the image feature vector and the initial text feature vector, and construct a loss function based on the first similarity; optimize the loss function through the gradient descent algorithm, update the value of the general text embedding vector until the loss function reaches a preset convergence condition, and obtain an optimized text embedding vector; decode the optimized text embedding vector to generate a text prompt with similar features to the target image.
[0083] Among them, the general text embedding vector refers to a low-dimensional dense vector obtained through the pre-training process and capable of representing text semantic information. These vectors usually come from deep learning models of large-scale text corpora, such as BERT, GPT, etc. In the graphic-text encoding model, the general text embedding vector is used as the initial representation of text information for subsequent matching and fusion with the image feature vector.
[0084] Among them, the text encoder is a component in the graphic-text encoding model (such as a shared CLIP encoder), which is responsible for converting the general text embedding vector into an initial text feature vector for subsequent generation of text prompts based on the initial text feature vector.
[0085] Among them, the initial text feature vector refers to the vector obtained after being processed by the text encoder and representing the high-level semantic information of the text prompt. It is used for calculating the similarity with the image feature vector to evaluate the correlation between the graphic and the text.
[0086] Among them, the first similarity refers to the similarity metric between the image feature vector and the initial text feature vector, which is used to evaluate the correlation or matching degree between the graphic and the text. It is used to construct a loss function to guide the model to optimize the graphic-text matching performance.
[0087] Among them, the gradient descent algorithm is an iterative method for optimizing the parameters of machine learning models. It minimizes the loss function by continuously adjusting the model parameters. It is used to optimize parameters such as text embedding vectors to minimize the loss function and improve the performance of image-text matching.
[0088] In one example, pre-trained general text embedding vectors are obtained. These vectors have been fully trained on a large amount of text data and can capture the basic semantic features of the text. Then, the general text embedding vectors are input into the text encoder part of the shared CLIP encoder. Through the neural network structure inside the encoder, an initial text feature vector is obtained. Subsequently, the first similarity between the image feature vector and the initial text feature vector is calculated. Based on the first similarity, a loss function is constructed, which aims to minimize the degree of mismatch between the image and text features. The loss function is optimized by the gradient descent algorithm, and the values of the general text embedding vectors are iteratively updated. In each iteration, each component of the embedding vector is adjusted according to the gradient information of the loss function, gradually approaching the optimal solution. When the loss function reaches the preset convergence condition (such as the loss value is less than a certain threshold or the number of iterations reaches the upper limit), the optimization stops, and the optimized text embedding vectors are obtained. Finally, the optimized text embedding vectors are decoded and converted back into text form to generate text prompts that are feature-similar to the target image. These text prompts can accurately reflect the main content and features of the target image.
[0089] The embodiments of this application can realize the preliminary extraction of text features by introducing pre-trained general text embedding vectors and combining with the text encoder of the image-text encoding model. On this basis, by calculating the similarity between the image feature vector and the initial text feature vector and constructing a loss function accordingly, the feature difference between the text and the image can be accurately measured. The gradient descent algorithm is used to optimize the loss function, so that the values of the general text embedding vectors continuously approach the optimal solution during the iteration process, thereby obtaining the optimized text embedding vectors. This process not only improves the alignment degree of the text and image features, but also significantly enhances the relevance between the text prompts and the target image. Finally, by decoding the optimized text embedding vectors, text prompts that are highly similar to the target image features can be generated. These prompts not only accurately reflect the main content of the image, but also have high readability and practicality.
[0090] In some optional implementation manners of this embodiment, the step of "optimizing the loss function by the gradient descent algorithm, updating the values of the text embedding vectors until the loss function reaches the preset convergence condition, and obtaining the optimized text embedding vectors" specifically includes the following steps:
[0091] Obtain the initial value of the general text embedding vector as the starting point of the gradient descent algorithm; take the general text embedding vector as the target text embedding vector, and calculate the loss value of the target text embedding vector according to the loss function; based on the loss value and the initial value, calculate the gradient of the loss function with respect to the target text embedding vector, and determine the direction and step size of the gradient descent of the target text embedding vector in the space of the loss function; update the value of the target text embedding vector according to the direction and step size to obtain a new text embedding vector; determine whether the loss value corresponding to the new text embedding vector meets the preset convergence condition; if the loss value corresponding to the new text embedding vector meets the convergence condition, it is determined that the loss function corresponding to the new text embedding vector meets the convergence condition, and the new text embedding vector is used as the optimized text embedding vector; if the loss value corresponding to the new text embedding vector does not meet the convergence condition, the new text embedding vector is used as the target text embedding vector, and return to execute the step of calculating the loss value until the loss value corresponding to the obtained text embedding vector meets the convergence condition, and the text embedding vector that meets the convergence condition is used as the optimized text embedding vector.
[0092] Among them, the initial value refers to the starting numerical value of the general text embedding vector obtained before optimizing by the gradient descent algorithm. This numerical value usually comes from a pre-trained model and is the basis for the iterative optimization of the algorithm. It represents the initial state of the general text embedding vector and is used as the starting point of the gradient descent to find the optimal solution.
[0093] Among them, the loss value of the target text embedding vector refers to the numerical value calculated according to the loss function in the current iteration round, which represents the degree of mismatch between the target text embedding vector and the image feature vector. It represents the distance between the current text embedding vector and the ideal position in the feature space and is used to measure the progress and direction of optimization. For example, in the text-image matching task, the smaller the loss value, the higher the similarity between the text embedding vector and the image feature vector, and the better the matching effect.
[0094] Among them, the gradient refers to the partial derivative of the loss function with respect to the target text embedding vector, which represents the rate of change of the loss value when the target text embedding vector changes in all directions in the space of the loss function. It represents the steepest upward direction of the loss function at the current point (the negative gradient direction is the direction of the fastest descent), and is used to guide the update direction of the target text embedding vector.
[0095] Among them, the direction refers to the specific path or trend of the update of the target text embedding vector in the gradient descent algorithm. Based on the calculation result of the gradient, it determines the direction in which the target text embedding vector should move in the space of the loss function to reduce the loss value. This direction is usually consistent with the negative direction of the gradient because the negative gradient direction is the direction of the fastest descent of the loss function.
[0096] Among them, the step size refers to the distance or stride by which the target text embedding vector moves along the gradient direction in each iteration of the gradient descent algorithm. It characterizes the convergence speed and stability of the algorithm during the optimization process and is an important parameter for adjusting the algorithm performance.
[0097] In one example, a pre-trained general text embedding vector is obtained as the initial value. This vector has been trained on a large amount of text data and has certain semantic representation capabilities. This initial value will serve as the starting point of the gradient descent algorithm for subsequent optimization processes. Then, this general text embedding vector is used as the target text embedding vector, and a loss function is used to measure the degree of mismatch between the target text embedding vector and the image feature vector. For example, a loss function is used to calculate the similarity between the target text embedding vector and the image feature vector and use it as the loss value. Then, based on the loss value and the initial value, the gradient of the loss function with respect to the target text embedding vector is calculated. The gradient represents the rate of change of the loss value when the target text embedding vector changes in various directions in the loss function space. By calculating the gradient, the direction and step size of the gradient descent of the target text embedding vector in the loss function space can be determined. Next, according to the direction and step size of the gradient descent, the value of the target text embedding vector is updated to obtain a new text embedding vector. This update process is achieved by adding a vector proportional to the gradient to the current value of the target text embedding vector, and the step size determines the magnitude of this vector. After the update, it is judged whether the loss value corresponding to the new text embedding vector meets the preset convergence condition. The convergence condition is usually a threshold. When the loss value is less than this threshold, it can be considered that the optimization process has converged and a satisfactory text embedding vector has been obtained. If the convergence condition is met, the new text embedding vector is used as the optimized text embedding vector. If the convergence condition is not met, the new text embedding vector is used as the target text embedding vector, and the step of calculating the loss value is returned to continue the optimization process until a text embedding vector that meets the convergence condition is obtained.
[0098] The embodiments of this application can ensure that there is a clear starting benchmark for the optimization process by setting the initial value of the general text embedding vector as the starting point of the gradient descent algorithm. Subsequently, the loss value of the target text embedding vector is accurately calculated using the loss function, providing a reliable basis for gradient calculation. By calculating the gradient based on the loss value and the initial value, the descent direction and step size in the loss function space are accurately determined, thus realizing the effective adjustment of the target text embedding vector. During the iterative update process, it is continuously judged whether the loss value corresponding to the new text embedding vector meets the preset convergence condition, ensuring the precise control of the optimization process. Once the convergence condition is met, that is, the optimized text embedding vector is obtained. This process not only improves the accuracy of the text embedding vector but also enhances its applicability and effect in subsequent processing tasks by continuously approaching the optimal solution.
[0099] In some alternative implementation manners of this embodiment, in step S204, a text encoding model is used to vectorize the text prompt to obtain a text feature vector of the text prompt, which specifically includes the following steps:
[0100] Perform word segmentation on the text prompt to obtain a sequence of prompt words; use a pre-trained word vector model to convert the sequence of prompt words into a sequence of word vectors; extract the text features of the sequence of word vectors through the text encoder of the text encoding model to obtain a text feature vector of the text prompt.
[0101] Among them, the sequence of prompt words is an ordered combination of words obtained by performing word segmentation on the input text prompt. It is used to be converted into a numerical sequence of word vectors through the word vector model for computer processing. For example, for the text prompt "beautiful scenery", its sequence of prompt words may be ["beautiful", "scenery"].
[0102] Among them, the word vector model is a pre-trained model that maps words to a high-dimensional vector space. This word vector model is trained with a large amount of text data to learn the distributed representation of each word, that is, to convert it into a vector with a specific dimension.
[0103] Among them, the sequence of word vectors is a set of vectors obtained by converting each word in the sequence of prompt words through the word vector model. Each vector in this sequence of word vectors corresponds to a word in the sequence of prompt words and maintains the original order of the words.
[0104] In an example, perform word segmentation on the input text prompt. For example, for the text prompt "by the sea at sunset", a rule-based or machine learning-based word segmentation algorithm can be used to segment it into the sequence of prompt words "sunset / by / sea". Then, use a pre-trained word vector model (such as Word2Vec, GloVe, or BERT, etc.) to convert each word in the sequence of prompt words into a corresponding word vector. These word vectors usually have high dimensionality and can capture the semantic relationships between words. For example, for the sequence of prompt words "sunset / by / sea", its corresponding sequence of word vectors may be a series of 128-dimensional or higher-dimensional vectors. Then, encode the sequence of word vectors through the text encoder of the shared CLIP encoder, and the text encoder can utilize the semantic information in the sequence of word vectors to extract a text feature vector that can represent the text content.
[0105] The embodiment of this application can effectively extract features of the text prompt by using a pre-trained word vector model and a text encoder. It not only improves the accuracy and efficiency of text feature extraction, but also can capture the semantic relationships between words, providing strong support for subsequent tasks.
[0106] In some alternative implementation manners of this embodiment, in step S205, inputting the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt may specifically include the following steps:
[0107] Input the text feature vector into a preset diffusion model, and generate an initial image through the multi-layer neural network of the diffusion model; according to the semantic features of the text prompt, obtain a reference image with a high degree of matching with the semantic features from a preset style image library; based on the reference image, optimize the style features of the initial image to obtain a style image semantically similar to the text prompt.
[0108] Among them, the diffusion model is a deep learning model, which is derived from the simulation of the thermodynamic diffusion process. It gradually introduces noise into the data or gradually removes noise from the noisy data through a multi-layer neural network to generate new data samples. The diffusion model receives the text feature vector as input and gradually generates an initial image that matches the text description through its internal multi-layer neural network structure. For example, when inputting the text feature vector describing "a golden wheat field", the diffusion model can generate an initial image containing elements of a golden wheat field.
[0109] Among them, the initial image is the image first generated by the diffusion model according to the input text feature vector. It is derived from the analysis and transformation of the text feature vector by the diffusion model and represents the initial visual manifestation of the text description.
[0110] Among them, the reference image is selected from a preset style image library and has a high degree of matching with the semantic features of the input text prompt. During the image style optimization process, the reference image is used as the source of style features to guide the style adjustment of the initial image.
[0111] In an example, input the text feature vector into a pre-trained diffusion model. The diffusion model is composed of a multi-layer neural network and can receive the text feature vector and generate an initial image. Although this initial image can reflect some content of the text prompt corresponding to the text feature vector, its style may be relatively single or lack personality. To enhance the style features of the image, a reference image that matches the semantic features can be retrieved from a preset style image library according to the semantic features of the text prompt. The style image library contains a large number of images with different style features, such as oil paintings, watercolor paintings, sketches, etc., ensuring that a style reference that fits the semantics of the text prompt can be found. Subsequently, optimize the style features of the initial image based on the retrieved reference image. This step can include various processing methods such as style transfer, color adjustment, and texture enhancement, and finally a style image that is both semantically similar to the text prompt and has unique style features can be obtained.
[0112] In one embodiment, after optimizing the style features of the initial image based on the reference image to obtain a style image semantically similar to the text prompt, it is also possible to determine whether the semantic similarity between the optimized image and the text prompt reaches a preset threshold; if it reaches the preset threshold, the optimized image is used as the style image semantically similar to the text prompt; if the semantic similarity between the optimized image and the text prompt does not reach the preset threshold, the optimized image is repeatedly input into the diffusion model for the next round of image generation and style feature optimization until the semantic similarity between the generated image and the text prompt reaches the preset threshold to obtain an iterative image, and the iterative image is used as the style image.
[0113] In the embodiment of the present application, the text feature vector can be transformed into an initial image by using a preset diffusion model, which ensures a high correlation between the initial construction of the image content and the text prompt. Subsequently, according to the semantic features of the text prompt, the reference images in the style image library are accurately screened to achieve an accurate match between the style and the content, enhancing the personalization and pertinence of image generation. Finally, by optimizing the style features of the initial image, not only the semantic essence of the text prompt is retained, but also the unique style of the selected reference image is skillfully incorporated, thus generating an image work that conforms to the text description and has a specific style, improving the accuracy of image generation.
[0114] In some optional implementation manners of this embodiment, after inputting the text feature vector into a preset diffusion model in step S205 to generate a style image semantically similar to the text prompt, the following steps may specifically be further included:
[0115] Use a text-image encoding model to encode the style image to obtain a style feature vector; calculate the second similarity between the text feature vector and the style feature vector; based on the second similarity, construct an alignment loss function, and based on the alignment loss function, optimize the text-image encoding model.
[0116] Among them, the style feature vector refers to the vector representation obtained by encoding the style image through a text-image encoding model (such as the image encoder part in the shared CLIP encoder). This style feature vector extracts and characterizes the style features of the image, such as unique visual attributes like color matching, texture patterns, and composition layouts.
[0117] Among them, the second similarity refers to the similarity index calculated between the text feature vector and the style feature vector after being processed by the text-image encoding model. This similarity measures the proximity in semantics and style between the style image generated from the text prompt and the original text prompt. By calculating the distance between the two in the vector space (such as cosine similarity), the performance of the text-image encoding model in transforming text into image style can be quantitatively evaluated.
[0118] Among them, the alignment loss function is a loss function used to optimize the image-text encoding model, which is constructed based on the second similarity between the text feature vector and the style feature vector. The alignment loss function aims to improve the accuracy of the image-text encoding model in the image-text alignment task by minimizing the inconsistency between the two. Specifically, the alignment loss function calculates the deviation between the feature vector of the actually generated style image and the expected text feature vector, and uses this as a basis to guide the training process of the image-text encoding model. During the training process, the image-text encoding model continuously adjusts its parameters to reduce this deviation.
[0119] In one example, to evaluate the consistency between the generated style image and the original text prompt, the image-text encoding model is used again to encode the style image to obtain the style feature vector. Then, the second similarity (such as cosine similarity) between the text feature vector corresponding to the text prompt and the style feature vector is calculated. Based on the calculated second similarity, an alignment loss function is constructed, and the image-text encoding model is optimized through the backpropagation algorithm, enabling it to more accurately capture the potential association between the image and the text. After multiple iterative trainings, the performance of the image-text encoding model in image feature extraction and text generation is significantly improved.
[0120] The embodiment of the present application can construct an alignment loss function by calculating the second similarity between the text feature vector and the feature vector of the generated style image. This alignment loss function, as an optimization tool, can feedback and guide the iterative improvement of the image-text encoding model, ensuring that the generated image features are closer to the semantic content of the original text prompt.
[0121] It should be emphasized that to further ensure the privacy and security of the above-mentioned target image, text prompt, and style image, the above-mentioned target image, text prompt, and style image can also be stored in a node of a blockchain.
[0122] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods, and each data block contains information about a batch of network transactions, used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0123] Embodiments of the present application can construct and optimize related models and networks based on artificial intelligence technologies, such as text-image encoding models, diffusion models, etc. Among them, an artificial intelligence (AI) model is the crystallization of theory and practice that simulates the decision-making process of human intelligence through algorithms and data analysis to solve complex problems, predict future trends, or automate tasks. These models utilize a large amount of historical data and real-time information and are trained and optimized through specific algorithm frameworks to achieve efficient, accurate, and reliable performance.
[0124] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), etc., or a random access memory (RAM), etc.
[0125] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time but can be executed at different times, and their execution order is not necessarily sequential but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0126] Further reference Figure 3 to Figure 2 as an implementation of the method shown above, an embodiment of an image generation device is provided in the present application. This device embodiment corresponds to the method embodiment shown in Figure 2 and can be specifically applied to various electronic devices.
[0127] As Figure 3 shown, the image generation device 400 in this embodiment includes: an acquisition module 401, a first encoding module 402, a word generation module 403, a vectorization module 404, and an image generation module 405. Among them:
[0128] The acquisition module 401 is used to acquire a target image;
[0129] The first encoding module 402 is configured to encode the target image using a preset graphic and text encoding model to obtain an image feature vector of the target image;
[0130] The word generation module 403 is configured to generate a text prompt similar to the target image feature according to the image feature vector through reverse optimization;
[0131] The vectorization module 404 is configured to vectorize the text prompt using the graphic and text encoding model to obtain a text feature vector of the text prompt;
[0132] The image generation module 405 is configured to input the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt.
[0133] The embodiment of the present application can achieve the reverse optimization from image to text prompt and the generation of style images based on the prompt through a series of carefully designed steps. Specifically, first, use a preset graphic and text encoding model to encode the target image into an image feature vector, which ensures the accurate capture of image information. Subsequently, through a reverse optimization strategy, generate a text prompt similar to the image feature, which enhances the semantic relevance between the text and the image. Further, use the graphic and text encoding model again to vectorize the generated text prompt to obtain a text feature vector. This model design with shared parameters not only saves computing and storage resources but also improves the execution efficiency of the model. Finally, input the text feature vector into a preset diffusion model to generate a style image semantically similar to the text prompt, which enhances the semantic accuracy of the generated content and improves the accuracy of image generation.
[0134] In one embodiment, the first encoding module 402 includes:
[0135] The first extraction sub-module is configured to extract features of the target image using the image encoder of the preset graphic and text encoding model to obtain multiple feature maps of the target image;
[0136] The pooling sub-module is configured to perform a pooling operation on the multiple feature maps to reduce the size of the multiple feature maps and extract key features of the multiple feature maps to obtain multiple pooled feature information;
[0137] The mapping sub-module is configured to map the multiple pooled feature information to a preset dimension to obtain an image feature vector of the target image.
[0138] The embodiment of the present application can deeply explore the multi-dimensional features in the target image through the picture encoder of the graphic coding model, and generate a rich set of feature maps, which describe the details and structure of the image in detail. Furthermore, the pooling operation is performed on multiple feature maps, which not only effectively reduces the size of the feature map and reduces the computational complexity, but also accurately extracts the key features in the image and enhances the robustness and representativeness of the feature information. This step ensures that the key features can still be stably captured even when the image size changes or noise interferes. Finally, multiple pooled feature information is mapped to a preset dimension, and the resulting image feature vector is highly condensed and rich in information.
[0139] In one embodiment, the word generation module 403 includes:
[0140] A vector acquisition submodule is used to obtain a pre-trained universal text embedding vector, input the universal text embedding vector into a text encoder of a graphic-text encoding model, and obtain an initial text feature vector;
[0141] A calculation submodule, used for calculating a first similarity between the image feature vector and the initial text feature vector, and constructing a loss function based on the first similarity;
[0142] An update submodule is used to optimize the loss function through a gradient descent algorithm, update the value of the general text embedding vector, until the loss function reaches a preset convergence condition, and obtain the optimized text embedding vector;
[0143] The decoding submodule is used to decode the optimized text embedding vector and generate text prompt words with similar characteristics to the target image.
[0144] The embodiment of the present application can realize the preliminary extraction of text features by introducing a pre-trained universal text embedding vector and combining the text encoder of the image-text encoding model. On this basis, by calculating the similarity between the image feature vector and the initial text feature vector and constructing a loss function accordingly, the feature difference between the text and the image can be accurately measured. The gradient descent algorithm is used to optimize the loss function so that the value of the universal text embedding vector continuously approaches the optimal solution during the iteration process, thereby obtaining the optimized text embedding vector. This process not only improves the alignment of text and image features, but also significantly enhances the correlation between text prompt words and target images. Finally, by decoding the optimized text embedding vector, text prompt words that are highly similar to the target image features can be generated. These prompt words not only accurately reflect the main content of the image, but also have high readability and practicality.
[0145] In one embodiment, the update sub-module is further configured to obtain the initial value of the general text embedding vector as the starting point of the gradient descent algorithm; use the general text embedding vector as the target text embedding vector, and calculate the loss value of the target text embedding vector according to the loss function; based on the loss value and the initial value, calculate the gradient of the loss function with respect to the target text embedding vector, and determine the direction and step size of the gradient descent of the target text embedding vector in the space of the loss function; according to the direction and step size, update the value of the target text embedding vector to obtain a new text embedding vector; determine whether the loss value corresponding to the new text embedding vector satisfies a preset convergence condition; if the loss value corresponding to the new text embedding vector satisfies the convergence condition, it is determined that the loss function corresponding to the new text embedding vector satisfies the convergence condition, and the new text embedding vector is used as the optimized text embedding vector; if the loss value corresponding to the new text embedding vector does not satisfy the convergence condition, the new text embedding vector is used as the target text embedding vector, and the step of calculating the loss value is returned until the loss value corresponding to the obtained text embedding vector satisfies the convergence condition, and the text embedding vector that satisfies the convergence condition is used as the optimized text embedding vector.
[0146] In the embodiment of the present application, by setting the initial value of the general text embedding vector as the starting point of the gradient descent algorithm, it is ensured that the optimization process has a clear starting benchmark. Subsequently, the loss value of the target text embedding vector is accurately calculated using the loss function, providing a reliable basis for gradient calculation. By calculating the gradient based on the loss value and the initial value, the descent direction and step size in the loss function space are accurately determined, thereby realizing the effective adjustment of the target text embedding vector. During the iterative update process, it is continuously determined whether the loss value corresponding to the new text embedding vector satisfies the preset convergence condition, ensuring the precise control of the optimization process. Once the convergence condition is met, the optimized text embedding vector is obtained. This process not only improves the accuracy of the text embedding vector but also enhances its applicability and effect in subsequent processing tasks by continuously approaching the optimal solution.
[0147] In one embodiment, the vectorization module 404 includes:
[0148] The word segmentation sub-module is configured to perform word segmentation on the text prompt to obtain a prompt word sequence;
[0149] The conversion sub-module is configured to use a pre-trained word vector model to convert the prompt word sequence into a word vector sequence;
[0150] The second extraction sub-module is configured to extract the text features of the word vector sequence through the text encoder of the image-text encoding model to obtain the text feature vector of the text prompt.
[0151] Embodiments of the present application can effectively extract features of text prompts by using a pre-trained word vector model and a text encoder. This not only improves the accuracy and efficiency of text feature extraction but also captures the semantic relationships between words, providing strong support for subsequent tasks.
[0152] In one embodiment, the image generation module 405 includes:
[0153] An input sub-module for inputting a text feature vector into a preset diffusion model and generating an initial image through the multi-layer neural network of the diffusion model;
[0154] An image acquisition sub-module for obtaining a reference image with a matching degree to the semantic feature from a preset style image library according to the semantic feature of the text prompt;
[0155] An optimization sub-module for optimizing the style feature of the initial image based on the reference image to obtain a style image similar in semantics to the text prompt.
[0156] Embodiments of the present application can convert a text feature vector into an initial image by using a preset diffusion model. This process ensures a high degree of relevance between the initial construction of the image content and the text prompt. Subsequently, according to the semantic feature of the text prompt, a reference image in the style image library is accurately selected, achieving an accurate match between style and content, enhancing the personalization and pertinence of image generation. Finally, by optimizing the style feature of the initial image, not only the semantic essence of the text prompt is retained, but also the unique style of the selected reference image is skillfully incorporated, thus generating an image work that conforms to the text description and has a specific style, improving the accuracy of image generation.
[0157] In one embodiment, the image generation device may further include:
[0158] A second encoding module for encoding the style image by using a text-image encoding model to obtain a style feature vector;
[0159] A calculation module for calculating a second similarity between the text feature vector and the style feature vector;
[0160] A construction module for constructing an alignment loss function based on the second similarity and optimizing the text-image encoding model based on the alignment loss function.
[0161] Embodiments of the present application can construct an alignment loss function by calculating the second similarity between the text feature vector and the feature vector of the generated style image. This alignment loss function, as an optimization tool, can feedback and guide the iterative improvement of the text-image encoding model to ensure that the generated image features are closer to the semantic content of the original text prompt.
[0162] To solve the above technical problems, an embodiment of the present application further provides a computer device. Specifically, please refer to Figure 4 , Figure 4 , which is the basic structural block diagram of the computer device in this embodiment.
[0163] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 6 with the memory 61, the processor 62, and the network interface 63 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0164] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, or other computing devices. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.
[0165] The memory 61 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the image generation method, etc. In addition, the memory 61 may also be used to temporarily store various data that have been output or will be output.
[0166] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the image generation method.
[0167] The network interface 63 may include a wireless network interface or a wired network interface, and the network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.
[0168] The embodiments of the present application can achieve the reverse optimization from images to text prompts and the generation of style images based on the prompts through a series of carefully designed steps. Specifically, first, a preset image-text encoding model is used to encode the target image into an image feature vector, which ensures the accurate capture of image information. Subsequently, through a reverse optimization strategy, text prompts similar to the image features are generated, which enhances the semantic correlation between the text and the image. Further, the generated text prompts are vectorized again using the image-text encoding model to obtain text feature vectors. This model design with shared parameters not only saves computing and storage resources but also improves the execution efficiency of the model. Finally, the text feature vectors are input into a preset diffusion model to generate style images semantically similar to the text prompts, which enhances the semantic accuracy of the generated content and improves the accuracy of image generation.
[0169] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that at least one processor executes the steps of the image generation method as described above.
[0170] The embodiments of the present application can achieve the reverse optimization from images to text prompts and the generation of style images based on the prompts through a series of carefully designed steps. Specifically, first, a preset image-text encoding model is used to encode the target image into an image feature vector, which ensures the accurate capture of image information. Subsequently, through a reverse optimization strategy, text prompts similar to the image features are generated, which enhances the semantic correlation between the text and the image. Further, the generated text prompts are vectorized again using the image-text encoding model to obtain text feature vectors. This model design with shared parameters not only saves computing and storage resources but also improves the execution efficiency of the model. Finally, the text feature vectors are input into a preset diffusion model to generate style images semantically similar to the text prompts, which enhances the semantic accuracy of the generated content and improves the accuracy of image generation.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0172] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the accompanying drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure made by using the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is similarly within the scope of patent protection of this application.
Claims
1. An image generation method, characterized in that: The steps include: Get the target image; Using a preset image-text coding model to encode the target image to obtain an image feature vector of the target image; According to the image feature vector, generating text prompt words similar to the target image feature through reverse optimization; Using the image-text coding model, vectorizing the text prompt word to obtain a text feature vector of the text prompt word; The text feature vector is input into a preset diffusion model to generate a style image that is semantically similar to the text prompt word.
2. The method according to claim 1, characterized in that The step of using a preset image-text coding model to encode the target image to obtain an image feature vector of the target image specifically includes: Using a picture encoder of a preset picture-text encoding model to extract features of the target image, and obtain multiple feature maps of the target image; Performing a pooling operation on the multiple feature maps to reduce the sizes of the multiple feature maps and extract key features of the multiple feature maps to obtain multiple pooled feature information; The multiple pooled feature information are mapped to a preset dimension to obtain an image feature vector of the target image.
3. The method according to claim 1, characterized in that The step of generating text prompt words similar to the target image features through reverse optimization according to the image feature vector specifically includes: Obtaining a pre-trained universal text embedding vector, and inputting the universal text embedding vector into a text encoder of the image-text encoding model to obtain an initial text feature vector; Calculating a first similarity between the image feature vector and the initial text feature vector, and constructing a loss function based on the first similarity; Optimizing the loss function by a gradient descent algorithm, updating the value of the general text embedding vector until the loss function reaches a preset convergence condition, and obtaining an optimized text embedding vector; The optimized text embedding vector is decoded to generate a text prompt word having features similar to those of the target image.
4. The method according to claim 3, characterized in that The step of optimizing the loss function by a gradient descent algorithm, updating the value of the text embedding vector, until the loss function reaches a preset convergence condition, and obtaining the optimized text embedding vector specifically includes: Obtaining an initial value of the universal text embedding vector as a starting point for a gradient descent algorithm; Using the universal text embedding vector as the target text embedding vector, and calculating the loss value of the target text embedding vector according to the loss function; Based on the loss value and the initial value, calculate the gradient of the loss function with respect to the target text embedding vector, and determine the direction and step size of the gradient descent of the target text embedding vector in the space of the loss function; According to the direction and the step size, the value of the target text embedding vector is updated to obtain a new text embedding vector; Determine whether the loss value corresponding to the new text embedding vector meets a preset convergence condition; If the loss value corresponding to the new text embedding vector satisfies the convergence condition, determining that the loss function corresponding to the new text embedding vector satisfies the convergence condition, and using the new text embedding vector as the optimized text embedding vector; If the loss value corresponding to the new text embedding vector does not meet the convergence condition, the new text embedding vector is used as the target text embedding vector, and the step of executing the loss value calculation is returned until the loss value corresponding to the obtained text embedding vector meets the convergence condition, and the text embedding vector that meets the convergence condition is used as the optimized text embedding vector.
5. The method according to claim 1, characterized in that The step of adopting the image-text coding model to vectorize the text prompt word to obtain the text feature vector of the text prompt word specifically includes: Performing word segmentation processing on the text prompt words to obtain a prompt word sequence; Using a pre-trained word vector model, converting the prompt word sequence into a word vector sequence; The text features of the word vector sequence are extracted through the text encoder of the image-text encoding model to obtain the text feature vector of the text prompt word.
6. The method according to claim 1, characterized in that The step of inputting the text feature vector into a preset diffusion model to generate a style image with semantic similarity to the text prompt word specifically includes: Inputting the text feature vector into a preset diffusion model, and generating an initial image through a multi-layer neural network of the diffusion model; According to the semantic features of the text prompt word, a reference image matching the semantic features is obtained from a preset style image library; Based on the reference image, the style features of the initial image are optimized to obtain a style image that is semantically similar to the text prompt word.
7. The method according to claim 1, characterized in that After the step of inputting the text feature vector into a preset diffusion model to generate a style image with semantic similarity to the text prompt word, the method further includes: Using the image-text encoding model, encoding the style image to obtain a style feature vector; Calculating a second similarity between the text feature vector and the style feature vector; Based on the second similarity, an alignment loss function is constructed, and based on the alignment loss function, the image-text encoding model is optimized.
8. An image generating device, characterized in that: include: An acquisition module, used for acquiring a target image; A first encoding module, used to encode the target image using a preset image-text encoding model to obtain an image feature vector of the target image; A word generation module, used for generating text prompt words similar to the target image features through reverse optimization according to the image feature vector; A vectorization module, used to vectorize the text prompt word by using the image-text coding model to obtain a text feature vector of the text prompt word; The image generation module is used to input the text feature vector into a preset diffusion model to generate a style image with semantics similar to the text prompt word.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the image generation method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the image generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Image generation method and device, equipment, storage medium and product
CN120997628A