Image generation method and device based on retrieval enhancement and multi-source feature fusion

By building a retrieval enhancement generation library and combining it with a large visual language model and an object detection model, the problem of insufficient fusion of image and text information is solved, and high-quality and diverse images are generated to meet the personalized needs of users.

CN120807714APending Publication Date: 2025-10-17BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510939083.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing image generation algorithms have shortcomings in fusing images and text information, making it difficult to generate high-quality and diverse images and unable to accurately reflect user expectations and needs.

Method used

Build a retrieval enhancement generation library containing image features and text description features, extract image and text features through the visual large language model and target detection model, perform multi-view feature attention processing and diffusion model fusion, and generate a target image that conforms to the reference image and text description.

Benefits of technology

It achieves the full integration of image and text information, improves the accuracy and quality of generated images, and meets the high-quality and diversified needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807714A_ABST
    Figure CN120807714A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device based on retrieval enhancement and multi-source feature fusion. The method comprises the steps that feature extraction is conducted on a reference image, similarity retrieval is conducted on image features in a retrieval enhancement generation library according to the features of the reference image, and a retrieval image is obtained; encoding the reference image and the retrieval image to obtain corresponding global features, and performing multi-view feature attention processing on the global features to obtain comprehensive image features; splitting the global text description and the local text description into a plurality of segmented descriptions according to a semantic structure, and retrieving and expanding the segmented descriptions by utilizing a retrieval enhancement generation library to obtain comprehensive text description features; and inputting the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and outputting a target image conforming to the reference image and the text description. According to the invention, the image and the text information can be fully fused, and the accuracy and the quality of the generated image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image generation, and particularly relates to an image generation method and device based on retrieval enhancement and multi-source feature fusion. BACKGROUND

[0002] In the digital era, image generation technology has been widely applied in artistic creation, design, virtual reality (VR) and augmented reality (AR) fields. With the increasing demand for high-quality and diversified image generation, how to generate images meeting the requirements according to specific inputs has become an important technical problem. Especially when the user inputs a text description and a reference image, how to generate an image highly matching the input has become the core goal of current technical development.

[0003] In the prior art, most image generation algorithms mainly rely on single image features or text descriptions, resulting in a lack of accurate semantic and visual combination in the image generation process, and it is difficult to meet the user's demand for high-quality and diversified images. In particular, the fusion of image and text information is insufficient, some methods excessively rely on image features and ignore the semantic content of the text description, resulting in the generated image failing to accurately reflect the user's expectations. While another part of the method excessively focuses on the text description and ignores the visual style and details of the image itself, the generated image lacks uniqueness and appeal.

[0004] In addition, existing algorithms often generate images only through a single image or text description provided by the user, and it is difficult to extract and fuse multi-dimensional image and text features. This way limits the accuracy and diversity of image generation, and the generated image may not meet the user's personalized needs or lack scene richness. Due to incomplete feature extraction, the generated image often lacks sufficient details and accuracy, and the overall scene construction is monotonous, making it difficult to present complex image content. SUMMARY

[0005] Therefore, the embodiments of the present application provide an image generation method and device based on retrieval enhancement and multi-source feature fusion to solve the problems of insufficient fusion of image and text information, lack of accuracy of generated images and poor image quality in the prior art.

[0006] In a first aspect, the embodiment of the present application provides an image generation method based on retrieval enhancement and multi-source feature fusion, comprising the following steps: constructing a retrieval enhancement generation library containing image features and text description features; performing feature extraction on an input reference image, and performing similarity retrieval on the image features in the retrieval enhancement generation library according to the features of the reference image to obtain a retrieval image; encoding the reference image and the retrieval image to obtain corresponding global features, and performing multi-view feature attention processing on the global features to obtain comprehensive image features; generating a global text description of the reference image by using a visual large language model, and identifying local objects in the reference image by using a target detection model to generate a corresponding local text description for each local object; splitting the global text description and the local text description into a plurality of segmented descriptions according to a semantic structure, and performing retrieval and expansion on the segmented descriptions by using the retrieval enhancement generation library to obtain comprehensive text description features; and inputting the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and outputting a target image conforming to the reference image and the text description.

[0007] In a second aspect, the embodiment of the present application provides an image generation device based on retrieval enhancement and multi-source feature fusion, comprising: a construction module configured to construct a retrieval enhancement generation library containing image features and text description features; an extraction module configured to perform feature extraction on an input reference image, and perform similarity retrieval on the image features in the retrieval enhancement generation library according to the features of the reference image to obtain a retrieval image; an encoding module configured to encode the reference image and the retrieval image to obtain corresponding global features, and perform multi-view feature attention processing on the global features to obtain comprehensive image features; a generation module configured to generate a global text description of the reference image by using a visual large language model, and identify local objects in the reference image by using a target detection model to generate a corresponding local text description for each local object; an expansion module configured to split the global text description and the local text description into a plurality of segmented descriptions according to a semantic structure, and perform retrieval and expansion on the segmented descriptions by using the retrieval enhancement generation library to obtain comprehensive text description features; and an output module configured to input the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and output a target image conforming to the reference image and the text description.

[0008] In a third aspect, the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above method.

[0009] The above at least one technical scheme adopted by the embodiment of the present application can achieve the following beneficial effects:

[0010] The retrieval enhancement generation library is constructed by constructing an image feature and a text description feature; feature extraction is performed on the input reference image, and similarity retrieval is performed on the image features in the retrieval enhancement generation library according to the features of the reference image to obtain a retrieval image; the reference image and the retrieval image are encoded to obtain corresponding global features, and the global features are subjected to multi-view feature attention processing to obtain comprehensive image features; a global text description of the reference image is generated by using a visual large language model, and a local object in the reference image is identified by using a target detection model, and a corresponding local text description is generated for each local object; the global text description and the local text description are split into a plurality of segmented descriptions according to a semantic structure, and the segmented descriptions are retrieved and expanded by using the retrieval enhancement generation library to obtain comprehensive text description features; the comprehensive text description features and the comprehensive image features are input into a diffusion model for feature fusion, and a target image conforming to the reference image and the text description is output. The present application can fully fuse image and text information, and improve the accuracy and image quality of the generated image. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0012] Figure 1 is the overall implementation flowchart of the image generation algorithm based on retrieval enhancement and multi-source feature fusion provided by the embodiments of the present application;

[0013] Figure 2 is the flowchart of the image generation method based on retrieval enhancement and multi-source feature fusion provided by the embodiments of the present application;

[0014] Figure 3 is the structural schematic diagram of the image generation device based on retrieval enhancement and multi-source feature fusion provided by the embodiments of the present application;

[0015] Figure 4 is the structural schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0016] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application, but it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.

[0017] In the digital age, the demand for image generation is growing, whether in the field of artistic creation, design, or in the application of virtual reality, augmented reality and other technologies, it is necessary to generate high-quality images that meet specific requirements based on specific inputs. This algorithm aims to solve a method that can combine image and text information for accurate image generation. In many practical application scenarios, users often want to generate specific images by inputting text descriptions and reference images, but existing methods have shortcomings in integrating image and text information, making it difficult to generate images that meet user expectations. Traditional image generation methods may generate relatively single images, lack of details and personalization, and cannot meet the user's demand for high-quality and diversified images. Therefore, a retrieval-enhanced image generation method is proposed to meet the demand for high-quality and diversified images.

[0018] Existing algorithms perform poorly in many aspects. First, the integration of image and text information is insufficient. In most existing algorithms, image and text cannot be closely combined. Either over-reliance on image features ignores text descriptions, resulting in generated images that do not match user expectations and cannot accurately convey specific concepts and emotions; or too much emphasis on text information ignores the style and content of the image, making the generated image lack visual appeal and uniqueness.

[0019] Second, existing algorithms usually only use user-provided images or text descriptions for image generation, making it difficult to quickly and accurately extract high-quality image or text features. This makes it difficult to have enough reference in the image generation process, greatly affecting the quality and efficiency of image generation. For example, generating images based on a single text description may generate inaccurate images due to biased text understanding; relying solely on a single reference image may be limited by the image itself and unable to meet the user's diverse needs. Finally, due to incomplete feature extraction, the accuracy and richness of the generated image are not enough, and the overall scene construction also lacks richness, and the picture may appear monotonous. At the same time, the images generated by existing algorithms lack personalization and are difficult to meet the user's pursuit of unique images. Therefore, a new image generation algorithm based on retrieval enhancement and multi-source feature fusion is needed to solve the above problems.

[0020] In view of the problems in the prior art, the present application proposes an image generation algorithm based on retrieval enhancement and multi-source feature fusion. The algorithm aims to build an integrated retrieval enhancement generation (RAG) library containing high-quality images and corresponding text description information. Through this integrated library, high-quality pictures that match the input reference image and text description can be generated more effectively. The algorithm integrates the relevant features of images and text through a series of key steps, and uses an advanced diffusion model to achieve high-quality image generation.

[0021] Before the technical solutions of the present application are described in detail, the overall implementation process of the image generation algorithm based on retrieval enhancement and multi-source feature fusion of the present application will be described in conjunction with the accompanying drawings. Figure 1 is the overall implementation process schematic diagram of the image generation algorithm based on retrieval enhancement and multi-source feature fusion provided by the embodiments of the present application. As shown in Figure 1 the specific implementation process of the image generation algorithm based on retrieval enhancement and multi-source feature fusion can include the following contents:

[0022] First, the input reference image is processed by downsampling to generate multiple reference images. These images will be sent to the image feature extraction module to extract the feature representation of the image using a deep learning-based feature extraction method.

[0023] Further, through the constructed retrieval enhancement generation library, the features of the reference image will be matched with the image features in the library, and the RAG (retrieval enhancement generation) module will be used for image retrieval to obtain images similar to the reference image.

[0024] Further, the features of the reference image and the retrieved image are encoded to obtain global features and local features of the image. These image features will be processed by the VIT (Vision Transformer) image encoder to obtain comprehensive image features. In addition, the text description is processed by a large language model (LLM), and the CoT (Chain of Thought) method is used to split the text into multiple segmented descriptions to enhance its semantic structure.

[0025] Further, through the target detection model (such as Grounding DINO), the local objects in the reference image are identified and corresponding local image descriptions are generated. These local descriptions will be further processed together with the global text description.

[0026] Further, the global text description and the local text description are retrieved and expanded by the RAG library to obtain comprehensive text description features. The comprehensive image features and the comprehensive text description features are then fused by a multi-view self-attention mechanism, a linear layer and an MLP (multi-layer perceptron) to obtain final fusion features.

[0027] Further, the fused features are input into a diffusion model for image generation. The diffusion model fuses the image features and the text description features through a step-by-step iterative process, and finally outputs a target image that meets the reference image and the text description.

[0028] The technical solutions of the present application will be described in detail below in combination with the drawings and specific embodiments.

[0029] Figure 2 is a flowchart of an image generation method based on retrieval enhancement and multi-source feature fusion provided by an embodiment of the present application. As shown in Figure 2 , the image generation method based on retrieval enhancement and multi-source feature fusion can specifically include:

[0030] S201, constructing a retrieval enhancement generation library containing image features and text description features;

[0031] S202, performing feature extraction on the input reference image, and performing similarity retrieval on the image features in the retrieval enhancement generation library according to the features of the reference image to obtain retrieval images;

[0032] S203, encoding the reference image and the retrieval images to obtain corresponding global features, and performing multi-view feature attention processing on the global features to obtain comprehensive image features;

[0033] S204, generating a global text description of the reference image using a visual large language model, and identifying local objects in the reference image using a target detection model, and generating a corresponding local text description for each local object;

[0034] S205, splitting the global text description and the local text description into a plurality of segmented descriptions according to a semantic structure, and retrieving and expanding the segmented descriptions using the retrieval enhancement generation library to obtain comprehensive text description features;

[0035] S206, inputting the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and outputting a target image that meets the reference image and the text description.

[0036] In some embodiments, constructing a retrieval enhancement generation library containing image features and text description features includes:

[0037] Collecting images and corresponding text descriptions, and pre-processing the images and the corresponding text descriptions;

[0038] The pre-processed image is feature extracted to obtain an image feature vector, and the image feature vector and the corresponding text description are stored in a retrieval enhancement generation library;

[0039] The pre-processed text description is initially encoded to obtain an initial feature representation, the initial feature representation is input into a text encoder for feature extraction to obtain a text feature vector, and the text feature vector and the corresponding text description and image are stored in a retrieval enhancement generation library.

[0040] Specifically, in the present embodiment, first, a variety of high-quality images and their corresponding text descriptions need to be collected extensively. The data sources can be professional image databases, art work websites, literature work illustrations, etc. In order to ensure the effectiveness and efficiency of subsequent processing, the collected images are first pre-processed.

[0041] The specific steps are as follows:

[0042] First, the collected images are subjected to size reduction operation, and common image scaling algorithms such as bilinear interpolation are used to ensure that the key information of the image is preserved as much as possible while reducing the size of the image. This step helps to reduce the computational complexity of the image, while ensuring that the main features of the image are not lost.

[0043] Then, the collected text descriptions are cleaned and standardized. The specific operations include removing noise, error information, unifying formats, such as converting text to lowercase letters, removing extra spaces and punctuation errors, etc. In addition, the description structure is also unified to ensure the consistency of each text description in format and content for subsequent processing.

[0044] Further, the pre-processed image will be input into a deep convolutional neural network (CNN) for feature extraction. In the present embodiment, a ResNet50 pre-trained model is selected for feature extraction, and the specific process is as follows:

[0045] First, the pre-processed image is input into the ResNet50 model, and the image is processed by multiple residual blocks. Each residual block contains multiple convolutional layers and a shortcut connection, allowing the network to learn deeper features of the image.

[0046] Next, after the image is subjected to convolution and pooling operations, the multi-level features of the image are mapped to a fixed-length embedding vector through a fully connected layer. The embedding vector of each image can fully represent the visual features of the image, including the structure, color, texture, etc. of the image.

[0047] Then, the extracted image embedding is stored together with the corresponding text description in the retrieval enhancement generation library for subsequent quick retrieval and image generation.

[0048] Further, for the text description, the embodiment adopts a large language model (LLM) and a CLIP text encoder for processing. The specific process is as follows:

[0049] First, each text description is input into the LLM for preliminary coding to generate an initial feature representation of the text. The LLM can capture the basic semantic information in the text and provide a high-dimensional feature space representation.

[0050] Next, the preliminarily coded text features are input into the CLIP text encoder. The CLIP uses its pre-trained model parameters to deeply encode the input text description. The CLIP can learn and extract semantic features of the text description, closely linking the relationship between the text and the image features, thereby obtaining a text embedding that can accurately represent the core semantics and feature information of the text description.

[0051] Then, the generated text embedding is stored together with the corresponding image and text description in the retrieval enhancement generation library to provide more accurate text guidance information for subsequent image generation.

[0052] In the embodiment, through the above image and text feature extraction process, a retrieval enhancement generation library containing image features and text description features is constructed. The main functions of the library are:

[0053] In the image generation process, based on the reference image and text description input by the user, the retrieval enhancement generation library can retrieve similar images and related text descriptions to the target image features through a similarity matching algorithm (such as cosine similarity), providing high-quality reference for subsequent image generation.

[0054] The library not only stores the basic feature representations of images and texts, but also can retrieve and expand the text description as needed to assist in generating image text guidance information that meets the requirements.

[0055] The embodiment details how to construct an efficient retrieval enhancement generation library through feature extraction and storage of images and text descriptions. With the support of the library, more accurate image generation can be achieved, and the diversity and personalization level of image generation are also improved. The technical solution of the embodiment provides a solid foundation for the image generation method based on retrieval enhancement and multi-source feature fusion.

[0056] In some embodiments, according to the features of the reference image, the image features in the retrieval enhancement generation library are retrieved for similarity, and a retrieval image is obtained, including:

[0057] According to the characteristics of the reference image, the image features in the retrieval enhancement generation library are retrieved using a similarity measurement algorithm;

[0058] According to the similarity between the characteristics of the reference image and the image features in the retrieval enhancement generation library, a plurality of images are screened out as retrieval images according to the similarity ranking.

[0059] Specifically, the embodiment mainly describes how to retrieve the similarity of image features in the constructed retrieval enhancement generation library according to the characteristics of the reference image, and screen out the most similar images as retrieval images. This process involves reference image feature extraction, similarity calculation and image screening, ensuring the accuracy and diversity of subsequent image generation.

[0060] First, for the reference image corresponding to the image to be generated, the same preprocessing operation as when constructing the image RAG library is performed. The specific steps are as follows:

[0061] The size reduction operation is performed on the reference image, and the same image scaling algorithm as when constructing the RAG library is used, such as bilinear interpolation. This operation ensures that the size and features of the reference image remain consistent during processing, so as to effectively match the images in the retrieval enhancement generation library.

[0062] The same feature extraction method as when constructing the library (such as a pre-trained convolutional neural network such as ResNet50 or VGG) is used to extract features from the reference image, generating an embedding vector for the image. This embedding vector can efficiently capture the visual features of the image, including color, texture, shape, etc., providing basic data for the subsequent retrieval process.

[0063] Further, in the image RAG library, pre-extracted image features are stored. Using the embedding vector of the reference image, similarity retrieval is performed in the library, and the specific steps are as follows:

[0064] According to the feature embedding of the reference image, a suitable similarity measurement algorithm is selected, such as cosine similarity, Euclidean distance, etc., to compare the image features stored in the library. These algorithms can calculate the similarity between the reference image and each image feature in the library, and determine their distance in the feature space.

[0065] According to the calculated similarity, the similarity values of all images in the library are sorted, for example, the top 3 images most similar to the reference image are screened out. These retrieved images have high similarity in visual content, style, structure, etc. with the reference image, and can provide effective reference for subsequent image generation.

[0066] Further, through the similarity retrieval and ranking process described above, 3 images most similar to the reference image are obtained from the image RAG library. These images will serve as input for subsequent image generation, to ensure that the generated target image is consistent with the reference image in terms of visual style and content, while integrating the details and diversity of the images.

[0067] These retrieved images will work together with the features of the reference image in subsequent steps, through further feature processing and fusion, to ultimately generate a target image that meets the user's needs.

[0068] This embodiment details how to perform similarity retrieval in the image RAG library based on the features of the reference image. By extracting features from the reference image and performing similarity calculation and screening with the images in the library, the most similar images are obtained, providing high-quality reference for subsequent image generation. This process ensures the accuracy, richness, and personalization of image generation, and is consistent with the construction process of the retrieval-enhanced generation library, ensuring the efficiency and consistency of the entire system.

[0069] In some embodiments, the reference image and the retrieved image are encoded to obtain corresponding global features, including:

[0070] The retrieved image and the reference image are divided into multiple image blocks using a visual transformer image encoder;

[0071] The image blocks are linearly embedded and position encoding information is added, and a multi-head self-attention mechanism and a multi-layer encoder are used to process the image blocks, output image feature representations, and generate corresponding classification labels;

[0072] A linear layer is used to transform the features of the classification labels, so that each classification label is mapped to a feature space dimension;

[0073] The classification labels after feature transformation are concatenated in order to form a comprehensive classification label feature representation, wherein the comprehensive classification label feature representation integrates the global features of the retrieved image and the reference image.

[0074] Specifically, this embodiment mainly describes how to use a visual transformer (VIT) image encoder to encode the reference image and the retrieved image to obtain their global features, and through a series of image processing and feature transformation, finally form a comprehensive classification label feature representation. The process includes image division, embedding, feature processing, and concatenation of global features.

[0075] First, for the input reference image and the retrieved image, a VIT (Vision Transformer) image encoder is used to process them. The specific steps are as follows:

[0076] The search image and the reference image are divided into multiple image patches, each containing local information of the image. Through this division, VIT can better capture local and global features in the image.

[0077] Linear embedding is performed on each image patch, converting the pixel information of the image patch into a fixed-dimensional feature vector while adding position encoding information. Position encoding can preserve the spatial relationship between image patches and provide position information for subsequent self-attention mechanisms.

[0078] Through multi-head self-attention mechanisms and multi-layer Transformer encoder blocks, VIT can capture the dependency and feature association between image patches from multiple perspectives, ultimately generating image feature representations that contain rich semantic and visual features. These feature representations include both low-level detailed information and high-level semantic information.

[0079] Through the encoder output of VIT, a corresponding CLS token (classification token) is also generated, which is used to summarize the global feature information of the image. The CLS token is an important part of VIT, which can provide global representation of the image and serve as input features for subsequent image generation or processing.

[0080] Further, for the 4 CLS tokens from the reference image and the search image (3 from the search image and 1 from the original image), further feature processing is performed. The specific process is as follows:

[0081] Each CLS token is transformed by an independent linear layer to map each CLS token to a feature space dimension suitable for subsequent fusion processing. This step ensures that the features of each CLS token can be unified and optimized for efficient fusion in subsequent steps.

[0082] The 4 CLS tokens after linear transformation are concatenated in a certain order to form a comprehensive CLS token feature representation. This feature representation contains global feature information from 3 search images and 1 reference image, integrating diverse image features. Through the concatenation operation, these global feature information is merged into a comprehensive feature representation, providing unified feature support for subsequent image generation and feature fusion.

[0083] The embodiment details how to process the reference image and the search image through the VIT image encoder, generate the global features of each image, and effectively fuse these features through the CLS token. Through linear embedding, position encoding, and multi-head self-attention mechanism, VIT can extract high-level semantic information of the image. Finally, through feature transformation and splicing, the four CLS tokens are merged into a comprehensive global feature representation, providing accurate input data for the subsequent image generation process. This process ensures efficient extraction and accurate fusion of image features, providing a solid foundation for subsequent image generation.

[0084] In some embodiments, the global features are subjected to multi-view feature attention processing to obtain comprehensive image features, including:

[0085] The multi-view feature attention mechanism is used to capture the relationship between global features from multiple perspectives, calculate the attention weights of global features under different perspectives, and perform deep learning according to the attention weights;

[0086] The feature results learned by the multi-view feature attention mechanism are input into a multi-layer perceptron, which uses the multi-layer perceptron to perform nonlinear transformation on the input features;

[0087] The features output by the multi-layer perceptron are spliced with the classification labels after feature transformation to obtain comprehensive image features.

[0088] Specifically, first, the four images (reference image and three search images) output from the VIT image encoder are subjected to image feature extraction, and the obtained image features are input into a multi-view self-attention (Multi-view self-attention) module. This module is used to capture the relationship between image features from multiple perspectives, and the specific process is as follows:

[0089] The multi-view self-attention module further excavates the potential relationship between image features by calculating the attention weights of image features under different perspectives. For example, images under different perspectives may exhibit similar styles, color distributions, or object structures. Through this mechanism, these potential similarities can be captured and the features can be weighted.

[0090] The module calculates the attention weights of each image feature under different perspectives through the self-attention mechanism, and performs deep learning on the image features according to these weights. This step can effectively extract and integrate the effective information in the image features, making the global features of the image more efficiently represented.

[0091] Further, after processing by the multi-view feature attention module, the image feature results are input into a multi-layer perceptron (MLP). The MLP processes the features in the following ways:

[0092] The MLP performs a non-linear transformation on the input image features through its multi-layer neural structure and activation function. This process helps to enhance the expressiveness of the features, uncover deep-level information in the image features, and improve the semantic understanding ability of the image.

[0093] Through the processing of the MLP, the expressiveness of the image features is enhanced, enabling them to more accurately convey the content and semantic information of the image during subsequent fusion processes.

[0094] Finally, the features output by the MLP are concatenated with the CLS token (classification token) features processed by the linear layer, resulting in comprehensive image features. The specific steps are as follows:

[0095] The CLS token represents the global image features and contains the overall information of the image. In this step, the image features processed by the MLP are concatenated with the CLS token features in a specific order to obtain a comprehensive image feature representation.

[0096] The final comprehensive image features will contain rich global-to-local feature information from the three search images and the original image. This comprehensive image feature will support the subsequent fusion with the text features, ensuring the accuracy and diversity of the generated images.

[0097] This embodiment processes image features through a multi-view feature attention mechanism, capturing relationships between images from multiple perspectives and combining MLP for feature transformation, ultimately generating a comprehensive image feature. This feature combines global and local features of reference images and search images, providing efficient feature support for subsequent image generation. This process effectively improves the accuracy and diversity of image generation, ensuring the quality and personalization of generated images.

[0098] In some embodiments, a target detection model is used to identify local objects in the reference image, and a corresponding local text description is generated for each local object, including:

[0099] A target detection model is used to identify local objects in the reference image, and a corresponding bounding box is generated for each local object. According to the bounding box, the local objects are extracted from the reference image to generate sub-images.

[0100] The sub-images are sequentially input into the visual large language model, and according to the visual features of the local objects in the sub-images, a local text description corresponding to each local object is generated.

[0101] Specifically, in this embodiment, the reference image is first processed using the Grounding DINO model. Grounding DINO is a deep learning-based object detection and localization model that can identify individual objects in an image and generate corresponding bounding boxes for each object. The specific operation is as follows:

[0102] The Grounding DINO model analyzes the reference image and identifies individual local objects within it. Each object is labeled by a bounding box, which accurately represents the specific location of each object in the image in terms of size and position.

[0103] According to each generated bounding box, a sub-image containing a single object is extracted from the reference image. These sub-images will be used for further analysis to generate detailed text descriptions.

[0104] Further, after obtaining the sub-images, the next step is to generate a text description corresponding to each sub-image using a visual large language model (VLM). VLM can handle cross-modal relationships between images and text and has strong visual understanding capabilities. The specific steps are as follows:

[0105] Each extracted sub-image is sequentially input into the visual large language model (VLM). VLM can understand the visual features in the image and generate corresponding text descriptions.

[0106] VLM analyzes the local objects in each sub-image and generates detailed text descriptions based on their visual features. The descriptions focus on the specific attributes, states, colors, shapes, etc. of the objects, providing a comprehensive understanding of each object in the image.

[0107] The local text descriptions generated by VLM can accurately reflect the details of each object in the image, supplementing the features of the image from a microscopic perspective. These local text descriptions can enrich the overall text representation of the reference image and provide more detailed information for subsequent image generation.

[0108] This embodiment not only generates a global text description using VLM, but also generates precise local text descriptions for each local object in the image. The global text description provides a macroscopic understanding of the entire image, while the local text description supplements the detailed information of the image from a microscopic perspective. Through the combination of these two parts, the image content can be better described, and sufficient text support can be provided for subsequent generation of target images that meet the requirements of reference images and text descriptions.

[0109] In some embodiments, the global text description and the local text description are split into several segmented descriptions according to the semantic structure, and the segmented descriptions are retrieved and expanded using a retrieval enhancement generation library to obtain comprehensive text description features, including:

[0110] In a large language model, the global text description and the local text description are split into several segment descriptions according to logical order and semantic association using the thought chain processing method;

[0111] Each segment description is input into a retrieval enhancement generation library, and a similarity matching algorithm is used to retrieve relevant text description features;

[0112] The retrieved relevant text description features are input into the large language model, and the segment descriptions are expanded according to the relevant text description features to obtain the expansion content of the segment descriptions, and the comprehensive text description features are generated according to the segment descriptions and the expansion content.

[0113] Specifically, first, the input global text description and local text description are input into a large language model (LLM) for processing, and the specific steps are as follows:

[0114] LLM processes through the thought chain (Chain of Thought, CoT) method. CoT aims to simulate human thinking logic and split complex text descriptions according to logical order and semantic association, and decompose long text into multiple segment descriptions. For example, for the text "There is a beautiful mountain peak by the lake, with green trees on the mountain, and the lake surface is rippling", CoT may split it into the following segment descriptions:

[0115] Description of the location and surrounding environment of the mountain peak

[0116] Description of the vegetation on the mountain peak

[0117] Description of the state of the lake surface

[0118] Through CoT, the generated segment descriptions are more organized and can better reflect the specific content in the text, making the text clearer and easier for subsequent retrieval and processing.

[0119] Further, after CoT decomposition, each segment description will be retrieved and expanded, and the specific steps are as follows:

[0120] Each segment description is input into the constructed retrieval enhancement generation library (RAG library) for retrieval. The RAG library contains image and text feature data and can perform similarity retrieval based on the input segment description. By using a similarity matching algorithm (such as cosine similarity or Euclidean distance), the most similar text description features to the segment description are retrieved from the library.

[0121] The retrieved relevant text descriptions are input back into the LLM, which expands the original segmented descriptions based on these reference information. By referring to the retrieved texts, the LLM can supplement each segmented description with more details, adjectives, and other information, further enriching the text content and making each segmented description more detailed and specific.

[0122] The expanded segmented descriptions are merged into a complete comprehensive text description. This description supplements the overall understanding of the reference image from both global and local perspectives, enabling more accurate communication of image generation requirements and providing more detailed text guidance for the image generation process.

[0123] This embodiment details how to use the CoT to split the global and local text descriptions into multiple segmented descriptions and retrieve and expand these segmented descriptions through the RAG library. Through multi-channel retrieval and LLM expansion, the details and accuracy of the text description can be significantly improved, resulting in a more detailed and comprehensive text description feature. This method can provide more accurate and rich text guidance, providing sufficient support for the subsequent image generation process and ensuring that the generated image meets the user's expectations in terms of content and style.

[0124] In some embodiments, this embodiment will also detail the splicing of the image global text description, image local text description, and text retrieval description, and further extracting the integrated text description features through the LLM and CLIP text encoder. This process involves the integration, feature extraction, and encoding of the text description, ultimately generating a comprehensive text description feature that accurately represents the image generation requirements.

[0125] First, the previously obtained image global text description, image local text description, and text retrieval description are integrated. The specific steps are as follows:

[0126] The global text description provides semantic information about the overall image, including the main scene, main subject, and other key content, and can reflect the image feature requirements from a macro perspective.

[0127] The local text description refines the specific objects in the image, describing the features, states, and other information of each local object in the image. These descriptions provide a supplement to the image details and enhance the accurate understanding of the image.

[0128] Through the relevant text descriptions obtained from the RAG, the content of the global and local descriptions is further enriched. These retrieval descriptions expand the original descriptions, supplementing more details and information to ensure the comprehensiveness of the text content.

[0129] In this stage, the three types of text descriptions are spliced in a certain order to form a complete comprehensive text description. This description integrates information from multiple dimensions, from macro to micro, from original input to search expansion, and can comprehensively and meticulously reflect the characteristics of the generated image.

[0130] Further, after obtaining the comprehensive text description, semantic analysis and feature extraction are performed on it.

[0131] The specific steps are as follows:

[0132] First, the spliced comprehensive text description is input into a large language model (LLM). LLM has strong language understanding and coding capabilities, and can perform deep semantic analysis on the input text. Through semantic analysis, LLM can understand the detailed information, relevance and context in the text, and generate integrated text semantic representation.

[0133] Next, the integrated text semantic representation output by LLM is input into the CLIP text encoder. The CLIP (Contrastive Language-Image Pretraining) text encoder is based on pre-trained model parameters and has strong text feature extraction capability, which can extract high-dimensional semantic information from text. Through the CLIP encoder, the semantic information of the comprehensive text description is further deeply coded to generate the feature representation of the text.

[0134] Then, the text feature representation output by the CLIP text encoder can accurately represent the core semantics and feature information of the entire text description. These features provide comprehensive and accurate text guidance for image generation requirements and will be fused with image features in subsequent steps.

[0135] This embodiment integrates the image global text description, local text description and text search description into a comprehensive text description. Then, the large language model (LLM) and CLIP text encoder are used to perform deep semantic analysis and feature extraction on the description to generate a comprehensive text description feature that accurately represents the image generation requirements. This feature will be used for the fusion of text and image features in the subsequent image generation process to ensure that the generated image accurately reflects the user's expectations.

[0136] In some embodiments, the comprehensive text description feature and the comprehensive image feature are input into a diffusion model for feature fusion, and a target image that meets the reference image and text description is output, including:

[0137] Generate an initial noise image as the input of the diffusion model, and use the initial noise image as the starting state of the target image;

[0138] In the diffusion model, multiple cross-attention modules are included. In the previous cross-attention module, the integrated text description features are fused with the input features of the current cross-attention module.

[0139] In the next cross-attention module, the fused features output by the previous cross-attention module are fused with the integrated image features.

[0140] After the diffusion model processes through each cross-attention module in turn, the target image that meets the reference image and text description is output.

[0141] Specifically, first, a random noise image is generated as the initial input of the diffusion model. The noise image is the starting state of the target image generation, and through the gradual fusion of text description features and image features, the noise image will gradually transform into a target image that meets the reference image and text description.

[0142] The diffusion model starts from a completely random noise image, which has no specific information as the initial input of the generation process. With the gradual fusion of features, the noise image gradually shows clear image structure.

[0143] The diffusion model adopts the Unet structure of Stable Diffusion, which is composed of multiple sub-modules, each containing two cross-attention mechanisms for fusing integrated text description features and integrated image features, respectively. The specific process is as follows:

[0144] First, the integrated text description features are fused with the input features of the current sub-module of the diffusion model. In this stage, the integrated text description features are used as key and value vectors, while the input features of the current sub-module are used as query vectors. Through cross-attention calculation, the text features and the current input features are preliminarily fused, making the initial features of the generated image gradually develop towards the direction of meeting the text description.

[0145] Next, the fused features of the previous cross-attention module are fused with the integrated image features. In this stage, the output of the previous cross-attention module is used as the query vector, while the integrated image features are used as the key and value vectors. Again through cross-attention calculation, the image features are further integrated, making the generated image meet the text requirements while also referencing the style, content, and other characteristics of the retrieved image.

[0146] The above two cross-attention modules are iterated multiple times in the diffusion model, and each iteration will gradually fuse the image features and make the generated image gradually clear and perfect. Each iteration will further optimize the image quality until the target image that meets the reference image and text description is generated.

[0147] Through the above iterative process, the diffusion model gradually transforms the initial noise map into the final target image. Each feature fusion in the iterative process further optimizes and enhances the visual content and text consistency of the image. Finally, after multiple iterations, the output image meets the requirements and serves as the final algorithm result.

[0148] This embodiment gradually fuses the comprehensive text description features and the comprehensive image features through multiple cross-attention modules in the diffusion model, and gradually transforms them into the target image that meets the reference image and the text description through iterative generation. Each iteration optimizes the content and details of the image by fusing features from different sources, ensuring that the generated image not only meets the requirements of the text description, but also retains the style and features of the reference image. Through this efficient feature fusion and iterative process, the final generated image has been significantly improved in accuracy, richness, and diversity.

[0149] According to the technical solutions provided in the above embodiments, the present application has at least the following advantages:

[0150] The present application proposes an image generation algorithm based on constructing a RAG library to enhance feature retrieval. An integrated RAG library that combines high-quality images and text descriptions is constructed. By constructing a RAG library that combines high-quality images and text descriptions, during the image and text retrieval stage, the library can quickly find similar reference images and corresponding text descriptions to the target image. This provides more comprehensive and accurate reference information for image and text generation, enhancing the features of the original image and text.

[0151] The present application proposes a deep fusion of multi-source features to improve image generation quality. With the help of VIT image encoder and multi-view self-attention module, rich global and local feature information is extracted from the retrieved images and the original image, providing a more comprehensive reference for image generation. On the other hand, the VLM generates global and local text descriptions of the image, and combines the target detection and positioning capabilities of the Grounding DINO model to enrich the text representation of the reference image from both macro and micro perspectives. At the same time, LLM CoT and multi-path retrieval and expansion methods are used to segment and expand the input text description, further improving the accuracy and richness of image generation, so that the generated image not only meets the text description requirements, but also references the style and content of similar images, thereby significantly improving the quality of image generation.

[0152] The following is an embodiment of the device of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0153] Figure 3FIG. 1 is a structural schematic diagram of an image generation device based on retrieval enhancement and multi-source feature fusion provided by an embodiment of the present application. As shown in FIG. 1, the image generation device based on retrieval enhancement and multi-source feature fusion comprises: Figure 3

[0154] a construction module 301 configured to construct a retrieval enhancement generation library containing image features and text description features;

[0155] an extraction module 302 configured to perform feature extraction on an input reference image, and perform similarity retrieval on the image features in the retrieval enhancement generation library according to the features of the reference image to obtain retrieval images;

[0156] an encoding module 303 configured to encode the reference image and the retrieval images to obtain corresponding global features, and perform multi-view feature attention processing on the global features to obtain comprehensive image features;

[0157] a generation module 304 configured to generate a global text description of the reference image by using a visual large language model, and identify local objects in the reference image by using a target detection model, and generate a corresponding local text description for each local object;

[0158] an expansion module 305 configured to split the global text description and the local text description into a plurality of segmented descriptions according to a semantic structure, and perform retrieval and expansion on the segmented descriptions by using the retrieval enhancement generation library to obtain comprehensive text description features;

[0159] an output module 306 configured to input the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and output a target image conforming to the reference image and the text description.

[0160] In some embodiments, Figure 3 the construction module 301 collects images and corresponding text descriptions, pre-processes the images and the corresponding text descriptions, performs feature extraction on the pre-processed images to obtain image feature vectors, stores the image feature vectors and the corresponding text descriptions in a retrieval enhancement generation library, and performs preliminary encoding on the pre-processed text descriptions to obtain initial feature representations, inputs the initial feature representations into a text encoder for feature extraction to obtain text feature vectors, and stores the text feature vectors and the corresponding text descriptions and images in the retrieval enhancement generation library.

[0161] In some embodiments, Figure 3 the extraction module 302 performs retrieval on the image features in the retrieval enhancement generation library according to the features of the reference image by using a similarity measurement algorithm, and according to the similarity between the features of the reference image and the image features in the retrieval enhancement generation library, filters a plurality of images as retrieval images according to similarity sorting.

[0162] ​In some embodiments, Figure 3 The encoding module 303 divides the retrieval image and the reference image into a plurality of image blocks by using a visual transformer image encoder; performs linear embedding on the image blocks and adds position encoding information; processes the image blocks by using a multi-head self-attention mechanism and a multi-layer encoder, outputs image feature representations, and generates corresponding classification labels; performs feature transformation on the classification labels by using a linear layer so as to map each classification label to a feature space dimension; and splices the classification labels after the feature transformation in sequence to form a comprehensive classification label feature representation, wherein the comprehensive classification label feature representation fuses global features of the retrieval image and the reference image.

[0163] In some embodiments, Figure 3 The encoding module 303 captures the relationship between global features from multiple perspectives by using a multi-perspective feature attention mechanism, calculates attention weights of the global features under different perspectives, and performs deep learning according to the attention weights; inputs feature results learned by the multi-perspective feature attention mechanism into a multi-layer perceptron, and performs nonlinear transformation on the input features by using the multi-layer perceptron; and splices the features output by the multi-layer perceptron and the classification labels after the feature transformation to obtain comprehensive image features.

[0164] In some embodiments, Figure 3 The generation module 304 identifies local objects in the reference image by using a target detection model, generates corresponding bounding boxes for each local object, extracts the local objects from the reference image according to the bounding boxes to generate sub-images, and inputs the sub-images into a visual large language model in sequence to generate local text descriptions corresponding to each local object according to visual features of the local objects in the sub-images.

[0165] In some embodiments, Figure 3 The expansion module 305 splits the global text description and the local text description into a plurality of segmented descriptions according to a logical order and semantic association by using a thought chain processing manner in the large language model; inputs the segmented descriptions into a retrieval enhancement generation library respectively, retrieves relevant text description features by using a similarity matching algorithm; inputs the retrieved relevant text description features into the large language model, expands the segmented descriptions according to the relevant text description features to obtain expansion contents of the segmented descriptions, and generates comprehensive text description features according to the segmented descriptions and the expansion contents.

[0166] In some embodiments, Figure 3The output module 306 generates an initial noise map as an input of the diffusion model, and takes the initial noise map as a starting state of the target image; the diffusion model comprises a plurality of cross-attention modules, in a previous cross-attention module, the integrated text description feature is fused with the input feature of the current cross-attention module; in a next cross-attention module, the fused feature output by the previous cross-attention module is fused with the integrated image feature; after the processing of each cross-attention module in the diffusion model in turn, a target image conforming to the reference image and the text description is output.

[0167] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0168] Figure 4 is a structural schematic diagram of an electronic device 4 provided by the embodiments of the present application. As shown in the figure, the electronic device 4 of the embodiments comprises a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. The processor 401 implements the steps in each of the above method embodiments when executing the computer program 403. Alternatively, the processor 401 implements the functions of each module / unit in each of the above device embodiments when executing the computer program 403. Figure 4

[0169] By way of example, the computer program 403 can be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 403 in the electronic device 4.

[0170] The electronic device 4 can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The electronic device 4 can include but is not limited to the processor 401 and the memory 402. Those skilled in the art can understand that the electronic device 4 can include more or fewer components than those shown, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc. Figure 4 The electronic device 4 is only an example and does not constitute a limitation on the electronic device 4, and can include more or fewer components than those shown, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc.

[0171] ​The processor 401 can be a central processing unit (CPU), or other general purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general purpose processor can be a microprocessor or the processor can be any conventional processor.

[0172] The memory 402 can be an internal storage unit of the electronic device 4, for example, a hard disk or a memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 can include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 can also be used to temporarily store data that has been output or will be output.

[0173] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the unit and module in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0174] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can refer to the relevant description of other embodiments.

[0175] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0176] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely schematic, for example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0177] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0178] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0179] The integrated modules / units, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The computer program can be executed by a processor to implement the steps of the above-mentioned various method embodiments. The computer program can include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc.

[0180] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An image generation method based on retrieval enhancement and multi-source feature fusion, characterized in that: include: Build a retrieval enhancement generation library containing image features and text description features; Extracting features from an input reference image, and performing similarity retrieval on image features in the retrieval enhancement generation library based on the features of the reference image to obtain a retrieval image; Encoding the reference image and the search image to obtain corresponding global features, and performing multi-view feature attention processing on the global features to obtain comprehensive image features; Generate a global text description of the reference image using a large visual language model, identify local objects in the reference image using an object detection model, and generate a corresponding local text description for each local object; Splitting the global text description and the local text description into a plurality of segmented descriptions according to semantic structure, and using the search enhancement generation library to search and expand the segmented descriptions to obtain comprehensive text description features; The comprehensive text description features and the comprehensive image features are input into a diffusion model for feature fusion, and a target image that conforms to the reference image and the text description is output.

2. The method according to claim 1, characterized in that The construction includes a retrieval enhancement generation library containing image features and text description features, including: Collecting images and corresponding text descriptions, and preprocessing the images and corresponding text descriptions; Performing feature extraction on the preprocessed image to obtain an image feature vector, and storing the image feature vector and the corresponding text description in the retrieval enhancement generation library; The preprocessed text description is preliminarily encoded to obtain an initial feature representation, the initial feature representation is input into a text encoder for feature extraction to obtain a text feature vector, and the text feature vector and the corresponding text description and image are stored in the retrieval enhancement generation library.

3. The method according to claim 1, characterized in that The step of performing similarity retrieval on image features in the retrieval enhancement generation library based on features of the reference image to obtain a retrieval image includes: Retrieving image features in the retrieval enhancement generation library using a similarity measurement algorithm based on features of the reference image; According to the similarity between the features of the reference image and the features of the images in the retrieval enhancement generation library, multiple images are screened out as retrieval images according to the similarity sorting.

4. The method according to claim 1, wherein The encoding of the reference image and the search image to obtain corresponding global features includes: dividing the search image and the reference image into a plurality of image blocks using the visual transformer image encoder; Linearly embed the image block and add position encoding information, process the image block using a multi-head self-attention mechanism and a multi-layer encoder, output image feature representation, and generate corresponding classification labels; Performing feature transformation on the classification labels using a linear layer so as to map each classification label to a feature space dimension; The classification labels after feature transformation are spliced ​​in order to form a feature representation of a comprehensive classification label, wherein the feature representation of the comprehensive classification label integrates the global features of the retrieval image and the reference image.

5. The method according to claim 4, characterized in that The performing multi-view feature attention processing on the global features to obtain comprehensive image features includes: Utilize the multi-view feature attention mechanism to capture the relationship between the global features from multiple perspectives, calculate the attention weights of the global features under different perspectives, and perform deep learning based on the attention weights; Inputting the feature results learned by the multi-view feature attention mechanism into a multi-layer perceptron, and using the multi-layer perceptron to perform nonlinear transformation on the input features; The features output by the multi-layer perceptron are concatenated with the classification labels after the feature transformation to obtain the comprehensive image features.

6. The method according to claim 1, characterized in that The method of using the object detection model to identify local objects in the reference image and generating a corresponding local text description for each local object includes: Identifying local objects in the reference image using an object detection model, generating a corresponding bounding box for each local object, and extracting the local object from the reference image based on the bounding box to generate a sub-image; The sub-images are sequentially input into a large visual language model, and a local text description corresponding to each local object is generated according to the visual features of the local objects in the sub-images.

7. The method according to claim 1, characterized in that The global text description and the local text description are split into a plurality of segmented descriptions according to the semantic structure, and the segmented descriptions are retrieved and expanded using the search enhancement generation library to obtain comprehensive text description features, including: Using a thought chain processing method in a large language model, the global text description and the local text description are split into a plurality of segmented descriptions according to a logical order and semantic association; Input each segment description into the search enhancement generation library, and use a similarity matching algorithm to retrieve relevant text description features; The retrieved relevant text description features are input into the large language model, the segment description is expanded according to the relevant text description features to obtain the expanded content of the segment description, and a comprehensive text description feature is generated according to the segment description and the expanded content.

8. The method according to claim 1, characterized in that The step of inputting the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and outputting a target image that conforms to the reference image and the text description, includes: generating an initial noise map as an input to a diffusion model, and using the initial noise map as a starting state of a target image; The diffusion model includes a plurality of cross-attention modules, and in a previous cross-attention module, the comprehensive text description feature is fused with the input feature of the current cross-attention module; In the next cross-attention module, the fusion feature output by the previous cross-attention module is fused with the comprehensive image feature; After being processed by each cross-attention module in the diffusion model in turn, a target image that conforms to the reference image and text description is output.

9. An image generation device based on retrieval enhancement and multi-source feature fusion, characterized in that: include: A construction module for constructing a retrieval enhancement generation library containing image features and text description features; An extraction module is used to extract features from an input reference image and perform similarity retrieval on image features in the retrieval enhancement generation library based on the features of the reference image to obtain a retrieval image; an encoding module, configured to encode the reference image and the search image to obtain corresponding global features, and perform multi-view feature attention processing on the global features to obtain comprehensive image features; a generation module, configured to generate a global text description of the reference image using a large visual language model, identify local objects in the reference image using an object detection model, and generate a corresponding local text description for each local object; An expansion module, configured to split the global text description and the local text description into a plurality of segmented descriptions according to semantic structures, and to retrieve and expand the segmented descriptions using the search enhancement generation library to obtain comprehensive text description features; The output module is used to input the comprehensive text description features and the comprehensive image features into a diffusion model for feature fusion, and output a target image that conforms to the reference image and the text description.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.