Image Generation Method, Apparatus, Electronic Device, and Storage Medium

Through image retrieval, feature extraction and generation processes, combined with image structured features and text features, the target image is generated using the SDXL model, which solves the problem of unrelated drawings caused by inaccurate description in the image library, and achieves high-quality image generation and diversity improvement.

CN117351116BActive Publication Date: 2025-07-08BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311338853.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-16
Publication Date
2025-07-08
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

In the prior art, some images in the image library are described as empty or inaccurate, resulting in unrelated pictures when matching pictures and texts, and the control of simple text and production pictures is weak, and the availability rate of generated images is poor.

Method used

Through the image retrieval, feature extraction and generation process, combined with image structured features and text features, the target image is generated using the SDXL model, including the image retrieval module, feature extraction module and image generation module, and the features are extracted using posture detection, line segment detection, depth detection and edge detection models, and images are generated by combining the Stable Diffusion neural network.

Benefits of technology

It improves the relevance and quality of generated images, increases the diversity and effect of the number of pictures, and improves the control ability and refinement effect of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351116B_ABST
    Figure CN117351116B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image generation method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence technology, and particularly to fields such as computer vision, intelligent search, and deep learning. The specific implementation solution is as follows: perform image retrieval according to the text content to obtain a target retrieval image related to the text content; extract features from the target retrieval image to obtain image structured features; and generate a target image according to the image structured features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to fields such as computer vision, intelligent search, deep learning, etc. Specifically, it relates to an image generation method, apparatus, electronic device, and storage medium. Background Art

[0002] As a rich media form of information, the graphic style can increase the information content of the text by adding images compared to information such as pure text articles and advertisements, making the text more vivid. The text with images is also more likely to arouse the interest of readers, thereby improving the click-through rate and monetization efficiency of the text. In this process, it is necessary to match high-quality pictures that meet the text relevance based on the given text. Summary of the Invention

[0003] The present disclosure provides an image generation method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided an image generation method, including: performing image retrieval according to text content to obtain a target retrieval image related to the text content; extracting features from the target retrieval image to obtain image structured features; and generating a target image according to the image structured features.

[0005] According to another aspect of the present disclosure, there is provided an image generation apparatus, including: an image retrieval module for performing image retrieval according to text content to obtain a target retrieval image related to the text content; a feature extraction module for extracting features from the target retrieval image to obtain image structured features; and an image generation module for generating a target image according to the image structured features.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image generation method of the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image generation method of the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, the computer program being stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the image generation method of the present disclosure.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 Schematically shows an exemplary system architecture to which the image generation method and apparatus according to embodiments of the present disclosure can be applied;

[0012] Figure 2 Schematically shows a flowchart of the image generation method according to embodiments of the present disclosure;

[0013] Figure 3 Schematically shows a schematic diagram of obtaining image structured features by feature extraction of a target retrieval image according to embodiments of the present disclosure;

[0014] Figure 4 Schematically shows an overall schematic diagram of generating a matching picture for text content according to embodiments of the present disclosure;

[0015] Figure 5 Schematically shows a structural diagram of the SDXL model according to embodiments of the present disclosure;

[0016] Figure 6 Schematically shows an application diagram of the image generation method according to embodiments of the present disclosure;

[0017] Figure 7 Schematically shows a block diagram of the image generation apparatus according to embodiments of the present disclosure; and

[0018] Figure 8 Shows a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.

[0021] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent is obtained.

[0022] When matching a picture to the text, it is usually completed by relying on existing images and matching the semantic similarity between the text description and the image description of the existing images.

[0023] The inventors found in the process of implementing the concept of the present disclosure that due to problems such as empty descriptions, inaccurate descriptions, and poor description value in a part of the images in the existing image library, relying on these descriptions for picture matching in the same modality will result in irrelevant picture matching. When adding the picture description as a medium in the process of picture-text matching, the picture description may only describe some attributes in the picture, which will cause a loss in the final relevance of the picture and text. In addition, the control ability of pure text-to-image is weak, and the availability of the generated images is poor.

[0024] Figure 1 An exemplary system architecture to which the image generation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0025] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the image generation method and apparatus can be applied may include a terminal device, but the terminal device can implement the image generation method and apparatus provided by the embodiments of the present disclosure without interacting with the server.

[0026] As Figure 1 shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0027] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (only examples).

[0028] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.

[0029] The server 105 can be a server that provides various services, such as a background management server (only as an example) that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (″Virtual Private Server″, or simply referred to as ″VPS″). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0030] It should be noted that the image generation method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the image generation device provided by the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0031] Alternatively, the image generation method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the image generation device provided by the embodiments of the present disclosure can generally be set in the server 105. The image generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0032] For example, when an image needs to be generated, the first terminal device 101, the second terminal device 102, and the third terminal device 103 can obtain the text content, and then send the obtained text content to the server 105. The server 105 performs image retrieval based on the text content to obtain a target retrieval image related to the text content, extracts features from the target retrieval image to obtain image structured features, and generates a target image based on the image structured features. Alternatively, a server or a server cluster capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105 analyzes the text content and realizes the generation of the target image.

[0033] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0034] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0035] As Figure 2 shown, the method includes operations S210 to S230.

[0036] In operation S210, image retrieval is performed based on the text content to obtain a target retrieval image related to the text content.

[0037] In operation S220, features are extracted from the target retrieval image to obtain image structured features.

[0038] In operation S230, a target image is generated based on the image structured features.

[0039] According to an embodiment of the present disclosure, after obtaining the text content, an image matching the semantic information can be retrieved from an image database based on the semantic information represented by the text content to obtain a target retrieval image. The image database can be an open-source database including an open-source image set, a copyright image database with obtained copyrights, etc., and is not limited thereto.

[0040] According to an embodiment of the present disclosure, the image structured features can represent different types of features respectively extracted for different types of elements in the image. For example, a skeleton diagram feature can be extracted for the human element in the image, a line feature can be extracted for the building element in the image, a depth feature can be extracted for the scene element in the image, an edge feature can be extracted for other elements in the image, etc. The obtained image structured features can include at least one of the following: skeleton diagram feature, line feature, depth feature, edge feature, etc., and is not limited thereto.

[0041] According to an embodiment of the present disclosure, the structural features of an image can be processed in combination with a base generation model to generate a target image. The target image can be used as an illustration for text content. The base generation model can include, for example, various text-to-image models, image-to-image models, etc., and is not limited thereto.

[0042] Through the above embodiments of the present disclosure, by combining the methods of image retrieval and image generation, features of multiple modalities such as text features and image features can be synchronously introduced as control conditions for image generation, and the generated target image can have high quality. In addition, compared with traditional matching schemes, this method can understand the picture requirements in more detail. When using the target image as an illustration for text content, it can effectively increase the number of illustrations, enhance the diversity of illustrations, and improve the illustration effect.

[0043] The following will further illustrate the Figure 2 method shown with specific embodiments.

[0044] According to an embodiment of the present disclosure, the above operation S210 may include: performing image retrieval according to the text content to obtain an initial retrieved image related to the text content; scoring the initial retrieved image according to a preset scoring rule to obtain a scoring value of the initial retrieved image; and screening a target retrieved image from the initial retrieved images according to the scoring value and a predefined screening condition.

[0045] According to an embodiment of the present disclosure, the preset scoring rule may include at least one of the following: scoring according to the semantic similarity between the text content and the initial retrieved image; scoring according to the aesthetics of the initial retrieved image, etc., and is not limited thereto. The judgment criteria for aesthetics can be custom-set according to actual business requirements and will not be elaborated here. The preset screening condition may include: the scoring value is greater than a first scoring value, the scoring system is less than a second scoring value, etc., where the first scoring value is greater than the second scoring value, and is not limited thereto.

[0046] According to an embodiment of the present disclosure, in the case where an image matching the semantic information of the text content is retrieved from the image database, these images can be first determined as the initial retrieved images. Then, according to the above preset scoring rule, the initial retrieved images can be scored to obtain the scoring value of each initial retrieved image. After that, a retrieved image whose scoring value meets the preset screening condition can be screened from the initial retrieved images as the target retrieved image.

[0047] Through the above embodiments of the present disclosure, by initially screening the retrieved images by combining the scoring values of the retrieved images, the quality of the obtained target retrieved images can be improved, which is conducive to improving the quality of the generated target images.

[0048] According to an embodiment of the present disclosure, the above operation S220 may include: classifying the target retrieval image to determine the target image category of the target retrieval image. According to the target feature extraction model adapted to the target image category, feature extraction is performed on the target retrieval image to obtain image structured features.

[0049] According to an embodiment of the present disclosure, the categories of the target retrieval image may include, for example: human category, building category, scene category, other category, etc. The target feature extraction model may be a full-category feature extraction model capable of extracting various features such as skeleton map features, line features, depth features, edge features, etc. For example, when the category of the target retrieval image is determined to be the human category, the scene category, and the building category, skeleton map features, line features, and depth features may be extracted based on this full-category feature extraction model.

[0050] According to an embodiment of the present disclosure, when the target retrieval image is determined to be an image including multiple element categories, a primary-secondary judgment may also be performed on the multiple element categories, and according to the primary-secondary judgment result, the weights of the model parameters for extracting features of different categories in the full-category feature extraction model are adjusted to obtain a target feature extraction model adapted to the target retrieval image.

[0051] For example, when the category of the target retrieval image is determined to be the human category and the scene category, a saliency detection model may be combined to detect the main element of the target retrieval image, for example, an element of the human category. In this case, a relatively large weight may be configured for the model parameters for extracting relevant features of elements of the human category in the full-category feature extraction model, and a relatively small weight may be configured for the model parameters for extracting relevant features of elements of the scene category in the full-category feature extraction model, so as to extract the features after feature weighting of different category elements according to the image content as the extracted image structured features.

[0052] It should be noted that based on the above classification result of the target retrieval image, relatively smaller weights may also be configured for the model parameters for extracting relevant features of elements of the building category and other category elements in the full-category feature extraction model. In some embodiments, for example, it may be configured to 0, and this is not limited thereto.

[0053] According to an embodiment of the present disclosure, depending on the category of the target retrieval image, the target feature extraction model may also include multiple independently existing feature extraction models such as a pose detection model, a line segment detection model, a depth detection model, an edge detection model, etc., and is not limited thereto. For example, the pose detection model may use openpose (a human pose recognition model), etc. The line segment detection model may use M-LSD (a real-time lightweight line segment detector for resource-constrained environments), for example. The depth detection model may use Depth (a neural network model for implementing depth detection), etc. The edge detection model may use canny, etc. The selection and structure of various models are not limited to those described above.

[0054] According to an embodiment of the present disclosure, when the target image category includes the human category and the target feature extraction model includes a pose detection model, feature extraction is performed on the target retrieval image according to the target feature extraction model adapted to the target image category, and the obtained image structured features may include: inputting the target retrieval image into the pose detection model to obtain bone graph features. Determining the bone graph features as the image structured features.

[0055] According to an embodiment of the present disclosure, when the target image category includes the building category and the target feature extraction model includes a line segment detection model, feature extraction is performed on the target retrieval image according to the target feature extraction model adapted to the target image category, and the obtained image structured features may include: inputting the target retrieval image into the line segment detection model to obtain line features. Determining the line features as the image structured features.

[0056] According to an embodiment of the present disclosure, when the target image category includes the scene category and the target feature extraction model includes a depth detection model, feature extraction is performed on the target retrieval image according to the target feature extraction model adapted to the target image category, and the obtained image structured features may include: inputting the target retrieval image into the depth detection model to obtain image depth features. Determining the image depth features as the image structured features.

[0057] According to an embodiment of the present disclosure, when the target retrieval image includes elements that do not belong to the task category, the scene category, or the building category, or the target retrieval image is determined to be of other categories, and the target feature extraction model includes an edge detection model, feature extraction is performed on the target retrieval image according to the target feature extraction model adapted to the target image category, and the obtained image structured features may include: inputting the target retrieval image into the edge detection model to obtain image edge features. Determining the image edge features as the image structured features.

[0058] Figure 3 Schematically shows a schematic diagram of obtaining image structured features by performing feature extraction on a target retrieval image according to an embodiment of the present disclosure.

[0059] As Figure 3 shown, after obtaining the target retrieval image 300, the target retrieval image 300 can be first input into the classification model 310 for image category classification. For example, the target retrieval image 300 can be classified into at least one of a person category retrieval image 311, a building category retrieval image 312, a scene category retrieval image 313, and an other category retrieval image 314. Then, the person category retrieval image 311 can be input into the openpose model 320 to obtain a skeletal map feature 321. The building category retrieval image 312 can be input into the M-LSD model 330 to obtain a line feature 332. The scene category retrieval image 313 can be input into the Depth model 340 to obtain an image depth feature 343. The other category retrieval image 314 can be input into the canny model 350 to obtain an edge feature 354. The skeletal map feature 321, the line feature 332, the image depth feature 343, and the edge feature 354 can all be used as image structured features 360.

[0060] It should be noted that the above openpose model 320, M-LSD model 330, Depth model 340, and canny model 350 can also be replaced with the aforementioned full-category feature extraction model, and image structured features can be extracted based on the foregoing embodiments, which will not be elaborated herein.

[0061] Through the above embodiments of the present disclosure, features adapted to different types of image elements can be respectively extracted for different types of image elements, and the obtained image structured features can more accurately and completely represent the features of each part of the content in the image, which is beneficial to improving the effect of the generated target image.

[0062] According to an embodiment of the present disclosure, the target retrieval image may include a target object. The above operation S220 may further include: performing target detection on the target retrieval image to obtain a target detection frame. According to the target detection frame, the target retrieval image is cropped to obtain a target object region image of the target object. Feature extraction is performed on the target object region image to obtain image structured features.

[0063] According to an embodiment of the present disclosure, the target object may represent the image main body in the target retrieval image. A target detection model or other saliency detection model can be used to detect the target object, and according to the target detection frame in the detection result, the target object is cropped out from the target detection image to obtain a target object region image. The image structured features may be features obtained by performing feature extraction on the cropped target object region image.

[0064] It should be noted that a lower limit value can be set corresponding to the size of the target detection box. Based on this lower limit value, the too-small target detection boxes obtained by detection can be directly deleted, and the image of the target object area corresponding to them is no longer obtained.

[0065] Through the above embodiments of the present disclosure, when the main body of the image is too small, the features of the main body of the image can be accurately extracted, which is conducive to centering the main body of the generated target image and further improving the quality of the generated image.

[0066] According to an embodiment of the present disclosure, the above operation S230 may include: generating a target image according to the image structured feature and the text feature of the text content.

[0067] According to an embodiment of the present disclosure, the image structured feature and the text feature may also be jointly processed in combination with a basic generation model to generate a target image. The basic generation model may include Stable Diffusion (a neural network model for generating realistic images), etc., and is not limited thereto. The Stable Diffusion may include a controlNet plugin.

[0068] In this embodiment, the image structured feature may be used as the input of the controlNet, and the text feature may be used as the input of the Stable Diffusion. Since ControlNet can control Stable Diffusion during decoding, during the process of processing the text feature based on Stable Diffusion, the image structured feature may be used as a control condition based on the controlNet plugin to perform text-to-image processing on the text feature to generate a target image.

[0069] It should be noted that the input of the Stable Diffusion is three-dimensional. During the process of using Stable Diffusion to generate a target image, at most only three types of image structured features can be received as control conditions. The process of using other basic generation models to generate a target image is not limited herein.

[0070] Through the above embodiments of the present disclosure, by combining various modal features such as image structured features and text features as control conditions, the generated target image can have a high degree of relevance to both the text content and the target retrieval image. When used as an illustration for the text content, the illustration effect can be effectively improved. In addition, using the image structured feature as a control condition can also improve the control ability of the basic generation model.

[0071] According to an embodiment of the present disclosure, before generating a target image based on text features according to image structured features and text content, the text content may first be preprocessed, for example, it may include: preprocessing the text content to obtain preprocessed content; extracting features from the preprocessed content to obtain text features.

[0072] According to an embodiment of the present disclosure, the preprocessing may include at least one of the following processing methods: text rewriting, data augmentation, risk control processing, etc., and is not limited thereto.

[0073] Through the above embodiments of the present disclosure, by performing the preprocessing process, the text semantics can be understood in more detail, the richness of the text content can be expanded, and further, the diversity of the accompanying pictures can be increased and the quality of the accompanying pictures can be improved.

[0074] Figure 4 Schematically shows an overall schematic diagram of generating an accompanying picture for text content according to an embodiment of the present disclosure.

[0075] As Figure 4 shown, there is Prompt 401, for example, "a female student with short hair is studying in the classroom". First, based on the text content of the Prompt 401, the initial retrieved image 411 can be retrieved from the image library 410. Then, the quality of the initial retrieved image 411 can be preferably selected to screen out the target retrieved image 412. The target retrieved image 412 can be directly input into the classification model 420, or after subject recognition and cropping of the target retrieved image 412, it can be input into the classification model 420 to obtain the target image category 421 of the target retrieved image. Combining with the target feature extraction model 430, features are extracted from the target retrieved image, and different types of features are weighted according to the target image category 421, and the image structured features 431 of the target retrieved image can be obtained, for example, it may include edge features 431_1, depth features 431_2, skeleton graph features 431_3, line features 431_4, etc. In this process, the Prompt 401 can also be preprocessed to obtain preprocessed content 402. By inputting the preprocessed content 402 and the image structured features 431 into the basic generation model 440, the target image 441 can be generated and used as the accompanying picture of the Prompt 401.

[0076] According to an embodiment of the present disclosure, the basic generation model 440 may use SDXL (Stable Diffusion XL) to perform the process of generating a target image according to image structured features, or according to text features of image structured features and text content. The SDXL model may include two text encoders and at least one diffusion model.

[0077] Figure 5Schematically shows a structural diagram of an SDXL model according to an embodiment of the present disclosure.

[0078] As Figure 5 shown, the SDXL model 500 may include a first text encoder 510, a second text encoder 520, a first diffusion model 530, and a second diffusion model 540.

[0079] According to an embodiment of the present disclosure, the first text encoder 510 may be a text encoder pre-trained using text-image pairs for text-image matching. The second text encoder 520 may be a text encoder pre-trained using text-text pairs of a text generation model. For example, the second text encoder 520 may select an open-source Chinese text generation encoder, and is not limited thereto. Both the first diffusion model 530 and the second diffusion model 540 may use the UNET structure. The second diffusion model 540 may also cascade a refiner (a kind of image-to-image model) 541 to improve the image generation quality.

[0080] According to an embodiment of the present disclosure, when training the SDXL model 500, additional conditional injection may also be adopted to improve data processing problems during training. For example, the features of the crop region of the object detection image may be input into the first diffusion model 530 as a control condition for training the SDXL model 500. Instead of the way of cropping first and then training, approximate multi-scale fine-tuning can increase the scale of training data and improve the control ability of the model.

[0081] According to an embodiment of the present disclosure, the above operation S230 may include: performing first feature encoding on the image structured features to obtain a first feature vector, where the first feature vector represents the encoding vector of the first part of the features in the image structured features. Performing second feature encoding on the image structured features to obtain a second feature vector, where the first feature vector represents the encoding vector of the second part of the features in the image structured features, and there are differences between the first part of the features and the second part of the features. Fusing the first feature vector and the second feature vector to obtain a first fused feature vector. Decoding the first fused feature vector to generate the target image.

[0082] According to an embodiment of the present disclosure, in combination with Figure 5As shown, since the first text encoder 510 and the second text encoder 520 are text encoding models trained based on different pre-training data, when the same feature is input into the first text encoder 510 and the second text encoder 520 respectively, two feature vectors with different dimensions can be obtained. For example, when the feature of the text "The vehicle is speeding in the desert and there is a cactus beside it" is input into the first text encoder 510, a feature vector encoded for texts such as "vehicle" and "desert" can be obtained. When it is input into the second text encoder 520, a feature vector encoded for "cactus", for example, can be obtained.

[0083] According to an embodiment of the present disclosure, in the process of generating a target image based on the image structured features, Figure 5 As shown, the image structured features can be first input into the first text encoder 510 and the second text encoder 520 respectively. Then, the first feature encoding can be performed based on the first text encoder 510 to obtain a first feature vector. And the second feature encoding can be performed based on the second text encoder 520 to obtain a second feature vector. After that, the first fusion feature vector obtained by fusing the first feature vector and the second feature vector can be input into the first diffusion model 530 for decoding to generate a target image.

[0084] According to an embodiment of the present disclosure, the above-mentioned generation of the target image based on the image structured features and the text features of the text may include: performing a third feature encoding on the image structured features and the text features to obtain a third feature vector, where the third feature vector represents the encoding vector of the third part of the features in the image structured features and the text features. Performing a fourth feature encoding on the image structured features and the text features to obtain a fourth feature vector, where the fourth feature vector represents the encoding vector of the fourth part of the features in the image structured features and the text features, and there are differences between the third part of the features and the fourth part of the features. Fusing the third feature vector and the fourth feature vector to obtain a second fusion feature vector. Decoding the second fusion feature vector to generate a target image.

[0085] According to an embodiment of the present disclosure, in the process of generating a target image based on the image structured features and the text features, Figure 5As shown, the image-text fusion features of the image structural features and text features can be input into the first text encoder 510 and the second text encoder 520 respectively. Then, the third feature encoding can be performed based on the first text encoder 510 to obtain the third feature vector. And the fourth feature encoding can be performed based on the second text encoder 520 to obtain the fourth feature vector. After that, the second fusion feature vector obtained by fusing the third feature vector and the fourth feature vector can be input into the first diffusion model 530 for decoding, and the target image can be generated.

[0086] Through the above embodiments of the present disclosure, while performing two-step feature encoding on the input features to extract more complete feature vectors, the dimension of the obtained feature vectors will also become larger. Based on the feature vectors with a larger dimension, the input features can be understood more thoroughly, and the generation effect can be further improved.

[0087] According to an embodiment of the present disclosure, the above decoding of the second fusion feature vector to generate the target image may include: decoding the second fusion feature vector to generate an initial image. Generating the target image according to the initial image and the text features.

[0088] Based on the above embodiments, in combination with Figure 5 As shown, after inputting the second fusion feature vector into the first diffusion model 530 for decoding, the generated image can first be used as the initial image. After that, the initial image can be input into the second diffusion model 540, and during the process of encoding and decoding the initial image by the refiner cascaded based on the second diffusion model 540, the text features can be introduced as the control vector for this encoding and decoding process to generate the target image.

[0089] Through the above embodiments of the present disclosure, by combining two diffusion models, the generated target image is more delicate, and the image generation quality can be further improved. In addition, compared with the lower version of the SD model (such as SD1.5), the model parameters representing the depth and width in the diffusion model of the SDXL model are both increased, making the overall diffusion model larger. When generating an image from a vector based on the diffusion model in SDXL, the image generation ability can be effectively enhanced, and the refinement effect of the image can be improved.

[0090] According to an embodiment of the present disclosure, after obtaining the target image, the above image generation method may further include: in response to receiving an instruction for the selected illustration size, cropping the target image according to the illustration size.

[0091] According to an embodiment of the present disclosure, the target image generated based on the above method may have a default image size. In the case where the user inputs text content and selects the desired illustration size, after obtaining the target image, the target image can be cropped according to the desired illustration size, and the cropped image can be used as the illustration for the text content.

[0092] Figure 6 FIG. schematically shows an application diagram of an image generation method according to an embodiment of the present disclosure.

[0093] As Figure 6 shown, in the page 600 for text-to-image generation, there are a promotion service module 610, a scene description module 620, a main subject option module 630, and a ratio / size option module 640. The type of service to be promoted can be filled in the promotion service module 610. For example, it can be filled with: tourism service → travel agency. The text content that requires an illustration can be filled in the scene description module 620. For example, it can be filled with: A vehicle is speeding in the desert, and there are cacti beside it. The category of the main subject of the image to be generated can be selected in the main subject option module 630. For example, it can be selected: desert scenery. The illustration size of the image to be generated can be selected in the ratio / size option module 640. For example, 1:1 can be selected. Through the foregoing configuration, an illustration 650 with a size of 1:1 can be obtained, for example.

[0094] It should be noted that the scene description module 620 and the main subject option module 630 can also be used alternatively, and this is not limited herein.

[0095] Through the above embodiments of the present disclosure, an illustration that better meets the business requirements can be obtained, improving the user experience.

[0096] Figure 7 FIG. schematically shows a block diagram of an image generation device according to an embodiment of the present disclosure.

[0097] As Figure 7 shown, the image generation device 700 includes an image retrieval module 710, a feature extraction module 720, and an image generation module 730.

[0098] The image retrieval module 710 is configured to perform image retrieval according to the text content to obtain a target retrieval image related to the text content.

[0099] The feature extraction module 720 is configured to perform feature extraction on the target retrieval image to obtain an image structured feature.

[0100] The image generation module 730 is configured to generate a target image according to the image structured feature.

[0101] According to an embodiment of the present disclosure, the image generation module includes a first feature encoding sub-module, a second feature encoding sub-module, a fusion sub-module, and a decoding sub-module.

[0102] The first feature encoding sub-module is configured to perform first feature encoding on the image structural features to obtain a first feature vector, where the first feature vector represents the encoded vector of the first part of the features in the image structural features.

[0103] The second feature encoding sub-module is configured to perform second feature encoding on the image structural features to obtain a second feature vector, where the first feature vector represents the encoded vector of the second part of the features in the image structural features, and there are differences between the first part of the features and the second part of the features.

[0104] The fusion sub-module is configured to fuse the first feature vector and the second feature vector to obtain a first fused feature vector.

[0105] The decoding sub-module is configured to decode the first fused feature vector to generate a target image.

[0106] According to an embodiment of the present disclosure, the image generation module includes an image generation sub-module.

[0107] The image generation sub-module is configured to generate a target image according to the image structural features and the text features of the text content.

[0108] According to an embodiment of the present disclosure, the image generation sub-module includes a third feature encoding unit, a fourth feature encoding unit, a fusion unit, and a decoding unit.

[0109] The third feature encoding unit is configured to perform third feature encoding on the image structural features and the text features to obtain a third feature vector, where the third feature vector represents the encoded vector of the third part of the features in the image structural features and the text features.

[0110] The fourth feature encoding unit is configured to perform fourth feature encoding on the image structural features and the text features to obtain a fourth feature vector, where the fourth feature vector represents the encoded vector of the fourth part of the features in the image structural features and the text features, and there are differences between the third part of the features and the fourth part of the features.

[0111] The fusion unit is configured to fuse the third feature vector and the fourth feature vector to obtain a second fused feature vector.

[0112] The decoding unit is configured to decode the second fused feature vector to generate a target image.

[0113] According to an embodiment of the present disclosure, the decoding unit includes a decoding sub-unit and an image generation sub-unit.

[0114] A decoding subunit for decoding the second fused feature vector to generate an initial image.

[0115] An image generation subunit for generating a target image according to the initial image and text features.

[0116] According to an embodiment of the present disclosure, the feature extraction module includes a classification sub-module and a first feature extraction sub-module.

[0117] The classification sub-module is used for classifying the target retrieval image to determine the target image category of the target retrieval image.

[0118] The first feature extraction sub-module is used for extracting features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain image structured features.

[0119] According to an embodiment of the present disclosure, the target image category includes a person category, and the target feature extraction model includes a pose detection model. The first feature extraction sub-module includes a skeleton map feature obtaining unit and a first definition unit.

[0120] The skeleton map feature obtaining unit is used for inputting the target retrieval image into the pose detection model to obtain skeleton map features.

[0121] The first definition unit is used for determining the skeleton map features as image structured features.

[0122] According to an embodiment of the present disclosure, the target image category includes a building category, and the target feature extraction model includes a line segment detection model. The first feature extraction sub-module includes a line feature obtaining unit and a second definition unit.

[0123] The line feature obtaining unit is used for inputting the target retrieval image into the line segment detection model to obtain line features.

[0124] The second definition unit is used for determining the line features as image structured features.

[0125] According to an embodiment of the present disclosure, the target image category includes a scene category, and the target feature extraction model includes a depth detection model. The first feature extraction sub-module includes a depth feature obtaining unit and a third definition unit.

[0126] The depth feature obtaining unit is used for inputting the target retrieval image into the depth detection model to obtain image depth features.

[0127] The third definition unit is used for determining the image depth features as image structured features.

[0128] According to an embodiment of the present disclosure, the target feature extraction model includes an edge detection model. The first feature extraction sub-module includes an edge feature obtaining unit and a fourth definition unit.

[0129] An edge feature acquisition unit for inputting a target retrieval image into an edge detection model to obtain image edge features.

[0130] A fourth definition unit for determining the image edge features as image structured features.

[0131] According to an embodiment of the present disclosure, the target retrieval image includes a target object. The feature extraction module includes a target detection sub-module, a cropping sub-module, and a second feature extraction sub-module.

[0132] The target detection sub-module for performing target detection on the target retrieval image to obtain a target detection box.

[0133] The cropping sub-module for cropping the target retrieval image according to the target detection box to obtain a target object region image of the target object.

[0134] The second feature extraction sub-module for extracting features from the target object region image to obtain image structured features.

[0135] According to an embodiment of the present disclosure, the image retrieval module includes an image retrieval sub-module, a scoring sub-module, and a screening sub-module.

[0136] The image retrieval sub-module for performing image retrieval according to the text content to obtain an initial retrieval image related to the text content.

[0137] The scoring sub-module for scoring the initial retrieval image according to a preset scoring rule to obtain a scoring value of the initial retrieval image.

[0138] The screening sub-module for screening the target retrieval image from the initial retrieval images according to the scoring value and predefined screening conditions.

[0139] According to an embodiment of the present disclosure, the image generation device further includes a preprocessing module and a text feature acquisition module.

[0140] The preprocessing module for preprocessing the text content to obtain preprocessed content.

[0141] The text feature acquisition module for extracting features from the preprocessed content to obtain text features.

[0142] According to an embodiment of the present disclosure, the image generation device further includes a cropping module.

[0143] The cropping module for, in response to receiving an instruction of a selected illustration size, cropping the target image according to the illustration size.

[0144] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0145] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image processing method of the present disclosure.

[0146] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the image processing method of the present disclosure.

[0147] According to an embodiment of the present disclosure, a computer program product includes a computer program, the computer program is stored on at least one of a readable storage medium and an electronic device, and the computer program implements the image processing method of the present disclosure when executed by a processor.

[0148] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0149] As Figure 8 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0150] Multiple components in device 800 are connected to the input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0151] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as an image processing method. For example, in some embodiments, the image processing method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the image processing method by any other suitable means (e.g., by means of firmware).

[0152] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0154] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0155] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0156] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0157] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0158] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0159] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An image generation method, comprising: Performing image retrieval according to the text content to obtain a target retrieval image related to the text content; Performing feature extraction on elements of different categories in the target retrieval image respectively according to a target feature extraction model to obtain image structured features of different categories, wherein the image structured features of different categories include at least two of a skeleton map feature, a line feature, a depth feature, and an edge feature, the weights of the model parameters for extracting elements of different categories in the target feature extraction model are different, and the weights of the model parameters are adjusted according to the priorities of multiple element categories included in the target retrieval image, and a target feature extraction model adapted to the target retrieval image is obtained by adjusting the weights of the model parameters; And Using the image structured features of different categories as control conditions to generate a target image.

2. The method according to claim 1, wherein The step of using the image structured features of different categories as control conditions to generate a target image includes: Performing first feature encoding on the image structured features to obtain a first feature vector, where the first feature vector represents an encoding vector of a first part of features in the image structured features; Performing second feature encoding on the image structured features to obtain a second feature vector, where the first feature vector represents an encoding vector of a second part of features in the image structured features, and the first part of features and the second part of features are different; Fusing the first feature vector and the second feature vector to obtain a first fused feature vector; and Decoding the first fused feature vector to generate the target image.

3. The method according to claim 1, wherein, The step of using the image structured features of different categories as control conditions to generate a target image includes: Generating the target image according to the image structured features and the text features of the text content.

4. The method according to claim 3, wherein, The step of generating the target image according to the image structured features and the text features of the text content includes: Performing third feature encoding on the image structured features and the text features to obtain a third feature vector, where the third feature vector represents an encoding vector of a third part of features in the image structured features and the text features; Performing fourth feature encoding on the image structured features and the text features to obtain a fourth feature vector, where the fourth feature vector represents an encoding vector of a fourth part of features in the image structured features and the text features, and the third part of features and the fourth part of features are different; Fusing the third feature vector and the fourth feature vector to obtain a second fused feature vector; and Decoding the second fused feature vector to generate the target image.

5. The method according to claim 4, wherein The step of decoding the second fused feature vector to generate the target image includes: Decoding the second fused feature vector to generate an initial image; and Generating the target image according to the initial image and the text features.

6. The method according to any one of claims 1-5, wherein The step of performing feature extraction on elements of different categories in the target retrieval image respectively according to a target feature extraction model to obtain image structured features of different categories includes: Classify the target retrieval image to determine the target image category of the target retrieval image; and Extract features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured features.

7. The method according to claim 6, wherein The target image category includes a person category, and the target feature extraction model includes a pose detection model; the extracting features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured features includes: Input the target retrieval image into the pose detection model to obtain skeleton map features; and Determine the skeleton map features as the image structured features.

8. The method according to claim 6, wherein, The target image category includes a building category, and the target feature extraction model includes a line segment detection model; the extracting features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured features includes: Input the target retrieval image into the line segment detection model to obtain line features; and Determine the line features as the image structured features.

9. The method according to claim 6, wherein, The target image category includes a scene category, and the target feature extraction model includes a depth detection model; the extracting features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured features includes: Input the target retrieval image into the depth detection model to obtain image depth features; and Determine the image depth features as the image structured features.

10. The method according to claim 6, wherein, The target feature extraction model includes an edge detection model; the extracting features from the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured features includes: Input the target retrieval image into the edge detection model to obtain image edge features; and Determine the image edge features as the image structured features.

11. According to the method according to any one of claims 1-10, wherein, The target retrieval image includes a target object; the extracting different types of image structured features from different types of elements in the target retrieval image according to a target feature extraction model includes: Perform target detection on the target retrieval image to obtain a target detection frame; Crop the target retrieval image according to the target detection frame to obtain a target object region image of the target object; and Extract features from the target object region image to obtain the image structured features.

12. The method according to any one of claims 1-11, wherein, The performing image retrieval according to the text content to obtain a target retrieval image related to the text content includes: Perform image retrieval according to the text content to obtain an initial retrieval image related to the text content; Score the initial retrieval image according to a preset scoring rule to obtain a score value of the initial retrieval image; and Screen the target retrieval image from the initial retrieval images according to the score value and predefined screening conditions.

13. The method according to any one of claims 3-5 further comprises: Before generating the target image according to the image structured features and the text features of the text content Preprocess the text content to obtain preprocessed content; and Extract features from the preprocessed content to obtain the text features.

14. The method according to any one of claims 1-13 further includes: In response to receiving an instruction for a selected illustration size, crop the target image according to the illustration size.

15. An image generation device includes: An image retrieval module configured to perform image retrieval according to text content to obtain a target retrieval image related to the text content; A feature extraction module configured to respectively extract features of different categories of elements in the target retrieval image according to a target feature extraction model to obtain image structure features of different categories, where the image structure features of different categories include at least two of skeleton graph features, line features, depth features, and edge features, weights of model parameters for extracting different categories of elements in the target feature extraction model are different, and the weights of the model parameters are adjusted according to the priorities of multiple element categories included in the target retrieval image, and a target feature extraction model adapted to the target retrieval image is obtained by adjusting the weights of the model parameters; and An image generation module configured to use the image structure features of different categories as control conditions to generate a target image.

16. The device according to claim 15, wherein, The image generation module includes: A first feature encoding sub-module configured to perform first feature encoding on the image structure features to obtain a first feature vector, where the first feature vector represents an encoding vector of a first part of the features in the image structure features; A second feature encoding sub-module configured to perform second feature encoding on the image structure features to obtain a second feature vector, where the first feature vector represents an encoding vector of a second part of the features in the image structure features, and the first part of the features is different from the second part of the features; A fusion sub-module configured to fuse the first feature vector and the second feature vector to obtain a first fused feature vector; and A decoding sub-module configured to decode the first fused feature vector to generate the target image.

17. The apparatus according to claim 15, wherein, The image generation module includes: An image generation sub-module configured to generate the target image according to the image structure features and the text features of the text content.

18. The device according to claim 17, wherein The image generation sub-module includes: A third feature encoding unit configured to perform third feature encoding on the image structure features and the text features to obtain a third feature vector, where the third feature vector represents an encoding vector of a third part of the features in the image structure features and the text features; A fourth feature encoding unit configured to perform fourth feature encoding on the image structure features and the text features to obtain a fourth feature vector, where the fourth feature vector represents an encoding vector of a fourth part of the features in the image structure features and the text features, and the third part of the features is different from the fourth part of the features; A fusion unit configured to fuse the third feature vector and the fourth feature vector to obtain a second fused feature vector; and A decoding unit, configured to decode the second fused feature vector to generate the target image.

19. The device according to claim 18, wherein, The decoding unit includes: A decoding subunit, configured to decode the second fused feature vector to generate an initial image; and An image generation subunit, configured to generate the target image according to the initial image and the text feature.

20. The device according to any one of claims 15 - 19, wherein, The feature extraction module includes: A classification sub-module, configured to classify the target retrieval image to determine the target image category of the target retrieval image; and A first feature extraction sub-module, configured to perform feature extraction on the target retrieval image according to a target feature extraction model adapted to the target image category to obtain the image structured feature.

21. The apparatus according to claim 20, wherein, The target image category includes a person category, and the target feature extraction model includes a pose detection model; the first feature extraction sub-module includes: A skeleton map feature obtaining unit, configured to input the target retrieval image into the pose detection model to obtain a skeleton map feature; and A first definition unit, configured to determine the skeleton map feature as the image structured feature.

22. The device according to claim 20, wherein The target image category includes a building category, and the target feature extraction model includes a line segment detection model; the first feature extraction sub-module includes: A line feature obtaining unit, configured to input the target retrieval image into the line segment detection model to obtain a line feature; and A second definition unit, configured to determine the line feature as the image structured feature.

23. The device according to claim 20, wherein, The target image category includes a scene category, and the target feature extraction model includes a depth detection model; the first feature extraction sub-module includes: A depth feature obtaining unit, configured to input the target retrieval image into the depth detection model to obtain an image depth feature; and A third definition unit, configured to determine the image depth feature as the image structured feature.

24. The device according to claim 20, wherein, The target feature extraction model includes an edge detection model; the first feature extraction sub-module includes: An edge feature obtaining unit, configured to input the target retrieval image into the edge detection model to obtain an image edge feature; and A fourth definition unit, configured to determine the image edge feature as the image structured feature.

25. The device according to any one of claims 15 - 24, wherein, The target retrieval image includes a target object; the feature extraction module includes: A target detection sub-module, configured to perform target detection on the target retrieval image to obtain a target detection box; A cropping sub-module, configured to crop the target retrieval image according to the target detection box to obtain a target object region image of the target object; and A second feature extraction sub-module, configured to perform feature extraction on the target object region image to obtain the image structured feature.

26. The apparatus according to any one of claims 15-25, wherein, The image retrieval module includes: An image retrieval sub-module, configured to perform image retrieval according to the text content to obtain an initial retrieval image related to the text content; A scoring sub-module, configured to score the initial retrieval image according to a preset scoring rule to obtain a scoring value of the initial retrieval image; and A screening sub-module, configured to screen the target retrieval image from the initial retrieval images according to the scoring value and a predefined screening condition.

27. The apparatus according to any one of claims 17-19, further comprising: a preprocessing module configured to preprocess the text content to obtain preprocessed content; and a text feature obtaining module configured to extract features from the preprocessed content to obtain the text features.

28. The apparatus according to any one of claims 15-27, further comprising: a cropping module configured to crop the target image according to the selected illustration size in response to receiving an instruction for the selected illustration size.

29. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-14.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

31. A computer program product, comprising a computer program stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Retrieval-based text-to-image generation with visual-semantic contrastive representation

    US20230260164A1