Image generation method and device, storage medium and electronic equipment

By extracting and fusing features of the reference image and target data information, and using a generative model to generate the target image, the problem of low efficiency in aesthetic design in the existing technology is solved, and efficient and accurate image generation is achieved.

CN120612380APending Publication Date: 2025-09-09LINGDI (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410263736.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing technology, the aesthetic design that integrates multiple elements is inefficient and mainly relies on manual creative output, resulting in low production efficiency and weak reusability.

Method used

The target image is generated by extracting features from the reference image and target data information and using the generative model to fuse the features.

Benefits of technology

The efficiency of image generation and the accuracy of generated content are improved, and the target image that meets the requirements is efficiently generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612380A_ABST
    Figure CN120612380A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, a storage medium and electronic equipment, and the method comprises the steps: carrying out the feature extraction of a reference image in a reference image set, and obtaining a first extraction result, and the reference image set comprises at least one reference image; feature extraction is carried out on target data information to obtain a second extraction result, and the target data information is used for determining the content of a target image; and generating the target image based on the first extraction result and the second extraction result. According to the method, feature extraction is performed on the reference image and the target data information for determining the content of the target image, and feature fusion is performed on the extracted reference image features and the extracted target data features, so that the target image meeting the requirements is generated, and the generation efficiency of the target image and the accuracy of the generated content are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a method, device, storage medium, and electronic device for generating an image. Background Art

[0002] With economic development, people are increasingly pursuing aesthetics and culture, such as fashion trends. To meet this growing demand, product design concepts and creative ideas need to be constantly updated. At the same time, with the diversification of the market, co-branded products are gradually entering the public eye. When designing co-branded products, it is often necessary to integrate the styles and characteristics of multiple brands to create new co-branded products.

[0003] In the existing technology, the realization of aesthetic design that requires the integration of multiple elements mainly relies on the creative output of designers and other staff. This implementation method is highly dependent on the creativity of the staff, resulting in relatively low efficiency in the production of related products.

[0004] There is currently no good solution to the above-mentioned problem of low product design efficiency. Summary of the Invention

[0005] In view of this, the present disclosure provides a method, device, storage medium and electronic device for image generation, which can efficiently generate images and avoid the problem of low image generation efficiency caused by relying on manual creative output to draw images.

[0006] According to a first aspect of an embodiment of the present disclosure, a method for generating an image is provided, the method comprising:

[0007] performing feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image;

[0008] Performing feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image;

[0009] The target image is generated based on the first extraction result and the second extraction result.

[0010] Optionally, performing feature extraction on a reference image in the reference image set to obtain a first extraction result includes:

[0011] Performing global feature extraction on the reference images to obtain global feature information of each reference image;

[0012] Performing local feature extraction on the reference images to obtain local feature information of each reference image;

[0013] The first extraction result is obtained based on the global feature information and the local feature information.

[0014] Optionally, performing feature extraction on the target data information to obtain a second extraction result includes:

[0015] Performing feature extraction on the target data information according to the data type using a multimodal understanding model to obtain a feature extraction result of at least one data type;

[0016] Feature fusion is performed on the feature extraction result of the at least one data type to obtain the second extraction result.

[0017] Optionally, generating the target image based on the first extraction result and the second extraction result includes:

[0018] Performing feature alignment on the global feature information in the first extraction result and the global feature information in the second extraction result to obtain a target global feature;

[0019] Performing feature alignment on the local feature information in the first extraction result and the local feature information in the second extraction result to obtain target local features;

[0020] The target image is generated based on the target global features and the target local features.

[0021] Optionally, generating the target image based on the first extraction result and the second extraction result includes:

[0022] Based on the first extraction result and the second extraction result, performing feature fusion on the first extraction result and the second extraction result through a generative model to obtain image data information;

[0023] The image data information is decoded to generate the target image.

[0024] Optionally, the method further includes:

[0025] Acquire a sample image set and sample data information, wherein the sample image set includes a reference sample image and a target sample image, and the sample data information is used to determine the content of the target sample image;

[0026] The generative model is trained based on the sample image set and the sample data information.

[0027] Optionally, the training of the generative model based on the sample image set and the sample data information includes:

[0028] Performing feature fusion on the features of the reference sample image and the features of the sample data information through a generative model to obtain sample image data information;

[0029] Decoding the sample image data information to generate a sample image;

[0030] determining a model loss based on a difference between the generated sample image and the target sample image;

[0031] The generative model is trained based on the model loss.

[0032] According to a second aspect of an embodiment of the present disclosure, there is provided an apparatus for generating an image, the apparatus comprising:

[0033] a first extraction unit, configured to perform feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image;

[0034] a second extraction unit, configured to perform feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image;

[0035] A generating unit is configured to generate the target image based on the first extraction result and the second extraction result.

[0036] Optionally, the first extraction unit is specifically configured to:

[0037] Performing global feature extraction on the reference images to obtain global feature information of each reference image;

[0038] Performing local feature extraction on the reference images to obtain local feature information of each reference image;

[0039] The first extraction result is obtained based on the global feature information and the local feature information.

[0040] Optionally, the second extraction unit is specifically configured to:

[0041] Performing feature extraction on the target data information according to the data type using a multimodal understanding model to obtain a feature extraction result of at least one data type;

[0042] Feature fusion is performed on the feature extraction result of the at least one data type to obtain the second extraction result.

[0043] Optional, generate unit, specifically used to:

[0044] Performing feature alignment on the global feature information in the first extraction result and the global feature information in the second extraction result to obtain a target global feature;

[0045] Performing feature alignment on the local feature information in the first extraction result and the local feature information in the second extraction result to obtain target local features;

[0046] The target image is generated based on the target global features and the target local features.

[0047] Optional, generate unit, specifically used to:

[0048] Based on the first extraction result and the second extraction result, performing feature fusion on the first extraction result and the second extraction result through a generative model to obtain image data information;

[0049] The image data information is decoded to generate the target image.

[0050] Optionally, the device further includes:

[0051] an acquiring unit, configured to acquire a sample image set and sample data information, wherein the sample image set includes a reference sample image and a target sample image, and the sample data information is used to determine the content of the target sample image;

[0052] A training unit is used to train the generative model based on the sample image set and the sample data information.

[0053] Optional training unit, specifically used for:

[0054] Performing feature fusion on the features of the reference sample image and the features of the sample data information through a fusion model generative model to obtain sample image coding information and image data information;

[0055] Decoding the sample image encoding information and image data information to generate a sample image;

[0056] determining a model loss based on a difference between the generated sample image and the target sample image;

[0057] The fusion model generative model is trained based on the model loss.

[0058] According to a third aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods described in the first aspect are implemented.

[0059] According to a fourth aspect of the embodiments of the present disclosure, there is provided an apparatus for generating an image, including:

[0060] processor;

[0061] a memory for storing processor-executable instructions;

[0062] The processor implements the steps of any one of the methods described in the first aspect above by running the executable instructions.

[0063] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0064] By extracting features from the reference image and the target data information that determines the target image content respectively, and fusing the extracted reference image features and target data features, a target image that meets the requirements is generated, thereby improving the generation efficiency of the target image and the accuracy of the generated content.

[0065] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0067] Figure 1 is a flow chart of a method for generating an image according to an exemplary embodiment of the present disclosure;

[0068] Figure 2 is a flow chart of a method for generating an image according to an exemplary embodiment of the present disclosure;

[0069] Figure 3 is a block diagram of an image generation device according to an exemplary embodiment of the present disclosure;

[0070] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device for image generation according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0071] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0072] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0073] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0074] With economic development, people are increasingly pursuing aesthetic culture, such as fashion trends. To meet people's growing cultural needs, product design concepts and creative ideas need to be constantly updated. At the same time, with the development of market diversification, co-branded products have gradually entered the public's field of vision. When designing co-branded products, it is usually necessary to integrate the product styles of multiple brands to create new co-branded products.

[0075] In existing technology, achieving aesthetic designs that require the integration of multiple elements primarily relies on designers manually searching for historical styles. Designers then spend a significant amount of time collecting and integrating inspiration based on their own experience and design requirements to create new works. This process is cumbersome for integrated design, inefficient, highly subjective, and lacks reusability.

[0076] Based on this, the present disclosure proposes a method for image generation, which extracts features from historical images through a machine learning model, and fuses historical feature information with new feature information to generate a new target image to achieve aesthetic design.

[0077] The following will explain the method of image generation, please refer to Figure 1 , Figure 1 This is a flow chart of a method for generating an image according to an exemplary embodiment of the present disclosure.

[0078] Step 102: Perform feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image.

[0079] First, a reference image set is determined for generating the target image. The reference images can be existing historical images. In some embodiments, the reference images can be images of clothing, backpacks, furniture, etc. Feature extraction is then performed on the reference images in the reference image set. The reference image set includes at least one reference image. If the reference image set includes multiple reference images, feature extraction is performed on each reference image one by one to obtain the first extraction result described above.

[0080] Step 104 : Perform feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image.

[0081] The target data information is information for determining the specific content of the target image. The target data information may be text information, image information, or video information. By performing feature extraction on the target data information, a second extraction result for generating the target image is obtained.

[0082] Step 106: Generate the target image based on the first extraction result and the second extraction result.

[0083] Based on the first and second extraction results, a generative model is used to generate a target image. In one embodiment, the reference image set contains two reference images, both of which are clothing images. The target data is textual information, such as floral patterns. Based on the first extraction result obtained from the reference image set and the second extraction result obtained from the textual information, a generative model is used to generate a clothing design image with floral patterns, i.e., the target image.

[0084] By extracting features from the reference image and the target data information that determines the content of the target image respectively, and fusing the extracted reference image features and target data features, a target image that fuses the reference image features and the target data information features is generated, thereby improving the generation efficiency of the target image and the accuracy of the generated content.

[0085] In one embodiment, the feature extraction of the reference images in the reference image set to obtain the first extraction result includes: performing global feature extraction on the reference images to obtain global feature information of each reference image; performing local feature extraction on the reference images to obtain local feature information of each reference image; and obtaining the first extraction result based on the global feature information and the local feature information.

[0086] By extracting global features from the reference images in the reference image set, global feature information of each reference image can be obtained. Global feature extraction refers to extracting representative, global features from the entire data set or image. These features are usually calculated on the entire data set or image, rather than based only on local areas or local information. In the field of computer vision, global feature extraction can be used to identify the overall features of an image, rather than just the features of a local area. In some embodiments of the present disclosure, global features may include information such as structural contours, shape geometry, and element matching in the image. For the extraction of global features, a variety of technical methods can be used, for example, by using a color histogram as a global feature, or by describing the shape information of the reference image, such as boundary features, contour features, etc., extracting a shape descriptor as a global feature, or by capturing the information of the reference image through the global average pooling or global maximum pooling layer of a convolutional neural network, mapping the entire image into a feature vector of a fixed length, and other methods to extract global feature information of the reference image.

[0087] Extract local feature information of the reference images in the reference image set to obtain local feature information of each reference image. Extracting local features of an image is an important task in computer vision, which can capture the details and structure of a local area in a reference image. In some embodiments of the present disclosure, local features may include information such as color, texture, and pattern of the reference image. The extraction of local features can be achieved through a variety of technical means, for example, based on the SITF (Scale Invariant Feature Transform) algorithm, by detecting key points in the reference image and calculating local feature descriptors around these key points to describe the local structure of the reference image, or based on the SURF (Speeded Up Robust Features) algorithm, by detecting points of interest in the reference image and calculating local feature descriptors around these points of interest to describe the local structure of the reference image, or by using some convolutional neural network models to slide convolution kernels in the reference image to extract local features, and other technical means to extract local feature information of the reference image.

[0088] Based on the global feature information and the local feature information of each reference image, a first extraction result of the reference image set may be determined.

[0089] By extracting the global feature information and local feature information of the reference image respectively, a comprehensive analysis of the reference image can be achieved, and the feature information of the reference image can be fully obtained, thereby improving the accuracy of subsequent generation of the target image based on the feature information of the reference image.

[0090] In one embodiment, the feature extraction of the target data information to obtain the second extraction result includes: performing feature extraction on the target data information according to the data type through a multimodal understanding model to obtain a feature extraction result of at least one data type; and performing feature fusion on the feature extraction result of the at least one data type to obtain the second extraction result.

[0091] To increase the flexibility of the image generation method disclosed herein, the target data information can be of multiple data types, such as text data, image data, or video data. To better extract feature information from the target data information, a multimodal understanding model can be introduced to perform feature extraction on the target data information. The multimodal understanding model can simultaneously perform feature extraction on target data information of different data types, and the extracted feature information of the different data types can be fused together to obtain a second extraction result. For example, the target data information may include text data and image data, where the text data is floral and the image data is a floral pattern image. The multimodal understanding model performs feature extraction on the text data and image data separately to obtain text features and image features. The text features and image features of the different data types are then fused to obtain a second extraction result representing the overall multimodal feature representation. For another example, the target data information may include text data and video data, where the text data is a national trend style and the video data is a national trend style clothing display video. The multimodal understanding model performs feature extraction on the text data and video data separately to obtain text features and video features. The text features and video features are then fused together to obtain a second extraction result. The above-mentioned feature fusion process can be achieved through simple concatenation, weighted summation, etc. In addition, more complex methods such as multimodal attention mechanism can also be used to perform weighted fusion of features according to the importance of different data types. The present disclosure does not limit the method of feature fusion.

[0092] In one embodiment of the present disclosure, the multimodal understanding model may be a CLIP model (Contrastive Language-Image Pre-training, a pre-training model based on contrastive text-image pairs).

[0093] The feature extraction of target data information is performed through a multimodal understanding model, thereby improving the accuracy of feature extraction of target data information.

[0094] In one embodiment, generating the target image based on the first extraction result and the second extraction result includes: performing feature alignment on the global feature information in the first extraction result and the global feature information in the second extraction result to obtain target global features; performing feature alignment on the local feature information in the first extraction result and the local feature information in the second extraction result to obtain target local features; and generating the target image based on the target global features and the target local features.

[0095] Global feature alignment is to align features within the entire image to ensure that globally similar information is properly corresponded. In the present disclosure, feature alignment of the global feature information in the first extraction result and the second extraction result can be achieved through a variety of methods. For example, through a feature matching algorithm (for example, a random sampling consensus algorithm), feature points in the first extraction result and the second extraction result are aligned to achieve alignment of the global feature information in the first extraction result and the global feature information in the second extraction result, or by calculating the similarity between the global feature information in the first extraction result and the second extraction result, feature alignment of the global feature information of the first extraction result and the second extraction result is performed according to the similarity value.

[0096] For the alignment of local features, the local feature descriptors in the first extraction result and the second extraction result can be matched and aligned by a feature descriptor matching algorithm (such as nearest neighbor matching, Han Ming distance matching, etc.), so as to achieve the alignment of the local feature information of the first extraction result and the second extraction result. The above-mentioned local feature descriptors are extracted during the feature extraction process. These descriptors can be used to represent the key points in the reference image and the target data information, and spatially correspond to the local structure of the reference image and the target data information. For the alignment of local features, local feature alignment can also be performed based on the color features of the local area in the reference image. By comparing the local color features of similar positions in different reference images, local areas with similar color distributions are found for alignment. For the method of local feature alignment, different feature alignment methods can also be selected according to the specific information of the extracted local features, such as texture information and pattern information.

[0097] Feature alignment aligns the reference images in the reference image set with the information used to determine the target image's content in the target data, improving the accuracy of the generated target image. For example, if the reference images are all clothing images, feature alignment can calculate the correspondence matrix between the clothing images in different reference images, thereby determining the association between the clothing images. Simultaneously, incorporating feature information from the target data effectively unifies the reference images and the target data.

[0098] In one embodiment, generating the target image based on the first extraction result and the second extraction result includes: based on the first extraction result and the second extraction result, performing feature fusion on the first extraction result and the second extraction result through a generative model to obtain image data information; decoding the image data information to generate the target image.

[0099] The first extraction result and the second extraction result are subjected to feature fusion by the generative model. When the generative model performs feature fusion on the first extraction result and the second extraction result, it is usually multi-scale feature fusion. Multi-scale feature fusion is the process of integrating feature information of different scales to obtain a more comprehensive and accurate feature representation. Features of different scales can capture structures and patterns of different levels and sizes in the reference image, so fusing these features can improve the performance of the algorithm. In the feature fusion process, by performing fusion operations such as series connection, addition, or weighted averaging on the scale features, the generative model can generate each pixel of the image based on the fused feature information, thereby obtaining image data information. In some embodiments of the present disclosure, the generative model can be a diffusion model, which can quickly and accurately generate the target image through the diffusion model, thereby improving the efficiency of generating the target image.

[0100] After obtaining the image data information, the image data information is decoded to generate the target image. In some embodiments of the present disclosure, the image data information is a tensor, which is composed of a set of regularly arranged numbers or functions, and can efficiently perform mathematical operations, thereby improving the efficiency of generating the target image.

[0101] In one embodiment, the method further includes: obtaining a sample image set and sample data information, wherein the sample image set includes a reference sample image and a target sample image, and the sample data information is used to determine the content of the target sample image; and training a generative model based on the sample image set and the sample data information.

[0102] Before generating a target image using a generative model, the model must be trained to improve its accuracy. To train a generative model, a sample image set and a sample data set must first be constructed. The sample image set must include reference sample images for generating images, as well as labeled target sample images. The sample data information is used to determine the content of the target sample images. The generative model is trained using the sample image set and sample data information, and its parameters are adjusted to improve its accuracy in generating target images.

[0103] In one embodiment, the training of the generative model based on the sample image set and the sample data information includes: performing feature fusion on features of the reference sample image and features of the sample data information through the generative model to obtain sample image data information; decoding the sample image data information to obtain a generated sample image; determining a model loss based on a difference between the generated sample image and the target sample image; and training the generative model based on the model loss.

[0104] Before training the generative model, the sample image set and sample data information need to be preprocessed. This preprocessing operation includes at least extracting global and local features from the reference sample images in the sample image set, extracting features from the sample data information using a multimodal understanding model, and then aligning the feature information of the reference sample images with the feature information of the sample data information. The aligned feature information is input into the generative model, which then generates sample image data information based on the feature information of the reference sample images and the feature information of the sample data information. The generated sample images are then obtained by image decoding of the sample image data information. Finally, the loss function of the generative model is determined based on the difference between the generated sample images and the target sample images in the sample image set. The generative model is then trained based on the loss function, and the parameters of the generative model are adjusted.

[0105] The sample image set and sample data information are used to train and adjust the parameters of the generative model, thereby improving the accuracy of the generative model in generating target images.

[0106] In one embodiment, the method for generating an image can refer to Figure 2 The flowchart shown. Figure 2 This is a flow chart of a method for generating an image according to an exemplary embodiment of the present disclosure.

[0107] In this embodiment, there are two reference images in the reference image set, namely image A and image B. First, the reference images are analyzed to extract feature information of the reference images. The feature information includes at least global feature information such as structural contour, shape geometry information, element matching, and local feature information such as color, texture, and pattern. At the same time, feature extraction is performed on the reference information through a multimodal understanding model. The reference information is used to determine the style and elements of the target image to be generated, which may be text information, image information, or video information.

[0108] The feature information of the reference image and the reference information are then aligned. This alignment process involves global and local alignment, effectively unifying the features of Image A, Image B, and the reference information, resulting in a target image that incorporates all three characteristics. This aligned feature information is then fed into a generative fusion model, where it is iteratively processed using a multi-layer network model to fuse and optimize the features of Image A, Image B, and the reference information, generating the target image's data information. Finally, image decoding is performed on the target image's data information to obtain the target image.

[0109] Through the image generation method described in the present disclosure, a target image that integrates reference image features and reference information features can be obtained. This image generation method can be applied to multiple scenarios, such as clothing design, home design, etc.

[0110] For the sake of simplicity, the aforementioned method embodiments are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited to the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously.

[0111] Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0112] Corresponding to the aforementioned application function implementation method embodiment, the present disclosure also provides an application function implementation device and a corresponding terminal embodiment.

[0113] Reference Figure 3 According to a block diagram of a device for image generation according to an exemplary embodiment, the device may include:

[0114] A first extraction unit 302 is configured to perform feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image;

[0115] A second extraction unit 304 is configured to perform feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image;

[0116] The generating unit 306 is configured to generate the target image based on the first extraction result and the second extraction result.

[0117] Optionally, the first extraction unit 302 is specifically configured to:

[0118] Performing global feature extraction on the reference images to obtain global feature information of each reference image;

[0119] Performing local feature extraction on the reference images to obtain local feature information of each reference image;

[0120] The first extraction result is obtained based on the global feature information and the local feature information.

[0121] Optionally, the second extraction unit 304 is specifically configured to:

[0122] Performing feature extraction on the target data information according to the data type using a multimodal understanding model to obtain a feature extraction result of at least one data type;

[0123] Feature fusion is performed on the feature extraction result of the at least one data type to obtain the second extraction result.

[0124] Optionally, the generating unit 306 is specifically configured to:

[0125] Performing feature alignment on the global feature information in the first extraction result and the global feature information in the second extraction result to obtain a target global feature;

[0126] Performing feature alignment on the local feature information in the first extraction result and the local feature information in the second extraction result to obtain target local features;

[0127] The target image is generated based on the target global features and the target local features.

[0128] Optionally, the generating unit 306 is specifically configured to:

[0129] Based on the first extraction result and the second extraction result, performing feature fusion on the first extraction result and the second extraction result through a generative model to obtain image data information;

[0130] The image data information is decoded to generate the target image.

[0131] Optionally, the device further includes:

[0132] an acquiring unit, configured to acquire a sample image set and sample data information, wherein the sample image set includes a reference sample image and a target sample image, and the sample data information is used to determine the content of the target sample image;

[0133] A training unit is used to train the generative model based on the sample image set and the sample data information.

[0134] Optional training unit, specifically used for:

[0135] Performing feature fusion on the features of the reference sample image and the features of the sample data information through a fusion model generative model to obtain sample image coding information and image data information;

[0136] Decoding the sample image encoding information and image data information to generate a sample image;

[0137] determining a model loss based on a difference between the generated sample image and the target sample image;

[0138] The fusion model generative model is trained based on the model loss.

[0139] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0140] Accordingly, in one aspect, an embodiment of the present disclosure provides an apparatus for generating an image, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to:

[0141] performing feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image;

[0142] Performing feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image;

[0143] The target image is generated based on the first extraction result and the second extraction result.

[0144] The present disclosure also provides Figure 4 The schematic structure diagram of the electronic device shown in FIG. Figure 4As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the image generation method provided by any embodiment of the present disclosure. Of course, in addition to software implementation, the present disclosure does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0145] The present disclosure also provides a computer-readable storage medium, which stores a computer program. The computer program can be used to execute the image generation method provided in any embodiment of the present specification.

[0146] The present disclosure also provides a computer program product, which can be used to execute the image generation method provided in any embodiment of the present specification.

[0147] Embodiments of the subject matter described in this disclosure may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or to control the operation of the data processing apparatus. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these.

[0148] The processes and logic flows described in this disclosure can be performed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output.

[0149] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to such mass storage devices to receive data from them or to transmit data to them, or both. However, a computer does not necessarily have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0150] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks.

[0151] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0152] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the inventions claimed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0153] It should be understood that the present disclosure is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of this specification is limited only by the appended claims.

[0154] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of this specification.

Claims

1. A method for generating an image, characterized in that: The method comprises: performing feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image; Performing feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image; The target image is generated based on the first extraction result and the second extraction result.

2. The method according to claim 1, characterized in that The step of extracting features from the reference image in the reference image set to obtain a first extraction result includes: Performing global feature extraction on the reference images to obtain global feature information of each reference image; Performing local feature extraction on the reference images to obtain local feature information of each reference image; The first extraction result is obtained based on the global feature information and the local feature information.

3. The method according to claim 1, characterized in that The feature extraction of the target data information to obtain a second extraction result includes: Performing feature extraction on the target data information according to the data type using a multimodal understanding model to obtain a feature extraction result of at least one data type; Feature fusion is performed on the feature extraction result of the at least one data type to obtain the second extraction result.

4. The method according to claim 1, wherein The generating the target image based on the first extraction result and the second extraction result includes: Performing feature alignment on the global feature information in the first extraction result and the global feature information in the second extraction result to obtain a target global feature; Performing feature alignment on the local feature information in the first extraction result and the local feature information in the second extraction result to obtain target local features; The target image is generated based on the target global features and the target local features.

5. The method according to claim 1, wherein The generating the target image based on the first extraction result and the second extraction result includes: Based on the first extraction result and the second extraction result, performing feature fusion on the first extraction result and the second extraction result through a generative model to obtain image data information; The image data information is decoded to generate the target image.

6. The method according to claim 5, characterized in that The method further comprises: Acquire a sample image set and sample data information, wherein the sample image set includes a reference sample image and a target sample image, and the sample data information is used to determine the content of the target sample image; The generative model is trained based on the sample image set and the sample data information.

7. The method according to claim 6, characterized in that The training of the generative model based on the sample image set and the sample data information includes: Performing feature fusion on the features of the reference sample image and the features of the sample data information through a generative model to obtain sample image data information; Decoding the sample image data information to generate a sample image; determining a model loss based on a difference between the generated sample image and the target sample image; The generative model is trained based on the model loss.

8. An image generating device, characterized in that: The device comprises: a first extraction unit, configured to perform feature extraction on a reference image in a reference image set to obtain a first extraction result, wherein the reference image set includes at least one reference image; a second extraction unit, configured to perform feature extraction on the target data information to obtain a second extraction result, wherein the target data information is used to determine the content of the target image; A generating unit is configured to generate the target image based on the first extraction result and the second extraction result.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An image generating device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 7 by running the executable instructions.