Image generation method and device, equipment and medium

By acquiring the description text and the original image and fusing the features to generate the target image, the problem that traditional image generation methods cannot meet the needs of users is solved, and the image is matched with the text and images is achieved, avoiding wasting hardware resources.

CN120107377APending Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311673719.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The images generated by traditional image generation methods often cannot meet the user's generation needs, resulting in wasted hardware resources.

Method used

By obtaining the description text and the original image, extracting the text and image features, fusing the category features with the image features, and generating the target image to match the description text and the original image.

Benefits of technology

The generated image is realized at the same time matched with the original image and description text specified by the user, meeting the user's generation needs, avoiding the waste of hardware resources, and improving the fidelity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107377A_ABST
    Figure CN120107377A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation method and device, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that a description text and a plurality of original images are acquired, the description text comprises a category identification text, and the original images comprise the face of a first object; extracting an original text feature of the description text and an image feature of the original image, wherein the original text feature comprises a category feature of the category identification text; fusing the category features with respective image features of the plurality of original images to obtain a plurality of fused features, and determining comprehensive features according to the fused features; replacing category features in the original text features with comprehensive features to obtain target text features; a target image is generated according to the target text features, the target image comprises the face of the second object, the face of the second object has the face feature of the face of the first object, and the image content of the target image is matched with the description content of the description text. By adopting the method, the waste of hardware resources for generating the image can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and more particularly to the field of image processing, and in particular to an image generation method, device, equipment and medium. Background Art

[0002] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology can be applied to scenarios where images are generated based on text. In traditional technology, images that match the description content of the description text are usually generated directly by inputting the description text.

[0003] However, the images generated by traditional image generation methods often fail to meet the generation requirements of users and the expectations of users, which results in a waste of hardware resources used to support image generation. Summary of the invention

[0004] Based on this, it is necessary to provide an image generation method, apparatus, device and medium that can avoid wasting hardware resources used to support image generation in order to address the above technical problems.

[0005] In a first aspect, the present application provides an image generation method, the method comprising:

[0006] Acquire description text and acquire at least one original image, wherein the description text includes category identification text, and any of the original images includes a face of a first object;

[0007] Extracting original text features of the description text and image features of each of the at least one original image, wherein the original text features include category features of the category identification text;

[0008] Fusing the category feature with the image features of the at least one original image to obtain at least one fused feature, and determining a comprehensive feature based on the at least one fused feature;

[0009] Replacing the category feature in the original text feature with the comprehensive feature to obtain the target text feature;

[0010] A target image is generated according to the target text features, the target image includes a face of a second object, the face of the second object has facial features of the face of the first object, and image content of the target image matches the description content of the description text.

[0011] In a second aspect, the present application provides an image generating device, the device comprising:

[0012] An acquisition module, used for acquiring description text and acquiring at least one original image, wherein the description text includes category identification text, and any of the original images includes a face of a first object;

[0013] An extraction module, used to extract original text features of the description text and image features of each of the at least one original image, wherein the original text features include category features of the category identification text;

[0014] a determination module, configured to fuse the category feature with the image features of the at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature;

[0015] A replacement module, used for replacing the category feature in the original text feature with the comprehensive feature to obtain a target text feature;

[0016] A generation module is used to generate a target image according to the target text features, wherein the target image includes a face of a second object, the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text.

[0017] In one embodiment, the generation module is also used to obtain noise image features extracted from a random noise image; refer to the target text features to reduce the noise image features to obtain reduced-noise image features; and generate a target image based on the reduced-noise image features.

[0018] In one embodiment, the generation module is also used to use the noise image features as the features to be denoised in the first round of denoising, use the first round as the current round, refer to the target text features, denoise the features to be denoised in the current round, and obtain the first intermediate features after the current round of denoising; use the first intermediate features as the features to be denoised in the next round of denoising, use the next round as the current round, return to the step of referring to the target text features, denoise the features to be denoised in the current round, and obtain the first intermediate features after the current round of denoising, and iteratively execute until the denoising stop condition is met, and use the first intermediate features after the last round of denoising as the denoised image features.

[0019] In one embodiment, the generation module is further used to perform a cross-attention operation on the target text features and the features to be denoised in this round through a preset cross-attention mechanism to obtain the first intermediate features after this round of denoising.

[0020] In one embodiment, the generation module is also used to use the features to be denoised in this round as the features to be operated for the first cross-attention operation, and the first time as this time, and through a preset cross-attention mechanism, perform a cross-attention operation on the target text features and the features to be operated this time, to obtain a second intermediate feature after the cross-attention operation; use the second intermediate feature as the features to be operated for the next cross-attention operation, and the next time as this time, return to the step of performing a cross-attention operation on the target text features and the features to be operated this time through the preset cross-attention mechanism, and obtain the second intermediate feature after the cross-attention operation, and iteratively execute until the operation stop condition is met, and use the second intermediate feature after the last round of cross-attention operation as the first intermediate feature after this round of denoising.

[0021] In one embodiment, the noise image feature is obtained by feature encoding the random noise image according to a preset feature encoding method; the generation module is also used to feature decode the denoised image feature according to a feature decoding method matching the feature encoding method to obtain a target image.

[0022] In one embodiment, there are multiple original images, and the determination module is further used to obtain the importance weights of each of the multiple original images; according to the importance weights of each of the multiple original images, the fusion features corresponding to each of the multiple original images are fused to obtain a comprehensive feature; wherein the similarity between the face of the second object and the face of the respective first object in the multiple original images is positively correlated with the importance weights of each of the multiple original images.

[0023] In one embodiment, the determination module is also used to weight the fusion features corresponding to each original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image; and to splice the weighted features corresponding to the multiple original images to obtain comprehensive features.

[0024] In one embodiment, the acquisition module is also used to acquire original description text, which includes original category identification text; when a text editing event for the original description text is triggered, the category identification text indicated by the text editing event is acquired; and the category identification text is used to replace the original category identification text in the original description text to obtain a description text.

[0025] In one embodiment, the extraction module is also used to input the description text into a trained text feature extraction network to extract the original text features of the description text through the text feature extraction network; input the at least one original image into a trained image feature extraction network respectively to extract the image features of each of the at least one original image through the image feature extraction network; the generation module is also used to input the target text features into a generation network to generate a target image through the generation network and based on the target text features, and the generation network is obtained by adversarial learning with the text feature extraction network and the image feature extraction network.

[0026] In one embodiment, the target image is generated by a trained image generation model, and the apparatus further comprises:

[0027] A training module is used to obtain sample description text and at least one sample image, wherein the sample description text includes category identification text, any of the sample images includes the face of a sample object, and the image content of the sample image matches the description content of the sample description text; input the sample description text and the at least one sample image into an image generation model to be trained so as to generate a predicted image through the image generation model to be trained; and train the image generation model to be trained according to the difference between the predicted image and the at least one sample image to obtain a trained image generation model.

[0028] In one embodiment, the training module is also used to obtain at least one original sample image, any of which contains the face of a sample object, and the image content of the original sample image matches the description content of the sample description text; obtain random noise; for each original sample image, fill the background area of ​​the original sample image except the sample object with the random noise to obtain a sample image corresponding to the original sample image.

[0029] In one embodiment, the first object is a first person, the second object is a second person, the category indicated by the category identification text is a person attribute category, and the category identification text is a person attribute category identification text.

[0030] In a third aspect, the present application provides a computer device including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps in the method embodiments of the present application are implemented.

[0031] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the method embodiments of the present application.

[0032] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the steps in the method embodiments of the present application when the computer program is executed by a processor.

[0033] The above-mentioned image generation method, device, equipment and medium obtain a description text and at least one original image, wherein the description text includes a category identification text, and any original image includes a face of a first object. The original text features of the description text and the image features of at least one original image are extracted, and the original text features include the category features of the category identification text. The category features are fused with the image features of at least one original image to obtain at least one fused feature, and a comprehensive feature is determined based on the at least one fused feature, and the category features in the original text features are replaced with the comprehensive features to obtain the target text features. A target image is generated based on the target text features, and the target image includes the face of a second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text. Compared with traditional image generation methods, the present application obtains description text and original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain fused features, and replaces the category features in the original text features with comprehensive features obtained based on the fused features to obtain target text features, and then generates a target image based on the target text features. The generated target image can match the original image and description text specified by the user at the same time, meeting the user's image generation needs. The generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A diagram showing an application environment of an image generation method in an embodiment;

[0035] Figure 2 is a schematic flow chart of an image generating method in one embodiment;

[0036] Figure 3 A schematic diagram of describing the composition of text and original text features in one embodiment;

[0037] Figure 4 A schematic diagram of the principle of generating a target image based on a generating network in one embodiment;

[0038] Figure 5 A schematic diagram of a principle of generating a first intermediate feature based on a noise reduction unit in one embodiment;

[0039] Figure 6 A model architecture diagram of an image generation model in one embodiment;

[0040] Figure 7 A schematic diagram showing a comparison of images generated by the image generation method of the present application and a traditional image generation method, respectively, with reference to a description text and an original image in an embodiment;

[0041] Figure 8 A schematic diagram showing a comparison of images generated by the image generation method of the present application and a traditional image generation method respectively with reference to a description text and an original image in another embodiment;

[0042] Fig. 9 A schematic diagram showing a comparison of images generated by the image generation method of the present application and a traditional image generation method respectively, with reference to a description text and an original image in another embodiment;

[0043] Fig.10 A schematic diagram showing a comparison of images generated by the image generation method of the present application and a traditional image generation method, respectively, with reference to a description text and two original images in an embodiment;

[0044] Fig.11 is a schematic flow chart of an image generating method in another embodiment;

[0045] Fig.12 is a structural block diagram of an image generating device in one embodiment;

[0046] Fig.13 is a structural block diagram of an image generating device in another embodiment;

[0047] Fig.14 is an internal structure diagram of a computer device in one embodiment;

[0048] Fig.15 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] The image generation method provided in this application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can be set up separately and can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other servers. Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptops, smart phones, tablet computers, vehicle-mounted terminals, intelligent voice interaction devices, aircraft, smart home appliances and portable wearable devices. Smart home appliances can be smart speakers, smart TVs and smart air conditioners. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, cloud security, host security and other network security services, CDN, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected by wired or wireless communication, and this application is not limited here.

[0051] The server 104 can obtain the description text and at least one original image from the terminal 102. The description text includes a category identification text, and any original image includes a face of a first object. The server 104 can extract the original text features of the description text and the image features of at least one original image. The original text features include the category features of the category identification text. The server 104 can fuse the category features with the image features of at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature, replace the category features in the original text features with the comprehensive feature, and obtain the target text feature. Furthermore, the server 104 can generate a target image based on the target text features. The target image includes the face of a second object. The face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text. It can be understood that this embodiment is not limited to this. It can be understood that Figure 1 The application scenarios are only for illustration and are not limited to this.

[0052] It should be noted that the image generation method in some embodiments of the present application uses artificial intelligence technology. For example, the original text features and image features are extracted using artificial intelligence technology, and the target image is also generated using artificial intelligence technology.

[0053] In one embodiment, Figure 2As shown, an image generation method is provided, which can be applied to a computer device, and the computer device can be a terminal or a server, that is, the method can be executed by the terminal or the server alone, or can be implemented through interaction between the terminal and the server. This embodiment takes the method applied to a computer device as an example for explanation, and includes the following steps:

[0054] Step 202, obtaining description text and obtaining at least one original image, wherein the description text includes category identification text, and any original image includes a face of a first object.

[0055] The description text is a text used to describe the image to be generated. The category identifier text is a text that carries a category identifier in the description text. The category identifier is used to identify the category to which the text belongs.

[0056] Specifically, the computer device may obtain a description text including a category identification text. It is understood that the category identification text is located in the description text, is one of the components of the description text, and is a subtext in the description text. Also, the computer device may obtain at least one original image. It is understood that the number of original images may be one or more. Each original image includes a face of a first object.

[0057] In one embodiment, the computer device may obtain the original description text, which includes the original category identification text. Furthermore, the computer device may directly use the original description text as the description text, and may directly use the original category identification text included in the original description text as the category identification text included in the description text.

[0058] Step 204: extracting original text features of the description text and image features of at least one original image, wherein the original text features include category features of the category identification text.

[0059] Among them, the original text feature is the text feature extracted from the description text. The category feature is the text feature extracted from the category identification text. The original text feature includes the category feature of the category identification text. It can be understood that the category feature is located in the original text feature, is one of the components of the original text feature, and belongs to the sub-feature of the original text feature.

[0060] Specifically, the computer device may extract the text features of the description text to obtain the original text features of the description text, wherein the original text features include the category features of the category identification text. Figure 3As shown, the description text includes multiple subtexts, namely, subtext 1, subtext 2, subtext 3, subtext 4, ..., subtext n, where n is a positive integer. Each subtext corresponds to its own subtext feature, namely, subtext feature 1, subtext feature 2, subtext feature 3, subtext feature 4, ..., subtext feature n. The original text feature is composed of the subtext features of the multiple subtexts. It can also be understood that the category identification text is one of the subtexts of the description text, and the subtext feature corresponding to the category identification text in the original text feature is the category feature. For example, subtext 2 is the category identification text in the description text, and the subtext feature 2 corresponding to subtext 2 in the original text feature is the category feature. And, the computer device can perform feature extraction on at least one original image respectively to obtain the image features of at least one original image respectively.

[0061] In one embodiment, the computer device may input the description text into a trained text feature extraction network to extract original text features of the description text through the text feature extraction network.

[0062] In one embodiment, the text feature extraction network may be a text encoder. The computer device may input the description text into a trained text encoder so that the text encoder performs feature encoding on the description text to obtain original text features of the description text.

[0063] In one embodiment, the computer device may input at least one original image into a trained image feature extraction network, respectively, so as to extract image features of each of the at least one original image through the image feature extraction network.

[0064] In one embodiment, the image feature extraction network is an image encoder. The computer device can input at least one original image into a trained image encoder to perform feature encoding on the at least one original image through the image encoder to obtain image features of the at least one original image.

[0065] Step 206: fuse the category feature with the image features of at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature.

[0066] Among them, the fusion feature is a feature obtained by fusing the category feature with the image feature. The comprehensive feature is a feature determined based on the fusion feature. It can be understood that each fusion feature in at least one fusion feature is referenced in the determination process of the comprehensive feature. The comprehensive feature has more details than a single fusion feature.

[0067] Specifically, the computer device may fuse the category feature with the image feature of at least one original image to obtain at least one fused feature. It is understood that for each original image, the computer device may fuse the category feature with the image feature of the original image to obtain a fused feature corresponding to the original image. Furthermore, the computer device may determine the comprehensive feature based on the fused features corresponding to at least one original image.

[0068] In one embodiment, for each image feature of the original image, the computer device may determine the feature dimension of the image feature for the original image according to the feature dimension of the category feature, wherein the feature dimension of the image feature for the original image is consistent with the feature dimension of the category feature. Furthermore, the computer device may fuse the image feature of the original image with the category feature of the consistent dimension to obtain the fused feature corresponding to the original image.

[0069] In one embodiment, when the number of original images is one, the computer device may directly use the fusion feature corresponding to the original image as the comprehensive feature.

[0070] In one embodiment, when there are multiple original images, the computer device may fuse the fusion features corresponding to the multiple original images to obtain a comprehensive feature.

[0071] Step 208: replace the category features in the original text features with comprehensive features to obtain target text features.

[0072] The target text feature is a feature obtained by replacing the category feature in the original text feature describing the text with the comprehensive feature.

[0073] Specifically, the computer device may replace the category features included in the original text features with the comprehensive features, and the features other than the category features in the original text features remain unchanged, so as to obtain the target text features.

[0074] In one embodiment, continue to refer to Figure 3 If subtext 2 in the description text is a category identification text and subtext feature 2 in the original text feature is a category feature, the computer device can replace subtext feature 2 in the original text feature with a comprehensive feature, and other subtext features in the original text feature except subtext feature 2 remain unchanged to obtain the target text feature.

[0075] Step 210, generating a target image according to the target text features, the target image including the face of the second object, the face of the second object having the facial features of the first object, and the image content of the target image matching the description content of the description text.

[0076] Specifically, the computer device can generate a target image according to the target text features, wherein the generated target image includes the face of the second object, and the face of the second object has the facial features of the face of the first object in the original image. It can be understood that the face of the second object in the target image is similar to the face of the first object in the original image. And the image content of the target image matches the description content of the description text. For example, if the description text is "a girl wearing a Santa hat", a girl wearing a Santa hat will be presented in the target image.

[0077] In one embodiment, the computer device may input the target text features into a trained generation network to generate a target image based on the target text features through the generation network.

[0078] In the above-mentioned image generation method, by obtaining a description text and obtaining at least one original image, the description text includes a category identification text, and any original image includes a face of a first object. The original text features of the description text and the image features of at least one original image are extracted, and the original text features include the category features of the category identification text. The category features are fused with the image features of at least one original image to obtain at least one fused feature, and a comprehensive feature is determined based on the at least one fused feature, and the category features in the original text features are replaced with the comprehensive features to obtain the target text features. A target image is generated based on the target text features, and the target image includes the face of a second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain the fusion features, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image based on the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meeting the user's image generation needs, and the generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0079] In one embodiment, generating a target image according to target text features includes: obtaining noise image features extracted from a random noise image; reducing noise on the noise image features with reference to the target text features to obtain reduced noise image features; and generating a target image according to the reduced noise image features.

[0080] Among them, the noise image feature is the image feature of the random noise image.

[0081] Specifically, the computer device may obtain a random noise image, and perform feature extraction on the random noise image to obtain noise image features of the random noise image. Furthermore, the computer device may reduce noise on the noise image features with reference to the target text features to obtain reduced noise image features, and generate a target image based on the reduced noise image features. It can be understood that there is a lot of noise in the random noise image, and the noise image features are reduced with reference to the target text features, and the target image generated based on the reduced noise image features has less noise than the random noise image.

[0082] In one embodiment, the computer device may refer to the target text features and perform multi-level noise reduction on the noise image features layer by layer to obtain noise reduction image features.

[0083] In one embodiment, the computer device may convolve the target text feature and the noise image feature to obtain the convolution feature. The computer device may use the convolution feature as the feature to be denoised in the first round of denoising, and use the first round as the current round to denoise the feature to be denoised in the current round to obtain the intermediate feature after the current round of denoising. Furthermore, the computer device may use the intermediate feature obtained in the current round of denoising as the feature to be denoised in the next round of denoising in the current round of denoising, and use the next round as the new current round, and return to the step of denoising the feature to be denoised in the current round to obtain the intermediate feature after the current round of denoising for iterative execution until the denoising stop condition is met, and the computer device may use the intermediate feature after the last round of denoising as the denoised image feature. It can be understood that the process of denoising the feature to be denoised in each round is the process of convolving the feature to be denoised.

[0084] In one embodiment, the computer device may refer to the target text feature and perform a single denoising on the noise image feature to obtain the denoised image feature.

[0085] In the above embodiment, by obtaining noise image features extracted from a random noise image, and denoising the noise image features with reference to target text features to obtain denoised image features, and then generating a target image based on the denoised image features, the image quality can be further improved, thereby further avoiding the waste of hardware resources used to support image generation.

[0086] In one embodiment, the noise image features are denoised with reference to the target text features to obtain the denoised image features, including: using the noise image features as the features to be denoised in the first round of denoising, taking the first round as the current round, denoising the features to be denoised in the current round with reference to the target text features, and obtaining the first intermediate features after the current round of denoising; using the first intermediate features as the features to be denoised in the next round of denoising, taking the next round as the current round, returning to refer to the target text features, denoising the features to be denoised in the current round, and obtaining the first intermediate features after the current round of denoising, the steps are iteratively performed until the denoising stop condition is met, and taking the first intermediate features after the last round of denoising as the denoised image features.

[0087] Specifically, the computer device may use the noise image features as the features to be denoised in the first round of denoising, take the first round as the current round, and refer to the target text features to denoise the features to be denoised in the current round to obtain the first intermediate features after the current round of denoising. Furthermore, the computer device may use the first intermediate features obtained in the current round of denoising as the features to be denoised in the next round of denoising in the current round of denoising, and refer to the next round as the new current round, return to refer to the target text features, denoise the features to be denoised in the current round, and obtain the first intermediate features after the current round of denoising. The steps are iteratively performed until the denoising stop condition is met, and the computer device may use the first intermediate features after the last round of denoising as the denoised image features.

[0088] In one embodiment, Figure 4 As shown, the target image is generated by a trained generative network. The generative network includes multiple denoising units connected layer by layer. The computer device can use the noise image feature as the feature to be denoised in the first round of denoising, take the first round as the current round, and refer to the target text feature, and use the denoising unit of the current round to denoise the feature to be denoised in the current round to obtain the first intermediate feature after the denoising in the current round. Furthermore, the computer device can use the first intermediate feature obtained by the current round of denoising as the feature to be denoised in the next round of denoising in the current round, and use the next round as the new current round, return to the reference target text feature, use the denoising unit of the current round to denoise the feature to be denoised in the current round, and obtain the first intermediate feature after the denoising in the current round. The steps are iteratively executed until the denoising stop condition is met, and the computer device can use the first intermediate feature after the last round of denoising, that is, the output of the last denoising unit in the generative network, as the denoised image feature.

[0089] In one embodiment, the computer device may convolve the target text feature with the feature to be denoised in this round to obtain the first intermediate feature after denoising in this round. It can be understood that the process of convolving the target text feature with the feature to be denoised in this round is the process of denoising the feature to be denoised in this round with reference to the target text feature.

[0090] In one embodiment, the noise reduction stopping condition may be that the number of noise reduction rounds reaches a preset number of noise reduction rounds.

[0091] In the above embodiment, multiple rounds of denoising are performed on the noise image features by referring to the target text features through multiple rounds of iterative denoising. Moreover, each round of denoising refers to the target text features, and when the denoising stop condition is met, the first intermediate feature after the last round of denoising is directly used as the denoised image feature, which can further improve the denoising effect, thereby further improving the quality of the generated image and further avoiding the waste of hardware resources used to support the generated image.

[0092] In one embodiment, the features to be denoised in this round are denoised with reference to the target text features to obtain the first intermediate features after this round of denoising, including: performing a cross-attention operation on the target text features and the features to be denoised in this round through a preset cross-attention mechanism to obtain the first intermediate features after this round of denoising.

[0093] In one embodiment, a preset cross-attention mechanism is used to perform multi-level cross-attention operations on the target text features and the features to be denoised in this round layer by layer to obtain the first intermediate features after this round of denoising.

[0094] In one embodiment, a single cross-attention operation is performed on the target text features and the features to be denoised in this round through a preset cross-attention mechanism to obtain the first intermediate features after this round of denoising.

[0095] In the above embodiment, in each round of denoising, the target text features and the features to be denoised are cross-attended through the cross-attention mechanism to obtain the first intermediate features after each round of denoising, which can improve the effect of each round of denoising, thereby further improving the quality of the generated image and further avoiding the waste of hardware resources used to support the generation of images.

[0096] In one embodiment, a cross-attention operation is performed on the target text features and the features to be operated in this round through a preset cross-attention mechanism to obtain the first intermediate features after this round of noise reduction, including: taking the features to be operated in this round as the features to be operated for the first cross-attention operation, taking the first time as this time, and performing a cross-attention operation on the target text features and the features to be operated this time through a preset cross-attention mechanism to obtain the second intermediate features after this cross-attention operation; taking the second intermediate features as the features to be operated for the next cross-attention operation, taking the next time as this time, returning to the step of performing a cross-attention operation on the target text features and the features to be operated this time through a preset cross-attention mechanism to obtain the second intermediate features after this cross-attention operation, and iteratively executing until the operation stop condition is met, and taking the second intermediate features after the last round of cross-attention operation as the first intermediate features after this round of noise reduction.

[0097] Specifically, the computer device may use the features to be denoised in this round as the features to be operated in the first cross-attention operation, and use the first time as this time, and perform a cross-attention operation on the target text features and the features to be operated in this time through a preset cross-attention mechanism to obtain the second intermediate features after this cross-attention operation. Furthermore, the computer device may use the second intermediate features obtained in this cross-attention operation as the features to be operated in the next cross-attention operation of this cross-attention operation, and use the next time as the new this time, and return to the step of performing a cross-attention operation on the target text features and the features to be operated in this time through a preset cross-attention mechanism to obtain the second intermediate features after this cross-attention operation for iterative execution until the operation stop condition is met, and the computer device may use the second intermediate features after the last round of cross-attention operations as the first intermediate features after this round of denoising.

[0098] In one embodiment, Figure 5 As shown, the target image is generated by a trained generative network. The generative network includes multiple denoising units connected layer by layer, and each denoising unit includes multiple cross-attention layers connected layer by layer. The computer device can use the features to be denoised in this round as the features to be operated in the first cross-attention operation, and use the first time as this time, and use the cross-attention mechanism preset in the cross-attention layer of this time to perform a cross-attention operation on the target text features and the features to be operated in this time to obtain the second intermediate features after the cross-attention operation. Furthermore, the computer device can use the second intermediate features obtained in this cross-attention operation as the features to be operated in the next cross-attention operation of this cross-attention operation, and use the next time as the new time, and return to the step of performing a cross-attention operation on the target text features and the features to be operated in this time through the cross-attention mechanism preset in the cross-attention layer of this time, and obtain the second intermediate features after the cross-attention operation of this time, and perform iterative execution until the operation stop condition is met, and the computer device can use the second intermediate features after the last round of cross-attention operation, that is, the output of the last cross-attention layer, as the first intermediate features after this round of denoising.

[0099] In one embodiment, a cross-attention operation is performed on the target text feature and the feature to be operated this time through a preset cross-attention mechanism to obtain a second intermediate feature after the cross-attention operation, which can be specifically implemented by the following formula:

[0100]

[0101]

[0102] in, For the noise reduction unit, The input of the denoising unit used in this round of denoising, namely the target text features and the features to be denoised in this round of denoising, is the target text feature, , , is the projection matrix, which is the network parameter of the denoising unit. is a criss-cross attention layer for performing the criss-cross attention operation, is the normalization function and d is the number of cross-attention layers in the denoising unit.

[0103] In one embodiment, the operation stop condition may be that the cross-attention operation rounds reach a preset cross-attention operation rounds.

[0104] In the above embodiment, in each round of denoising process, multiple rounds of cross-attention operations are performed on the denoised features and the target text features, which can improve the effect of each round of denoising, thereby further improving the quality of the generated image and further avoiding the waste of hardware resources used to support the generation of images.

[0105] In one embodiment, the noise image feature is obtained by feature encoding the random noise image according to a preset feature encoding method; generating a target image according to the denoised image feature includes: feature decoding the denoised image feature according to a feature decoding method matching the feature encoding method to obtain the target image.

[0106] Specifically, the computer device may acquire a random noise image, and perform feature encoding on the random noise image according to a preset feature encoding method to obtain noise image features. Furthermore, the computer device may reduce noise on the noise image features with reference to the target text features to obtain reduced noise image features, and perform feature decoding on the reduced noise image features according to a feature decoding method matching the feature encoding method to obtain the target image.

[0107] In the above embodiment, the random noise image is feature encoded according to a preset feature encoding method to obtain the noise image feature. Then, the noise reduction image feature is feature decoded according to a feature decoding method matching the feature encoding method to obtain the target image, thereby further improving the quality of the generated image and further avoiding the waste of hardware resources used to support the generated image.

[0108] In one embodiment, there are multiple original images, and a comprehensive feature is determined based on at least one fusion feature, including: obtaining the importance weights of each of the multiple original images; fusing the fusion features corresponding to each of the multiple original images based on the importance weights of each of the multiple original images to obtain a comprehensive feature; wherein the similarity between the face of the second object and the face of the first object in each of the multiple original images is positively correlated with the importance weights of each of the multiple original images.

[0109] Specifically, when there are multiple original images, the computer device can obtain the importance weights of each of the multiple original images, and fuse the corresponding fusion features of each of the multiple original images according to the importance weights of each of the multiple original images to obtain comprehensive features. Furthermore, the computer device can replace the category features in the original text features with comprehensive features to obtain target text features, and generate a target image according to the target text features. The target image includes the face of the second object, and the face of the second object has the facial features of the face of the first object. Moreover, the similarity between the face of the second object in the target image and the face of the first object in each of the multiple original images is positively correlated with the importance weights of each of the multiple original images. And, the image content of the target image matches the description content of the description text.

[0110] In one embodiment, for each original image, the computer device may weight the fusion features corresponding to the original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image. Furthermore, the computer device may convolve the weighted features corresponding to the multiple original images to obtain the comprehensive features.

[0111] In the above embodiment, by obtaining the importance weights of each of the multiple original images, and according to the importance weights of each of the multiple original images, the fusion features corresponding to each of the multiple original images are fused to obtain the comprehensive features. Then, the target image is generated based on the comprehensive features, and the similarity between the face of the second object in the target image and the face of the first object in each of the multiple original images is positively correlated with the importance weights of each of the multiple original images. In this way, by customizing the importance weights of each of the multiple original images, the facial features of the second object in the generated target image can be flexibly controlled, thereby further meeting the image generation needs of the user and further avoiding the waste of hardware resources used to support image generation.

[0112] In one embodiment, according to the importance weights of each of the multiple original images, the fused features corresponding to the multiple original images are fused to obtain a comprehensive feature, including: for each original image, according to the importance weight of the original image, the fused features corresponding to the original image are weighted to obtain the weighted features corresponding to the original image; the weighted features corresponding to the multiple original images are spliced ​​to obtain a comprehensive feature.

[0113] Specifically, for each original image, the computer device may weight the fusion features corresponding to the original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image. Furthermore, the computer device may splice the weighted features corresponding to the multiple original images to obtain the comprehensive features. It can be understood that the spliced ​​comprehensive features have richer features than a single weighted feature, while also retaining the weighted features corresponding to the multiple original images.

[0114] In the above embodiment, for each original image, the fusion features corresponding to the original image are weighted according to the importance weight of the original image to obtain the weighted features corresponding to the original image, and the weighted features corresponding to multiple original images are spliced ​​to obtain comprehensive features. Compared with a single weighted feature, the comprehensive feature can have richer features while retaining the weighted features corresponding to multiple original images, thereby further meeting the user's image generation needs and further avoiding the waste of hardware resources used to support image generation.

[0115] In one embodiment, obtaining description text includes: obtaining original description text, the original description text including original category identification text; when a text editing event for the original description text is triggered, obtaining the category identification text indicated by the text editing event; and replacing the original category identification text in the original description text with the category identification text to obtain the description text.

[0116] The original category identification text is the category identification text contained in the original description text.

[0117] Specifically, the computer device may obtain the original description text, wherein the original description text includes the original category identification text. The computer device may monitor the text editing event in real time. When it is detected that the text editing event for the original description text is triggered, the computer device may obtain the category identification text indicated by the text editing event, and replace the original category identification text in the original description text with the category identification text indicated by the text editing event to obtain the description text.

[0118] For example, if the original category identification text is "a boy wearing a Santa hat", where "boy" is the original category identification text, and the category identification text indicated by the text editing event is "girl", the computer device can replace the "boy" in "a boy wearing a Santa hat" with "girl" to obtain the description text, namely "a girl wearing a Santa hat".

[0119] In the above embodiment, when a text editing event is triggered for the original description text, the category identification text indicated by the text editing event is obtained, and the category identification text is used to replace the original category identification text in the original description text to obtain the description text. The corresponding target image can be quickly generated by simply replacing the category identification text in the description text, thereby improving the image generation efficiency.

[0120] In one embodiment, extracting original text features of a description text and image features of each of at least one original image comprises: inputting the description text into a trained text feature extraction network to extract original text features of the description text through the text feature extraction network; inputting at least one original image into a trained image feature extraction network to extract image features of each of at least one original image through the image feature extraction network; generating a target image according to target text features, comprises: inputting the target text features into a generation network to generate a target image through the generation network and based on the target text features, wherein the generation network is obtained by adversarial learning with the text feature extraction network and the image feature extraction network.

[0121] Specifically, the computer device may obtain a description text and at least one original image, wherein the description text includes a category identification text, and any original image includes a face of a first object. The computer device may input the description text into a trained text feature extraction network to extract the original text features of the description text through the text feature extraction network, and input at least one original image into a trained image feature extraction network to extract the image features of each of the at least one original image through the image feature extraction network. The computer device may fuse the category features with the image features of each of the at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature, and replace the category features in the original text features with the comprehensive feature to obtain the target text feature. Furthermore, the computer device may input the target text feature into a generation network to generate a target image through the generation network and based on the target text feature. The generated target image includes the face of the second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text. Wherein, the generation network is obtained by adversarial learning with the text feature extraction network and the image feature extraction network.

[0122] In one embodiment, the trained image generation model includes a text feature extraction network, an image feature extraction network, and a generation network. It can be understood that the original text features of the description text are extracted from the description text by the text feature extraction network in the trained image generation model. The image features of the original image are extracted from the original image by the image feature extraction network in the trained image generation model. The target image is generated based on the target text features by the generation network in the trained image generation model.

[0123] In the above embodiment, by inputting the description text into a trained text feature extraction network and extracting the original text features of the description text, the accuracy of the original text features can be improved. By inputting at least one original image into a trained image feature extraction network and extracting the image features of at least one original image, the accuracy of the image features can be improved. And, by inputting the target text features into a generation network to generate a target image, the image quality can be further improved, thereby further meeting the image generation needs of the user and further avoiding the waste of hardware resources used to support image generation.

[0124] In one embodiment, the target image is generated by a trained image generation model, and the method also includes a model training step, which includes: obtaining a sample description text and obtaining at least one sample image, wherein the sample description text includes a category identification text, any sample image includes the face of a sample object, and the image content of the sample image matches the description content of the sample description text; inputting the sample description text and at least one sample image into the image generation model to be trained to generate a predicted image through the image generation model to be trained; and training the image generation model to be trained according to the difference between the predicted image and at least one sample image to obtain a trained image generation model.

[0125] Specifically, the computer device may obtain a sample description text and at least one sample image, wherein the sample description text includes a category identification text, any sample image includes the face of a sample object, and the image content of the sample image matches the description content of the sample description text. It can be understood that the category identification text set to which the category identification text included in the sample description text used in the model training phase belongs is consistent with the category identification text set to which the category identification text included in the description text used in the model use phase belongs. The computer device may input the sample description text and at least one sample image into the image generation model to be trained to generate a predicted image through the image generation model to be trained. The computer device may determine a loss value based on the difference between the predicted image and at least one sample image, and iteratively train the image generation model to be trained in the direction of reducing the loss value until the iteration stop condition is met to obtain a trained image generation model.

[0126] In one embodiment, the iteration stopping condition may be that the number of training iterations reaches a preset number of training iterations, or that the loss value reaches a preset loss value.

[0127] In one embodiment, the computer device may obtain at least one original sample image, any original sample image includes a face of a sample object, and the image content of the original sample image matches the description content of the sample description text. The computer device may directly use the at least one original sample image as the at least one sample image.

[0128] In the above embodiment, by inputting sample description text and at least one sample image into the image generation model to be trained, so as to generate a predicted image through the image generation model to be trained, and training the image generation model to be trained according to the difference between the predicted image and at least one sample image, to obtain the trained image generation model, the generation capability of the trained image generation model can be improved, thereby further meeting the image generation needs of users and further avoiding the waste of hardware resources used to support image generation.

[0129] In one embodiment, obtaining at least one sample image includes: obtaining at least one original sample image, any original sample image containing a face of a sample object, and the image content of the original sample image matches the description content of the sample description text; obtaining random noise; for each original sample image, filling the background area of ​​the original sample image except the sample object with random noise to obtain a sample image corresponding to the original sample image.

[0130] Specifically, the computer device may obtain at least one original sample image, any original sample image contains a face of a sample object, and the image content of the original sample image matches the description content of the sample description text. The computer device may obtain random noise, and for each original sample image, the computer device may determine the area in the original sample image other than the sample object to obtain the background area, and fill the background area in the original sample image with random noise to obtain the sample image corresponding to the original sample image.

[0131] In the above embodiment, for each original sample image, the background area except the sample object in the original sample image is filled with random noise to obtain a sample image corresponding to the original sample image, and the model is trained through the obtained sample image, so that the characteristics of the sample object can be better learned during the model training process, which can further enhance the generation capability of the image generation model, thereby further meeting the image generation needs of users and further avoiding the waste of hardware resources used to support image generation.

[0132] In one embodiment, the first object is a first person, the second object is a second person, the category indicated by the category identification text is a person attribute category, and the category identification text is a person attribute category identification text.

[0133] Specifically, the computer device may obtain a description text and at least one original image, the description text includes a character attribute category identification text, and any original image includes a face of a first character. The computer device may extract the original text features of the description text and the image features of at least one original image, the original text features include the category features of the character attribute category identification text. The computer device may fuse the category features with the image features of at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature, and replace the category features in the original text features with the comprehensive feature to obtain the target text feature. The computer device may generate a target image based on the target text feature, the target image includes the face of a second character, the face of the second character has the facial features of the face of the first character, and the image content of the target image matches the description content of the description text.

[0134] In one embodiment, the character attribute category includes at least one of gender or age. For example, the character attribute category identification text can be at least one of "man", "woman", "boy" or "girl".

[0135] In the above embodiment, by obtaining a description text and obtaining at least one original image, the description text includes a character attribute category identification text, and any original image includes a face of a first character. The original text features of the description text and the image features of at least one original image are extracted, and the original text features include the category features of the character attribute category identification text. The category features are fused with the image features of at least one original image to obtain at least one fused feature, and a comprehensive feature is determined based on the at least one fused feature, and the category features in the original text features are replaced with the comprehensive features to obtain the target text features. A target image is generated based on the target text features, and the target image includes the face of a second character, and the face of the second character has the facial features of the face of the first character, and the image content of the target image matches the description content of the description text. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the character attribute category identification text in the description text to obtain the fusion features, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image based on the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meeting the user's image generation needs, and the generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0136] In one embodiment, Figure 6As shown, the target image is generated by a trained image generation model, and the image generation model includes a text feature extraction network, an image feature extraction network and a generation network. Specifically, the computer device can obtain a description text and at least one original image, wherein the description text includes a category identification text, and any original image includes a face of a first object. The computer device can input the description text into the text feature extraction network to extract the original text features of the description text through the text feature extraction network, and input at least one original image into the image feature extraction network respectively to extract the image features of each of the at least one original image through the image feature extraction network. The computer device can fuse the category feature with the image features of each of the at least one original image respectively to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature, and replace the category feature in the original text feature with the comprehensive feature to obtain the target text feature. Furthermore, the computer device can input the target text feature into the generation network to generate a target image through the generation network and based on the target text feature. The generated target image includes the face of the second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text.

[0137] Understandable, continue to refer to Figure 6 During the training of the image generation model, the sample image after the background area other than the sample object is filled with random noise can be used for training. This allows the image generation model to pay more attention to the sample object during the training process, so that the trained image generation model has better image generation capabilities, which can further meet the image generation needs of users and further avoid the waste of hardware resources used to support image generation. It should be noted that during the inference process, the trained image generation model can directly use the original image after the background area other than the first object is filled with random noise to generate an image to obtain a target image.

[0138] In one embodiment, in order to further illustrate that the image generation method of the present application is superior to the traditional image generation method, four different traditional methods and the image generation method of the present application are used to generate target images respectively. Figure 7As shown, description text 1 is "a man in a space suit", description text 2 is "a man wearing a Santa hat", description text 3 is "a woman sitting on the beach with a purple sunset", description text 4 is "a woman with red hair", and description text 5 is "a woman wearing a doctoral hat". The number of original images is one. The images generated by the four traditional image generation methods often cannot meet the generation needs of users, the fidelity of the images is low, and the generated images cannot meet the user's expectations. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain the fusion feature, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion feature to obtain the target text feature, and then generates the target image according to the target text feature, so that the generated target image can match the original image and description text specified by the user at the same time, meet the user's generation needs for the image, and the generated image meets the user's expectations. Moreover, since the image features of the original image are fused in the target text feature, the target image generated based on the target text feature is similar to the original image, which improves the fidelity of the image.

[0139] In one embodiment, in order to further illustrate that the image generation method of the present application is superior to the traditional image generation method, two different traditional methods and the image generation method of the present application are used to generate target images respectively, such as Figure 8 As shown, the description texts are "a man sitting in front of a computer programming" and "a man driving a spaceship" respectively. The number of original images is one. The images generated by the two traditional image generation methods often cannot meet the generation needs of users, the image fidelity is low, and the generated images cannot meet user expectations. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain a fusion feature, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image based on the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meet the user's image generation needs, and the generated image meets the user's expectations. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0140] In one embodiment, in order to further illustrate that the image generation method of the present application is superior to the traditional image generation method, two different traditional methods and the image generation method of the present application are used to generate target images respectively, such as Fig. 9As shown, the description texts are "a woman smiling at the camera" and "a girl wearing a Santa hat", respectively. The number of original images is one. The images generated by the two traditional image generation methods often cannot meet the generation needs of users, the image fidelity is low, and the generated images cannot meet user expectations. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain a fusion feature, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image based on the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meet the user's image generation needs, and the generated image meets the user's expectations. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0141] In one embodiment, in order to further illustrate that the image generation method of the present application is superior to the traditional image generation method, two different traditional methods and the image generation method of the present application are used to generate target images respectively, such as Fig.10 As shown, the description texts are "a man holding a bottle of red wine", "a woman frowning at the camera" and "a man in a space suit". The number of original images is two. The images generated by the two traditional image generation methods often cannot meet the generation needs of users, the image fidelity is low, and the generated images cannot meet user expectations. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain a fusion feature, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image according to the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meet the user's image generation needs, and the generated image meets the user's expectations. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0142] like Fig.11 As shown, in one embodiment, an image generation method is provided, which can be applied to a computer device, which can be a terminal or a server, that is, the method can be executed by the terminal or the server alone, or can be implemented through interaction between the terminal and the server. This embodiment is described by taking the method applied to a computer device as an example, and the method specifically includes the following steps:

[0143] Step 1102, obtaining the original description text, where the original description text includes the original category identification text.

[0144] Step 1104: when a text editing event for the original description text is triggered, the category identification text indicated by the text editing event is obtained.

[0145] Step 1106, the original category identification text in the original description text is replaced with the category identification text to obtain a description text, wherein the description text includes the category identification text.

[0146] Step 1108: Acquire multiple original images, any original image containing a face of a first object.

[0147] Step 1110 , extracting original text features of the description text and image features of each of the multiple original images, wherein the original text features include category features of the category identification text.

[0148] Step 1112, the category feature is fused with the image features of each of the multiple original images to obtain multiple fused features.

[0149] Step 1114, obtaining the importance weight of each of the multiple original images.

[0150] Step 1116: for each original image, weight the fusion features corresponding to the original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image.

[0151] Step 1118, concatenating the weighted features corresponding to the multiple original images to obtain a comprehensive feature, wherein the similarity between the face of the second object and the face of the first object in each of the multiple original images is positively correlated with the importance weights of the multiple original images.

[0152] Step 1120, replacing the category features in the original text features with the comprehensive features to obtain the target text features.

[0153] Step 1122, obtaining noise image features obtained by performing feature encoding on the random noise image according to a preset feature encoding method.

[0154] In step 1124, the noise image features are used as the features to be denoised in the first round of denoising, and the first round is used as this round. Through a preset cross-attention mechanism, the target text features and the features to be denoised in this round are cross-attentioned to obtain the first intermediate features after this round of denoising.

[0155] Step 1126, taking the first intermediate feature as the feature to be denoised for the next round of denoising, taking the next round as this round, returning to the preset cross-attention mechanism, performing a cross-attention operation on the target text feature and the feature to be denoised in this round, and obtaining the first intermediate feature after this round of denoising, and performing the step of iteratively executing until the denoising stop condition is met, and taking the first intermediate feature after the last round of denoising as the denoised image feature.

[0156] Step 1128, feature decoding is performed on the denoised image features according to a feature decoding method that matches the feature encoding method to obtain a target image, wherein the target image includes the face of the second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text.

[0157] The present application also provides an application scenario, which applies the above-mentioned image generation method. Specifically, the image generation method can be applied to the scenario of human image generation. It can be understood that the first object is a first person, the second object is a second person, the category indicated by the category identification text is a person attribute category, and the category identification text is a person attribute category identification text. Specifically, the computer device can obtain an original description text, and the original description text contains the original person attribute category identification text. When a text editing event for the original description text is triggered, the person attribute category identification text indicated by the text editing event is obtained. The original person attribute category identification text in the original description text is replaced with the person attribute category identification text to obtain a description text, and the description text contains the person attribute category identification text. A plurality of original images are obtained, and any of the original images contains a face of a first person. The original text features of the description text and the image features of each of the plurality of original images are extracted, and the original text features contain the category features of the person attribute category identification text. The category features are respectively fused with the image features of each of the plurality of original images to obtain a plurality of fused features. The importance weights of each of the plurality of original images are obtained. For each original image, the fusion features corresponding to the original image are weighted according to the importance weight of the original image to obtain the weighted features corresponding to the original image. The weighted features corresponding to the multiple original images are concatenated to obtain the comprehensive features, wherein the similarity between the face of the second person and the face of the first person in each of the multiple original images is positively correlated with the importance weight of each of the multiple original images. The category features in the original text features are replaced with the comprehensive features to obtain the target text features.

[0158] The computer device can obtain the noise image features obtained by feature encoding the random noise image according to the preset feature encoding method. The noise image features are used as the features to be denoised in the first round of denoising, and the first round is used as the current round. Through the preset cross-attention mechanism, the target text features and the features to be denoised in the current round are cross-attention operated to obtain the first intermediate features after the current round of denoising. The first intermediate features are used as the features to be denoised in the next round of denoising, and the next round is used as the current round. The step of performing the cross-attention operation on the target text features and the features to be denoised in the current round through the preset cross-attention mechanism to obtain the first intermediate features after the current round of denoising is iteratively executed until the denoising stop condition is met, and the first intermediate features after the last round of denoising are used as the denoised image features. According to the feature decoding method matching the feature encoding method, the denoised image features are feature decoded to obtain the target image, the target image includes the face of the second person, the face of the second person has the facial features of the face of the first person, and the image content of the target image matches the description content of the description text.

[0159] Applying the image generation method of the present application to the scene of human image generation can achieve the following beneficial effects: by obtaining the description text and the original image, the image features of the original image are fused with the category features of the character attribute category identification text in the description text to obtain the fusion features, and the category features in the original text features are replaced with the comprehensive features obtained based on the fusion features to obtain the target text features, and then the target image is generated based on the target text features, so that the generated target image can match the original image and the description text specified by the user at the same time, meeting the user's image generation requirements, and the generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0160] The present application also provides an application scenario, which applies the above-mentioned image generation method. Specifically, the image generation method can be applied to the scene of animal image generation. It can be understood that the first object is a first animal, the second object is a second animal, the category indicated by the category identification text is an animal attribute category, and the category identification text is an animal attribute category identification text. Specifically, the computer device can obtain an original description text, and the original description text contains the original animal attribute category identification text. When a text editing event for the original description text is triggered, the animal attribute category identification text indicated by the text editing event is obtained. The original animal attribute category identification text in the original description text is replaced with the animal attribute category identification text to obtain a description text, and the description text contains the animal attribute category identification text. A plurality of original images are obtained, and any original image contains a face of a first animal. The original text features of the description text and the image features of each of the plurality of original images are extracted, and the original text features contain the category features of the animal attribute category identification text. The category features are respectively fused with the image features of each of the plurality of original images to obtain a plurality of fused features. The importance weights of each of the plurality of original images are obtained. For each original image, the fusion features corresponding to the original image are weighted according to the importance weight of the original image to obtain the weighted features corresponding to the original image. The weighted features corresponding to the multiple original images are concatenated to obtain the comprehensive features, wherein the similarity between the face of the second animal and the face of the first animal in the multiple original images is positively correlated with the importance weight of the multiple original images. The category features in the original text features are replaced with the comprehensive features to obtain the target text features.

[0161] The computer device can obtain the noise image features obtained by feature encoding the random noise image according to the preset feature encoding method. The noise image features are used as the features to be denoised in the first round of denoising, and the first round is used as the current round. Through the preset cross-attention mechanism, the target text features and the features to be denoised in the current round are cross-attention operated to obtain the first intermediate features after the current round of denoising. The first intermediate features are used as the features to be denoised in the next round of denoising, and the next round is used as the current round. The step of performing the cross-attention operation on the target text features and the features to be denoised in the current round through the preset cross-attention mechanism to obtain the first intermediate features after the current round of denoising is iteratively executed until the denoising stop condition is met, and the first intermediate features after the last round of denoising are used as the denoised image features. According to the feature decoding method matching the feature encoding method, the denoised image features are feature decoded to obtain the target image, the target image includes the face of the second animal, the face of the second animal has the facial features of the face of the first animal, and the image content of the target image matches the description content of the description text.

[0162] Applying the image generation method of the present application to the scene of animal image generation can achieve the following beneficial effects: by obtaining the description text and the original image, the image features of the original image are fused with the category features of the animal attribute category identification text in the description text to obtain the fused features, and the category features in the original text features are replaced with the comprehensive features obtained based on the fused features to obtain the target text features, and then the target image is generated based on the target text features, so that the generated target image can be matched with the original image and the description text specified by the user at the same time, meeting the user's image generation requirements, and the generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. Moreover, since the target text features are fused with the image features of the original image, the target image generated based on the target text features is similar to the original image, thereby improving the fidelity of the image.

[0163] It should be understood that, although each step in the flow chart of the above-mentioned embodiments is shown in order, these steps are not necessarily performed in order. Unless there is a clear explanation in this article, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned embodiments may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in order, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0164] In one embodiment, Fig.12 As shown, an image generating device 1200 is provided, and the device specifically includes:

[0165] The acquisition module 1202 is used to acquire a description text and a plurality of original images, wherein the description text includes a category identification text, and any original image includes a face of a first object;

[0166] An extraction module 1204 is used to extract original text features of the description text and image features of each of the multiple original images, wherein the original text features include category features of the category identification text;

[0167] A determination module 1206 is used to fuse the category feature with the image features of each of the multiple original images to obtain multiple fused features, and determine the comprehensive feature based on the multiple fused features;

[0168] A replacement module 1208 is used to replace the category features in the original text features with the comprehensive features to obtain the target text features;

[0169] The generating module 1210 is used to generate a target image according to the target text features, wherein the target image includes the face of the second object, the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text.

[0170] In one embodiment, the generation module 1210 is also used to obtain noise image features extracted from the random noise image; reduce noise on the noise image features with reference to the target text features to obtain reduced noise image features; and generate the target image based on the reduced noise image features.

[0171] In one embodiment, the generation module 1210 is also used to use the noise image features as the features to be denoised in the first round of denoising, use the first round as the current round, refer to the target text features, denoise the features to be denoised in the current round, and obtain the first intermediate features after the current round of denoising; use the first intermediate features as the features to be denoised in the next round of denoising, use the next round as the current round, return to refer to the target text features, denoise the features to be denoised in the current round, and obtain the first intermediate features after the current round of denoising, and the steps are iteratively performed until the denoising stop condition is met, and use the first intermediate features after the last round of denoising as the denoised image features.

[0172] In one embodiment, the generation module 1210 is further used to perform a cross-attention operation on the target text features and the features to be denoised in this round through a preset cross-attention mechanism to obtain the first intermediate features after denoising in this round.

[0173] In one embodiment, the generation module 1210 is also used to use the features to be denoised in this round as the features to be operated for the first cross-attention operation, and the first time as this time, and perform a cross-attention operation on the target text features and the features to be operated this time through a preset cross-attention mechanism to obtain a second intermediate feature after the cross-attention operation; use the second intermediate feature as the features to be operated for the next cross-attention operation, and the next time as this time, and return to the step of performing a cross-attention operation on the target text features and the features to be operated this time through a preset cross-attention mechanism to obtain the second intermediate feature after the cross-attention operation, and iteratively execute until the operation stop condition is met, and use the second intermediate feature after the last round of cross-attention operation as the first intermediate feature after this round of denoising.

[0174] In one embodiment, the noise image features are obtained by feature encoding the random noise image according to a preset feature encoding method; the generation module 1210 is also used to feature decode the noise reduction image features according to a feature decoding method matching the feature encoding method to obtain the target image.

[0175] In one embodiment, there are multiple original images, and the determination module 1206 is further used to obtain the importance weights of each of the multiple original images; according to the importance weights of each of the multiple original images, the fusion features corresponding to each of the multiple original images are fused to obtain a comprehensive feature; wherein the similarity between the face of the second object and the face of the first object in each of the multiple original images is positively correlated with the importance weights of each of the multiple original images.

[0176] In one embodiment, the determination module 1206 is also used to weight the fusion features corresponding to each original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image; and to splice the weighted features corresponding to multiple original images to obtain comprehensive features.

[0177] In one embodiment, the acquisition module 1202 is also used to obtain the original description text, which includes the original category identification text; when a text editing event for the original description text is triggered, the category identification text indicated by the text editing event is obtained; and the category identification text is used to replace the original category identification text in the original description text to obtain the description text.

[0178] In one embodiment, the extraction module 1204 is also used to input the description text into a trained text feature extraction network to extract the original text features of the description text through the text feature extraction network; input multiple original images into the trained image feature extraction network respectively to extract the image features of each of the multiple original images through the image feature extraction network; the generation module 1210 is also used to input the target text features into the generation network to generate the target image through the generation network and based on the target text features, and the generation network is obtained by adversarial learning with the text feature extraction network and the image feature extraction network.

[0179] In one embodiment, the target image is generated by a trained image generation model, such as Fig.13 As shown, the image generating device 1200 further includes:

[0180] The training module 1212 is used to obtain a sample description text and a plurality of sample images, wherein the sample description text includes a category identification text, any sample image includes the face of a sample object, and the image content of the sample image matches the description content of the sample description text; the sample description text and the plurality of sample images are input into the image generation model to be trained so as to generate a predicted image through the image generation model to be trained; and the image generation model to be trained is trained according to the difference between the predicted image and the plurality of sample images to obtain a trained image generation model.

[0181] In one embodiment, the training module 1212 is also used to obtain multiple original sample images, any original sample image contains the face of a sample object, and the image content of the original sample image matches the description content of the sample description text; obtain random noise; for each original sample image, fill the background area of ​​the original sample image except the sample object with random noise to obtain a sample image corresponding to the original sample image.

[0182] In one embodiment, the first object is a first person, the second object is a second person, the category indicated by the category identification text is a person attribute category, and the category identification text is a person attribute category identification text.

[0183] The image generation device obtains a description text and a plurality of original images, wherein the description text includes a category identification text, and any of the original images includes a face of a first object. The original text features of the description text and the image features of each of the plurality of original images are extracted, wherein the original text features include the category features of the category identification text. The category features are fused with the image features of each of the plurality of original images to obtain a plurality of fused features, and a comprehensive feature is determined based on the plurality of fused features, and the category features in the original text features are replaced with the comprehensive features to obtain the target text features. A target image is generated based on the target text features, wherein the target image includes the face of a second object, and the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text. Compared with the traditional image generation method, the present application obtains the description text and the original image, fuses the image features of the original image with the category features of the category identification text in the description text to obtain the fusion features, and replaces the category features in the original text features with the comprehensive features obtained based on the fusion features to obtain the target text features, and then generates the target image based on the target text features, so that the generated target image can match the original image and description text specified by the user at the same time, meeting the user's image generation needs, and the generated image meets the user's expectations, thereby avoiding the waste of hardware resources used to support image generation. Moreover, since the image features of the original image are fused in the target text features, the target image generated based on the target text features is similar to the original image, which improves the fidelity of the image.

[0184] Each module in the above-mentioned image generation device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module above.

[0185] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.14 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image generation method is implemented.

[0186] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.15 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an image generation method is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.

[0187] Those skilled in the art will understand that Fig.14 and Fig.15The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0188] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0189] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0190] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0191] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0192] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0193] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0194] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for generating an image, It is characterized in that The method comprises: Acquire description text and acquire at least one original image, wherein the description text includes category identification text, and any of the original images includes a face of a first object; Extracting original text features of the description text and image features of each of the at least one original image, wherein the original text features include category features of the category identification text; Fusing the category feature with the image features of the at least one original image to obtain at least one fused feature, and determining a comprehensive feature based on the at least one fused feature; Replacing the category feature in the original text feature with the comprehensive feature to obtain the target text feature; A target image is generated according to the target text features, the target image includes a face of a second object, the face of the second object has facial features of the face of the first object, and image content of the target image matches the description content of the description text.

2. The method according to claim 1, It is characterized in that The step of generating a target image according to the target text feature comprises: Obtaining noise image features extracted from the random noise image; Referring to the target text feature, denoising the noise image feature to obtain a denoised image feature; A target image is generated according to the denoised image features.

3. The method according to claim 2, It is characterized in that The step of reducing the noise image feature by referring to the target text feature to obtain the reduced noise image feature comprises: The noise image features are used as the features to be denoised in the first round of denoising, the first round is used as the current round, and the features to be denoised in the current round are denoised with reference to the target text features to obtain the first intermediate features after the current round of denoising; The first intermediate feature is used as the feature to be denoised in the next round of denoising, and the next round is used as the current round. Return to the reference to the target text feature, denoise the feature to be denoised in the current round, and obtain the first intermediate feature after the current round of denoising. The step is iteratively performed until the denoising stop condition is met, and the first intermediate feature after the last round of denoising is used as the denoised image feature.

4. The method according to claim 3, It is characterized in that The step of denoising the features to be denoised in this round by referring to the target text features to obtain the first intermediate features after denoising in this round includes: Through a preset cross-attention mechanism, the target text features and the features to be denoised in this round are cross-attended to obtain the first intermediate features after this round of denoising.

5. The method according to claim 4, It is characterized in that The cross-attention operation is performed on the target text features and the features to be denoised in this round by a preset cross-attention mechanism to obtain the first intermediate features after denoising in this round, including: The features to be denoised in this round are used as the features to be operated in the first cross-attention operation, and the first time is used as this time. Through the preset cross-attention mechanism, the target text features and the features to be operated this time are cross-attention operated to obtain the second intermediate features after this cross-attention operation; The second intermediate feature is used as the feature to be operated for the next cross-attention operation, and the next time is used as this time. The preset cross-attention mechanism is returned to, and the target text feature and the feature to be operated this time are cross-attention operated. The step of obtaining the second intermediate feature after this cross-attention operation is iteratively executed until the operation stop condition is met, and the second intermediate feature after the last round of cross-attention operation is used as the first intermediate feature after this round of denoising.

6. The method according to claim 2, It is characterized in that The noise image feature is obtained by performing feature encoding on the random noise image according to a preset feature encoding method; The step of generating a target image according to the noise reduction image features comprises: The denoised image features are feature decoded according to a feature decoding method that matches the feature encoding method to obtain a target image.

7. The method according to any one of claims 1 to 6, It is characterized in that The number of the original images is multiple, and the determining of the comprehensive feature according to the at least one fusion feature includes: Obtain the importance weights of multiple original images; According to the importance weights of the multiple original images, the fusion features corresponding to the multiple original images are fused to obtain a comprehensive feature; The similarity between the face of the second object and the faces of the first objects in the multiple original images is positively correlated with the importance weights of the multiple original images.

8. The method according to claim 7, It is characterized in that The step of fusing the fusion features corresponding to the multiple original images according to their respective importance weights to obtain comprehensive features includes: For each original image, weighting the fusion features corresponding to the original image according to the importance weight of the original image to obtain the weighted features corresponding to the original image; The weighted features corresponding to the multiple original images are concatenated to obtain comprehensive features.

9. The method according to any one of claims 1 to 6, It is characterized in that The obtaining of the description text includes: Obtaining original description text, wherein the original description text includes original category identification text; When a text editing event for the original description text is triggered, obtaining a category identification text indicated by the text editing event; The category identification text is used to replace the original category identification text in the original description text to obtain a description text.

10. The method according to any one of claims 1 to 6, It is characterized in that The extracting of original text features of the description text and image features of each of the at least one original image comprises: Inputting the description text into a trained text feature extraction network to extract original text features of the description text through the text feature extraction network; Inputting the at least one original image into a trained image feature extraction network respectively, so as to extract image features of each of the at least one original image through the image feature extraction network; The step of generating a target image according to the target text feature comprises: The target text features are input into a generation network to generate a target image through the generation network and based on the target text features. The generation network is obtained by adversarial learning with the text feature extraction network and the image feature extraction network.

11. The method according to any one of claims 1 to 6, It is characterized in that The target image is generated by a trained image generation model, and the method further includes a model training step, which includes: Acquire a sample description text and acquire at least one sample image, wherein the sample description text includes a category identification text, any of the sample images includes a face of a sample object, and the image content of the sample image matches the description content of the sample description text; Inputting the sample description text and the at least one sample image into the image generation model to be trained, so as to generate a predicted image through the image generation model to be trained; The image generation model to be trained is trained according to the difference between the predicted image and the at least one sample image to obtain a trained image generation model.

12. The method according to claim 11, It is characterized in that The acquiring of at least one sample image comprises: Acquire at least one original sample image, any of the original sample images contains a face of a sample object, and the image content of the original sample image matches the description content of the sample description text; Get random noise; For each original sample image, the background area except the sample object in the original sample image is filled with the random noise to obtain a sample image corresponding to the original sample image.

13. The method according to any one of claims 1 to 6, It is characterized in that The first object is a first person, the second object is a second person, the category indicated by the category identification text is a person attribute category, and the category identification text is a person attribute category identification text.

14. An image generating device, It is characterized in that The device comprises: An acquisition module, used for acquiring description text and acquiring at least one original image, wherein the description text includes category identification text, and any of the original images includes a face of a first object; An extraction module, used to extract original text features of the description text and image features of each of the at least one original image, wherein the original text features include category features of the category identification text; a determination module, configured to fuse the category feature with the image features of the at least one original image to obtain at least one fused feature, and determine a comprehensive feature based on the at least one fused feature; A replacement module, used for replacing the category feature in the original text feature with the comprehensive feature to obtain a target text feature; A generation module is used to generate a target image according to the target text features, wherein the target image includes a face of a second object, the face of the second object has the facial features of the face of the first object, and the image content of the target image matches the description content of the description text.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.

16. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

17. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.