Image generation model processing method and device, image generation method and device and computer equipment

By extracting image and text features and updating text features to generate personalized images, the problem of requiring multiple photos training in the prior art is solved, and efficient personalized image generation is achieved.

CN120259708APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410015489.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When generating personalized images with different expressions, different backgrounds, and different styles in the prior art, users need to input multiple photos in advance for training and learning, resulting in low image generation processing efficiency.

Method used

By acquiring sample images and description text, the image generation model to be trained extracts image features and text features respectively, updates text features to generate personalized images, and updates the model based on personalized images and image generation features to directly generate personalized images corresponding to the image object.

Benefits of technology

The processing efficiency of image generation is improved, and the operation process of pre-entering multiple images for training and learning is eliminated. It can directly generate corresponding personalized images based on the input image and text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259708A_ABST
    Figure CN120259708A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation model processing method and device, an image generation method and device, computer equipment, a storage medium and a computer program product. The method relates to an artificial intelligence technology, and comprises the following steps: obtaining a sample image and a description text for an image object in the sample image; respectively extracting image features of the sample image and text features of the description text through a to-be-trained image generation model; updating an object text feature representing the image object in the text feature through the image feature to obtain an updated text feature; obtaining an image generation feature according to the updated text feature, and generating a personalized image corresponding to the sample image and the description text according to the image generation feature; and updating a to-be-trained image generation model based on the personalized image and the sub-features representing the image object in the image generation features to obtain a trained image generation model. By adopting the method, the processing efficiency of image generation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to an image generation model processing method, apparatus, computer device, storage medium, and computer program product, as well as an image generation method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] Generative Artificial Intelligence (AIGC, Artificial Intelligence Generated Content) refers to a technology method of artificial intelligence based on generative adversarial networks, large pre-trained models, etc. Through the learning and recognition of existing data, it generates relevant content with appropriate generalization ability. Its core idea is to use artificial intelligence algorithms to generate content with certain creativity and quality. AIGC is an important symbol of the new era of artificial intelligence. Image generation is an important application of AIGC. Specifically, it can generate images that match specified text or images based on AIGC, such as personalized images with different expressions, backgrounds, and styles.

[0003] However, currently, when a user generates personalized images with different expressions, backgrounds, and styles, generally, the user needs to input multiple photos in advance for training and learning, which reduces the processing efficiency of image generation. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an image generation model processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the processing efficiency of image generation, as well as an image generation method, apparatus, computer device, computer-readable storage medium, and computer program product.

[0005] On the one hand, the present application provides an image generation model processing method. The method includes:

[0006] Obtain a sample image and a description text for an image object in the sample image;

[0007] Extract, through an image generation model to be trained, the image features of the sample image and the text features of the description text respectively;

[0008] Update the object text features representing the image object in the text features with the image features to obtain updated text features;

[0009] Obtain image generation features according to the updated text features, and generate a personalized image corresponding to the sample image and the description text according to the image generation features;

[0010] Update the image generation model to be trained based on the sub-features that characterize the image object in the personalized image and the image generation features, and obtain the trained image generation model; the trained image generation model is used to generate corresponding personalized images according to the input images and texts.

[0011] On the other hand, the present application also provides an image generation model processing device. The device includes:

[0012] A sample data acquisition module, configured to acquire sample images and descriptive texts for the image objects in the sample images;

[0013] A sample feature extraction module, configured to respectively extract the image features of the sample images and the text features of the descriptive texts through the image generation model to be trained;

[0014] A text feature update module, configured to update the object text features that characterize the image object in the text features through the image features, and obtain the updated text features;

[0015] An image generation processing module, configured to obtain image generation features according to the updated text features, and generate personalized images corresponding to the sample images and the descriptive texts according to the image generation features;

[0016] A model update module, configured to update the image generation model to be trained based on the sub-features that characterize the image object in the personalized image and the image generation features, and obtain the trained image generation model; the trained image generation model is used to generate corresponding personalized images according to the input images and texts.

[0017] On the other hand, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above image generation model processing method are implemented.

[0018] On the other hand, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above image generation model processing method are implemented.

[0019] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above image generation model processing method are implemented.

[0020] The above image generation model processing method, device, computer equipment, storage medium and computer program product extract the image features of sample images and the text features of descriptive texts respectively through the image generation model to be trained, update the object text features representing image objects in the text features through the image features, generate personalized images corresponding to the sample images and descriptive texts according to the image generation features obtained from the updated text features, and update the image generation model to be trained based on the personalized images and the sub-features representing image objects in the image generation features, so as to obtain the trained image generation model. By using the image features to update the object text features representing image objects in the text features to obtain image generation features, and updating and training the image generation model to be trained based on the personalized images and the sub-features representing image objects in the image generation features, the image generation model can directly generate personalized images corresponding to the image objects, eliminating the operation process of pre-inputting multiple images for training and learning, and can directly generate corresponding personalized images according to the input images and texts, thereby improving the processing efficiency of image generation.

[0021] On the one hand, the present application provides an image generation method. The method includes:

[0022] Obtain a reference image and a descriptive text for an image object in the reference image;

[0023] Input the reference image and the descriptive text into the image generation model to obtain a personalized image output by the image generation model and corresponding to the reference image and the descriptive text;

[0024] Wherein, the image generation model is obtained through the above image generation model processing method.

[0025] On the other hand, the present application also provides an image generation device. The device includes:

[0026] A reference data acquisition module for obtaining a reference image and a descriptive text for an image object in the reference image;

[0027] A reference data processing module for inputting the reference image and the descriptive text into the image generation model to obtain a personalized image output by the image generation model and corresponding to the reference image and the descriptive text;

[0028] Wherein, the image generation model is obtained through the above image generation model processing method.

[0029] On the other hand, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above image generation method are implemented.

[0030] On the other hand, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above image generation method are implemented.

[0031] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above image generation method are implemented.

[0032] For the above image generation method, device, computer device, storage medium, and computer program product, by inputting a reference image and a reference text into an image generation model, a personalized image corresponding to the reference image and the description text output by the image generation model is obtained. The image generation model can directly generate a corresponding personalized image according to the input reference image and reference text, eliminating the operation process of pre-inputting multiple images for training and learning, thereby improving the processing efficiency of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 It is an application environment diagram of an image generation model processing and an image generation method in an embodiment;

[0035] Figure 2 It is a flowchart of an image generation model processing method in an embodiment;

[0036] Figure 3 It is a flowchart block diagram of an image generation model processing method in an embodiment;

[0037] Figure 4 It is a flowchart of a method for determining a matching relationship in an embodiment;

[0038] Figure 5 It is a flowchart of an image generation method in an embodiment;

[0039] Figure 6 It is a flowchart block diagram of an image generation method in an embodiment;

[0040] Figure 7 It is a schematic block diagram of fusing to obtain a personalized image in an embodiment;

[0041] Figure 8Schematic diagram of a user generating a personalized picture in an embodiment;

[0042] Figure 9 Schematic diagram of a man wearing glasses in an embodiment;

[0043] Figure 10 For Figure 9 Schematic diagram of the segmentation result;

[0044] Figure 11 For Figure 9 Schematic diagram of calculating the similarity of the people in;

[0045] Figure 12 Schematic diagram of performing text feature update processing in an embodiment;

[0046] Figure 13 Schematic diagram of a woman in an embodiment;

[0047] Figure 14 For Figure 13 Schematic diagram of generating various personalized pictures;

[0048] Figure 15 Structural block diagram of an image generation model processing device in an embodiment;

[0049] Figure 16 Structural block diagram of an image generation device in an embodiment;

[0050] Figure 17 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] The image generation model processing method provided by the embodiments of the present application can be applied to, such as Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can capture different sample images for different image objects and generate corresponding description texts for the image objects in the sample images. The terminal 102 can send the sample images and description texts to the server 104 through the network. For the obtained sample images and description texts, the server 104 extracts the image features of the sample images and the text features of the description texts respectively through the image generation model to be trained. The server 104 updates the object text features representing the image objects in the text features through the image features, generates personalized images corresponding to the sample images and description texts according to the image generation features obtained from the updated text features. The server 104 updates the image generation model to be trained based on the personalized images and the sub-features representing the image objects in the image generation features, and obtains the trained image generation model. The trained image generation model can generate corresponding personalized images according to the input images and texts, specifically generate corresponding personalized images according to the images and texts input by the terminal 102, and feedback the generated personalized images to the terminal 102.

[0053] The image generation method provided by the embodiments of this application can be applied to, for example, Figure 1 the application environment shown. The user can upload a reference image to the server 104 through the terminal 102 and select a description text for the image object in the reference image. The server 104 can input the reference image and reference text into the pre-trained image generation model, obtain the personalized image output by the image generation model corresponding to the reference image and description text, and feedback the personalized image output by the image generation model to the terminal 102.

[0054] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0055] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable them to have functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0056] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing image processing to make the images processed by the computer more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT (Vision Transformer), V-MOE, MAE (Masked Autoencoders), etc. can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0057] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, has evolved from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0058] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0059] The pre-training model is the latest development result of deep learning, integrating the above technologies. The pre-training model (PTM), also known as the foundation model or large model, refers to a deep neural network (DNN) with large parameters. It is trained on a large amount of unlabeled data, and uses the function approximation ability of the large-parameter DNN to extract common features from the data. Through technologies such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also be used as a general interface connecting multiple specific task models.

[0060] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0061] The solution provided by the embodiments of this application involves technologies such as computer vision technology, natural language processing, and machine learning in artificial intelligence, and will be specifically described through the following embodiments.

[0062] In an exemplary embodiment, as Figure 2 shown, a method for processing an image generation model is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of this application, taking this method applied to Figure 1 the server in it as an example for illustration, it includes the following steps 202 to step 210. Among them:

[0063] Step 202, obtain a sample image and a description text for the image object in the sample image.

[0064] Among them, the image generation model is an artificial intelligence model used to generate an image that meets the conditions, and can specifically generate a corresponding image according to the specified image and / or text. The sample image can include different image objects, and the image objects can include at least one of various types of objects such as animals, plants, and human figures. The description text is the description information for the image objects included in the sample image. For example, the description text can be "a child wearing a red dress".

[0065] Specifically, the server can obtain the sample image and the description text for the image object in the sample image. The description text can be extracted from the sample image. For example, it can be extracted from the sample image through an image description extraction model. The sample image can include at least one image object, such as at least human figures or animals and plants, etc. The description text is used to describe at least one image object in the sample image, and can specifically describe the appearance, actions, states, etc. of at least one image object. The image generation model to be trained by the server is used to generate a corresponding personalized image for the image object in the sample image. For example, it can generate a personalized image that includes the image object and has a different style, expression, or environment from the sample image.

[0066] Step 204, respectively extract the image features of the sample image and the text features of the description text through the image generation model to be trained.

[0067] Among them, the image generation model can be an artificial intelligence network model constructed based on machine learning. Specifically, it can be obtained based on at least one of generation models such as the Stable Diffusion base model, the Imagen model, and the Dalle model. It can also be constructed based on various artificial neural network algorithms, such as at least one of the swin-transformer algorithm, the ViT algorithm, the DNN algorithm, and the CNN (Convolutional Neural Networks) algorithm. The image feature is a feature extracted from the sample image to represent the image characteristics of the sample image; the text feature is a feature extracted from the description text and is used to represent the characteristics of the description text. The image feature and the text feature are extracted by the image generation model to be trained. Specifically, they can be extracted by different feature extraction networks in the image generation model to be trained for the sample image and the description text respectively.

[0068] Optionally, the server can obtain the image generation model to be trained and perform feature extraction on the sample image and the description text respectively through the image generation model to be trained, so as to extract the image feature of the sample image and the text feature of the description text. In a specific application, the image generation model to be trained can include feature extraction networks for different modal information, so that corresponding feature extraction can be performed through the feature extraction networks of the same modal type. For example, image feature extraction can be performed through the image feature extraction network, and text feature extraction can be performed through the text feature extraction network.

[0069] Step 206: Update the object text feature representing the image object in the text feature with the image feature to obtain the updated text feature.

[0070] Among them, the description text may include description statements for the image object, specifically including at least one description word, that is, the description text can be obtained from a statement composed of at least one description word. The text features of the description text can be obtained according to the combined features of the respective description words included. For example, for the description text "a flexible girl", the description words may include "a", "flexible", and "girl", and the text features of the description text can be obtained according to the combined features of each description word. If the respective features of "a", "flexible", and "girl" are F1, F2, and F3 respectively, then the text features of the description text can be "F1+F2+F3", that is, it can be the sequential concatenation of the respective features F1, F2, and F3 of each description word. The object text feature is the feature in the text features that characterizes the image object, that is, the object text feature is the feature of the description word in the description text that matches the image object, specifically, it can be the feature of the descriptive entity word in the description text that matches the image object, and the descriptive entity word can be the named entity corresponding to the image object. For example, in the description text "a flexible girl wearing a baseball cap", the descriptive entity word can be "girl".

[0071] Exemplarily, the server can determine the object text feature that characterizes the image object from the text features. Specifically, the server can analyze the text features to determine the object text feature that characterizes the characteristics of the image object. In a specific implementation, the server can match the image object in the sample image with each description word in the description text to determine the descriptive entity word in the description text used to describe the image object, and determine the feature corresponding to the descriptive entity word from the text features, so as to obtain the object text feature that characterizes the image object. In a specific application, when the server extracts the text features of the description text through the image generation model to be trained, the description text can be divided to divide the description text into at least one description word, and the text features of each description word are extracted in sequence to obtain the features of each description word, and the features of each description word are concatenated in the order of the description words in the description text to obtain the text features of the description text.

[0072] The server updates the object text feature that characterizes the image object through the image features, so as to update the text features and obtain the updated text features. Specifically, the server can fuse the image features with the object text features through the image generation model to be trained, and replace the object text feature in the text features with the fused result, that is, use the fused result as the new object text feature, so as to update the text features and obtain the updated text features.

[0073] Step 208, obtain the image generation features according to the updated text features, and generate a personalized image corresponding to the sample image and the description text according to the image generation features.

[0074] Among them, the image generation feature is a feature used to generate personalized images. The image generation feature is obtained based on the updated text feature, and specifically, it can be obtained by performing further feature extraction processing on the updated text feature. The personalized image corresponds to the sample image and the description text. Specifically, the personalized image matches the image object in the sample image, such as it can include the image object or an object similar to the image object. The personalized image also conforms to the description in the description text, such as conforming to various descriptions of expressions, actions, styles, environments, etc. in the description text.

[0075] Specifically, the server can obtain the image generation feature according to the updated text feature. For example, it can perform feature extraction through a to-be-trained image generation model based on the updated text feature. Specifically, it can perform feature extraction based on the attention mechanism to obtain the image generation feature. The server generates a personalized image according to the image generation feature. Specifically, it can perform image decoding through the to-be-trained image generation model according to the image generation feature, so as to generate a personalized image corresponding to the sample image and the description text.

[0076] Step 210, update the to-be-trained image generation model based on the personalized image and the sub-feature in the image generation feature that represents the image object, and obtain the trained image generation model; the trained image generation model is used to generate a corresponding personalized image according to the input image and text.

[0077] Among them, the sub-feature refers to the feature in the image generation feature that represents the image object, that is, in the image generation feature, the image object in the sample image is represented by the sub-feature. Optionally, the server can update the to-be-trained image generation model based on the personalized image and the sub-feature in the image generation feature that represents the image object. Specifically, it can determine the loss based on the personalized image and the sub-feature in the image generation feature that represents the image object, and update the to-be-trained model parameters in the to-be-trained image generation model through the determined loss. For example, it can adjust the numerical values of the to-be-trained model parameters in the to-be-trained image generation model according to the determined loss, so as to update the to-be-trained image generation model and obtain the image generation model after this training. The sub-feature can be determined by the server from the image generation feature. Specifically, the server can parse the image generation feature to determine the sub-feature that represents the image object from the image generation feature, and update the to-be-trained image generation model in combination with the personalized image.

[0078] After obtaining the image generation model after this training, the server can return to obtain the next sample image and continue training the image generation model after this training until the training end condition is met. For example, when the training reaches a preset number of times, the loss reaches the convergence condition, or the output of the image generation model reaches the accuracy condition, it can be considered that the training end condition is met. The server can end the training of the image generation model and obtain the trained image generation model. The trained image generation model can generate images based on the input images and texts and output personalized images corresponding to the input images and texts.

[0079] In a specific application, such as Figure 3 As shown, when training the image generation model, the server can obtain the sample image and the description text for the image object in the sample image. The description text can specifically be obtained by extracting the description of the image object in the sample image. The server can input the sample image and the description text into the image generation model to be trained this time, so that the image generation model can respectively extract the image features of the sample image and the text features of the description text. Based on the image generation model, the server updates the object text features in the text features through the image features to obtain the updated text features. The image generation model performs feature processing based on the updated text features. For example, it can perform feature extraction based on the attention mechanism to obtain the image generation features. The image generation model can perform image decoding based on the image generation features to achieve image generation and output the image generation result, thereby obtaining personalized images. The server can determine the sub-features representing the image object from the image generation features and update the image generation model to be trained based on the sub-features and the personalized images, thereby obtaining the image generation model after this training. The server can continue the next model training until the training ends to obtain an image generation model that can generate corresponding personalized images according to the input images and texts.

[0080] In the above image generation model processing method, the image features of the sample image and the text features of the description text are respectively extracted by the image generation model to be trained. The object text features representing the image object in the text features are updated by the image features, and personalized images corresponding to the sample image and the description text are generated based on the image generation features obtained from the updated text features. Then, the image generation model to be trained is updated based on the personalized images and the sub-features representing the image object in the image generation features, and the trained image generation model is obtained. By using the image features to update the object text features representing the image object in the text features to obtain image generation features, and updating and training the image generation model to be trained based on the personalized images and the sub-features representing the image object in the image generation features, the image generation model can directly generate personalized images corresponding to the image object, eliminating the operation process of pre-inputting multiple images for training and learning. It can directly generate corresponding personalized images according to the input images and texts, thereby improving the processing efficiency of image generation.

[0081] In an exemplary embodiment, updating the object text features representing the image object in the text features by the image features to obtain the updated text features includes: determining the object text features representing the image object from the text features; fusing the image features and the object text features to obtain an object fusion feature; and replacing the object text features in the text features with the object fusion feature to obtain the updated text features.

[0082] Among them, the text features can be obtained by combining the respective features of each description word in the description text, that is, the text features include the respective features of each description word in the description text. The object text features are the features representing the image object in the text features, that is, the object text features are the features corresponding to the description words describing the image object in the description text. The object fusion feature is obtained by fusing the image features and the object text features, and specifically can be obtained by weighted fusion or splicing of the image features and the object text features.

[0083] Specifically, the server determines the object text features from the text features. Specifically, it can determine the descriptive words for describing the image object from the descriptive text and determine the features corresponding to the descriptive words from the text features. The features corresponding to the descriptive words in the text features can be used as the object text features representing the image object. In a specific implementation, the server can pre-match the image objects in the sample images with each descriptive word in the descriptive text, so as to determine the descriptive words for describing the image object from the descriptive text. The server can determine the object text features representing the image object from the text features based on the descriptive words. The server fuses the image features with the object text features to obtain object fusion features. Specifically, the server can splice the image features and the object text features to obtain splicing features, and the server can perform dimensionality reduction processing on the obtained splicing features to obtain object fusion features with the same dimension as the object text features. In addition, the server can also perform weighted fusion on the image features and the object text features according to the preset fusion weights to obtain object fusion features. Among them, the fusion weights can be flexibly set according to actual needs. The server updates the text features through the object fusion features. Specifically, it replaces the object text features in the text features with the object fusion features to obtain the updated text features. The object fusion features further fuse the image features. By replacing the object text features in the text features with the object fusion features to update the text features, the image features can be effectively introduced into the text features, which is beneficial to improving the feature expression ability of the text features for the sample images and the image objects.

[0084] In this embodiment, the server fuses the image features with the object text features representing the image object in the text features, and replaces the object text features in the text features with the object fusion features obtained by the fusion, so as to realize the update of the text features, which can improve the feature expression ability of the text features for the sample images and the image objects, enabling the image generation model to fully learn the correlation between the image objects and the descriptive text in the sample images during the training process. Thus, the image generation model can directly generate personalized images corresponding to the image objects, eliminating the need for the operation process of pre-inputting multiple images for training and learning, and directly generating corresponding personalized images according to the input images and texts, thereby improving the processing efficiency of image generation.

[0085] In an exemplary embodiment, determining the object text features representing the image object from the text features includes: obtaining the matching relationship between the image object in the sample image and the descriptive entity words in the descriptive text; determining the object descriptive entity words for describing the image object from the descriptive text according to the matching relationship; and determining the object text features representing the image object from the text features according to the object descriptive entity words.

[0086] Among them, the matching relationship records the corresponding relationship between the image objects in the sample image and the description entity words in the description text. The description entity words can include the entity words or phrases in the description text used to describe each object in the sample image. Specifically, they can be named entities for each object, such as various entity words like "man", "woman", "boy", "big tree", "lawn", "airplane", "ship", "windbreaker", etc., and can also include various entity phrases like "red shirt", "black cap", "nine-point pants", "waterproof windbreaker", etc. The object description entity word is the description entity word in the description text used to describe the image object, that is, it is considered that the object description entity word is the description entity word in the description text that matches the image object. When generating a personalized image for the sample image and the description text, the generated personalized image needs to conform to the object description entity word first, so that the generated personalized image can be more similar to the image object.

[0087] Optionally, the server can obtain the matching relationship between the image objects in the sample image and the description entity words in the description text. This matching relationship can be determined by the server through graphic-text matching processing for the sample image and the description text in advance. Based on this matching relationship, the server determines the object description entity words used to describe the image objects from the description text. Specifically, based on this matching relationship, the server filters out the description entity words that match the image objects from the description text and determines the description entity words that match the image objects as the object description entity words for describing the image objects. In a specific application, different image objects can be matched with different description entity words, so different object description entity words can be determined for different image objects. The server determines the object text features corresponding to the object description entity words from the text features. Specifically, the server can determine the features representing the object description entity words in the text features as the object text features, and the object text features can be used to represent the image objects.

[0088] In this embodiment, the server determines the object description entity words for describing the image objects from the description text according to the matching relationship between the image objects in the sample image and the description entity words in the description text, and determines the object text features from the text features according to the object description entity words. It can accurately determine the object text features from the text features and update them, so as to improve the feature expression ability of the text features for the sample image and the image objects.

[0089] In an exemplary embodiment, as Figure 4 shown, the image generation model processing method further includes the processing of determining the matching relationship, specifically including:

[0090] Step 402, segment the sample image into image blocks including different image entities, and divide each description entity word from the description text.

[0091] Among them, an image entity refers to various entities included in a sample image, such as entities like animals, plants, people, physical objects, etc. An image object is the image entity in the sample image for which image generation processing is to be performed. That is, in the sample image, in addition to the image object for which a personalized image is to be generated, other types of image entities may also be included. An image block is an image region obtained by segmenting the sample image. Specifically, when the sample image is segmented according to image entities, different image blocks can correspond to different image entities, that is, each image block can include different image entities. A descriptive entity word is an entity vocabulary or phrase in the descriptive text used to describe each image entity in the sample image. Different image entities can be described by different descriptive entity words, so different image entities can be determined by dividing different descriptive entity words.

[0092] Exemplarily, the server can segment the sample image. Specifically, the sample image can be segmented according to different image entities, so as to obtain image blocks including different image entities. There can be a one-to-one correspondence between the image blocks and the image entities, that is, different image blocks correspond to different image entities. In a specific application, the server can segment the sample image through a pre-trained object recognition model to obtain image blocks including different image entities. The server performs text division on the descriptive text to obtain each descriptive entity word. Specifically, the server can be based on an entity annotation model, such as can divide each descriptive entity word from the descriptive text based on the spacy model.

[0093] Step 404: Perform image-text matching between the image blocks and the descriptive entity words to obtain an image-text matching result.

[0094] Among them, the image-text matching result is obtained by performing image-text matching between the image blocks and the descriptive entity words. Specifically, it can include the feature similarity between the image blocks and the descriptive entity words. Through the image-text matching result, the association relationship between each image block and each descriptive entity word can be determined.

[0095] Specifically, for each image block obtained by segmentation, the server can perform image-text matching between the targeted image block and each descriptive entity word respectively. Specifically, it can determine the feature similarity between the targeted image block and each descriptive entity word respectively, so as to obtain the image-text matching result. The server traverses each image block to obtain the image-text matching results between each image block and each descriptive entity word.

[0096] Step 406: According to the image-text matching result, determine the matching relationship between the image object in the sample image and the descriptive entity word in the descriptive text.

[0097] Among them, the matching relationship records the corresponding relationship between the image objects in the sample image and the descriptive entity words in the descriptive text. Optionally, the server can determine the object descriptive entity words that match the image objects in the sample image from the descriptive entity words in the descriptive text based on the graphic-text matching results between each image block and each descriptive entity word, so as to obtain the matching relationship between the image objects and the descriptive entity words in the descriptive text. In specific applications, the server can perform screening based on the graphic-text matching results between each image block and each descriptive entity word. Specifically, it can be screened based on the greedy algorithm. For example, it can be screened in the order of the feature similarity values in the graphic-text matching results from large to small to determine the matching relationship between the image objects and the descriptive entity words in the descriptive text.

[0098] In this embodiment, the server divides the sample image into image blocks and divides the descriptive entity words from the descriptive text. By performing graphic-text matching between the image blocks and the descriptive entity words, the matching relationship between the image objects in the sample image and the descriptive entity words in the descriptive text can be accurately determined according to the graphic-text matching results. Thus, based on the matching relationship, the object descriptive entity words can be accurately determined for feature update, which is beneficial to improving the feature expression ability of the text features for the sample image and the image objects.

[0099] In an exemplary embodiment, performing graphic-text matching between the image blocks and the descriptive entity words to obtain the graphic-text matching results includes: matching the image block features of the image blocks with the entity word features of the descriptive entity words to obtain a first matching result; matching the image entity labels associated with the image blocks with the entity word features of the descriptive entity words to obtain a second matching result; and obtaining the graphic-text matching results according to the first matching result and the second matching result.

[0100] Among them, the image block features are the features extracted for the image blocks, the entity word features are the features extracted for the descriptive entity words, the first matching result is the matching result obtained by matching the image block features with the entity word features, and specifically may include the feature similarity between the image block features and the entity word features, such as at least one of various similarities such as cosine similarity, Euclidean distance, Manhattan distance, and log-likelihood similarity. The image entity label is the label of the image entity corresponding to the image block, and can be obtained specifically when the sample image is segmented. The image block is an image area in the sample image, and the pixel features of the image entity are carried therein; the image entity label is the classification information of the image entity corresponding to the image block, and the category features of the image entity are carried therein. The second matching result is the matching result obtained by matching the image entity label with the entity word features, and specifically may include the feature similarity between the image entity label and the entity word features.

[0101] Optionally, when performing image-text matching between each image patch and each descriptive entity word, for each image patch, the server can perform image-text matching between the image patch and each descriptive entity word respectively to obtain the image-text matching results between the image patch and each descriptive entity word, and by traversing each image patch, the image-text matching results between each image patch and each descriptive entity word can be obtained. Specifically, for the targeted image patch, the server can perform feature matching between the image patch features of the image patch and the entity word features of the descriptive entity word to obtain a first matching result. In specific implementation, the server can perform feature extraction for each image patch and descriptive entity word respectively to obtain the image patch features of the image patch and the entity word features of the descriptive entity word. The server can calculate the feature similarity between the image patch features of the image patch and the entity word features of the descriptive entity word, such as calculating the cosine similarity, to obtain the first matching result.

[0102] The server can perform feature matching between the image entity label associated with the image patch and the entity word features of the descriptive entity word. Specifically, it can calculate the feature similarity between the label features of the image entity label and the entity word features, such as calculating the cosine similarity, to obtain a second matching result. In a specific application, the server can determine the image entity label associated with the image patch, extract the label features for the image entity label, and the server performs feature matching between the label features and the entity word features to obtain the second matching result. The server obtains the image-text matching result based on the first matching result and the second matching result. Specifically, the server can fuse the first matching result and the second matching result, such as performing weighted fusion to obtain the image-text matching result.

[0103] In this embodiment, the server matches the image patch features of the image patch and the image entity label associated with the image patch with the entity word features of the descriptive entity word respectively, and obtains the image-text matching result based on the first matching result and the second matching result obtained from the respective matches. By comprehensively performing image-text matching between the image patch and the image entity label, the accuracy of image-text matching can be improved, the accuracy of the matching relationship can be ensured, and it is beneficial to improve the feature expression ability of the text features for the sample image and the image object.

[0104] In an exemplary embodiment, obtaining image generation features based on the updated text features and generating a personalized image corresponding to the sample image and the descriptive text includes: based on the attention mechanism, obtaining attention features according to the updated text features and the image features; performing image decoding according to the attention features to generate a personalized image corresponding to the sample image and the descriptive text.

[0105] Among them, the attention mechanism is a method in deep learning that mimics the human visual and cognitive systems. It allows the neural network to focus on relevant parts when processing input data. By introducing the attention mechanism, the neural network can automatically learn and selectively focus on important information in the input, improving the performance and generalization ability of the model. The attention mechanism can allocate computing resources to more important tasks in the case of limited computing power, while solving the problem of information overload. By introducing the attention mechanism, focusing on the information that is more critical to the current task among numerous input information, reducing the attention to other information, and even filtering out irrelevant information, the problem of information overload can be solved, and the efficiency and accuracy of task processing can be improved. The attention feature is an image generation feature obtained based on the updated text feature and image feature through the attention mechanism. By using the attention feature as the image generation feature and performing image decoding on the attention feature, a personalized image corresponding to the sample image and the description text can be generated.

[0106] Exemplarily, the server can perform feature processing on the updated text feature and image feature based on the attention mechanism. For example, it can perform feature extraction on the updated text feature and image feature based on self-attention mechanism, cross-attention mechanism, multi-head self-attention mechanism, channel attention mechanism, spatial attention mechanism, etc., to obtain the attention feature. The server can perform image decoding according to the attention feature. Specifically, it can perform image decoding on the attention feature through the image decoding network in the image generation model to be trained, to obtain a personalized image corresponding to the sample image and the description text.

[0107] In this embodiment, the server obtains the attention feature based on the attention mechanism according to the updated text feature and image feature, and generates a personalized image by performing image decoding on the attention feature, which can enable the image generation model to learn important information in the updated text feature and image feature during the training process, thereby ensuring the image quality of the personalized image generated by the image generation model.

[0108] In an exemplary embodiment, obtaining the attention feature based on the attention mechanism according to the updated text feature and image feature includes: determining attention parameters based on the cross-attention mechanism according to the updated text feature and image feature; obtaining weight parameters according to the query parameter and key parameter in the attention parameters; and obtaining the attention feature according to the weight parameters and the value parameter in the attention parameters.

[0109] Among them, the cross-attention mechanism is an extended form based on the traditional attention mechanism, which further considers the correlation between different input sequences, such as the correlation between images and texts. By introducing the cross-attention mechanism, the model can better capture the correlation information between different inputs, thereby improving the performance of the model. In the image generation task, the model needs to generate personalized images related to the image and / or text content. After introducing the cross-attention mechanism, the model can consider the correlation information between the image and the text while generating personalized images, so as to generate more accurate and higher-quality personalized images. The attention parameters are intermediate parameters for constructing attention features, which can specifically include query parameters (Query, Q), key parameters (Key, K), and value parameters (Value, V). The weight parameters are calculated based on the query parameters and key parameters in the attention parameters, and the attention features are calculated based on the weight parameters and the value parameters in the attention parameters.

[0110] Specifically, the server can determine the attention parameters based on the cross-attention mechanism according to the updated text features and image features. The attention parameters can include query parameters, key parameters, and value parameters. Among them, the query parameters can be obtained by the server according to the updated text features and query intermediate parameters, the value parameters can be obtained by the server according to the updated text features and value intermediate parameters, and the key parameters can be obtained by the server according to the image features and key intermediate parameters. The query intermediate parameters, value intermediate parameters, and key intermediate parameters are trainable model parameters in the image generation model to be trained. The server further obtains the weight parameters according to the query parameters and key parameters in the attention parameters. Specifically, the weight parameters can be obtained by performing mapping processing on the query parameters and key parameters, and the weight parameters can specifically include a tensor matrix. The server obtains the attention features according to the weight parameters and the value parameters in the attention parameters. For example, the attention features can be obtained by multiplying the weight parameters and the value parameters.

[0111] In this embodiment, the server obtains the attention parameters based on the cross-attention mechanism according to the updated text features and image features, and constructs the attention features based on the attention parameters, which can enable the image generation model to learn the important information in the updated text features and image features during the training process, thereby ensuring the image quality of the personalized images generated by the image generation model.

[0112] In an exemplary embodiment, based on the personalized image and the sub-features representing the image object in the image generation features, the image generation model to be trained is updated to obtain a trained image generation model, including: determining a first loss according to the personalized image and the sample image; determining a second loss according to the sub-features representing the image object in the image generation features and the object image patch; the object image patch is the image patch including the image object in the sample image; obtaining an image generation loss based on the first loss and the second loss, and updating the image generation model to be trained through the image generation loss to obtain a trained image generation model.

[0113] Among them, the first loss is determined based on the personalized image, and specifically can be determined according to the difference between the personalized image and the sample image. The second loss is determined based on the sub-features representing the image object in the image generation features, and specifically can be determined according to the difference between the sub-features and the object image patch. The object image patch is the image patch including the image object in the sample image, that is, the region in the object image patch represents the region where the image object is located in the sample image. The image generation loss is obtained according to the first loss and the second loss, and specifically can be obtained by fusing the first loss and the second loss. Through the image generation loss, the image generation model to be trained can be updated, so as to realize the training and updating process of the image generation model.

[0114] Exemplarily, the server can determine the first loss based on the personalized image and the sample image, specifically obtained from the pixel differences between the same positions in the personalized image and the sample image. The first loss can specifically include at least one of Mean Square Error (MSE), 0-1 loss, Mean Absolute Error Loss, Logarithmic Loss, Exp-Loss, Hinge Loss, or Cross-Entropy Loss Function. The server can determine the second loss based on the sub-features representing the image object in the image generation features and the object image block. Specifically, the server can determine the sub-features representing the image object from the image generation features, such as determining the sub-features representing the image object from the updated text features. The server determines the object image block including the image object from each image block obtained after segmenting the sample image. The server determines the second loss based on the sub-features and the object image block, specifically calculating the second loss based on the sub-features and the object image block according to various loss forms. The server obtains the image generation loss based on the first loss and the second loss, such as performing weighted fusion on the first loss and the second loss to obtain the image generation loss. The server updates the image generation model to be trained according to the image generation loss, specifically adjusting each model parameter in the image generation model to be trained, thereby achieving the current training update of the image generation model and obtaining the trained image generation model.

[0115] In this embodiment, the server determines the first loss based on the personalized image and the sample image, determines the second loss based on the sub-features representing the image object in the image generation features and the object image block, and comprehensively obtains the image generation loss based on the first loss and the second loss, and updates the model through the image generation loss, thereby introducing the difference between the sub-features representing the image object and the object image block into the image generation loss, ensuring the training effect of the image generation model, and thus ensuring the image quality of the personalized image generated by the image generation model.

[0116] In an exemplary embodiment, obtaining a sample image and a description text for the image object in the sample image includes: obtaining a sample image including the image object and extracting the description text for the image object based on the sample image.

[0117] Among them, the sample image includes the image object, and the description text is the text content describing the image object in the sample image. Specifically, the server can obtain a sample image including the image object and generate a text description for the sample image, thereby extracting the description text for the image object.

[0118] Further, through the image generation model to be trained, the image features of the sample image and the text features of the description text are respectively extracted, including: through the image feature extraction network in the image generation model to be trained, the image features of the sample image are extracted; through the text feature extraction network in the image generation model to be trained, the text features of the description text are extracted.

[0119] Among them, the image generation model to be trained includes an image feature extraction network and a text feature extraction network to respectively extract features for different modal contents. Optionally, the server can extract features for the sample image through the image feature extraction network in the image generation model to be trained to obtain the image features of the sample image. The server can extract features for the description text through the text feature extraction network in the image generation model to be trained to obtain the text features of the description text.

[0120] In this embodiment, the server can generate an accurate description text based on the sample image. By using the image feature extraction network and the text feature extraction network in the image generation model to be trained, the image features of the sample image and the text features of the description text are respectively extracted, and the image features and text features can be accurately extracted for image generation processing, which is beneficial to improving the image quality of the personalized images generated by the image generation model.

[0121] In an exemplary embodiment, as Figure 5 shown, an image generation method is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server as an example for illustration, it includes the following steps 502 to step 504. Among them:

[0122] Step 502, obtain a reference image and a description text for the image object in the reference image.

[0123] Among them, the reference image is an image for generating a personalized image, that is, a corresponding personalized image is generated by referring to this reference image. The image object is an entity object included in the reference image, and specifically can include at least one of various types of objects such as animals, plants, human figures, and physical objects. The description text is description information for the image object in the reference image. The description text can be generated by performing text description based on the reference image, or can be configured by the user for the reference image. For example, it can be configured by the user selecting from candidate texts for the reference image, or can be directly configured by the user editing the reference image.

[0124] Specifically, the server can obtain a reference image and descriptive text for an image object in the reference image, and thus perform personalized image generation processing based on the reference image and the descriptive text.

[0125] Step 504: Input the reference image and the descriptive text into an image generation model to obtain a personalized image output by the image generation model and corresponding to the reference image and the descriptive text; wherein, the image generation model is obtained through an image generation model processing method.

[0126] Among them, the image generation model is obtained based on the image generation model processing method. The image generation model can generate a corresponding personalized image according to the input image and text. Exemplarily, the server can query a pre-trained image generation model, input the reference image and the descriptive text into the image generation model, and the image generation model performs personalized image generation processing on the reference image and the descriptive text, so as to output a personalized image corresponding to the reference image and the descriptive text. In a specific application, as Figure 6 shown, the server can obtain a reference image and descriptive text for an image object in the reference image, and input the reference image and the descriptive text into a pre-trained image generation model. The image generation model performs image generation processing on the reference image and the descriptive text, and outputs a personalized image corresponding to the reference image and the descriptive text.

[0127] In the above image generation method, by inputting the reference image and the reference text into the image generation model, a personalized image output by the image generation model and corresponding to the reference image and the descriptive text is obtained. The image generation model can directly generate a corresponding personalized image according to the input reference image and reference text, and the operation process of pre-inputting multiple images for training and learning can be omitted, thereby improving the processing efficiency of image generation.

[0128] In an exemplary embodiment, inputting the reference image and the descriptive text into the image generation model to obtain a personalized image output by the image generation model and corresponding to the reference image and the descriptive text includes: inputting the reference image and the descriptive text into the image generation model, and generating a first personalized image by the image generation model based on the reference image and the descriptive text; generating a second personalized image by the image generation model based on the reference image; generating a third personalized image by the image generation model based on the descriptive text; and performing weighted fusion on the first personalized image, the second personalized image, and the third personalized image to obtain a personalized image corresponding to the reference image and the descriptive text.

[0129] Among them, the first personalized image is the image generation result obtained by the image generation model based on the reference image and the description text; the second personalized image is the image generation result obtained by the image generation model only based on the reference image; the third personalized image is the image generation result obtained by the image generation model only based on the description text. The personalized image finally output by the image generation model is obtained by weighted fusion of the first personalized image, the second personalized image, and the third personalized image.

[0130] Optionally, the server can input the reference image and the description text into the image generation model to perform image generation processing by the image generation model. Specifically, the server can perform image generation based on the reference image and the description text through the image generation model. For example, image decoding can be performed to obtain the first personalized image; the server decodes the reference image through the image generation model to generate the second personalized image; the server decodes the description text through the image generation model to generate the third personalized image. The server can determine a preset weighting weight and perform weighted fusion of the first personalized image, the second personalized image, and the third personalized image according to the weighting weight to obtain a personalized image corresponding to the reference image and the description text. The weighting weight can be flexibly set according to empirical values.

[0131] In a specific application, such as Figure 7 As shown, when the server inputs the reference image and the description text into the image generation model, the image generation model can comprehensively generate the first personalized image based on the reference image and the description text, the image generation model can separately generate the second personalized image based on the reference image, and the image generation model can also separately generate the third personalized image based on the description text. The server obtains a personalized image corresponding to the reference image and the description text by fusing the first personalized image, the second personalized image, and the third personalized image.

[0132] In this embodiment, the server performs image generation based on the reference image, the description text, the reference image, and the description text respectively through the image generation model, and performs weighted fusion of the separately generated first personalized image, second personalized image, and third personalized image to obtain the required personalized image, which can avoid the generated personalized image being too biased towards the description text or too biased towards the reference image, and improve the image quality of the personalized image.

[0133] This application also provides an application scenario that applies the above image generation model processing method and image generation method. Specifically, the application of the image generation model processing method and image generation method in this application scenario is as follows:

[0134] Due to the rise of AIGC, in the information flow scenario, a technology for generating personalized human faces has emerged, and there are two mainstream methods. One is to directly use multiple different photos of the same person and specifically train a person lora (Low-Rank Adaptation of Large Language Models) for this person to control the similarity between the generated person and the input person. The other is to use multiple different photos of the same person, specifically train a person lora sub-model for this person, and then train multiple different style lora sub-models for various different styles such as Hanfu and ID photos. Finally, by combining the person lora sub-model and the style lora sub-model, personalized human face images with different styles are jointly generated, and both of the above two methods are completed based on the Stable Diffusion (AI painting generation tool) base model.

[0135] However, both of these methods have certain deficiencies. Since both methods are based on the technology of lora sub-models, for each person, 20 different-angle photos of this person are required, and a lora sub-model needs to be trained specifically for them. This will lead to three problems. Problem one, the user needs to collect 20 different-angle photos of themselves, which brings a certain usage threshold to the user. Problem two, a lora model needs to be trained for each user. When the number of users is tens of millions or even more, the time cost and storage cost of training the model will increase exponentially. Problem three, when the person lora and the style lora are used together, they will affect each other, and it is impossible to ensure that the generated result simultaneously meets the specified style and the specified person.

[0136] Based on this, the image generation model processing method and the image generation method provided in this embodiment involve a face personalization generation scheme based on fine-grained text-image matching and face localization. Based on the Stable Diffusion base model, a base model with the ability to generate personalized human faces is directly trained. Without the need for training, only one photo of a person is required to generate the corresponding personalized human face photo, which can be applied in the information flow scenario. The image generation model processing method and the image generation method provided in this embodiment are suitable for a face personalization generation model for large-scale applications. The user only needs to provide one photo, and the model can generate personalized pictures to reduce the user's usage threshold. Moreover, for any user, there is no need to train a model specifically for them, but personalized pictures can be directly generated for them. In addition, whether it is the style or the person, the base model can be directly used for control without the need for a separate sub-model for control.

[0137] Specifically, for the image generation model processing method and image generation method provided in this embodiment, data is first prepared. Based on the open-source person image dataset, the extraction of text descriptions, the fine-grained matching of text and images are completed to construct the data required for training. Then, the model is trained. Based on the prepared data, the person in the located image is positioned and a corresponding objective function is designed to learn the model. Then, model inference is performed. To prevent overfitting, conditional delay strategy and weight adjustment strategy are applied in the model inference stage. Finally, model application is realized, and the model can be applied in the information flow scenario. As Figure 8 shown, the user uploads a photo, selects a specific text from the pre-set texts, and inputs the photo and the text into the image generation model together. The image generation model is used to predict the user's personalized picture. When the user publishes content, the user selects the desired picture from the personalized pictures generated by the image generation model, and finally publishes the generated personalized picture to the information flow scenario. The published content supports distribution at the consumer side and can be consumed by other users.

[0138] Specifically, for data preparation, based on the open-source person image dataset, the extraction of text descriptions, the fine-grained matching of text and images are completed to construct the data required for training. First, a person dataset is collected. As Figure 9 shown, a total of 70,000 photos containing people are collected. The screening conditions for the person photos are upper body photos, and it is required that the photos have high clarity and rich person details. For example, facial wrinkles and beards are clearly visible. Such data is beneficial for the model to learn the distinguishable details of people. Further, text descriptions are extracted. Specifically, the Blip-2 model can be used to extract the image caption (text description of the picture). For example, for Figure 9 the extracted image caption is "a man wearing a checkered shirt and sunglasses sitting down" (a man sitting down wearing a checkered shirt and sunglasses). Further, noun phrases are extracted. Specifically, the spacy model can be used to extract noun phrases from the text description. For example, the noun phrases extracted from the image caption are 'a man', 'a checkered shirt','sunglasses', a total of three noun phrases. Further, the original picture is segmented. Specifically, the mask2former model can be used to perform panoramic segmentation on the original picture to obtain all the entities in the picture. The segmentation result is as Figure 10As shown, a total of 5 segmentation results are obtained. Each segmentation result consists of two parts. The first part is the segmented pixel points. For example, the area of the filled points in the figure corresponds to the person (the male with glasses) in the original image. The second part is the segmentation label. For example, the area of the filled points in the figure corresponds to the "person-0" label. In addition, in the segmentation result for Figure 9 also includes labels such as "fence-merged-0", "tie-0", "tree-merged-0", and "ceiling-merged-0", and each label corresponds to Figure 9 different image regions in

[0139] Furthermore, fine-grained text-image matching processing. Specifically, the Sentence Transformer model and the OpenCLIP model can be used to calculate the similarity. First, a Cartesian product is performed on the determined 3 noun phrases and 5 segmentation results, and 15 relationship pairs can be matched. Then, for each relationship pair, two similarities are calculated. The Sentence Transformer model is used to extract the vectors of the noun phrase and the segmentation label and calculate the cosine similarity of these two vectors. The OpenCLIP model is used to extract the vectors of the noun phrase and the segmentation result (obtain the pixel values of the original image from the original image using the segmented pixel points) and calculate the cosine similarity between them. Finally, the value obtained by multiplying these two cosine similarities is used as the final similarity between the noun phrase and the segmentation result. Taking the segmentation result of "a man" and "this male in the figure" as an example, the process of calculating the similarity is as Figure 11 shown, calculating the similarity between 'a man' and 'person', and calculating the similarity between 'a man' and the pixel region of 'this male in the figure'.

[0140] Cosine similarity measures the similarity between two vectors by measuring the cosine value of the angle between them. The cosine value of a 0-degree angle is 1, and the cosine value of any other angle is not greater than 1; and its minimum value is -1. Thus, the cosine value of the angle between two vectors determines whether the two vectors generally point in the same direction. When the two vectors have the same direction, the value of the cosine similarity is 1; when the angle between the two vectors is 90°, the value of the cosine similarity is 0; when the two vectors point in exactly opposite directions, the value of the cosine similarity is -1. This result is independent of the length of the vectors and only related to the direction of the vectors. Cosine similarity is usually used in positive space, so the value given is between 0 and 1. The calculation formula of cosine similarity is shown in the following formula (1),

[0141] (1)

[0142] Among them, A and B are two n-dimensional vectors. A is [A1, A2, ..., An], and B is [B1, B2, ..., Bn].

[0143] After the above similarity calculation, 15 similarity matching relationships are obtained. Further, the greedy algorithm is used to obtain the final fine-grained matching result of text and image. Specifically, the maximum similarity is taken from the 15 similarities to obtain the first matching relationship. The matched noun and the image segmentation result will not participate in the subsequent process. Then, the maximum similarity is continuously selected from the remaining similarities to obtain the second matching relationship, and so on, until all the noun phrases obtained in the third step are matched to the segmentation results in the image.

[0144] Finally, 70,000 images can be obtained. Each image has a corresponding text description (image caption), as well as the matching relationship between the people in the image and the nouns in the text description.

[0145] Furthermore, train the model. Based on the prepared data, locate the people in the image and design a corresponding objective function to learn the model. Specifically, take Stable Diffusion as the base model and input the above-prepared data into the model. First, use the text encoder to extract the vector of the text description. Then, use the image encoder to extract the vector of the image. Using the matching relationship between the people in the image and the nouns in the text description extracted in the data preparation stage, splice the image vector with the corresponding noun vector. As Figure 12 shown, specifically, the image vector can be directly spliced with the vector of the token "man", and then input into a MLP (Multilayer Perceptron) to reduce the dimension of the spliced vector to half of the original, and the vectors of other words remain unchanged. Finally, the new vector of each word is obtained.

[0146] Furthermore, locate the people and design a suitable learning objective. The final vectors of each word obtained together form a 77×768 matrix. This vector matrix is used as a conditional input into the cross-attention mechanism of the Stable Diffusion model. Denote this matrix as c, and denote the other input of the attention mechanism, that is, the image vector, as z. z and c are respectively multiplied by , to obtain the Q, K, V matrices, as shown in the following formula (2) specifically,

[0147] (2)

[0148] Among them, 、 and are trainable parameter matrices in the image generation model, and the Q, K, and V matrices are query parameter, key parameter, and value parameter matrices respectively.

[0149] The weight tensor A can be obtained based on the Q and K matrices, as specifically shown in the following formula (3):

[0150] (3)

[0151] Among them, the size of the tensor A is h×w×n, representing n matrices of size h×w. is the length of the vector. In this embodiment, n is equal to 77. Taking the input "a man wearing a checkered shirt and sunglasses sittingdown" as an example, there are a total of 10 word vectors. If there are less than 77, they are padded with empty characters to 77 word vectors, and "man" corresponds to the second word vector. Therefore, the second matrix of size h×w in the tensor A corresponds to the concept of the word "man". Then, it is desired that the second matrix of size h×w in the tensor A only pays attention to this male in the input image, that is, "person" in the segmentation result. Therefore, an L1 loss can be calculated between the second matrix of size h×w in the tensor A and the segmentation result "person" to force this weight matrix to learn this male in the image and ignore other word concepts. Finally, the learned weight matrix A is multiplied by the V matrix to obtain the result of cross-attention, as specifically shown in the following formula (4):

[0152] (4)

[0153] is the result of cross-attention, and a corresponding personalized picture is generated through the result of cross-attention.

[0154] In this embodiment, after special processing of the input image and text data, they are input into the model. At the same time, an additional L1 loss is calculated for the cross-attention mechanism and the segmentation result of the original image as a learning objective to optimize the model parameters.

[0155] In model inference, to prevent overfitting, a weight adjustment strategy is proposed and applied in the model inference stage. To prevent overfitting, a weight adjustment strategy is proposed to adjust the weights of the image and text respectively, so as to adjust the generated result and avoid the generated result being too biased towards the text or the picture. As specifically shown in the following formula (5):

[0156] (5)

[0157] Among them, i represents an image, and f represents text. represents the result obtained by simultaneously inputting an image and input text. represents the result obtained by only inputting an image, that is, when all text is set to empty. represents the result obtained by only inputting text, that is, when the image is set to empty. α and β are hyperparameters. β is used to control the weights of inputting text alone and inputting an image alone, and α is used to control the weights of simultaneously inputting an image and text and inputting image and text alone. is the finally obtained personalized picture.

[0158] In specific applications, it can be applied to the information flow scenario. Without the need for training, only by inputting a picture and selecting pre-set text, a series of personalized pictures can be generated. For example, Figure 13 as shown, input a photo of a woman and select any text: A smiling woman, A wry smile woman, A cry loudly woman, A grievance woman, A Growl woman, and corresponding pictures can be generated respectively, such as Figure 14 shown, personalized pictures corresponding to A smiling woman, A wry smile woman, A cry loudly woman, A grievance woman, and A Growl woman can be generated respectively.

[0159] In this embodiment, for any user, personalized photos can be generated without training, which can reduce the technical cost of the image generation model; and for any user, only one photo is needed to generate personalized photos, reducing the user's usage threshold. In addition, it can be used for both the generation of users' expressions and for users to generate their own personalized photos in the information flow scenario, with a wide range of applicable scenarios.

[0160] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0161] Based on the same inventive concept, the embodiments of the present application also provide an image generation model processing device for implementing the above-mentioned image generation model processing method and an image generation device for implementing the above-mentioned image generation method. The implementation solutions provided by the device for solving problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the image generation model processing device provided below can refer to the limitations on the image generation model processing device method in the above text, and the specific limitations in one or more embodiments of the image generation device provided below can refer to the limitations on the image generation method in the above text, which will not be elaborated here.

[0162] In an exemplary embodiment, as Figure 15 shown, an image generation model processing device 1500 is provided, including: a sample data acquisition module 1502, a sample feature extraction module 1504, a text feature update module 1506, an image generation processing module 1508, and a model update module 1510, where:

[0163] The sample data acquisition module 1502 is configured to acquire a sample image and a description text for an image object in the sample image;

[0164] The sample feature extraction module 1504 is configured to respectively extract the image feature of the sample image and the text feature of the description text through the image generation model to be trained;

[0165] The text feature update module 1506 is configured to update the object text feature representing the image object in the text feature through the image feature to obtain the updated text feature;

[0166] The image generation processing module 1508 is configured to obtain an image generation feature according to the updated text feature, and generate a personalized image corresponding to the sample image and the description text according to the image generation feature;

[0167] A model update module 1510 is configured to update an image generation model to be trained based on a personalized image and sub-features representing an image object in the image generation features, so as to obtain a trained image generation model; the trained image generation model is used to generate a corresponding personalized image according to the input image and text.

[0168] In one embodiment, the text feature update module 1506 is further configured to determine object text features representing the image object from the text features; fuse the image features and the object text features to obtain object fusion features; and replace the object text features in the text features with the object fusion features to obtain updated text features.

[0169] In one embodiment, the text feature update module 1506 is further configured to obtain a matching relationship between the image object in the sample image and the description entity word in the description text; determine, according to the matching relationship, an object description entity word for describing the image object from the description text; and determine, according to the object description entity word, object text features representing the image object from the text features.

[0170] In one embodiment, a matching relationship determination module is further included, configured to segment the sample image into image blocks including different image entities, and divide each description entity word from the description text; perform image-text matching on the image blocks and the description entity words to obtain an image-text matching result; and determine the matching relationship between the image object in the sample image and the description entity word in the description text according to the image-text matching result.

[0171] In one embodiment, the matching relationship determination module is further configured to match the image block features of the image blocks with the entity word features of the description entity words to obtain a first matching result; match the image entity labels associated with the image blocks with the entity word features of the description entity words to obtain a second matching result; and obtain the image-text matching result according to the first matching result and the second matching result.

[0172] In one embodiment, the image generation processing module 1508 is further configured to obtain attention features based on an attention mechanism according to the updated text features and image features; and perform image decoding according to the attention features to generate a personalized image corresponding to the sample image and the description text.

[0173] In one embodiment, the image generation processing module 1508 is further configured to determine attention parameters based on a cross-attention mechanism according to the updated text features and image features; obtain weight parameters according to the query parameters and key parameters in the attention parameters; and obtain attention features according to the weight parameters and the value parameters in the attention parameters.

[0174] In one embodiment, the model update module 1510 is further configured to determine a first loss according to the personalized image and the sample image; determine a second loss according to the sub-feature representing the image object in the image generation feature and the object image block, where the object image block is an image block including the image object in the sample image; obtain an image generation loss based on the first loss and the second loss, and update the image generation model to be trained through the image generation loss to obtain a trained image generation model.

[0175] In one embodiment, the sample data acquisition module 1502 is further configured to acquire a sample image including an image object, and extract a description text for the image object based on the sample image; the sample feature extraction module 1504 is further configured to extract the image feature of the sample image through the image feature extraction network in the image generation model to be trained; and extract the text feature of the description text through the text feature extraction network in the image generation model to be trained.

[0176] In an exemplary embodiment, as Figure 16 shown, an image generation device 1600 is provided, including: a reference data acquisition module 1602 and a reference data processing module 1604, where:

[0177] The reference data acquisition module 1602 is configured to acquire a reference image and a description text for the image object in the reference image;

[0178] The reference data processing module 1604 is configured to input the reference image and the description text into the image generation model to obtain a personalized image output by the image generation model corresponding to the reference image and the description text; where the image generation model is obtained through the above image generation model processing method.

[0179] In one embodiment, the reference data processing module 1604 is further configured to input the reference image and the description text into the image generation model, generate a first personalized image by the image generation model based on the reference image and the description text; generate a second personalized image by the image generation model based on the reference image; generate a third personalized image by the image generation model based on the description text; and perform weighted fusion on the first personalized image, the second personalized image, and the third personalized image to obtain a personalized image corresponding to the reference image and the description text.

[0180] Each module in the above image generation model processing device or image generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0181] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as shown in Figure 17 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image generation model data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an image generation model processing or an image generation method.

[0182] Those skilled in the art can understand that Figure 17 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0183] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0184] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0185] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0188] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0189] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for processing an image generation model, characterized in that, The method includes: Obtaining a sample image and descriptive text for an image object in the sample image; Respectively extracting image features of the sample image and text features of the descriptive text through an image generation model to be trained; Updating object text features in the text features that represent the image object through the image features to obtain updated text features; Obtaining image generation features according to the updated text features, and generating a personalized image corresponding to the sample image and the descriptive text according to the image generation features; Updating the image generation model to be trained based on the personalized image and sub-features in the image generation features that represent the image object, and obtaining a trained image generation model; the trained image generation model is used to generate a corresponding personalized image according to the input image and text.

2. The method according to claim 1, wherein The updating the object text features in the text features that represent the image object through the image features to obtain updated text features includes: Determining object text features representing the image object from the text features; Fusing the image features and the object text features to obtain object fusion features; Replacing the object text features in the text features with the object fusion features to obtain updated text features.

3. The method according to claim 2, wherein The determining object text features representing the image object from the text features includes: Obtaining a matching relationship between the image object in the sample image and descriptive entity words in the descriptive text; Determining object descriptive entity words for describing the image object from the descriptive text according to the matching relationship; Determining object text features representing the image object from the text features according to the object descriptive entity words.

4. The method according to claim 3, wherein The method further includes: Segmenting the sample image into image blocks including different image entities, and dividing each descriptive entity word from the descriptive text; Performing graphic-text matching between the image blocks and the descriptive entity words to obtain a graphic-text matching result; Determining a matching relationship between the image object in the sample image and the descriptive entity words in the descriptive text according to the graphic-text matching result.

5. The method according to claim 4, wherein The performing graphic-text matching between the image blocks and the descriptive entity words to obtain a graphic-text matching result includes: Matching image block features of the image blocks with entity word features of the descriptive entity words to obtain a first matching result; Matching image entity labels associated with the image blocks with entity word features of the descriptive entity words to obtain a second matching result; Obtaining a graphic-text matching result according to the first matching result and the second matching result.

6. The method according to claim 1, wherein The obtaining image generation features according to the updated text features, and generating a personalized image corresponding to the sample image and the descriptive text according to the image generation features includes: Based on an attention mechanism, obtaining attention features according to the updated text features and the image features; Performing image decoding according to the attention features to generate a personalized image corresponding to the sample image and the descriptive text.

7. The method according to claim 6, wherein Based on the attention mechanism, obtaining attention features according to the updated text features and the image features, including: Based on the cross-attention mechanism, determining attention parameters according to the updated text features and the image features; Obtaining weight parameters according to the query parameter and the key parameter in the attention parameters; Obtaining attention features according to the weight parameters and the value parameter in the attention parameters.

8. The method according to claim 1, characterized in that Updating the image generation model to be trained based on the personalized image and the sub-features representing the image object in the image generation features, and obtaining the trained image generation model, including: Determining a first loss according to the personalized image and the sample image; Determining a second loss according to the sub-features representing the image object in the image generation features and the object image patch; the object image patch is the image patch including the image object in the sample image; Obtaining an image generation loss based on the first loss and the second loss, and updating the image generation model to be trained through the image generation loss to obtain the trained image generation model.

9. The method according to any one of claims 1 to 8, characterized in that The obtaining of the sample image and the description text for the image object in the sample image includes: Obtaining a sample image including an image object, and extracting the description text for the image object based on the sample image; Extracting the image features of the sample image and the text features of the description text respectively through the image generation model to be trained, including: Extracting the image features of the sample image through the image feature extraction network in the image generation model to be trained; Extracting the text features of the description text through the text feature extraction network in the image generation model to be trained.

10. An image generation method, characterized in that, The method includes: Obtaining a reference image and the description text for the image object in the reference image; Inputting the reference image and the description text into the image generation model, and obtaining the personalized image corresponding to the reference image and the description text output by the image generation model; Wherein, the image generation model is obtained by the image generation model processing method according to any one of claims 1 to 9.

11. The method according to claim 10, wherein The inputting the reference image and the description text into the image generation model, and obtaining the personalized image corresponding to the reference image and the description text output by the image generation model includes: Inputting the reference image and the description text into the image generation model, and generating a first personalized image by the image generation model based on the reference image and the description text; Generating a second personalized image by the image generation model based on the reference image; Generating a third personalized image by the image generation model based on the description text; Performing weighted fusion on the first personalized image, the second personalized image and the third personalized image to obtain the personalized image corresponding to the reference image and the description text.

12. An image generation model processing device, characterized in that, The device includes: A sample data acquisition module, configured to acquire a sample image and the description text for the image object in the sample image; A sample feature extraction module, configured to respectively extract the image feature of the sample image and the text feature of the description text through an image generation model to be trained; A text feature update module, configured to update the object text feature representing the image object in the text feature through the image feature to obtain an updated text feature; An image generation processing module, configured to obtain an image generation feature according to the updated text feature, and generate a personalized image corresponding to the sample image and the description text according to the image generation feature; A model update module, configured to update the image generation model to be trained based on the personalized image and the sub-feature representing the image object in the image generation feature, and obtain a trained image generation model; the trained image generation model is configured to generate a corresponding personalized image according to the input image and text.

13. An image generation device, characterized in that, The apparatus includes: A reference data acquisition module, configured to acquire a reference image and a description text for an image object in the reference image; A reference data processing module, configured to input the reference image and the description text into an image generation model, and obtain a personalized image output by the image generation model and corresponding to the reference image and the description text; Wherein, the image generation model is obtained by the image generation model processing method according to any one of claims 1 to 9.

14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 or claims 10 to 11 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 or claims 10 to 11 are implemented.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 or claims 10 to 11 are implemented.