Multi-person image generation method and device, equipment and storage medium
By generating the base map and performing human body and face detection, binding the portrait image and only repainting the face, the problems of low quality of multi-person image generation and difficulty in automatic matching in the prior art are solved, and high-quality multi-person image generation and automatic character matching are achieved.
Patent Information
- Application Number
- CN202411754006.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, when generating multi-person images, the image generation quality is not high and it is difficult to automatically match characters that need to be replaced, resulting in frequent manual manual operations of manual checksums.
By receiving multiple portrait images and scene description words input by the user, the base map is generated and the human body and face detection is performed, the portrait images and detection results are bound, and only the face is redrawn to ensure that the quality of the base map generation and the matching of characters are consistent.
The image quality generated by multi-person images is improved, the need for manual checksum operation is reduced, and the function of automatic matching and replacement of characters is realized.
Smart Images

Figure CN119941924A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for generating multi-person images. Background Art
[0002] Existing technology cannot generate multiple people at a time. The general solution for generating multiple people is as follows:
[0003] Solution 1: Manual face swap:
[0004] It is necessary to generate some images according to the scene description word prompt, manually check whether they meet the requirements, and then manually cut out a person's face and replace it with the face of the target portrait character (the image of the person to be generated). When this solution is used to generate multiple images at a time (for example, portraits character_1 and character_2), it is necessary to manually check whether the generated images meet the requirements, and manually confirm which person in the generated image should be replaced with character_1 and which person should be replaced with character_2.
[0005] Solution 2: Automatic generation
[0006] It is necessary to generate an image according to the prompt, and then use dino+sam (open set target detection algorithm and segmentation algorithm) to segment the human body separately. For each human body, use the instantID (single person artistic portrait generation) algorithm to redraw it. This solution needs to fix the random number seed of the first generated image, and use this seed to initialize when redrawing to ensure that the background remains unchanged except for the person being redrawn. The problem with this solution is that the segmentation method of dino+sam can easily cause target segmentation errors, and it is difficult to distinguish between two subjects of the same gender. Furthermore, since the second redrawing of each person depends on the seed generated for the first time, this requires that the basic Wensheng graph model used in the two generations (the first image generated according to the prompt and the second image redrawn according to the segmentation result) must be consistent. Since there are many variants of the Wensheng graph basic model, different variants may have different structures and different generation effects. The newer variants have higher quality images generated by a given prompt and more meet the prompt description (for example, gender accuracy, number of people accuracy, body integrity of people, etc.). However, the instantID algorithm used in the redraw is a plug-in on the Vincent graph base model, which is related to the model structure of the Vincent graph. The update speed of this plug-in is not that fast, which means that if you use a newer Vincent graph base model variant, you cannot find the corresponding instantID plug-in. This will limit the image generation quality of the overall solution.
[0007] Therefore, there is an urgent need for a multi-person image generation method with higher image generation quality. Summary of the invention
[0008] The present disclosure provides a method, apparatus, device and storage medium for generating a multi-person image.
[0009] According to a first aspect of the present disclosure, a method for generating a multi-person image is provided. The method comprises:
[0010] Receive multiple portrait images and scene description words input by the user;
[0011] generating a base map based on the scene description words;
[0012] Performing human body detection and face detection on the base image to obtain a human body frame and a face frame;
[0013] Binding the human body frame, the face frame and the portrait image;
[0014] According to the face images in the bound portrait images, the face frame area images in the corresponding base image are redrawn to obtain a multi-person image.
[0015] According to the above aspects and any possible implementation, an implementation is further provided.
[0016] The scene description words are used to describe the scene of the multi-person image that the user expects to generate;
[0017] The generating a base map based on the scene description words comprises:
[0018] The scene description words are input into the text graph model to obtain a base map.
[0019] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the method further includes:
[0020] A quality inspection is performed on the face frame, and if the quality inspection result does not meet a preset condition, a base map is regenerated according to the scene description word.
[0021] According to the above aspects and any possible implementation, an implementation is further provided.
[0022] The performing quality inspection on the face frame, and if the quality inspection result does not meet a preset condition, regenerating a base map according to the scene description word, includes:
[0023] If the face frame is empty, regenerating the base image according to the scene description words;
[0024] If the number of face frames is less than the number of portrait images input by the user, regenerating a base image according to the scene description words;
[0025] If the number of face frames is greater than or equal to the number of portrait images input by the user, then after screening the face frames according to the area size of the face region and the confidence of the face frames, if the number of remaining face frames is less than the number of portrait images input by the user, the base map is regenerated according to the scene description words.
[0026] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the method further includes:
[0027] According to the number of face frames that meet the preset conditions in the final quality detection results, the number of human frames is determined as the target number;
[0028] Performing human body detection on the base image to obtain a human body frame also includes:
[0029] It is determined whether the number of human body frames obtained by human body detection is equal to the target number; if it is less than the target number, the detection is continued.
[0030] According to the above aspects and any possible implementation, an implementation is further provided.
[0031] The portrait image carries a gender identifier of the portrait;
[0032] The step of binding the human body frame, the face frame and the portrait image comprises:
[0033] Inputting the image in the base map corresponding to the human body frame into a pre-trained gender recognition model to obtain the gender identification of the human body frame;
[0034] Binding the human body frame with the portrait image according to the gender identifier of the human body frame and the gender identifier of the portrait;
[0035] Sort the face frames in descending order according to their area sizes, take the first n face frames as target face frames, and randomly match the target face frames with the human body frames one by one;
[0036] Wherein, n=the number of targets.
[0037] According to the above aspects and any possible implementation, an implementation is further provided.
[0038] The face frame area image in the corresponding base image is redrawn according to the face image in the bound portrait image to obtain a multi-person image, including:
[0039] Perform mask processing on the face frame area image in the corresponding base image to obtain a mask image;
[0040] The face image in the portrait image bound to the face frame and the mask image are passed into the inpainting+instantID algorithm to redraw the face area in the base image;
[0041] When all the face images in the portrait image are redrawn into the base image, a multi-person image is obtained.
[0042] According to a second aspect of the present disclosure, a multi-person image generation device is provided. The device comprises:
[0043] A receiving module, used for receiving multiple portrait images and scene description words input by a user;
[0044] A generating module, used for generating a base map based on the scene description words;
[0045] A detection module, used to perform human body detection and face detection on the base image to obtain a human body frame and a face frame;
[0046] A binding module, used for binding the human body frame, the face frame and the portrait image;
[0047] The redrawing module is used to redraw the face frame area image of the corresponding base image according to the face image in the bound portrait image to obtain a multi-person image.
[0048] According to a third aspect of the present disclosure, an electronic device is provided, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described above is implemented.
[0049] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0050] The embodiments of the present disclosure provide a method, device, equipment and storage medium for generating a multi-person image. A basic Wensheng map model is used to generate a basic base map according to the prompt scene prompt word, and the basic base map is processed. The purpose is to confirm whether the generated base map meets the requirements in the first stage, and then the people generated in the base map are matched. The purpose is to automatically match the character portrait image that needs to be replaced to the generated base map. Then, according to the matching results obtained, each character portrait image is replaced with the person in the base map. This solution decouples the two models of base map generation and redrawing generation, so that the base map can be generated using the most advanced Wensheng map basic model, ensuring the quality of base map generation, for example, the normal structure of human body limbs and normal aesthetic light. The additional matching method used in this solution ensures that the character and the person in the base map are consistent. This solution only redraws the face, ensuring the normal structure of the base map limbs and the beauty of the original image.
[0051] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0053] Figure 1 A flowchart of a method for generating a multi-person image according to an embodiment of the present disclosure is shown;
[0054] Figure 2 A block diagram of a multi-person image generating device according to an embodiment of the present disclosure is shown;
[0055] Figure 3 A schematic block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0057] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0058] The present disclosure adopts the technical route of generating a base map → matching a person → only changing the face. In this technical route, the base map generation model and the face changing model are further decoupled, so that the base map model can adopt the latest Wensheng graph model architecture, ensuring the beauty of the image and the normality of the human body structure.
[0059] Figure 1 A flowchart of a method 100 for generating a multi-person image according to an embodiment of the present disclosure is shown. The method 100 includes:
[0060] Step 110: receiving multiple portrait images and scene description words input by the user.
[0061] In some embodiments, the scene description word prompt is used to describe the scene of the multi-person image that the user expects to generate. Multi-person generation can meet many group photo scenes, such as wedding dress scenes, etc. For example, a portrait of male y and a portrait of female x need to generate a costume wedding photo of male y and female x.
[0062] Step 120: Generate a base map based on the scene description words.
[0063] In some embodiments, generating a base map based on the scene description words includes: inputting the scene description words into a text graph model to obtain a base map. For example, using a FLUX text graph model.
[0064] In some embodiments, after each prompt is input, the base map generation module generates a picture from a random noise, but this picture may not meet the user's requirements. For example, the generated face is too small, the generated face has too many sides, and the back of the person is generated. These pictures whose faces do not meet the requirements, although they meet the description of the prompt for the picture, are not conducive to inserting specific characters. Therefore, it is necessary to screen out the base map that is conducive to inserting the face of a specific character. If it does not meet the requirements, it is regenerated. By decoupling the generation of the base map and the redrawing of the face, the quality control of the base map generation part can be achieved, ensuring the generation quality of the base map.
[0065] Step 130, performing human body detection and face detection on the base image to obtain a human body frame and a face frame.
[0066] In some embodiments, the face frame is quality-checked, and if the quality check result does not meet the preset conditions, the base map is regenerated according to the scene description words. The face frame is quality-checked, and if the quality check result does not meet the preset conditions, the base map is regenerated according to the scene description words, including: if the face frame is empty, the base map is regenerated according to the scene description words; if the number of face frames is less than the number of portrait images input by the user, the base map is regenerated according to the scene description words; if the number of face frames is greater than or equal to the number of portrait images input by the user, after screening the face frames according to the area size of the face area and the confidence of the face frame, if the number of remaining face frames is less than the number of portrait images input by the user, the base map is regenerated according to the scene description words. For example, in the above-mentioned generation of the ancient wedding photo of male y and female x, that is, the generation of a two-person image, (a) when there is no face, it is considered not to meet the requirements and the base map is regenerated; (b) according to the descending order of the face area detected by the detector, the faces with an area smaller than a fixed pre-made thr_area (generally 64*64) and a confidence level smaller than a fixed pre-made thr_prob (generally 0.6, the confidence level is a parameter automatically given by the face detector when performing face recognition) are deleted. The number of remaining faces is less than 2 (this value is determined according to the specific number of people in the multi-person image to be generated), which is considered not to meet the requirements and the base map is regenerated. (c) After the above (b) is completed, the first 3 faces are taken in the order of descending order of face area. If there are less than 3, all are taken (theoretically 2, because less than 2 are filtered out and regenerated in step b). Of course, the above example is limited to the case where the embodiment is to generate a two-person image. If it is a group photo of more than 3 people, it is not limited by this example.
[0067] In some embodiments, according to the number of face frames that meet preset conditions in the final quality detection result, the number of human body frames is determined as the target number; human body detection is performed on the base image to obtain human body frames, and it also includes: judging whether the number of human body frames obtained by human body detection is equal to the target number; if it is less than the target number, continue the detection. For example, in the case of generating a photo of two people in the above example, assuming that the number of face frames is 3, then there must be a corresponding number of human body frames in the base image. Using an open set detection model such as dino, the given text is "person", so all the detection frames of person can be detected, that is, the human body frames are obtained.
[0068] Step 140: Bind the human body frame, the face frame and the portrait image.
[0069] In some embodiments, the portrait image carries a portrait gender identifier; the binding of the human body frame, the face frame and the portrait image includes: inputting the image in the base image corresponding to the human body frame into a pre-trained gender recognition model to obtain the gender identifier of the human body frame; binding the human body frame to the portrait image according to the gender identifier of the human body frame and the portrait gender identifier; sorting the face frames in descending order according to their area sizes, taking the first n face frames as target face frames, and randomly matching the target face frames with the human body frames one by one; wherein n=the number of targets.
[0070] In some embodiments, a pre-trained gender recognition model can obtain a gender label by inputting the image in the base map corresponding to the human body frame, thereby determining the gender of the human body frame, and then bind the human body frame with the same gender to the portrait image according to the gender identifier carried by the portrait image.
[0071] In some embodiments, the pre-trained gender recognition model is a clip model, wherein the clip model includes an image feature extraction part and a text feature extraction part, and the image in each human body frame is clipped out, and the image feature extraction part of the clip model is used to extract the features of the clipped image, and the text feature extraction part of the clip model is used to extract the text features of the two words "boy, man, gentleman, male" and "woman, lady, girl, female". Then, training is performed, and subsequently, only the image of the human body frame in the base map needs to be input to obtain the gender identification.
[0072] In some embodiments, the portrait image male y may be bound to the human body frame and the face frame in the base image first, and then the face of the portrait image female x may be redrawn after the face is redrawn. At this time, each time the human body frame and the face frame are matched, it is only necessary to calculate the face frame with the largest area and assign it to the corresponding human body frame. Of course, each time the match is completed, the face frame and the human body frame are considered to be occupied. For example, the IoU of the face frame and the human body frame is calculated, and the face frame with the largest IoU is assigned to the human body frame. It is also possible to calculate the area of the face frame, which is not specifically limited here.
[0073] In some embodiments, the portrait image, the face frame and the body frame may be completely matched before the face is redrawn one by one and / or synchronously. For example, the face frames are sorted in descending order according to the size of the area, the first n face frames are taken as the target face frames, and the target face frames are randomly matched with the body frames one by one; wherein n = the number of targets. You can also refer to the same method as described above, that is, calculate the IoU of the face frame and the body frame, and match them one by one. Note that those that have been matched are considered occupied and cannot participate in the next match. The priority is given to face frames with larger areas so that a higher-definition face image can be obtained when the face is redrawn later.
[0074] In this way, the binding of the human frame and the portrait image, and the binding of the human frame and the face frame are completed, that is, the binding of the portrait image and the face frame is indirectly realized, and the binding relationship is: portrait image-human frame-face frame. This solution ensures that the character target in the character portrait image and the base image is consistent through face detection, human gender recognition, etc. Only the face is redrawn, ensuring the normal structure of the base image limbs and the beauty of the original image.
[0075] In some embodiments, if the gender labels of the identified human body frames are, for example, male 1 and male 2, but the user needs to generate ancient costume wedding photos of male y and female x, then the gender cannot be covered by the portrait image. At this time, the user needs to be prompted to enter a more accurate scene prompt word, or in the process of generating the base map, the gender identifier carried by the portrait image uploaded by the user is integrated to assist in the generation of a more accurate base map.
[0076] Step 150 , redrawing the face frame area image in the corresponding base image according to the face images in the bound portrait images to obtain a multi-person image.
[0077] In some embodiments, the facial images in the bound portrait images are facially redrawn on the face frame area images in the corresponding base image to obtain a multi-person image, including: masking the face frame area images in the corresponding base image to obtain a mask image; passing the face images in the portrait images bound to the face frame and the mask image into the inpainting+instantID algorithm to redraw the face area in the base image; when the face images in all the portrait images are redrawn into the base image, a multi-person image is obtained.
[0078] In some embodiments, the sam segmentation model is used to cut out the face area in the base image and perform mask processing, that is, to identify which area in the base image is to be redrawn in the subsequent facial redrawing, and the mask image and the corresponding face in the character portrait image are passed into the inpainting+instantID algorithm to redraw the face area in the base image.
[0079] Among them, the inpainting algorithm, also known as image processing technology, is an algorithm used to repair missing parts in an image. Its purpose is to fill in damaged or missing areas in the image so that the image looks natural and complete. The inpainting algorithm has a wide range of applications in the fields of image processing and computer vision, such as repairing old photos, removing unwanted objects in images, and filling in occlusions in videos. instantID is an AI-driven image generation tool. Through deep learning and neural network technology, it provides users with an innovative character image creation experience, and is committed to converting the text descriptions entered by users into vivid and realistic character visual artworks while ensuring the integrity of the identity information of the characters in the image. instantID is a powerful solution based on a diffusion model. The designed plug-and-play module can skillfully handle various styles of image personalization using only a single facial image while ensuring high fidelity.
[0080] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0081] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.
[0082] Figure 2 FIG. 2 is a block diagram of a multi-person image generating apparatus 200 according to an embodiment of the present disclosure.
[0083] like Figure 2 As shown, the device 200 includes:
[0084] A receiving module 210 is used to receive multiple portrait images and scene description words input by a user;
[0085] A generating module 220, configured to generate a base map based on the scene description words;
[0086] A detection module 230, configured to perform human body detection and face detection on the base image to obtain a human body frame and a face frame;
[0087] A binding module 240, used to bind the human body frame, the face frame and the portrait image;
[0088] The redrawing module 250 is used to redraw the face frame area image of the corresponding base image according to the face images in the bound portrait images to obtain a multi-person image.
[0089] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0090] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0091] Figure 3 A schematic block diagram of an electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0092] The electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a ROM 302 or a computer program loaded from a storage unit 308 into a RAM 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0093] A number of components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0094] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as method 100. For example, in some embodiments, the method 100 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform the method 100 in any other appropriate manner (e.g., by means of firmware).
[0095] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0096] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0097] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0099] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0100] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0101] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0102] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for generating a multi-person image, characterized in that: include: Receive multiple portrait images and scene description words input by the user; generating a base map based on the scene description words; Performing human body detection and face detection on the base image to obtain a human body frame and a face frame; Binding the human body frame, the face frame and the portrait image; According to the face images in the bound portrait images, the face frame area images in the corresponding base image are redrawn to obtain a multi-person image.
2. The method according to claim 1, characterized in that: The scene description words are used to describe the scene of the multi-person image that the user expects to generate; The generating a base map based on the scene description words comprises: The scene description words are input into the text graph model to obtain a base map.
3. The method according to claim 1, characterized in that The method further comprises: A quality inspection is performed on the face frame, and if the quality inspection result does not meet a preset condition, a base map is regenerated according to the scene description word.
4. The method according to claim 3, characterized in that The performing quality inspection on the face frame, and if the quality inspection result does not meet a preset condition, regenerating a base map according to the scene description word, includes: If the face frame is empty, regenerating the base image according to the scene description words; If the number of face frames is less than the number of portrait images input by the user, regenerating a base image according to the scene description words; If the number of face frames is greater than or equal to the number of portrait images input by the user, then after screening the face frames according to the area size of the face region and the confidence of the face frames, if the number of remaining face frames is less than the number of portrait images input by the user, the base map is regenerated according to the scene description words.
5. The method according to claim 4, characterized in that The method further comprises: According to the number of face frames that meet the preset conditions in the final quality detection results, the number of human frames is determined as the target number; Performing human body detection on the base image to obtain a human body frame also includes: It is determined whether the number of human body frames obtained by human body detection is equal to the target number; if it is less than the target number, the detection is continued.
6. The method according to claim 5, characterized in that The portrait image carries a gender identifier of the portrait; The step of binding the human body frame, the face frame and the portrait image comprises: Inputting the image in the base map corresponding to the human body frame into a pre-trained gender recognition model to obtain the gender identification of the human body frame; Binding the human body frame with the portrait image according to the gender identifier of the human body frame and the gender identifier of the portrait; Sort the face frames in descending order according to their area sizes, take the first n face frames as target face frames, and randomly match the target face frames with the human body frames one by one; Wherein, n=the number of targets.
7. The method according to claim 6, characterized in that The face frame area image in the corresponding base image is redrawn according to the face image in the bound portrait image to obtain a multi-person image, including: Perform mask processing on the face frame area image in the corresponding base image to obtain a mask image; The face image in the portrait image bound to the face frame and the mask image are passed into the inpainting+instantID algorithm to redraw the face area in the base image; When all the face images in the portrait image are redrawn into the base image, a multi-person image is obtained.
8. A multi-person image generation device, characterized in that: include: A receiving module, used for receiving multiple portrait images and scene description words input by a user; A generating module, used for generating a base map based on the scene description words; A detection module, used to perform human body detection and face detection on the base image to obtain a human body frame and a face frame; A binding module, used for binding the human body frame, the face frame and the portrait image; The redrawing module is used to redraw the face frame area image of the corresponding base image according to the face image in the bound portrait image to obtain a multi-person image.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.