Multi-view portrait generation method and device, computer equipment and storage medium

By combining text generation and multimodal large language model, the character portraits and postures in the generation process are optimized, and the problems of controllability and consistency of traditional multi-view portrait generation methods are solved, and efficient and controllable multi-view portrait generation is achieved.

CN120495465APending Publication Date: 2025-08-15SHANGHAI YIQIUCHUN CULTURE MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463298.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The traditional multi-view portrait generation method has obvious limitations in controllability, intelligence and multi-view consistency, and the generation efficiency is inefficient.

Method used

By obtaining the initial text input by the user, using the text generation model for amplification, combining the character portrait generation model and the multimodal large language model to generate the perspective requirements text, clothing text, character posture text and scene text, and finally generating the complete image in the multi-view portrait generation model, and introducing the portrait optimization model and attitude control module to improve image quality and consistency.

Benefits of technology

It realizes user-controllable multi-view portrait generation, improves the controllability and image quality of the generated results, and ensures image consistency and visual effects at different perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495465A_ABST
    Figure CN120495465A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-view portrait generation method and device, computer equipment and a storage medium. The multi-view portrait generation method comprises the following steps: acquiring an initial text input by a user and used for multi-view portrait generation; the initial text is input into a text generation model for text content amplification processing, and an amplified text is obtained; the amplified text comprises a portrait text; inputting the amplified text into a human portrait generation model for human portrait generation to obtain an initial human portrait; inputting the initial text, the amplified text and the initial portrait into a multi-modal large language model, and generating a view angle requirement text, a clothing text, a figure posture text and a scene text through the multi-modal large language model; and inputting the image and the initial portrait into a multi-view portrait generation model for image generation to obtain a complete image of the initial portrait under multiple different views. By adopting the method, the controllability of the generated result and the image quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a multi-perspective portrait generation method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the rapid development of image generation models and technologies in the field of artificial intelligence, portrait generation has become an important research hotspot in image generation. Among them, multi-view portrait generation has attracted much attention due to its wide application in virtual fitting, film and television production, and other fields.

[0003] Traditional technologies use image generation models to generate multi-perspective portraits. However, the generation scheme of traditional technologies is carried out in steps, and the control of multimodal information (such as text, images, and posture) is separate. As a result, the generation results still have obvious limitations in controllability, intelligence, and multi-perspective consistency. At the same time, the efficiency of multiple generation and debugging is relatively low. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, computer equipment and storage medium that can quickly generate multi-perspective portraits with high consistency to address the above technical problems.

[0005] In a first aspect, the present application provides a multi-perspective portrait generation method, the method comprising:

[0006] Obtaining the initial text input by the user for generating multi-view portraits;

[0007] Inputting the initial text into a text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes portrait text;

[0008] Inputting the amplified text into a character portrait generation model to generate a character portrait, thereby obtaining an initial character portrait;

[0009] Inputting the initial text, the augmented text, and the initial portrait into a multimodal large language model, and generating a viewing angle requirement text, a clothing text, a character posture text, and a scene text through the multimodal large language model;

[0010] The initial portrait, the viewing angle requirement text, the clothing text, the character posture text and the scene text are input into a multi-view portrait generation model for image generation to obtain complete images of the initial portrait at multiple different viewing angles.

[0011] In one embodiment, after inputting the portrait text into a character portrait generation model to generate a character portrait and obtaining an initial character portrait, the method further includes:

[0012] Performing text processing on the initial portrait to obtain initial portrait text features;

[0013] Comparing the initial portrait text feature with the portrait text to obtain a similarity between the initial portrait text feature and the portrait text;

[0014] Determining whether the similarity is greater than or equal to a preset threshold;

[0015] If the similarity is less than a preset threshold, the portrait text and the initial portrait are input into a portrait optimization model, and the initial portrait is optimized by the portrait optimization model to improve the similarity between the initial portrait text features corresponding to the initial portrait and the portrait text;

[0016] If the similarity is greater than or equal to a preset threshold, the initial portrait is not processed.

[0017] In one embodiment, the portrait text includes facial expression description text, hair texture description text and clothing pattern description text, and the portrait optimization model includes an attention mechanism module, which extracts the facial expression, hair texture and clothing pattern of the initial portrait through the attention mechanism module, optimizes the facial expression based on the facial expression description text, optimizes the hair texture based on the hair texture description text, and optimizes the clothing pattern based on the clothing pattern description text.

[0018] In one embodiment, the multi-view portrait generation model includes a posture control module, and the character posture text and the initial portrait are input into the posture control module for posture generation to obtain target posture information that matches the skeleton key points of the initial portrait.

[0019] In one embodiment, the character portrait generation model is trained by the following steps:

[0020] Based on a pre-trained image generation model with general image generation capabilities, a portrait generation model is obtained by training using a collection of portrait images with annotated text.

[0021] The text annotations include age, gender, skin color, clothing, posture and lighting scene.

[0022] In one embodiment, before obtaining the initial text input by the user for generating the multi-view portrait, the method further includes:

[0023] A text input interface is displayed, which includes a portrait text input area, a clothing text input area, and a background text input area.

[0024] In a second aspect, the present application further provides a multi-perspective portrait generation device, the device comprising:

[0025] An acquisition module, used to obtain initial text input by a user for generating a multi-view portrait;

[0026] A first text generation module is configured to input the initial text into a text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes a portrait text;

[0027] A portrait generation module, configured to input the portrait text into a character portrait generation model to generate a character portrait and obtain an initial character portrait;

[0028] a second text generation module, configured to input the initial text, the augmented text, and the initial portrait into a multimodal large language model, and generate a viewing angle requirement text, a clothing text, a character posture text, and a scene text through the multimodal large language model;

[0029] The multi-view image generation module is used to input the initial portrait, the view requirement text, the clothing text, the character posture text and the scene text into the multi-view portrait generation model to generate an image, thereby obtaining a complete image of the initial portrait at multiple different view angles.

[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned multi-perspective portrait generation when executing the computer program.

[0031] In a fourth aspect, the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned multi-perspective portrait generation are implemented.

[0032] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps in the above-mentioned multi-view portrait generation when executed by a processor.

[0033] In the above-mentioned multi-perspective portrait generation method, device, computer device and storage medium, the initial text input by the user is first obtained, and the initial text is amplified to obtain the amplified text. The portrait text in the amplified text is used to generate the character portrait to obtain the initial portrait. Then, the initial text, the amplified text and the initial portrait are input into the multimodal large language model to generate the perspective requirement text, the clothing text, the character posture text and the scene text. In this way, the perspective requirement text, the clothing text, the character posture text and the scene text can be used to generate a multi-perspective scene. Based on the initial portrait, the multi-perspective portrait generation model can generate a complete image of different perspectives in a specific scene. In the multi-perspective portrait generation method provided by the present invention, the initial text, the portrait text and the initial portrait are simultaneously input into the multimodal large language model for generation. The resulting perspective requirement text, clothing text, the character posture text and the scene text have a higher coupling strength with the initial portrait and a higher matching degree with the initial text input by the user, thereby realizing user-controllable multi-perspective portrait image generation and effectively improving the controllability and image quality of the generated results. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 A diagram showing an application environment of a multi-view portrait generation method according to an embodiment;

[0036] Figure 2 1 is a flow chart of a multi-view portrait generation method according to an embodiment;

[0037] Figure 3 is a schematic flow chart of a multi-view portrait generation method according to another embodiment;

[0038] Figure 4 is a structural block diagram of a multi-view portrait generation device in one embodiment;

[0039] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0041] The multi-view portrait generation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.

[0042] In an exemplary embodiment, Figure 2 As shown, a multi-view portrait generation method is provided, which is applied to Figure 1 The server 104 in the example is used as an example to illustrate the process, including the following steps 202 to 210. Among them:

[0043] Step 202: Obtain the initial text input by the user for generating a multi-view portrait.

[0044] For example, the initial text input by the user is a customized requirement of the user, and serves as a prompt for the subsequent image generation model. The initial text includes keywords used in the image generation model. Alternatively, the user can input the initial text on the terminal 102, and the server 104 obtains the text information input by the user on the terminal 102.

[0045] Exemplarily, the initial text input by the user may include at least one of portrait text, clothing text, and background text.

[0046] Step 204 : Input the initial text into the text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes the portrait text.

[0047] Exemplarily, the text generation model may be a large language model. After the initial text is input into the large language model, the initial text may be amplified by the large language model to obtain an amplified text with greater information content.

[0048] For example, the text generation model can be a conversational pre-trained autoregressive model using a Transformer architecture. The text generation model augments the initial text input by the user, using the initial text as a semantic guide or constraint. This means that the augmented text will not contradict the initial text and will only augment the initial text.

[0049] Exemplarily, the text generation model is preset so that the portrait text in the amplified text includes at least a text description of the region, a text description of gender, a text description of age, a text description of skin color, a text description of facial features, a text description of hairstyle, and a text description of facial expression, so as to provide semantic guidance or restriction conditions for the subsequent image generation model to perform the generation and correction process.

[0050] Step 206: Input the amplified text into a character portrait generation model to generate a character portrait, thereby obtaining an initial character portrait.

[0051] Exemplarily, the character portrait generation model generates an image based on the portrait text in the augmented text.

[0052] Exemplarily, the character portrait generation model performs a multi-step denoising process based on the portrait text to generate an initial portrait that matches the user's needs.

[0053] In step 208 , the initial text, the augmented text, and the initial portrait are input into a multimodal large language model, and the multimodal large language model is used to generate a viewing angle requirement text, a clothing text, a character posture text, and a scene text.

[0054] For example, to form a complete multi-view image, not only the person's portrait is required, but also their clothing, posture from different perspectives, and the scene they are in. In this step, the multimodal large language model performs semantic optimization on the initial text, the augmented text, and the initial portrait to obtain the perspective requirement text, clothing text, character posture text, and scene text.

[0055] For example, a large multimodal language model encodes text input and image input to generate text features and image features, respectively. These features are then mapped to a fused multimodal feature set to generate viewpoint requirement text, clothing text, character pose text, and scene text. The text features include the original text and augmented text, while the image features include the original portrait.

[0056] In step 210 , the initial portrait, the view requirement text, the clothing text, the character posture text, and the scene text are input into a multi-view portrait generation model for image generation, thereby obtaining complete images of the initial portrait at multiple different view angles.

[0057] For example, the multi-view portrait generation model is trained through multiple sample sets. Specifically, the training process is to stitch images of the same person taken from different viewpoints in the same scene into a composite image for training, so that the model can capture and integrate the appearance consistency and spatial structure information from different viewpoints, ensuring that the generated portrait is both consistent and meets the visual diversity requirements from various viewpoints.

[0058] The proposed controllable multi-view portrait generation method based on multimodal collaboration utilizes a large language model, multimodal collaboration, and a specialized multi-view portrait generation model. By injecting cross-view consistency conditions into the generation process, the method ensures the coherence and uniformity of facial features, clothing details, background style, and lighting effects across all viewpoints. Multi-view splicing training and customized prompts are used to treat images of the same person from different viewpoints as a whole, enabling the simultaneous generation of multiple highly consistent images. This avoids the repeated generation of images using traditional methods and improves efficiency.

[0059] In the above-mentioned multi-perspective portrait generation method, the initial text input by the user is first obtained, and the initial text is amplified to obtain the amplified text. The portrait text in the amplified text is used to generate the character portrait to obtain the initial portrait. Then, the initial text, the amplified text and the initial portrait are input into the multimodal large language model to generate the perspective requirement text, the clothing text, the character posture text and the scene text. In this way, the perspective requirement text, the clothing text, the character posture text and the scene text can be used to generate a multi-perspective scene. Based on the initial portrait, the multi-perspective portrait generation model can generate a complete image of different perspectives in a specific scene. In the multi-perspective portrait generation method provided by the present invention, the initial text, the portrait text and the initial portrait are simultaneously input into the multimodal large language model for generation. The coupling strength between the perspective requirement text, the clothing text, the character posture text and the scene text obtained in this way and the initial portrait is higher, and the matching degree with the initial text input by the user is higher, thereby realizing user-controllable multi-perspective portrait image generation, effectively improving the controllability and image quality of the generated results.

[0060] In one embodiment, after step 206 inputs the portrait text into the character portrait generation model to generate the character portrait and obtains the initial portrait, the multi-view portrait generation method further includes: textually processing the initial portrait to obtain initial portrait text features; comparing the initial portrait text features with the portrait text to obtain a similarity between the initial portrait text features and the portrait text; determining whether the similarity is greater than or equal to a preset threshold; if the similarity is less than the preset threshold, inputting the portrait text and the initial portrait into the portrait optimization model, and optimizing the initial portrait through the portrait optimization model to improve the similarity between the initial portrait text features corresponding to the initial portrait and the portrait text; if the similarity is greater than or equal to the preset threshold, not processing the initial portrait.

[0061] For example, the initial portrait serves as the basis for the final multi-view portrait. It is necessary to determine whether the initial portrait needs to be optimized. The initial portrait generated by the portrait generation model is converted into initial portrait text features to form text information for easy comparison with the portrait text. Alternatively, the feature extraction model can be used to extract the visual feature vector of the initial portrait to form the initial portrait text features.

[0062] For example, the similarity ratio can be calculated using the following formula:

[0063]

[0064] Among them, V image is the initial portrait text feature, and V text is the initial text.

[0065] Optionally, the preset threshold may be 0.8, and when the similarity is greater than or equal to 0.8, it is considered that no optimization is required.

[0066] When the similarity is less than 0.8, the initial portrait needs to be optimized and input into the portrait optimization model. The portrait optimization model adjusts the local features of the initial portrait and reduces the difference between the initial portrait text features and the portrait text, so that the initial portrait is more consistent with the initial text entered by the user.

[0067] In one embodiment, the portrait text includes facial expression description text, hair texture description text and clothing pattern description text. The portrait optimization model includes an attention mechanism module, which extracts the facial expression, hair texture and clothing pattern of the initial portrait through the attention mechanism module. The portrait optimization model optimizes the facial expression based on the facial expression description text, optimizes the hair texture based on the hair texture description text, and optimizes the clothing pattern based on the clothing pattern description text.

[0068] For example, the attention mechanism enables the portrait optimization model to focus more on facial expressions, hair texture, and clothing patterns when performing calculations.

[0069] For example, the attention mechanism module first focuses on the facial region of the subject in the initial portrait based on the input facial expression description text, extracting the emotional changes, expression details, and dynamic changes in the face. By comparing the differences between the initial facial expression in the portrait and the text description, this module automatically adjusts and optimizes the facial expression to better match the emotion or expression in the text description.

[0070] For example, based on the hair texture description text, the portrait optimization model further optimizes the hair details. The attention mechanism module focuses on the hair region, extracting and enhancing the hair texture, shape, and texture features consistent with the text description, and adjusts the hair details to achieve the hair effect required by the text. This optimization process ensures that the appearance of the subject's hair closely matches the texture style presented in the text description.

[0071] For example, the portrait optimization model also optimizes a person's clothing design based on the clothing pattern description text. Using the attention mechanism module, the system precisely adjusts the clothing in the portrait based on the pattern style, color matching, and clothing details in the description text, ensuring that the pattern, color, and style of the clothing meet the requirements of the text description and are consistent with the overall image of the person.

[0072] In one embodiment, the multi-view portrait generation model includes a posture control module. The character posture text and the initial portrait are input into the posture control module for posture generation to obtain target posture information that matches the skeleton key points of the initial portrait.

[0073] Exemplarily, the posture control module is a neural network structure that can use posture for precise guidance based on the image generation model. By introducing additional conditional branches in the multi-step denoising process or key layers of the diffusion model to impose strong constraints on the generation process, controllable generation of posture, shape or layout can be achieved.

[0074] Exemplarily, the posture control module may be a neural network structure such as ControlNet or T2I-Adapter.

[0075] Exemplarily, the posture control module can preset or receive target posture data provided by the user, including two-dimensional skeleton key points. Based on this, it can generate target posture information corresponding to the initial portrait, which is convenient for providing to the multi-view portrait generation model for further posture synthesis to ensure that the generated result meets the expected posture structure.

[0076] In one embodiment, a character portrait generation model is trained through the following steps: based on a pre-trained image generation model with general image generation capabilities, a portrait image set with annotated text is used for training to obtain a character portrait generation model; wherein the text annotations include age, gender, skin color, clothing, posture, and lighting scene.

[0077] For example, pre-trained image generation models can be based on mainstream models such as StabilityAI's StableDiffusion or Academia Sinica's CogView. These pre-trained models already possess certain image generation capabilities, but lack specific details or fail to fully meet specific artistic requirements. Therefore, further training is required.

[0078] The images used for training are primarily from professional portrait photography, covering a wide range of ages, genders, skin tones, clothing, poses, and lighting scenarios. This provides the model with ample image detail to enhance the diversity and realism of its generated images. Furthermore, each image is accompanied by precise annotated text, providing important semantic guidance for the model and ensuring that the generated portraits meet the specified visual standards.

[0079] The core goal of fine-tuning training is to enhance the model's expertise in portrait image generation. Building on the general image generation capabilities already possessed by the pre-trained model, the model undergoes further learning, focusing on high-precision features relevant to portraits. This ensures that generated portraits are not only realistic but also exhibit natural expressions, hair texture, and clothing details. Through fine-tuning training, the model is able to better capture these details, generating portraits that are more in line with realistic photographic styles, or adapting them to a specific style based on specific needs.

[0080] In one embodiment, before step 202 of obtaining the initial text input by the user for generating the multi-view portrait, the method further includes: displaying a text input interface, the text input interface including a portrait text input area, a clothing text input area, and a background text input area.

[0081] For example, the user can enter or select a portrait customization text in the portrait text input area, can enter or select a clothing requirement text in the clothing text input area, and can enter or select a background requirement text in the background text input area as the initial text. Then, the server 104 can obtain the initial text entered by the user.

[0082] Optionally, portrait customization requirements include region, gender, age, skin color, facial features, hairstyle, and expression.

[0083] Optionally, clothing requirements include tops, bottoms, jumpsuits, accessories, and bags.

[0084] Optionally, the background requirement is a single text description, such as indoor, outdoor, studio, etc.

[0085] In one embodiment, see Figure 3 ,The multi-view portrait generation method includes the following steps:

[0086] Step 302: The user selects portrait customization parameters, multi-view scenes, and clothing.

[0087] The initial text input by the user is "Asian, female, 25 years old, brown eyes, brown curly hair".

[0088] Step 304: The large language model automatically refines and optimizes the portrait to generate prompt words.

[0089] Among them, after obtaining the initial text input by the user, it is input into the text generation model to obtain the augmented text: "25-year-old Chinese girl, her eyes are dark brown, exuding a charming glow. They are expressive, giving her a charming and intelligent demeanor. Her hair is warm brown, loosely wavy, and cut to complement her face shape, enhancing her overall appearance. Her skin looks smooth and healthy, with a subtle glow, making her radiant. Her facial features embody the grace and elegance of Chinese aesthetics, blending traditional Asian beauty with modern sophistication. Professional photography ensures that every detail is captured accurately. The balance of light and shadow creates a visually appealing image, and the balanced exposure highlights her features without overexposing any part of her face."

[0090] Step 306: The character portrait generation model generates a character portrait according to the prompt word.

[0091] The amplified text is input into a character portrait generation model to obtain an initial portrait.

[0092] Step 308: Determine whether the portrait needs to be optimized.

[0093] Step 310: If the portrait needs to be optimized, the portrait details are enhanced to improve the authenticity through the character portrait optimization model.

[0094] In step 312 , if the portrait does not need to be optimized, the large language model generates multi-perspective prompt words according to the multi-perspective template and the user input.

[0095] The initial text, augmented text, and initial portrait are input into a multimodal large language model to generate detailed description text including perspective requirement text, clothing text, character posture text, and scene text: "This picture shows a female model. She has long, dark brown hair that is straight and smooth. Her expression is neutral, and her eyes are looking softly forward. She is wearing a light blue off-shoulder dress with thin straps and a ruffled neckline. The top of the dress is fitted, with pleats at the waist, and the hem is a multi-layered flared shape. The background is lush green palm trees, simple modern buildings, and natural sunlight shining through the trees.

[0096] The images are divided into four columns and a single row, showing four different poses. The following is a description of each pose:

[0097] (1) She stood straight, facing the camera, with her hands hanging naturally at her sides, showing the overall fit of the skirt and the flared skirt.

[0098] (2) She is at a third angle, with one hand resting lightly on her hip, emphasizing the drape and fluidity of the skirt.

[0099] (3) She is photographed sideways with her head slightly turned towards the camera, highlighting the ruffled neckline of her dress and her naturally falling hair.

[0100] (4) She turned her back to the camera, revealing the details of the dress’s back, including the spaghetti straps and gathered waist, showcasing the overall silhouette.”

[0101] In step 314 , the multi-view portrait generation model generates a consistent multi-view character image based on the prompt words and the portrait.

[0102] In this step, the multi-view portrait generation model receives the prompt words output from step 312 (including key information such as the character's posture, clothing details, environmental background, and multi-view requirements) and the previously generated initial portrait, and generates complete images of the same person from different perspectives.

[0103] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0104] Based on the same inventive concept, embodiments of the present application further provide a multi-perspective portrait generation device for implementing the multi-perspective portrait generation method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations of one or more of the following multi-perspective portrait generation device embodiments can be found in the above-described limitations of the multi-perspective portrait generation method and are not further elaborated here.

[0105] In an exemplary embodiment, Figure 4 As shown, a multi-view portrait generation device 400 is provided, comprising: an acquisition module 402, a first text generation module 404, a portrait generation module 406, a second text generation module 408 and a multi-view image generation module 410, wherein:

[0106] The acquisition module 402 is configured to acquire initial text input by a user for generating a multi-view portrait.

[0107] The first text generation module 404 is configured to input the initial text into a text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes a portrait text.

[0108] The portrait generation module 406 is used to input the amplified text into the character portrait generation model to generate a character portrait and obtain an initial character portrait.

[0109] The second text generation module 408 is used to input the initial text, the augmented text and the initial portrait into the multimodal large language model, and generate the perspective requirement text, the clothing text, the character posture text and the scene text through the multimodal large language model.

[0110] The multi-view image generation module 410 is used to input the initial portrait, view requirement text, clothing text, character posture text and scene text into the multi-view portrait generation model to generate images, thereby obtaining complete images of the initial portrait at multiple different viewpoints.

[0111] In one embodiment, the multi-view portrait generation device 400 also includes a portrait optimization module, which is used to: textually process the initial portrait to obtain initial portrait text features; compare the initial portrait text features with the portrait text to obtain the similarity between the initial portrait text features and the portrait text; determine whether the similarity is greater than or equal to a preset threshold; if the similarity is less than the preset threshold, input the portrait text and the initial portrait into a portrait optimization model, and optimize the initial portrait through the portrait optimization model to improve the similarity between the initial portrait text features corresponding to the initial portrait and the portrait text; if the similarity is greater than or equal to the preset threshold, do not process the initial portrait.

[0112] In one embodiment, the portrait text includes facial expression description text, hair texture description text and clothing pattern description text, and the portrait optimization model includes an attention mechanism module, which extracts the facial expression, hair texture and clothing pattern of the initial portrait through the attention mechanism module. The portrait optimization model optimizes the facial expression based on the facial expression description text, optimizes the hair texture based on the hair texture description text, and optimizes the clothing pattern based on the clothing pattern description text.

[0113] In one embodiment, the multi-view portrait generation model includes a posture control module, and the character posture text and the initial portrait are input into the posture control module for posture generation to obtain target posture information that matches the skeleton key points of the initial portrait.

[0114] In one embodiment, a character portrait generation model is trained through the following steps: based on a pre-trained image generation model with general image generation capabilities, a portrait picture set with annotated text is used for training to obtain a character portrait generation model; wherein the text annotations include age, gender, skin color, clothing, posture, and lighting scene.

[0115] In one embodiment, the multi-view portrait generation device 400 further includes a display module for:

[0116] Display the text input interface, which includes portrait text input area, clothing text input area and background text input area.

[0117] Each module in the multi-view portrait generation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0118] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store image data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multi-perspective portrait generation method is implemented.

[0119] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0120] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0122] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0123] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0124] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multi-view portrait generation method, characterized in that: The method comprises: Obtaining the initial text input by the user for generating multi-view portraits; Inputting the initial text into a text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes portrait text; Inputting the amplified text into a character portrait generation model to generate a character portrait, thereby obtaining an initial character portrait; Inputting the initial text, the augmented text, and the initial portrait into a multimodal large language model, and generating a viewing angle requirement text, a clothing text, a character posture text, and a scene text through the multimodal large language model; The initial portrait, the viewing angle requirement text, the clothing text, the character posture text and the scene text are input into a multi-view portrait generation model for image generation to obtain complete images of the initial portrait at multiple different viewing angles.

2. The method according to claim 1, characterized in that After inputting the portrait text into the character portrait generation model to generate the character portrait and obtaining the initial character portrait, the method further includes: Performing text processing on the initial portrait to obtain initial portrait text features; Comparing the initial portrait text feature with the portrait text to obtain a similarity between the initial portrait text feature and the portrait text; Determining whether the similarity is greater than or equal to a preset threshold; If the similarity is less than a preset threshold, the portrait text and the initial portrait are input into a portrait optimization model, and the initial portrait is optimized by the portrait optimization model to improve the similarity between the initial portrait text features corresponding to the initial portrait and the portrait text; If the similarity is greater than or equal to a preset threshold, the initial portrait is not processed.

3. The method according to claim 2, characterized in that The portrait text includes facial expression description text, hair texture description text and clothing pattern description text. The portrait optimization model includes an attention mechanism module, which extracts the facial expression, hair texture and clothing pattern of the initial portrait through the attention mechanism module. The portrait optimization model optimizes the facial expression based on the facial expression description text, optimizes the hair texture based on the hair texture description text, and optimizes the clothing pattern based on the clothing pattern description text.

4. The method according to claim 1, wherein The multi-view portrait generation model includes a posture control module. The character posture text and the initial portrait are input into the posture control module for posture generation to obtain target posture information that matches the skeleton key points of the initial portrait.

5. The method according to claim 1, wherein The character portrait generation model is trained by the following steps: Based on a pre-trained image generation model with general image generation capabilities, a portrait generation model is obtained by training using a collection of portrait images with annotated text. The text annotations include age, gender, skin color, clothing, posture and lighting scene.

6. The method according to claim 1, characterized in that Before obtaining the initial text input by the user for generating the multi-view portrait, the method further includes: A text input interface is displayed, which includes a portrait text input area, a clothing text input area, and a background text input area.

7. A multi-view portrait generation device, characterized in that: The device comprises: An acquisition module, used to obtain initial text input by a user for generating a multi-view portrait; A first text generation module is configured to input the initial text into a text generation model to perform text content amplification processing to obtain an amplified text based on the initial text; the amplified text includes a portrait text; A portrait generation module, configured to input the amplified text into a character portrait generation model to generate a character portrait and obtain an initial character portrait; a second text generation module, configured to input the initial text, the augmented text, and the initial portrait into a multimodal large language model, and generate a viewing angle requirement text, a clothing text, a character posture text, and a scene text through the multimodal large language model; The multi-view image generation module is used to input the initial portrait, the view requirement text, the clothing text, the character posture text and the scene text into the multi-view portrait generation model to generate an image, thereby obtaining a complete image of the initial portrait at multiple different view angles.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.