Face consistency multi-angle lens video generation method and device
The integration of LoRA-tuned image models, video synthesis, and LLM-driven text processing in the method addresses inefficiencies in generating multi-angle videos, improving efficiency and realism through automated lip-syncing and seamless transitions.
Patent Information
- Application Number
- CN202510658703.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art has a single lens angle when generating oral videos, lacks natural switching effects of multiple angles, and the generation process is scattered, and a unified automation solution is lacking. The lip-driven is incoherent, making it difficult to meet diversified and personalized needs.
LoRA fine-tuning technology is used to generate multi-angle and multi-pose image sequences, and frame-by-frame generation and style consistency processing is performed by combining the image-generating video model. The LLM large language model is used to generate copy scripts that conform to the speech broadcast logic, and the natural transition of multi-angle lenses is achieved through lip-driven and video stitching technology.
It realizes the efficient generation of personalized oral videos of multi-angle lenses, improves the authenticity and dynamic expression of the video, reduces the need for manual intervention, and is suitable for customized applications in various vertical fields.
Smart Images

Figure CN120321472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-angle lens video generation, and particularly to a method and device for generating a multi-angle lens video with facial consistency. Background Art
[0002] In today's digital content creation field, the application of virtual human technology is becoming increasingly widespread. Especially in scenarios such as short videos, online education, and virtual customer service, the demand for voice-over videos is continuously increasing. However, the existing voice-over video generation technologies still face many challenges in meeting diversification, personalization, and efficient production. Traditional voice-over video production usually relies on live recording, which is not only costly, time-consuming, and laborious, but also difficult to achieve large-scale personalized customization. With the development of artificial intelligence technology, 2D / 3D modeling and animation-driven technologies have gradually been introduced into the synthesis of voice-over videos. These technologies can generate virtual human voice-over videos with certain dynamic effects through computer graphics methods. However, these methods often require complex manual operations when generating multi-angle lenses, lacking an automated process, resulting in obvious defects in the realism, continuity, personalization, lens diversification, and naturalness of the generated video content.
[0003] In recent years, generative artificial intelligence technology has made remarkable progress, especially in the field of image generation. For example, the LoRA (Low-Rank Adaptation) technology injects low-rank matrices into the generative AI model and only trains these matrices while keeping the main model parameters frozen, thus achieving efficient fine-tuning of the image generation model. This technology performs excellently in enhancing the personalization and customization of image generation, but its application is mainly concentrated on static image generation, and the support for dynamic video generation is still insufficient.
[0004] In the field of video generation, some technologies attempt to enhance the dynamic expression ability by generating short videos from static image sequences. For example, models such as WAN2.1 and Hunyuan can synthesize short video clips from multiple static images, but most of these technologies are limited to the generation of single-angle lenses and lack an automated process for multi-angle lens synthesis. In addition, when generating voice-over videos, existing technologies usually need to separately process multiple links such as image generation, video conversion, copywriting generation, and lip movement driving, lacking a unified automated process. This decentralized processing method is not only inefficient but also likely to result in insufficient visual and auditory coherence of the generated video.
[0005] In addition, the existing voice-over video generation technologies also have obvious deficiencies in lip movement driving. Although some tools can match the audio signal with the lip movement of the character, it is often difficult to ensure the natural connection of the lip movement during multi-angle lens switching. This makes the generated video look rigid visually and lack a sense of reality.
[0006] Generally speaking, the following main problems exist in the prior art when generating speaking videos: First, the camera angles are single, making it difficult to generate natural switching effects from multiple angles; second, the generation process is scattered, lacking a unified automated solution, resulting in low efficiency; third, the coherence of lip movement driving is insufficient, affecting the realism of the video; fourth, the overall generation effect is difficult to meet diverse and personalized needs. These problems severely limit the wide application of speaking video generation technology.
[0007] In view of this, the present application is proposed. Summary of the Invention
[0008] The present invention provides a method and device for generating a multi-angle video with consistent facial features, which can at least partially improve the above problems.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A method for generating a multi-angle video with consistent facial features, comprising:
[0011] Obtain a multi-angle image set to be processed and text data, and based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning processing on the multi-angle image set to generate a multi-angle and multi-pose image sequence of the same person;
[0012] Call a video generation model from images to perform frame-by-frame generation and style unification processing on the multi-angle and multi-pose image sequence, and synthesize multiple short video segments;
[0013] Obtain a preset theme or keyword, and use a large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the logic of voice broadcast;
[0014] Based on the copywriting script, drive the generation of the lip shape of the person, and fuse and splice multiple short video segments to generate a final multi-angle video.
[0015] The present invention also provides a device for generating a multi-angle video with consistent facial features, comprising:
[0016] A fine-tuning unit, configured to obtain a multi-angle image set to be processed and text data, and based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning processing on the multi-angle image set to generate a multi-angle and multi-pose image sequence of the same person;
[0017] A segment synthesis unit, configured to call a video generation model from images to perform frame-by-frame generation and style unification processing on the multi-angle and multi-pose image sequence, and synthesize multiple short video segments;
[0018] A copywriting generation unit for obtaining a preset theme or keyword, and using a large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the logic of voice broadcast;
[0019] A fusion unit for driving the generation of a character's lip shape based on the copywriting script, and fusing and splicing multiple short video clips to generate a final multi-angle shot video.
[0020] In summary, the method for generating a multi-angle shot video with face consistency efficiently generates a personalized oral broadcast video with multi-angle shots. First, the image generation model is optimized using the advanced LoRA fine-tuning technology, which can generate a diverse image sequence of the same person at different angles and poses from a single input image. This process not only improves the personalization degree of the generated images but also significantly enhances the model's ability to capture specific character features. Subsequently, through the image-to-video technology, the generated multi-angle image sequence is converted into coherent short video clips, further enriching the dynamic expressiveness of the video content. At the same time, combined with the powerful text generation ability of the large language model (LLM), an oral broadcast copywriting script highly matching the video content is quickly generated according to the preset theme or keyword. These copywritings are not only logically clear and natural in language but also can be flexibly adjusted in style and tone according to different scenarios. Finally, through the lip shape driving and video splicing technology, the generated copywriting is accurately matched with the character image, and the character's mouth movement is driven to achieve a natural transition between multi-angle shots. The entire process is highly automated, greatly reducing the need for manual intervention and significantly improving the efficiency and quality of video generation.
[0021] The innovation of this method lies in seamlessly integrating multiple links such as image generation, video synthesis, copywriting creation, and lip shape driving, forming a set of efficient, flexible, and scalable automated video generation methods and devices. It is not only applicable to multiple vertical fields such as education, marketing, and social networking but also can be customized according to different application scenarios and user needs, with broad application prospects and significant practical value. Description of the Drawings
[0022] Figure 1 is a schematic flowchart of the method for generating a multi-angle shot video with face consistency provided by the first embodiment of the present invention;
[0023] Figure 2 is a flowchart block diagram of the method for generating a multi-angle shot video with face consistency provided by the first embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of the modules of the device for generating a multi-angle shot video with face consistency provided by the second embodiment of the present invention. Detailed Embodiments
[0025] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0026] Referring Figure 1 、 Figure 2 As shown in [references not provided in the original, assumed to be some figures], the first embodiment of the present invention discloses a method for generating a multi-angle face-consistent lens video, which can be executed by a multi-angle face-consistent lens video generation device (hereinafter referred to as the generation device), specifically, by one or more processors in the generation device to implement the following method:
[0027] S1. Obtain a multi-angle image set and text data to be processed. Based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning on the multi-angle image set to generate a multi-angle and multi-pose image sequence of the same person;
[0028] Preferably, the text data includes a clear angle instruction, a style instruction, and a content instruction, and the multi-angle image set includes images of the same person from multiple different angles.
[0029] Specifically, step S1 includes: training the Flux model using a preset PEFT open-source framework and selecting multiple hyperparameters as the rank of LoRa;
[0030] Use the multi-angle image set as the input of the model, obtain the output results of the model under different rank values, and compare them;
[0031] Select the image with the best effect from multiple output results as the final fine-tuning model result to obtain a multi-angle and multi-pose image sequence of the same person.
[0032] The core of the present invention is to implement an efficient automated process for generating personalized on-camera videos with multi-angle lenses. The multi-angle face-consistent lens video generation method aims to solve problems such as single lenses, scattered processes, inconsistent lip movement driving, and excessive manual intervention in the prior art.
[0033] In this embodiment, first, a multi-angle image set and text data to be processed are obtained. These multi-angle image sets include images of the same person from multiple different angles, such as front, side, top view, etc. These images provide the basic materials for generating a multi-angle and multi-pose image sequence later. The text data contains clear angle instructions, style instructions, and content instructions, which are used to guide the fine-tuning process of the image generation model to ensure that the generated images meet specific requirements and styles.
[0034] After obtaining these input data, a pre-trained image generation model is called and subjected to LoRA (Low-Rank Adaptation) fine-tuning. LoRA is an efficient fine-tuning technique that achieves rapid adjustment of the model by injecting low-rank matrices into certain linear layers of the model and only training these matrices, while keeping the parameters of the main model frozen; it does not directly fine-tune all the parameters of the large model. This method not only improves the adaptability of the model but also significantly reduces the training cost and time.
[0035] Specifically, in the implementation process, a preset PEFT open-source framework is used to train the Flux model. The PEFT framework provides strong technical support for LoRA fine-tuning, enabling the model to quickly adapt to different input data and text instructions. During the training process, multiple hyperparameters are selected as the ranks of LoRA, such as 4, 6, 8, etc. By comparing the model output results under different rank values, the model with the best effect is selected as the final fine-tuned model result.
[0036] In addition, when using the multi-angle image set as the input of the model, corresponding output results will be generated according to different rank values. And these output results are compared and evaluated in detail (such as: performance comparison, efficiency comparison, image quality comparison, resource cost comparison, scene applicability comparison, etc.), and the best image sequence is selected from multiple candidate results. This process not only ensures the quality and diversity of the generated images but also further improves the performance and efficiency of the model by optimizing the selection of ranks. Based on this, an image sequence of the same person with multiple angles and postures is successfully generated. These image sequences not only have a high degree of realism and naturalness visually but also can flexibly adjust the style and content according to text instructions to meet the diverse needs in different scenarios.
[0037] S2, call the image-to-video model to perform frame-by-frame generation and style consistency processing on the multi-angle and multi-posture image sequence, and synthesize multiple short video segments;
[0038] Preferably, each of the short video segments is obtained by splicing multiple video segments from different angles.
[0039] Specifically, in this embodiment, the image-to-video model receives the multi-angle and multi-posture image sequence after LoRA fine-tuning as the input. These image sequences contain rich visual information of the same person at different angles and postures, providing a solid foundation for generating diverse video content. The model converts static images into dynamic video frames in a frame-by-frame manner, while ensuring the natural and smooth transition between each frame and avoiding the abruptness of the video content.
[0040] During the frame-by-frame generation process, the image-to-video model also performs style consistency processing. This processing step is crucial because it ensures that the generated video clips maintain a consistent visual style, even if these clips are composed of images taken from different angles. Style consistency processing involves adjusting visual elements such as the color, lighting, and texture of the images, so that the final generated video clips have a unified visual style, enhancing the overall visual perception and professionalism of the video.
[0041] It should be noted that each short video clip is obtained by splicing multiple video clips from different angles; for example, the initially input image sequence contains images from three angles: the front, left, and right. At this time, three videos from the front, left, and right angles are generated accordingly. Subsequently, by splicing the three independent videos, a short video clip composed of multiple video clips from different angles can be obtained. This multi-angle splicing strategy not only enriches the visual content of the video but also enhances the realism and dynamic expressiveness of the video. By skillfully switching different shooting angles, the video can better simulate the visual changes in real scenes, enabling the audience to feel a more three-dimensional and vivid visual effect. For example, in a video of a person's oral broadcast, the front shot of the person can be shown first, then switched to the side shot, and finally back to the front shot. This angle switching can better attract the audience's attention and enhance the attractiveness of the video.
[0042] During the implementation process, the image-to-video model will automatically select appropriate splicing points according to preset rules and algorithms to ensure seamless transition when splicing video clips from different angles. This process not only relies on the model's understanding of the image content but also involves grasping the video rhythm and narrative logic. Through intelligent splicing algorithms, the model can maximize the advantages of multi-angle images while maintaining the coherence of the video, creating short video clips with visual impact and narrative coherence. Based on this, this method not only realizes the efficient conversion from static images to dynamic videos but also significantly improves the quality and visual effect of the videos through multi-angle splicing and style consistency processing.
[0043] S3. Obtain a preset theme or keyword, and use the large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the voice broadcast logic;
[0044] Specifically, in this embodiment, a preset theme or keyword is obtained. These themes and keywords are the core of the video content, and they can be specific product introductions, educational content, news reports, or any other information that needs to be conveyed through oral broadcasts. For example, if the theme of the video is "healthy diet", the keywords may include "balanced nutrition", "fruits and vegetables", "low sugar and low fat", etc. These information provide clear directions and key points for copywriting generation.
[0045] Subsequently, the LLM (Large Language Model) is called to generate copywriting based on these preset topics or keywords. With its powerful natural language processing capabilities, the LLM can generate high-quality copywriting scripts according to the input topics and keywords. During the generation process, the model will consider the structure, logical coherence, and language style of the copywriting to ensure that the generated copywriting is not only rich in content and accurate in information but also has good readability and attractiveness. The copywriting scripts generated by the LLM can not only accurately revolve around the preset topics but also flexibly adjust the style and tone according to different scenarios. For example, if it is a video targeting a young audience, the copywriting can adopt a relaxed and humorous style; if it is a professional lecture, the copywriting will be more formal and rigorous. This flexibility enables the generated copywriting to meet the needs of multiple application scenarios.
[0046] To further enhance the applicability of the copywriting, the logic and characteristics of voice broadcasting are incorporated during the generation of the copywriting. This means that the copywriting not only needs to be smooth and natural in written language but also suitable for oral expression. For example, the copywriting will avoid using overly complex sentence patterns and professional jargon but adopt a concise and clear language style to ensure that the information can be clearly conveyed during oral broadcasting. In addition, the copywriting will be optimized according to the duration and rhythm of the video to ensure the compactness and integrity of the content.
[0047] In practical applications, the generated copywriting scripts can be directly used for the voice-over link of the video. Since the copywriting has been optimized according to the logic of voice broadcasting, during the recording process, the anchor or the voice synthesis system can broadcast more smoothly, reducing the situation of recording interruptions or repeated recordings caused by copywriting problems. This not only improves the efficiency of video production but also enhances the quality and professionalism of the final video.
[0048] Among them, during this generation process, the corresponding prompt words can be as follows:
[0049] """
[0050] You are a world-class podcast writer who has worked as a ghostwriter for people like Joe Rogan, Lex Fridman, Ben Shapiro, and Tim Ferriss. Through your writing, their shows have become lively and engaging.
[0051] In this special scenario, we assume that you are actually writing every word they say, and this content is directly received by them and conveyed to the audience. With this unique talent, you have won multiple podcast awards, proving your outstanding status in this field.
[0052] Your task is to adapt the uploaded text and conversation scenarios into an engaging dialogue. You need to verbatim record the exchanges between the two speakers, including all of Speaker 2's short responses such as "right", "correct", and "yes" to maintain the authenticity of the conversation. Although the discussion may occasionally deviate from the topic, it should generally revolve around the core topic.
[0053] Remember, Speaker 2 is new to this topic, so you should skillfully incorporate real-life stories and analogies into the conversation to help explain concepts and deepen understanding. At the same time, this speaker should maintain the coherence of the conversation by asking follow-up questions and can show great interest or confusion when asking questions, thus leading to more interesting discussions.
[0054] Ensure that Speaker 2's interjections are not limited to simple confirmations but can also be unexpected or interesting comments. There should be natural breaks and impromptu interactions throughout the conversation to make the listeners feel the atmosphere of a real conversation.
[0055] The final result should be like an actual recorded podcast episode, capturing every detail as comprehensively as possible. Welcome the listeners with a super interesting overview and ensure that the entire conversation is extremely engaging, almost like clickbait.
[0056] At the beginning of the conversation, Speaker 1 starts the topic. Instead of providing a separate episode title or chapter title, let Speaker 1 naturally determine the theme in their speech. Strictly speaking, this should be a conversation, not a one-sided narrative. The conversation language must be in Chinese.
[0057] Scene:
[0058] Face-to-face communication
[0059] Text:
[0060] {input_text}
[0061] """
[0062] Through this implementation step, this method achieves an efficient transformation from the theme and keywords to a high-quality copywriting script. The application of the LLM large language model not only ensures the diversity and adaptability of the copywriting but also significantly improves the practicality of the copywriting and the final effect of the video by optimizing the voice broadcast logic. This innovative copywriting generation process provides a solid content foundation for the automated voice-over video generation, further promoting the development of virtual human video generation technology.
[0063] S4. Drive the generation of the character's lip shape based on the copywriting script, and fuse and splice multiple short video clips to generate the final multi-angle shot video.
[0064] Specifically, in this embodiment, a copywriting script generated by a large language model (LLM) is received. This script has been optimized to ensure that its content is suitable for voice broadcast and can accurately convey a preset theme or keyword. The voice conversion of the copywriting script is achieved through text-to-speech technology, and the generated voice file will serve as the basis for subsequent lip movement driving. Immediately afterwards, the lip movement driving module is called. The core function of this module is to drive the lip movements of a character based on the generated voice file. Lip movement driving technology uses deep learning algorithms to analyze the pronunciation features in the voice signal and map them to the lip movements of the character. This process requires precise matching of the rhythm and pronunciation of the voice to ensure that the lip movements of the character are natural, realistic, and synchronized with the voice.
[0065] During the lip movement driving process, lip animations corresponding to the voice content of the copywriting script are generated. These lip animations are generated based on a multi-angle and multi-pose image sequence of the character, ensuring that the lip movements of the character can be naturally and smoothly presented from different angles. For example, when the character switches from a front view to a side view, the details and angles of the lip movements will be adjusted accordingly to maintain visual coherence. At the same time, multiple short video clips are fused and spliced. These short video clips are generated by an image-to-video model, and each clip contains character shots from different angles. To generate the final multi-angle shot video, these clips need to be intelligently spliced. During the splicing process, the switching rhythm of the shots, the action coherence of the character, and the overall narrative logic of the video are considered. Through advanced video editing algorithms, seamless shot switching can be achieved, avoiding the abruptness of the video content.
[0066] During the splicing process, the audio and video tracks of the video are also synchronized. The precise alignment of the lip animation and the voice file is the key to ensuring the video quality. Through timeline alignment technology, it is ensured that the lip movements are completely synchronized with the voice content in terms of time, thus creating a natural and smooth voice-over effect. In addition, visual elements such as the color and lighting of the video are optimized to further enhance the overall visual perception of the video. Finally, a complete multi-angle shot video is generated. This video not only contains rich character expressions and actions but also enhances the visual attractiveness and realism of the video through the switching of multi-angle shots; the conversion angle here is like a cut shot. Because the facial consistency is maintained, it is like there are three camera positions shooting the same person. Through lip movement driving technology, the lip movements of the character are perfectly synchronized with the copywriting voice, making the video content more vivid and natural.
[0067] In summary, the multi-angle lens video generation method for face consistency is used to generate personalized mouthpiece videos with multi-angle lenses, aiming to solve many problems existing in the prior art, such as single lens, scattered processes, discontinuous lip driving, and excessive manual intervention.
[0068] Specifically, first, optimize the image generation model through the LoRA fine-tuning technology to generate a multi-angle and multi-pose image sequence of the same person from a single photo. This process not only improves the personalization degree of the generated images but also significantly enhances the model's ability to capture specific person features. Subsequently, use the image-to-video model to generate each frame of these image sequences and perform style consistency processing to synthesize multiple short video segments. Each short video segment is composed of multiple video segments with different angles, and the dynamic expressiveness and visual richness of the video are enhanced through multi-angle switching.
[0069] At the same time, generate a copywriting script that conforms to the voice broadcast logic through the LLM (Large Language Model). These copywriting scripts are not only rich in content and clear in logic but also can flexibly adjust the style and tone according to different scenarios to ensure a high degree of matching with the video content. Finally, through lip driving and video splicing technology, accurately match the generated copywriting with the character images and drive the lip movements of the characters to achieve a natural transition between multi-angle lenses and generate a smooth and coherent complete mouthpiece video.
[0070] Compared with the prior art, the beneficial effects of the multi-angle lens video generation method for face consistency include: 1. By LoRA fine-tuning and multi-angle image generation, the problem of single lens in traditional technologies is solved, and the realism and visual attraction of the video are significantly improved. 2. Integrate multiple links such as image generation, video synthesis, copywriting creation, and lip driving into an integrated process, significantly improving the efficiency of video generation and reducing the need for manual intervention. 3. Through the natural transition of multi-angle lenses and lip driving technology, the problem of discontinuous lip movements in the prior art is solved, making the generated video more natural and smooth. 4. Strong scalability, applicable to multiple vertical fields such as education, marketing, and social networking, and can be customized according to different application scenarios and user needs.
[0071] Please refer to Figure 3 , the second embodiment of the present invention provides a multi-angle lens video generation device for face consistency, which includes:
[0072] A fine-tuning unit 101, configured to obtain a multi-angle image set and text data to be processed, and based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning processing on the multi-angle image set to generate a multi-angle and multi-pose image sequence of the same person;
[0073] The segment synthesis unit 102 is used to call the image-to-video model to perform frame-by-frame generation and style consistency processing on the multi-angle and multi-pose image sequence, and synthesize multiple short video segments;
[0074] The copywriting generation unit 103 is used to obtain a preset theme or keyword, and use the large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the logic of voice broadcast;
[0075] The fusion unit 104 is used to drive the generation of the lip shape of the character based on the copywriting script, and fuse and splice multiple short video segments to generate the final multi-angle shot video.
[0076] Preferably, the text data includes a clear angle instruction, a style instruction, and a content instruction, and the multi-angle image set includes images of the same person from multiple different angles.
[0077] Preferably, the fine-tuning unit 101 specifically includes: training the Flux model using a preset PEFT open-source framework, and selecting multiple hyperparameters as the rank of LoRa;
[0078] Taking the multi-angle image set as the input of the model, obtaining the output results of the model under different rank values, and making comparisons;
[0079] Selecting the image with the best effect from multiple output results as the final fine-tuning model result to obtain a multi-angle and multi-pose image sequence of the same person.
[0080] Preferably, each of the short video segments is spliced from multiple video segments with different angles.
[0081] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A method for generating a multi-angle lens video with face consistency, characterized in that Including: Obtain a multi-angle image set and text data to be processed. Based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning on the multi-angle image set, and generate a multi-angle and multi-pose image sequence of the same person; Call a image-to-video model to perform frame-by-frame generation and style consistency processing on the multi-angle and multi-pose image sequence, and synthesize multiple short video segments; Obtain a preset theme or keyword, and use a large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the logic of voice broadcast; Drive the generation of the lip shape of the person based on the copywriting script, and fuse and splice multiple short video segments to generate the final multi-angle shot video.
2. The method for generating a multi-angle lens video of face consistency according to claim 1, wherein, The text data includes clear angle instructions, style instructions, and content instructions, and the multi-angle image set includes images of the same person from multiple different angles.
3. The method for generating a multi-angle lens video of face consistency according to claim 1, wherein Calling a pre-trained image generation model to perform LoRA fine-tuning on the multi-angle image set to generate a multi-angle and multi-pose image sequence of the same person specifically includes: Use a preset PEFT open-source framework to train the Flux model, and select multiple hyperparameters as the rank of LoRa; Use the multi-angle image set as the input of the model to obtain the output results of the model under different rank values, and compare them; Select the image with the best effect from multiple output results as the final fine-tuning model result to obtain a multi-angle and multi-pose image sequence of the same person.
4. The method for generating a multi-angle face consistency lens video according to claim 1, wherein Each short video segment is spliced by multiple video segments from different angles.
5. A multi-angle lens video generation device for face consistency, characterized in that, Including: A fine-tuning unit for obtaining a multi-angle image set and text data to be processed. Based on the text data, call a pre-trained image generation model to perform LoRA fine-tuning on the multi-angle image set, and generate a multi-angle and multi-pose image sequence of the same person; A segment synthesis unit for calling an image-to-video model to perform frame-by-frame generation and style consistency processing on the multi-angle and multi-pose image sequence, and synthesizing multiple short video segments; A copywriting generation unit for obtaining a preset theme or keyword, and using a large language model (LLM) to perform copywriting generation processing on the preset theme or keyword to obtain a copywriting script that conforms to the logic of voice broadcast; A fusion unit for driving the generation of the lip shape of the person based on the copywriting script, and fusing and splicing multiple short video segments to generate the final multi-angle shot video.
6. The multi-angle lens video generation device for face consistency according to claim 5, wherein The text data includes clear angle instructions, style instructions, and content instructions, and the multi-angle image set includes images of the same person from multiple different angles.
7. The multi-angle lens video generation device for face consistency according to claim 5, wherein The fine-tuning unit specifically includes: Use a preset PEFT open-source framework to train the Flux model, and select multiple hyperparameters as the rank of LoRa; Use the multi-angle image set as the input of the model to obtain the output results of the model under different rank values, and compare them; Select the image with the best effect from multiple output results as the final fine-tuning model result to obtain a multi-angle and multi-pose image sequence of the same person.
8. The face consistency multi-angle lens video generation device according to claim 5, characterized in that, Each short video segment is spliced by multiple video segments from different angles.
Citation Information
Cited By
AI video output system and method combined with text image model
CN121442166A