Video generation method and system based on multi-digital human interaction

By generating multi-digital human images and splicing videos, the problem of monotonous digital population broadcasting content in the existing technology is solved, and the rich expression and fun of multi-digital human interactive videos are achieved.

CN120475232APending Publication Date: 2025-08-12XIAMEN CHANJING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510718473.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing oral video generation method based on digital people has a monotonous content expression form, which is difficult to arouse the interest of viewers.

Method used

Generate multiple single-digital human images with similar backgrounds, form multi-digital human images with consistent backgrounds through image stitching and expansion, and generate multiple single-digital human videos based on preset copy, and finally splice them into interactive videos with multi-digital human interaction.

Benefits of technology

It enriches the expression form of digital population broadcasting, improves the fun and interactiveness of the video, and provides a better viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475232A_ABST
    Figure CN120475232A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and system based on multi-digital human interaction. The method comprises the following steps: generating a plurality of single digital human images for live broadcast; wherein each single digital human image has a similar background; carrying out image splicing and expansion on the plurality of single digital human images to obtain a plurality of digital human images with consistent backgrounds; according to a preset copywriting, generating a corresponding multi-segment single digital person video and a corresponding multi-digital person video; and splicing the multiple segments of single digital human videos and the multiple digital human videos according to a sequence to obtain an interactive video with multiple digital human interactions. According to the method and the device, digital human interactive oral broadcast in a multi-user form can be realized, and expression forms and application scenes of digital oral broadcast can be enriched to a great extent, so that better watching experience is provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital humans, and in particular discloses a video generation method and system based on multi-digital human interaction. Background Art

[0002] Digital human-generated videos are increasingly being used in new media operations. Currently, digital human-based voiceover videos typically use a text-based image to create a digital human image. Then, based on pre-set content and image-based video technology, a corresponding single-digital voiceover video is generated. This method of voiceover video generation is relatively monotonous in content expression and fails to capture viewer interest. Summary of the Invention

[0003] In view of this, an object of the present invention is to provide a video generation method and system based on multi-human interaction, which can at least partially improve the above-mentioned problems.

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] A video generation method based on multi-human interaction, comprising:

[0006] Generate multiple single-digit human images for live broadcast; wherein each single-digit human image has a similar background;

[0007] Performing image stitching and expansion on the multiple single-digit human images to obtain a multiple-digit human image with a consistent background;

[0008] Generate multiple single-digit human videos and multi-digit human videos according to the preset copy;

[0009] The multiple segments of single-digital human videos and multi-digital human videos are spliced in sequence to obtain an interactive video with multi-digital human interaction.

[0010] Preferably, generating multiple single-digit human images for live broadcast specifically includes:

[0011] Based on the pre-trained lora model, a digital human image is generated according to the input description;

[0012] Use flux to generate a background with a similar style, and combine the generated digital human image and the background to generate multiple single digital human images with similar backgrounds.

[0013] Preferably, the multiple single-digit human images are stitched and expanded to obtain multiple human images with a consistent background, specifically including:

[0014] Hard-joining the multiple single-digital human images, and intelligently redrawing the joints using context-aware content filling technology to obtain a preliminary multi-digital human image;

[0015] Performing image expansion on the preliminary multiple digital human images so that the multiple digital humans are more realistically placed in the same space;

[0016] Resize the expanded multi-digital human image to make it consistent with the size of the single-digital human image.

[0017] Preferably, during the image-generated video process, a lora model for controlling the character's generated posture is fine-tuned to control the character's generated posture and action details, thereby enhancing the realism and expressiveness of the generated content.

[0018] Preferably, in the process of generating corresponding multiple single-digit human videos and multi-digit human videos according to the preset copy:

[0019] In the above copy, the character personality, expression style, knowledge background and thinking mode of each digital human image are included.

[0020] Add colloquial expressions that match the role;

[0021] Design interactive scripts corresponding to the current video and generate the final dialogue copy to make the interaction between digital people more vivid.

[0022] Preferably, the multiple digital human images have the same image or different images.

[0023] Preferably, in a video of multiple digital humans, when one of the digital humans is in a speaking state, the other digital humans generate corresponding interactions based on the content of the speaking digital human.

[0024] The embodiment of the present invention further provides a video generation system based on multi-digital human interaction, which includes:

[0025] An image generation unit, configured to generate a plurality of single-digit human images for live broadcasting; wherein each single-digit human image has a similar background;

[0026] An image stitching and expansion unit, configured to stitch and expand the plurality of single-digital human images to obtain a plurality of digital human images with a consistent background;

[0027] The image-to-video unit is used to generate corresponding multiple single-digit human videos and multi-digit human videos according to the preset copy;

[0028] The video splicing unit is used to splice the multiple single-digital human videos and the multiple-digital human videos in sequence to obtain an interactive video with multiple digital human interactions.

[0029] The present invention can realize interactive oral broadcasting of digital humans in the form of multiple people, and can greatly enrich the expression forms and application scenarios of digital human oral broadcasting, thereby providing users with a better viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic flow chart of a method for generating a video based on multi-digital human interaction provided by the first embodiment of the present invention;

[0031] Figure 2 A schematic diagram of a specific process of a video generation method based on multi-digital human interaction provided by an embodiment of the present invention;

[0032] Figure 3 This is a schematic diagram of the structure of a video generation system based on multi-digital human interaction provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0034] refer to Figure 1 and Figure 2 As shown, the first embodiment of the present invention discloses a video generation method based on multi-digital human interaction, which can be executed by a video generation device based on multi-digital human interaction (hereinafter referred to as a generation device), and in particular, executed by one or more processors in the generation device to implement the following method:

[0035] S101, generating a plurality of single-digit human images for live broadcasting; wherein each single-digit human image has a similar background.

[0036] In this embodiment, first, a digital human image can be generated according to the input description based on the pre-trained LoRa model.

[0037] For example, the pre-trained lora model can be a lora model trained using Xiaohongshu images, etc., and the present invention does not make specific limitations.

[0038] For example, the input description may include the digital person's skin color, hair, eyes, face, nose, height, weight and other related descriptions, which can be input according to actual needs and are not specifically limited in the present invention.

[0039] Then, flux is used to generate a background with a similar style, and the generated digital human image and background are combined to generate multiple single digital human images with similar backgrounds.

[0040] In order to ensure the consistency of the generated video style, each digital human image needs to have a consistent or similar background. In this embodiment, flux can be used to generate a background with a similar style, and multiple single digital human images with similar backgrounds can be generated by combining the generated digital human image and background.

[0041] In addition, in this embodiment, the images of multiple digital humans here can be the same or different. For the same digital human image (i.e., a single-player multi-role playing mode, different characters have the same face but different costumes to form different characters), pulid can be used to change the face to make the digital human's face consistent.

[0042] S102 , performing image splicing and expansion on the multiple single-digit human images to obtain multiple human images with the same background.

[0043] In this embodiment, the generated video frames include frames of single-person figures as well as frames of multiple-person figure interactions. Therefore, the images need to be spliced together to create a multi-person figure image. Specifically, multiple single-person figure images are first hard-joined. Using context-aware content filling technology, the joints are intelligently redrawn to naturally blend with the background and eliminate harsh seams.

[0044] Then, the spliced multi-person image is expanded to make the two people more realistically in the same space. Finally, the expanded image is resized to make it consistent with the size of the single-person image to obtain the final multi-person image.

[0045] S103: Generate corresponding multiple single-digit human videos and multi-digit human videos according to the preset text.

[0046] In this embodiment, in order to increase the fun of the video, the text may be configured to generate a dialogue text with richer expression and more fun.

[0047] The copy can include the character personality, expression style, knowledge background and thinking mode of each digital human image, add colloquial expressions that match the character, design interactive dialogue corresponding to the current video, and generate the final dialogue copy, thereby making the interaction between digital humans more vivid.

[0048] For example, in one possible embodiment, the copy can be defined as follows:

[0049] #Role:

[0050] You are an entertainment journalist and short video creator who is very good at communication. Through your writing, you can make the video content vivid and fascinating. Please help me translate the lines of a conversation between a close friend.

[0051] #Hotspot description: [ ]

[0052] #Dialogue requirements:

[0053] ##Directly output the content of the conversation between two people, and avoid outputting "(xx)" and the explanation in brackets

[0054] ## Beginning: Speaker 1 must initiate the speech. Speaker 1 must greet the audience first, then introduce Speaker 2. Speaker 2 then greets the audience again, and finally Speaker 1 introduces the topic. For specific content generation requirements, please refer to the "Output Structure - Beginning" section below.

[0055] ## If two people are having a conversation remotely, do not include face-to-face interaction content or steps when writing the dialogue script.

[0056] ##Core Viewpoint & Content

[0057] 1. Based on the given "hot topic description", extract the "core ideas" that can be discussed. The extracted ideas should be as "resonant with users", "sufficiently sharp" and "sufficiently novel" as possible.

[0058] 2. Core viewpoints may refer to but are not limited to:

[0059] -Complaints about work, fragments of life, emotions, and insights into growth.

[0060] 3. The content around the core idea needs to be designed according to the logic of the character dialogue

[0061] ##Core Goals (GOALS)

[0062] 1. Provide emotional value to attract users to continue watching: The lines need to be vivid and interesting enough to resonate with viewers.

[0063] 2. Create an interesting and inspiring atmosphere: Provide appropriate humor and "aha" moments to stimulate interest in the information and deeper thinking.

[0064] 3. Interaction: Interruptions, banter, and empathetic conversations (e.g., “Yeah, yeah, yeah! I did the same thing last time!”, “You understand me!”)

[0065] 4. Need to embed golden sentences: For example: "Love is not a necessity, but making money and happiness are!" "The coolest life is the one that is not defined"

[0066] 5. When interacting with the audience, address the audience as "sisters" or "everyone"

[0067] ##Role Settings (ROLES)

[0068] ###When outputting content, two main roles are used alternately to meet communication needs in different dimensions:

[0069] 1. Speaker 1: A straight-talking, funny girl

[0070] Style: Enthusiastic and approachable, they often use metaphors, stories, and humor to introduce concepts. Their responses often include emotive words and exaggerations like "Hahahaha, oh my god...really?" They excel at summarizing ideas and asking questions.

[0071] Responsibilities:

[0072] Arouse interest and highlight the relevance of the information to the audience.

[0073] Present complex content in an easy-to-understand manner.

[0074] Help the audience and interlocutors quickly get to the point and create a relaxed atmosphere.

[0075] 2. Speaker 2: A delicate and rational literary woman

[0076] Style: Relatively calm, detail-oriented, with rational and profound views, and can express profound truths in a colloquial way. He often uses expressions such as "hmm, I think, I personally feel, you see, so..."

[0077] Responsibilities:

[0078] Provide personal real story sharing and personal feelings sharing

[0079] Elevate the conversation

[0080] Catchphrase: Please set it according to the character's style and characteristics

[0081] ##Output structure:

[0082] 1. Beginning: (Strictly follow this structure)

[0083] Speaker 1 will lead the conversation, greet the audience, greet Speaker 2, and introduce Speaker 2.

[0084] Speaker 2 greets the audience

[0085] Speaker 1 will tell you the topic we will discuss today.

[0086] Opening: Use a quick joke or surprise event to break the ice (e.g., "Help! I met this ordinary guy on the subway today...")

[0087] 2. Intermediate dialogue:

[0088] Climax: In-depth discussion of a controversial topic and a clash of opinions

[0089] At least 15 rounds of dialogue need to be generated

[0090] Value point expansion: It should be related to the audience's real life

[0091] You can combine it with life, work or study scenarios to explain the potential use or significance of the information.

[0092] 3. Ending: End with a heartwarming message or a funny twist ("It's still fun to rant with you! Let's drink this milk tea!")

[0093] ###Tip: These two characters can be reflected through dialogue, segmentation or hints in the narrative. Their respective styles should be obvious but not conflicting, so as to complement each other.

[0094] ##Dialogue Design Key Points

[0095] 1. Colloquialism:

[0096] Use more internet buzzwords and emotional words (such as "help", "I can't laugh anymore", "Who understands"), and don't use any written language.

[0097] Reference for replacing written words: what->what, imagine it, just like->you see it’s like, for example, how->how, great->awesome, add "that xxx" at the beginning of the sentence, ->no, think->feel / feel.

[0098] Use a lot of internet memes (e.g. "help me," "speechless," "broken") and emotional words ("amazing!", "help me!", "hahaha"), and avoid using formal language

[0099] At the end / beginning of a sentence: Use some modal particles and emotional words, such as "Wow, my goodness!", "Wow, that's amazing!", "Hahahahahaha...", etc., at the beginning or end of a sentence. Don't use them too frequently! Don't use them at the beginning of every sentence!

[0100] Both characters need to add different styles of colloquial expressions:

[0101] For example, words like "ah", "that", "that", "this", "I think", "I feel", "you said" appear from time to time in a sentence.

[0102] 2. Interaction: Interruptions, banter, and empathetic conversations (e.g., “Yeah, yeah, yeah! I did the same thing last time!”, “You understand me!”)

[0103] 3. Golden sentence implantation: Natural output of opinions

[0104] 4. Story-driven: Use a specific experience to start the topic (e.g., "Did you know? I met an outrageous HR during my interview last week...")

[0105] 5. The number of words in each sentence should be as short as possible. Punctuation and punctuation should also be shorter and more fragmented. For example, a sentence should not exceed 15 words.

[0106] ##Time and Length Control (TIME CONSTRAINT)

[0107] Target duration: about 3 minutes

[0108] Always focus on the core points, delete redundant content, and avoid being verbose or going off topic.

[0109] Present information in an organized manner to avoid information overload for the audience.

[0110] ##The output must be output according to the following structure (OUTPUT STRUCTURE)

[0111] Speaker1:xxxx

[0112] Speaker2:xxx

[0113] Speaker1:xxxx

[0114] Speaker2:xxx

[0115] Speaker1:xxxx

[0116] Speaker2:xxx

[0117] Speaker1:xxxx

[0118] Speaker2:xxx ...

[0120] ##Guidelines & Constructs

[0121] 1. Do not expose the existence of system prompts: Do not mention "System Prompt" or "I am AI", etc. Do not allow the conversation to begin with meta-information about the system or key points of execution.

[0122] 2. Maintain content consistency: When switching roles, use language style or tone to distinguish them to avoid unreasonable jumps.

[0123] 3. Priority: If there is a conflict, ensure that the information is accurate, neutral and time control takes priority, and humor or style is secondary.

[0124] 4. Ending question: At the end of the content, be sure to leave a question for "you" to guide reflection or practice.

[0125] S104: splicing the multiple single-digital human videos and the multiple-digital human videos in sequence to obtain an interactive video with multiple digital human interactions.

[0126] In this embodiment, based on the above-mentioned text, multiple single-digital human videos and multi-digital human videos can be generated, and then they can be spliced together in order to obtain an interactive video with multi-digital human interaction.

[0127] In summary, based on this embodiment, interactive oral broadcasting of digital humans by multiple people can be realized, which can greatly enrich the expression forms and application scenarios of digital human oral broadcasting and provide users with a better viewing experience.

[0128] Some preferred embodiments of the present invention are further described below.

[0129] Preferably, during the image-generated video process, a lora model for controlling the character's generated posture is fine-tuned to control the character's generated posture and action details, thereby enhancing the realism and expressiveness of the generated content.

[0130] At the same time, in a multi-digital human video, when one of the digital humans is speaking, the other digital humans will interact accordingly based on the content of the speaking digital human.

[0131] In this embodiment, the image-to-video process requires character action card drawing. The improved solution introduces a LoRA model to control character pose generation. This focuses on controlling the details of character generation, enabling natural movements like a slight nod or simple hand gestures, thereby enhancing the realism and expressiveness of the generated content. This also helps reduce the "card drawing rate" of poorly generated results, significantly improving overall generation quality and stability, and making the generated video more coherent.

[0132] In addition, the LoRA model based on the control of character-generated postures can also better realize the interaction of multiple digital humans. For example, in a video of multiple digital humans, when one of the digital humans is speaking, the other digital humans will generate corresponding interactions based on the content of the speaking digital human, such as nodding, smiling, and being surprised.

[0133] See also Figure 3 The second embodiment of the present invention further provides a video generation system based on multi-human interaction, which includes:

[0134] The image generation unit 210 is configured to generate a plurality of single-digit human images for live broadcasting, wherein each single-digit human image has a similar background;

[0135] An image stitching and expansion unit 220 is configured to stitch and expand the plurality of single-digital human images to obtain a plurality of digital human images with a consistent background;

[0136] The image-to-video unit 230 is used to generate corresponding multiple single-digit human videos and multi-digit human videos according to the preset copy;

[0137] The video splicing unit 240 is configured to splice the multiple single-digital human videos and the multiple-digital human videos in sequence to obtain an interactive video with multiple digital human interactions.

[0138] The third embodiment of the present invention further provides a video generation device based on multi-digital human interaction, which includes a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement the video generation method based on multi-digital human interaction as described above.

[0139] The fourth embodiment of the present invention further provides a computer-readable storage medium storing a computer program. The computer program can be executed by the processor to implement the above-mentioned method for generating a video based on multiple digital human interactions.

[0140] For example, the computer programs described in the third and fourth embodiments of the present invention may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the error correction device. For example, the system described in the second embodiment of the present invention may be used.

[0141] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the error correction method, and utilizes various interfaces and lines to connect the various parts of the method for achieving high-confidence intelligent live broadcast response.

[0142] The memory can be used to store the computer program and / or module, and the processor implements various functions of an error correction method by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0143] Wherein, if the implemented module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of each of the above-mentioned method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0144] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0145] The above are only preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention.

Claims

1. A video generation method based on multi-human interaction, characterized in that: include: Generate multiple single-digit human images for live broadcast; wherein each single-digit human image has a similar background; Performing image stitching and expansion on the multiple single-digit human images to obtain a multiple-digit human image with a consistent background; Generate multiple single-digit human videos and multi-digit human videos according to the preset copy; The multiple segments of single-digital human videos and multi-digital human videos are spliced in sequence to obtain an interactive video with multi-digital human interaction.

2. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: Generating multiple single-digit human images for live broadcast specifically includes: Based on the pre-trained lora model, a digital human image is generated according to the input description; Use flux to generate a background with a similar style, and combine the generated digital human image and the background to generate multiple single digital human images with similar backgrounds.

3. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: Performing image stitching and expansion on the multiple single-digit human images to obtain multiple human images with consistent backgrounds specifically includes: Hard-joining the multiple single-digital human images, and intelligently redrawing the joints using context-aware content filling technology to obtain a preliminary multi-digital human image; Performing image expansion on the preliminary multiple digital human images so that the multiple digital humans are more realistically placed in the same space; Resize the expanded multi-digital human image to make it consistent with the size of the single-digital human image.

4. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: During the image-generated video process, a lora model that controls the character's generated posture is fine-tuned to control the character's generated posture and movement details, thereby enhancing the realism and expressiveness of the generated content.

5. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: In the process of generating multiple single-digit and multi-digit human videos according to the preset text: In the above copy, the character personality, expression style, knowledge background and thinking mode of each digital human image are included. Add colloquial expressions that match the role; Design interactive scripts corresponding to the current video and generate the final dialogue copy to make the interaction between digital people more vivid.

6. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: The multiple digital human images have the same image or different images.

7. The video generation method based on multi-digital human interaction according to claim 1, characterized in that: In a multi-digital human video, when one of the digital humans is speaking, the other digital humans will interact accordingly based on the content of the speaking digital human.

8. A video generation system based on multi-digital human interaction, characterized in that: include: An image generation unit, configured to generate a plurality of single-digit human images for live broadcasting; wherein each single-digit human image has a similar background; An image stitching and expansion unit, configured to stitch and expand the plurality of single-digital human images to obtain a plurality of digital human images with a consistent background; The image-to-video unit is used to generate corresponding multiple single-digit human videos and multi-digit human videos according to the preset copy; The video splicing unit is used to splice the multiple single-digital human videos and the multiple-digital human videos in sequence to obtain an interactive video with multiple digital human interactions.