Video generation method and device based on large model, intelligent agent, equipment, medium and product
By combining temporal alignment and the collaborative processing of a large video generation model, the problem of inconsistency between lip movements and body movements in digital human video generation is solved, enabling automated control and efficient generation of object changes in the target video.
Patent Information
- Application Number
- CN202511317862.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for generating digital human videos suffer from problems such as inconsistencies between audio and body movements, representational limitations, and difficulty in learning the connection between lip movements and body movements, resulting in unnatural and inaccurate videos when editing lip movements or body movements separately.
By temporally aligning the reference content with the first video segment, at least two reference video frames are identified. Then, a large video generation model is used to process the reference content and video frames to generate a second video segment that matches the reference change process, thus achieving the splicing of the target video.
It improves the consistency between the change process of the target object in the target video and the reference content, enhances the generation efficiency and the naturalness of the video, and ensures that the change process of the target object area matches the change process of the reference.
Smart Images

Figure CN121000950A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model, etc., which can be applied to the scenarios of digital human, content generation based on artificial intelligence, etc., and more particularly to a large model based video generation method, device, agent, equipment, medium and product. BACKGROUND
[0002] Digital human (i.e., Digital Human) refers to a digitalized figure existing in a non-physical world, created and used by computer means, and having multiple human characteristics (e.g., appearance characteristics, human performance ability, interaction ability, etc.).
[0003] Digital human has been widely used in various industries, such as entertainment, education, e-commerce. With the gradual improvement of industry demand and user experience, the demand for digital human video in various industries is also increasing. SUMMARY
[0004] The present disclosure provides a large model based video generation method, device, agent, equipment, medium and product.
[0005] According to one aspect of the present disclosure, a large model based video generation method is provided, comprising: determining at least two reference video frames by performing time sequence alignment on reference content and a first video segment, wherein the first video segment is extracted from a reference video according to the reference content, and the reference content indicates a reference change process related to a target object in the reference video; and splicing the first video segment and a second video segment to obtain a target video, wherein the second video segment is obtained by processing the reference content and the at least two reference video frames using a video generation large model, and a change process of at least one region of the target object in the target video matches the reference change process.
[0006] According to another aspect of the present disclosure, a large model based video generation device is provided, comprising: a time sequence alignment module configured to determine at least two reference video frames by performing time sequence alignment on reference content and a first video segment, wherein the first video segment is extracted from a reference video according to the reference content, and the reference content indicates a reference change process related to a target object in the reference video; and a video generation module configured to splice the first video segment and a second video segment to obtain a target video, wherein the second video segment is obtained by processing the reference content and the at least two reference video frames using a video generation large model, and a change process of at least one region of the target object in the target video matches the reference change process.
[0007] According to another aspect of the present disclosure, there is provided a large model-based agent, comprising:
[0008] an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a target large model based on the target task, execute a large model-based video generation method by calling the target large model, and obtain output information, the target large model comprising at least one of a visual language large model, a video generation large model, and a video evaluation large model; and an output module configured to output the output information obtained by the processing module.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, comprising: one or more processors; a memory configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.
[0010] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the steps of the method.
[0011] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the steps of the method.
[0012] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0013] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0014] Figure 1 A system architecture to which a large model-based video generation method can be applied according to an embodiment of the present disclosure is schematically illustrated;
[0015] Figure 2 A flowchart of a large model-based video generation method according to an embodiment of the present disclosure is schematically illustrated;
[0016] Figure 3 An example schematic diagram of a process of obtaining a first video clip according to an embodiment of the present disclosure is schematically illustrated;
[0017] Figure 4An example schematic diagram of a process of determining at least two reference video frames by time-aligning the reference content with the first video segment according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 5 An example schematic diagram of a process of obtaining the second video segment according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 6 An example schematic diagram of a process of obtaining the second video segment in a case where the reference video includes at least two candidate objects according to an embodiment of the present disclosure is schematically shown;
[0020] Figure 7 An example schematic diagram of a process of generating a video based on a large model according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 8 A block diagram of a device for generating a video based on a large model according to an embodiment of the present disclosure is schematically shown;
[0022] Figure 9 A structural block diagram of an agent of a large model according to an embodiment of the present disclosure is schematically shown; and
[0023] Figure 10 A block diagram of an electronic device adapted to implement a method of generating a video based on a large model according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0024] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It should be understood, however, that the description which follows is merely exemplary and is not intended to limit the scope of the present disclosure. In the following detailed description of the embodiments of the present disclosure, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that the embodiments of the present disclosure can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present disclosure.
[0025] The terms used herein are merely used to describe specific embodiments and are not intended to limit the present disclosure. The terms "include", "comprise", and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present specification, and should not be interpreted in an idealized or excessively formal manner.
[0027] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally to be understood that the expression is to be interpreted in the same manner as a list follows the expression, such as "at least one of A, B, and C" should be interpreted as including at least one of A or B or C, but not limited to the entire set of A and B and C.
[0028] The method for generating a digital human video can include at least one of a method for editing lip movement alone, a method for editing body movement alone, and a method for editing lip movement and body movement simultaneously.
[0029] However, since the method for editing lip movement alone can only perform lip editing, there can be a case where audio and body movements in the original base video do not match; the method for editing body movement alone is limited by the representation itself, and has limited expression capability for local areas that need to be precisely driven, such as the face and hands; and the method for editing lip movement and body movement simultaneously generally uses a unified human face and body 3D representation, and there is a "many-to-many" relationship between the audio and the representation, making it difficult to learn the relationship between them, resulting in stiff and unnatural body movements and inaccurate lip movements.
[0030] To this end, an embodiment of the disclosure proposes a video generation scheme based on a large model. For example, by performing time sequence alignment on the reference content and the first video segment, at least two reference video frames are determined, wherein the first video segment is extracted from the reference video according to the reference content, and the reference content indicates a reference change process related to a target object in the reference video; the first video segment and the second video segment are spliced to obtain a target video, wherein the second video segment is obtained by processing the reference content and at least two reference video frames using a video generation large model, and the change process of at least one region of the target object in the target video matches the reference change process.
[0031] According to an embodiment of the disclosure, the first video segment is intercepted from the reference video based on the provided reference content, at least two reference video frames are determined by performing time sequence alignment on the reference content and the first video segment, the second video segment is obtained by processing the reference content and at least two reference video frames using a video generation large model, and the first video segment and the second video segment are spliced to obtain a target video. This time sequence alignment and video generation large model collaborative processing method can control the change process of the target object in the generated target video, automatically generate a target video that matches the reference content, improve the consistency of the change process of at least one region of the target object in the target video with the reference change process indicated by the reference content, and improve the generation efficiency of the target video and the degree of consistency with the reference content.
[0032] The collection, storage, use, processing, transmission, provision and disclosure of the user personal information in the technical solution of the present application comply with relevant laws and regulations and do not violate public order and good customs.
[0033] In the technical solution of the present application, the authorization or consent of the user is obtained before the user personal information is acquired or collected.
[0034] Figure 1 The system architecture to which the large model-based video generation method according to the embodiments of the present disclosure can be applied is schematically shown. It should be noted that, Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0035] As Figure 1 shown, the system architecture 100 according to the embodiments can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0036] The user can use at least one of the first terminal device 101, the second terminal device 102 and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102 and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0037] The first terminal device 101, the second terminal device 102 and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.
[0038] The server 105 can be a server providing various services, such as a background management server supporting the website browsed by the user using the first terminal device 101, the second terminal device 102 and the third terminal device 103 (only as an example). The background management server can analyze and process the received user requests and other data, and feed back the processing results (such as web pages, information or data generated according to the user requests, etc.) to the terminal device.
[0039] It should be noted that the large model-based video generation method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the large model-based video generation apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The large model-based video generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the large model-based video generation apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0040] Alternatively, the large model-based video generation method provided by the embodiments of the present disclosure can also be executed by the first terminal device 101, the second terminal device 102 or the third terminal device 103, or can also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103. Correspondingly, the large model-based video generation apparatus provided by the embodiments of the present disclosure can also be arranged in the first terminal device 101, the second terminal device 102 or the third terminal device 103, or can also be arranged in other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103.
[0041] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above system is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.
[0042] It should be noted that the serial numbers of the various operations in the following method are only used to represent the operations for description, and should not be regarded as representing the execution sequence of the various operations. Unless explicitly stated, the method does not need to be executed in the order shown.
[0043] Figure 2 A flowchart of a large model-based video generation method according to an embodiment of the present disclosure is schematically shown.
[0044] As Figure 2 The large model-based video generation method 200 includes operations S210-S220.
[0045] At operation S210, at least two reference video frames are determined by time sequence alignment of the reference content and the first video segment, wherein the first video segment is extracted from a reference video according to the reference content, and the reference content indicates a reference change process related to a target object in the reference video.
[0046] In operation S220, the first video segment and the second video segment are spliced to obtain a target video, wherein the second video segment is obtained by processing the reference content and the at least two reference video frames by using the video generation large model, and a change process of at least one region of the target object in the target video matches the reference change process.
[0047] The reference content refers to content provided by a user to guide a video generation process. The reference content indicates a reference change process related to a target object in a reference video, that is, a change process that the target object should present in a target video expected to be generated. The target object can refer to a subject in the target video, for example, a person, a commodity, a background, and the like. In an embodiment of the present disclosure, before obtaining the reference content and the reference video, consent or authorization of the user can be obtained. For example, before operation S210, a request for obtaining the reference content and the reference video can be sent to the user. In a case where the user agrees or authorizes that the reference content and the reference video can be obtained, operation S210 is performed.
[0048] In one example, the reference content can include at least one of a reference text and a first reference audio. For example, in a case where the large model-based video generation method provided by the present disclosure is applied to a live broadcast scenario, the reference text can include a script describing that “the host picks up a commodity and smiles”, and the first reference audio can include a voice obtained by processing the script through text-to-speech (TTS) conversion.
[0049] The first video segment refers to one or more video paragraphs extracted from the reference video according to the reference content. In one example, the first video segment can be extracted from the reference video according to scene semantic information of the reference content. For example, continuing to take the script describing that “the host picks up a commodity and smiles” as the reference content, the first video segment can be a segment in which the host is picking up the commodity in the reference video.
[0050] The time sequence alignment refers to a process of matching the reference content and the first video segment in time. In one example, an audio-visual alignment algorithm can be used to align by combining speech recognition and visual action detection; alternatively, a multi-modal alignment network based on an attention mechanism can be used to align the reference content and the first video segment; alternatively, a dynamic time warping (DTW) algorithm can be used to align speech and video features; alternatively, rough alignment can be performed through human key point detection and speech keyword extraction, and then fine-tuning can be performed.
[0051] For example, according to the reference content, three first video clips are extracted from the reference video, that is, first video clip 1, first video clip 2, and first video clip 3. Before the time alignment, the three first video clips can be arranged in time sequence or not. By time aligning the reference content and the three first video clips, the three first video clips can be arranged in time sequence to facilitate the subsequent extraction of reference video frames.
[0052] After arranging the first video clips in time sequence, at least two reference video frames can be determined according to the sequence relationship between the first video clips. The reference video frame refers to a key frame in the first video clip, which is used to indicate the start and end state of the second video clip to be generated. After determining at least two reference video frames, the second video clip for splicing generated by the video generation large model according to the reference content and the reference video frame can be used to fill the gap between the first video clips or expand the video content.
[0053] It should be noted that for each two adjacent first video clips, they can be regarded as a combination, and the combination corresponds to at least two reference video frames, and at least one of the reference video frames is extracted from the preceding video clip and at least one is extracted from the following video clip.
[0054] For example, continuing with the three first video clips, the time sequence relationship can be first video clip 1->first video clip 2->first video clip 3. In this case, the first video clip 1 and the first video clip 2 can be regarded as combination 1, and the first video clip 2 and the first video clip 3 can be regarded as combination 2.
[0055] For combination 1, the first video clip 1 is the preceding video clip, the first video clip 2 is the following video clip, and the purpose is to generate the second video clip 1 between the first video clip 1 and the first video clip 2. In this case, the reference video frame 1 can be extracted from the first video clip 1, the reference video frame 2 can be extracted from the first video clip 2, and the reference video frame 1 can be used as the first frame of the subsequently generated second video clip 1, and the reference video frame 2 can be used as the tail frame of the subsequently generated second video clip 1.
[0056] For combination 2, the first video clip 2 is the preceding video clip, the first video clip 2 is the following video clip, and the purpose is to generate the second video clip 2 between the first video clip 2 and the first video clip 3. In this case, the reference video frame 3 can be extracted from the first video clip 2, the reference video frame 4 can be extracted from the first video clip 3, and the reference video frame 3 can be used as the first frame of the subsequently generated second video clip 2, and the reference video frame 4 can be used as the tail frame of the subsequently generated second video clip 2.
[0057] After the second video clip between each first video clip is generated, the first video clip and the second video clip can be spliced to obtain the target video. Splicing refers to the process of seamlessly splicing the existing first video clip and the generated second video clip on the time axis to form a coherent target video. In an example, the video generation large model can be used to directly splice after generating the second video clip and ensure quality through a quality inspection cycle; alternatively, optical flow estimation and frame interpolation techniques can be used to smooth the transition boundary; alternatively, a timing consistency loss can be introduced to train the generation model to ensure motion coherence between video clips; alternatively, an adversarial training can be used to enhance the realism of each video clip to avoid splicing traces.
[0058] In an example, the region can include at least one of a body region of the target object, a lip region of the target object, and an article display region, and the reference change process can include a change process corresponding to the body region of the target object, a change process corresponding to the lip region of the target object, a change process corresponding to the article display region, and the like, which are merely examples. Based on this, the target video can be a video in which the change process of each region matches the reference change process.
[0059] In an embodiment, the large model-based video generation method provided by the present disclosure can be applied to a live broadcast scene, in which case the target object includes a digital human, and a live broadcast marketing video of an anchor can be generated based on the large model-based video generation method provided by the embodiments of the present disclosure. For example, the digital human can be an anchor in a live broadcast scene, and the live broadcast scene can include, but is not limited to, e-commerce live broadcast, game live broadcast, reality show live broadcast, education lecture live broadcast, news broadcast live broadcast, and the like.
[0060] In another embodiment, the large model-based video generation method provided by the present disclosure can also be applied to any application scenario such as animation production, film and television product production, and metaverse scene construction, and the large model-based video generation method provided by the embodiments of the present disclosure is not limited to a specific application scenario.
[0061] According to an embodiment of the present disclosure, the first video segment is intercepted in the reference video based on the provided reference content, at least two reference video frames are determined by time sequence alignment of the reference content and the first video segment, the second video segment is obtained by processing the reference content and the at least two reference video frames by using the video generation large model, and the target video is obtained by splicing the first video segment and the second video segment. The cooperative processing mode of time sequence alignment and the video generation large model realizes the control of the change process of the target object in the generated target video, can automatically generate the target video matched with the reference content, improves the consistency of the change process of at least one region of the target object in the target video with the reference change process indicated by the reference content, and improves the generation efficiency of the target video and the consistency with the reference content.
[0062] Figure 3 An example schematic diagram of the obtaining process of the first video segment according to an embodiment of the present disclosure is schematically shown.
[0063] As Figure 3 shown, in the obtaining process 300 of the first video segment, the scene semantic analysis of the reference content 310 is performed by using the visual language large model M310 to obtain the scene semantic information 320 for describing the reference change process. The visual language large model M310 refers to a large-scale pre-training deep learning model capable of processing and understanding images and texts simultaneously, which can convert visual content into language description or understand visual content according to language description.
[0064] The scene semantic analysis refers to the process of in-depth understanding of the reference content 310 by using the visual language large model M310 to analyze the rich scene semantic information 320 such as actions, events, objects, emotions, intentions, etc. contained therein. For example, in a live broadcast scene, scene semantic analysis is performed on the above script to analyze the information such as objects (e.g., anchor, audience, mobile phone, and articles), actions (e.g., greeting, picking up, and showing), states (e.g., smiling and new model), and intentions (e.g., introducing products).
[0065] The scene semantic information 320 obtained in this way refers to the structured or unstructured semantic representation obtained after scene semantic analysis, which is a summary and description of the deep meaning of the reference content and can be used to guide subsequent video retrieval and extraction. For example, the scene semantic information 320 can be: “the person (anchor) faces the camera, the expression is a smile, and the waving action (greeting) is performed. Scene two: the person's hand interacts with the mobile phone on the table, and the picking up action is performed. Scene three: the person holds the mobile phone to the chest, and the screen is towards the camera (showing)”.
[0066] The way in which the visual language large model M310 generates the scene semantic information 320 can include a way based on an instructional prompt, a way based on serialized information extraction, a way based on multi-round question and answer analysis, and the like. The way based on the instructional prompt refers to designing a detailed text prompt (i.e., Prompt) to guide the visual language large model M310 to perform analysis. The way based on the serialized information extraction refers to combining the visual language large model M310 with information extraction technology, and requiring the visual language large model M310 to output the analysis result in a fixed JSON format. The way based on the multi-round question and answer analysis refers to using an iterative strategy, first letting the visual language large model M310 perform preliminary analysis, then proposing a follow-up prompt according to the result of the preliminary analysis, and obtaining deeper and more detailed scene semantic information through multiple rounds of interaction.
[0067] After obtaining the scene semantic information 320, the scene semantic information 320 can be used to extract a segment matching the scene semantic information 320 from the reference video 330, to obtain first video segments 341, …, and first video segments 34N, where N is a positive integer. The extraction method of the first video segments can be configured according to actual business requirements, and is not limited herein.
[0068] In one example, the reference video 330 can be uniformly sampled into video frames, each video frame can be encoded into a first feature vector using the visual language large model M310, and the scene semantic information 320 can be encoded into a second feature vector. By calculating the cosine similarity between the first feature vector and the second feature vector, all key frames with high similarity can be found, and these key frames can be spliced into the first video segments through time sequence analysis. In another example, a pre-trained action recognition model can be used to perform time sequence action detection on the reference video 330, to directly output the start and end time stamps of the action occurrence, thereby extracting the first video segments corresponding to the scene semantic information 320.
[0069] According to embodiments of the present disclosure, by introducing a visual language large model to perform deep scene semantic analysis on reference content, first video segments that are highly matched in semantics with the reference content can be accurately automatically retrieved and extracted from the reference video 330, thereby improving the automation degree, semantic understanding accuracy, and processing efficiency in the video preprocessing stage.
[0070] Figure 4 An example schematic diagram of a process of determining at least two reference video frames by time sequence alignment of the reference content and the first video segments according to embodiments of the present disclosure is schematically shown.
[0071] As Figure 4As shown, in the process 400 of determining the at least two reference video frames, the reference content can include at least one of a reference text and a first reference audio 410. The reference text is a script, a dialogue or a content outline described in natural language, indicating events and dialogues that should occur in the target video to be generated. The first reference audio 410 is audio data automatically synthesized from the reference text by text-to-speech technology. The first reference audio 410 or a second reference audio generated based on the reference text can be time-aligned with the first video clips to determine the timing relationship of the plurality of first video clips.
[0072] Time alignment refers to a process of establishing a time correspondence between the first reference audio 410 or the second reference audio and the first video clips, i.e., determining which first video clip a certain dialogue or a certain sentence corresponds to, and the start and end time in the first video clip. The timing relationship is the arrangement order and duration relationship of the plurality of first video clips on the timeline.
[0073] For example, by time-aligning the first reference audio 410 with the first video clip 421, the first video clip 422 and the first video clip 423, the timing relationship of the plurality of first video clips can be determined as first video clip 422 -> first video clip 421 -> first video clip 423, and the first video clip 422 lasts for 3 seconds, the first video clip 421 lasts for 5 seconds, and the first video clip 423 lasts for 4 seconds.
[0074] According to an embodiment of the present disclosure, by time-aligning the reference audio with the plurality of first video clips, the audio-visual timing correspondence is automatically established, and the reference video frames for video generation are intelligently determined accordingly, ensuring that the actions, lip shapes and voice content in the finally generated target video are completely synchronized on the timeline, and improving the automation degree, timing accuracy and overall fluency of video generation.
[0075] After obtaining the timing relationship, at least two reference video frames can be determined in each two adjacent first video clips according to the timing relationship. In one embodiment, the tail video frame of the preceding video clip and the head video frame of the following video clip can be determined as the reference video frames.
[0076] For each two adjacent first video clips, the preceding video clip refers to the first video clip in front of the timing relationship, and the following video clip refers to the first video clip behind the timing relationship. The tail video frame refers to the last frame of image of the preceding video clip in time, and the head video frame refers to the first frame of image of the following video clip in time.
[0077] For example, for the combination of the first video segment 422 and the first video segment 421, the first video segment 422 is the preceding video segment and the second video segment 421 is the following video segment, based on which, the tail video frame 431_1 in the first video segment 422 and the head video frame 432_1 in the first video segment 421 can be determined as the reference video frame 440.
[0078] Alternatively, for the combination of the first video segment 421 and the first video segment 423, the first video segment 421 is the preceding video segment and the first video segment 423 is the following video segment, based on which, the tail video frame 433_1 in the first video segment 421 and the head video frame 434_1 in the first video segment 423 can be determined as the reference video frame 440.
[0079] Alternatively, the median frame of the last N frames of the first video segment 422 can also be taken as a "virtual" tail video frame, and similarly, the "virtual" head video frame can be obtained by processing the first N frames of the second video segment 421, so as to improve the stability of the reference frame.
[0080] In another embodiment, the first video frame with the largest data amount or the highest image quality in the preceding video segment and the second video frame with the largest data amount or the highest image quality in the following video segment can be determined as the reference video frame. The data amount is an index for measuring the richness of image information, and the frame with a large data amount usually contains more details and changes, and is more informative. The highest image quality is an index for measuring the visual performance of the image, for example, the image quality can be evaluated from multiple dimensions, including but not limited to definition, noise level, exposure accuracy, artifact-free, etc.
[0081] For example, for the combination of the first video segment 422 and the first video segment 421, the first video segment 422 is the preceding video segment and the second video segment 421 is the following video segment, based on which, the first video frame 431_2 with the largest data amount or the highest image quality in the first video segment 422 and the second video frame 432_2 with the largest data amount or the highest image quality in the first video segment 421 can be determined as the reference video frame 440.
[0082] Alternatively, for the combination of the first video segment 421 and the first video segment 423, the first video segment 421 is the preceding video segment and the first video segment 423 is the following video segment, based on which, the first video frame 433_2 with the largest data amount or the highest image quality in the first video segment 421 and the head video frame 434_2 with the largest data amount or the highest image quality in the first video segment 423 can be determined as the reference video frame 440.
[0083] According to an embodiment of the present disclosure, by providing multiple key frame determination strategies, the reference for transition generation between video clips is ensured, and by introducing the filtering dimension of data volume or image quality, the selected reference video frames have rich information content and high-quality visual performance, thereby providing more stable and clearer visual conditions for subsequent video generation large models, improving the coherence and authenticity of the generated transition clips, and helping to improve the accuracy of the generated target video.
[0084] Figure 5 An example schematic diagram of the obtaining process of the second video clip according to an embodiment of the present disclosure is schematically shown.
[0085] As shown in the obtaining process 500 of the second video clip, the reference content (such as the first reference audio 510) can be used to guide the visual language large model M510 to understand the content of the reference video frame 521 and the reference video frame 522, so as to guide the visual language large model M510 to understand the semantic relationship between the reference video frame 521 and the reference video frame 522 through the reference content, and obtain the description text 530. Figure 5 The description text 530 is a natural language text representing the intermediate change process between the reference video frame 521 and the reference video frame 522. For example, "the right hand of the person slowly lifts up and holds the cup on the table, and the background remains unchanged". The intermediate change process refers to the visual dynamic change that occurs between the reference video frame 521 and the reference video frame 522, such as including character actions, object movements, scene transformations, etc.
[0086] After obtaining the description text 530, the video generation large model M520 can be used to process the description text 530, the reference video frame 521 and the reference video frame 522, to obtain the second video clip 550 that conforms to the description text 530 and is visually coherent. For example, the input description text is "the hand slowly lifts up and holds the cup", the picture content of the reference video frame 521 is that the hand starts to lift up, and the picture content of the reference video frame 522 is that the hand holds the cup, and the second video clip 550 showing the intermediate action can be generated.
[0087] According to an embodiment of the present disclosure, by introducing the visual language large model to understand the dynamic change process between the reference video frames and describe the text, and combining the video generation large model to generate high-quality second video clips, the semantic consistency, action naturalness and overall expressiveness of the generated video are improved, which is conducive to realizing the automatic and integrated generation process from the original material to the high-expressiveness target video.
[0088]
[0089] In one embodiment, the process of processing the description text 530, the reference video frame 521 and the reference video frame 522 by the video generation large model M520 to obtain the second video clip 540 can include: processing the description text 530, the reference video frame 521 and the reference video frame 522 by the video generation large model M520 to obtain a plurality of candidate video clips 540.
[0090] The video generation large model M520 is a model for quality evaluation of the generated candidate video clip 540. For example, the video generation large model M520 can be a classification model based on CNN or Transformer structure, outputting a quality score or a defect label.
[0091] The candidate video clip 540 is a plurality of candidate video clips generated by the video generation large model M520 for the same set of input description text 530 and reference video frame, and the quality can be different. For example: 5 candidate video clips 540 of "picking up a cup" are generated, of which 3 are natural actions, and 2 have hand deformation or position jump of the cup.
[0092] In one example, the possibility of improving the final output quality can be increased by increasing the diversity of the plurality of candidate video clips 540. For example, a diffusion model can be used to generate video clips corresponding to a plurality of random seeds, and each seed generates a candidate video clip 540; alternatively, candidate video clips 540 of different styles can be generated by conditional control (such as optical flow, depth map); alternatively, a multi-scale generation strategy can be introduced, first generating a low-resolution candidate video clip 540, and then super-resolution or refinement generation is performed on the high-score candidate video clip 540.
[0093] For the obtained plurality of candidate video clips 540, the video evaluation large model M530 can be used to evaluate the quality of the plurality of candidate video clips 540 to obtain the second video clip 550. Quality evaluation refers to the process of automatically scoring or classifying candidate video clips 540 to filter out clips that meet the quality requirements.
[0094] In one example, the quality evaluation method can include using a classification model to perform binary classification judgment on common problems (such as artifacts, deformation, and physical irrationality), and the scores are integrated. Alternatively, contrastive learning or reinforcement learning can be introduced to train the video evaluation large model M530 to simulate human aesthetic preferences and perform more fine-grained quality sorting. Alternatively, low-level features such as optical flow consistency, image clarity, and semantic alignment, and high-level semantic features can be combined for multi-dimensional evaluation, and a comprehensive quality evaluation index can be constructed.
[0095] According to an embodiment of the present disclosure, by generating a plurality of candidate video clips and combining an automated quality evaluation mechanism, the reliability, naturalness and usability of the generated video clips are improved. By using the video evaluation large model to screen a plurality of candidate video clips, it is ensured that the final obtained second video clip meets the physical laws and semantic requirements in terms of character action, commodity performance, scene rationality, etc., and the automatic, iterative and high-quality generation of high-expression video content is realized.
[0096] In one embodiment, the quality evaluation of a plurality of candidate video clips 540 by the video evaluation large model M530 to obtain a second video clip 550 can include operation S510. The video evaluation large model M530 can be trained using samples with pre-labeled sample quality evaluation values.
[0097] For each candidate video clip 540, the video evaluation large model M530 can be used to evaluate the quality of the candidate video clip 540 to obtain a quality evaluation value 550. After obtaining the quality evaluation value 550, operation S510 is performed.
[0098] In operation S510, is the quality evaluation value 550 satisfied the preset quality condition? The preset quality condition can be that the quality evaluation value is greater than or equal to a preset quality threshold. The quality evaluation value 550 is the quality quantitative output of the video evaluation large model M530 on the candidate video clip 540, which can be a score, a probability value or a binary classification result. For example: a score of 0.6 (full score of 1.0), or a pass / fail label.
[0099] If yes, the candidate video clip 540 corresponding to the quality evaluation value 550 can be determined as the second video clip 560. If no, the following operations can be repeatedly performed until the updated quality evaluation value satisfies the preset quality condition: using the video generation large model M520 to reprocess the description text 530, the reference video frame 521 and the reference video frame 522 to obtain an updated candidate video clip; using the video evaluation large model M530 to evaluate the quality of the updated candidate video clip to obtain an updated quality evaluation value. The candidate video clip 540 whose updated quality evaluation value satisfies the preset quality condition is determined as the second video clip 560.
[0100] When the quality of the candidate video clip does not meet the standard, the generation-evaluation-generation cycle is automatically restarted to gradually optimize the output quality through multiple iterations. For example, taking a preset quality threshold of 0.8 as an example, the quality evaluation value of the first generated candidate video clip 540 is 0.6, which is lower than the preset quality threshold, so the generation-evaluation-generation cycle can be restarted until the updated quality evaluation value is greater than or equal to 0.8.
[0101] In the process of regenerating the candidate video clip 540, the same video generation large model M520 can be used but the random seed is adjusted to generate diverse candidate video clips 540; alternatively, the input conditions of the video generation large model M520 can be adjusted according to the last quality evaluation to guide the generation of more reasonable candidate video clips 540; alternatively, a reinforcement learning strategy can be introduced, taking the last quality evaluation value as a reward signal, and dynamically optimizing the parameters of the video generation large model M520.
[0102] According to an embodiment of the present disclosure, by introducing an iterative generation and evaluation mechanism, the automatic quality optimization and reliability improvement of the video clip are realized, and candidate videos can be continuously generated and screened until a clip meeting the preset quality condition is output, improving the usability, authenticity and expressiveness of the generated video, reducing the cost of manual intervention, and realizing efficient, automatic and closed-loop production of high expressiveness video content.
[0103] In one embodiment, the quality evaluation value can be obtained based on quality evaluation items, and the quality evaluation items include at least one of the following: an object-related quality evaluation item, an item-related quality evaluation item, and a scene-related quality evaluation item.
[0104] The object-related quality evaluation item can evaluate the performance of the characters in the video clip. For example, the object-related quality evaluation item can include at least one of the following: a lip movement accuracy item, a body movement naturalness item, an expression reasonableness item, an identity consistency item, etc. The lip movement accuracy item is used to evaluate whether the lip shape matches the audio. The body movement naturalness item is used to evaluate whether the movement conforms to human body kinematics, and whether there is unreasonable joint bending or shaking. The expression reasonableness item is used to evaluate whether the expression is consistent with the speech content and scene context. The identity consistency item is used to evaluate whether the generated facial features of the character are consistent with the target character in the original material, and whether there is identity drift.
[0105] The item-related quality evaluation item can evaluate the goods or other handheld or interactive items appearing in the video clip. For example, the item-related quality evaluation item can include at least one of the following: an item appearance integrity item, an item clarity item, an interaction reasonableness item, etc. The item appearance integrity item is used to evaluate whether the goods are deformed, distorted, torn or have artifacts. The item clarity item is used to evaluate whether the goods are clear and identifiable, and whether there is blur. The interaction reasonableness item is used to evaluate whether the interaction between the character and the item is reasonable, for example, whether the holding posture is correct, and whether the item appears or disappears in the air.
[0106] The scene-related quality evaluation item can evaluate the background environment and global changes in the video segment. The scene-related quality evaluation item can include at least one of the following: a background consistency item, a physical law compliance item, a semantic rationality item, and the like. The background consistency item is used to evaluate whether the background of the generated segment is continuous with the background of the first and last frames, and whether there is flickering or unreasonable jump. The physical law compliance item is used to evaluate whether the light and shadow changes, object motion trajectories, and the like in the scene comply with the physical law, for example, a free-falling object should accelerate downward, rather than at a constant speed. The semantic rationality (whether the overall changes of the scene are consistent with the process described in the description text.
[0107] According to an embodiment of the present disclosure, by establishing a multi-dimensional and fine-grained video quality evaluation system, the quality control of the generated video segment is expanded from a single overall score to a special evaluation for different semantic elements such as objects, items, scenes, etc., which helps to improve the pertinence, accuracy and explainability of quality detection, ensures that the finally generated target video meets high standard requirements in terms of detail performance, semantic consistency and overall perception, and improves the reliability and practicality of automatic video production in high requirement scenes such as live streaming and goods selling.
[0108] In one embodiment, if the user has additional requirements for the expressiveness of the video, or has additional design adjustments for more complex two-person / multi-person videos, the description text 530 or the reference video frames can be changed based on an interactive interface. For example, the reference video frame 521, the reference video frame 522 and the description text 530 can be displayed in the interactive interface. The interactive interface refers to an interface for showing the user with automatically generated intermediate results (such as reference video frames, description texts), and receiving editing input from the user.
[0109] For example, the reference video frame 521, the reference video frame 522 and the description text 530 can be displayed side by side in the interactive interface. Alternatively, a timeline view can be provided, the reference video frame 521 and the reference video frame 522 are placed at both ends of the timeline, and the description text 530 is decomposed into multiple action labels and arranged in time sequence on the track below the timeline, so that the user can more intuitively understand the action timing. Alternatively, a comparison view can be provided to simultaneously display the original unedited version and the user-edited version, facilitating the user to compare the changes.
[0110] In response to detecting an editing operation on at least one of the reference video frame 521, the reference video frame 522 and the description text 530 via the interactive interface, the description text 530 or the reference video frame 521 or the reference video frame 522 is updated to obtain an updated description text or an updated reference video frame. The editing operation refers to a modification instruction issued by the user through the interactive interface, which is used to adjust the content of the reference video frame or the semantics of the description text to change the final generated video content.
[0111] For example, the editing operation can include directly uploading a reference video frame, selecting other reference video frames, and directly modifying the textual content of the description text 530. Alternatively, a drop-down menu can be provided to recommend relevant vocabulary based on a large language model to speed up the editing process and maintain semantic reasonableness.
[0112] After obtaining the updated description text or the updated reference video frame, the updated description text or the updated reference video frame can be processed by the video generation large model M520 to obtain the second video segment 550.
[0113] According to an embodiment of the present disclosure, by providing a visual interactive interface, users can directly participate in and intervene in the process of video generation, combining automatic generation with user subjective intention, which helps to improve flexibility and controllability in processing complex, personalized or high expressiveness scenarios. In this process, users can guide the model to generate high-quality video segments that meet expectations through intuitive editing and correction, thereby meeting the needs of professional video production for detailed fine-tuning on the basis of ensuring automation efficiency.
[0114] Figure 6 An example schematic diagram of the obtaining process of the second video segment is schematically shown in the case where the reference video includes at least two candidate objects according to an embodiment of the present disclosure.
[0115] As Figure 6 shown, in the obtaining process 600 of the second video segment, the reference video includes at least two candidate objects. The candidate object is a different individual in the reference video that needs to be identified and distinguished. The target object is the main object that needs to be highlighted or edited in a certain time period, which is usually the current speaker or the main behavior initiator.
[0116] Taking the obtained reference video frame 621 and the reference video frame 622 as an example, the reference video frame 621 and the reference video frame 622 each include a candidate object 651 and a candidate object 652. The following will describe how to determine the target object from the candidate object 651 and the candidate object 652.
[0117] In one embodiment, the target object can be determined in the candidate object 651 and the candidate object 652 according to the voiceprint features of the candidate object 651 and the candidate object 652 based on audio data corresponding to the reference video. The audio data is an audio signal synchronized with the reference video. The voiceprint feature is a speech parameter capable of representing the speaker's biological characteristics and behavior characteristics, which has individual uniqueness and can be used for identity recognition.
[0118] In this case, speaker recognition is performed based on the audio data, that is, the voiceprint features are extracted from the audio data, and are compared with the reference voiceprint features of the candidate objects 651 and 652 respectively, and the candidate object with the highest similarity is determined as the target object, that is, the speaker in the current time period.
[0119] In another embodiment, the target object can be determined in at least one of the following ways: determining the target object from the at least two candidate objects according to the participation proportion of each candidate object in the reference content. The participation proportion can be used to measure the dominance or participation of each candidate object in the video content in a specific time period. For example, in one minute, the participation proportion of the candidate object 651 is 70%, and the participation proportion of the candidate object 652 is 30%.
[0120] In this case, by analyzing the reference text or the first reference audio in the reference content, the length, frequency or semantic importance of the lines associated with each candidate object are counted, and the one with the highest proportion is determined as the target object.
[0121] After the target object is determined, the reference content (such as the first reference audio 610) can be used to guide the visual language large model M610 to understand the content of the reference video frames 621 and 622, so as to guide the visual language large model M610 to understand the semantic relationship between the reference video frames 621 and 622 through the reference content, and obtain the description text 630.
[0122] In one embodiment, for the candidate objects other than the target object in the at least two candidate objects, the visual language large model can be guided according to the reference content to understand the content of the reference video frames 621 and 622 in this case, and the change process of at least one region of the other candidate objects is constrained to avoid the phenomenon that the non-speaker generated action is inconsistent with the scene, and the description text 630 is obtained.
[0123] The other candidate objects refer to the character objects that appear in the reference video frames but are not in a dominant position except for the target object. The constraint is to limit or avoid the description of the active or large action of the non-target object by means of instructions or constraint conditions when generating the description text.
[0124] In one example, a constraint instruction can be explicitly added in the text prompt (base prompt) of the input visual language large model M610, such as "when generating the description, please ensure that other objects only maintain basic poses or smiles, and do not generate active actions such as raising hands or walking". Alternatively, before inputting the reference video frame into the visual language large model M610, a segmentation model or other object recognition technology can be used to generate a visual mask in the image area where the other candidate objects are located, and the visual mask is input into the visual language large model M610 together with the reference video frame to indicate that the model should not change significantly in these areas, thereby guiding the visual language large model M610 to ignore or constrain the changes in these areas when generating the description text 630.
[0125] According to an embodiment of the present disclosure, by introducing behavior constraints for non-target objects in the video generation scene of multiple objects, the non-target objects can be intelligently identified according to the reference content and guided to actively constrain the behavior changes when the visual language large model generates the description text, ensuring that the non-target objects remain natural and static or only perform minor actions consistent with the scene logic in the generated video segment, thereby improving the rationality, focus, and overall visual impression of the multi-person interactive video.
[0126] After obtaining the description text 630, the description text 630, the reference video frame 621, and the reference video frame 622 can be processed by the video generation large model M620 to obtain a second video segment 640 that is consistent with the description text 630 and visually coherent.
[0127] According to an embodiment of the present disclosure, for complex scenes of multi-person interaction, by providing multiple automated and accurate target object determination mechanisms, combining voiceprint recognition of the audio modality and participation analysis of the content modality, the target object currently speaking or dominating the interaction can be reliably identified, thereby ensuring that the subsequent video generation process can focus on the correct target object, which helps to improve the semantic accuracy, interaction rationality, and overall visual impression of the generated video.
[0128] In one embodiment, the reference change process can include at least one of the following: an action change process of the target object, an article change process with the target object, a foreground change process of the picture, and a background change process of the picture.
[0129] The action change process of the target object refers to the changes in the motion, posture, or expression of the body parts of the target object in the target video. For example, the right hand of the host is lifted from the side of the body to the chest, the fingers change from stretching to a holding state, the head slightly nods, and the face shows a smile.
[0130] The process of change of objects in relation to the target object refers to the changes in the state, position, and posture of the object that interacts with the target object in the target video. For example, a mobile phone on a table changes from a static state with the screen facing up to a state where it is picked up, the screen is turned towards the camera, and it is slightly shaken in the air.
[0131] The foreground change in a video refers to changes occurring in objects or elements located between the target object and the background, but not at the core of the interaction. For example, a hand reaches into the frame from outside, hands over a microphone, and then leaves the frame. The background change in a video refers to changes occurring in the underlying environment of the target video. For example, a PowerPoint presentation on the background screen flips from the first slide to the second.
[0132] According to embodiments of this disclosure, by decomposing the dynamic changes between video frames into multiple independent and fine-grained dimensions such as target object actions, object interactions, foreground and background, the visual language big model can generate more structured and accurate descriptive text, thereby providing more explicit and controllable generation conditions for the video generation big model. This helps to improve the realism, semantic consistency and physical rationality of the generated video, ensuring natural character movements, accurate object interactions and coherent scene transitions, ultimately producing high-quality and highly expressive digital human video content.
[0133] Figure 7 The illustration shows an example schematic diagram of a large-model-based video generation process according to an embodiment of the present disclosure.
[0134] like Figure 7 As shown, in the large-model-based video generation process 700, the visual language large model M710 is used to perform scene semantic analysis on the reference content 710 to obtain scene semantic information 720 used to describe the reference change process. The visual language large model M710 refers to a large-scale pre-trained deep learning model that can simultaneously process and understand images and text, and can convert visual content into language descriptions, or understand visual content based on language descriptions.
[0135] After obtaining scene semantic information 720, segments matching scene semantic information 720 can be extracted from reference video 730 based on scene semantic information 720 to obtain first video segment 741 and first video segment 742. For the combination of first video segment 741 and first video segment 742, taking first video segment 741 as the preceding video segment and first video segment 742 as the following video segment as an example, reference video frame 751 is extracted from first video segment 741 and reference video frame 752 is extracted from first video segment 742.
[0136] Taking the reference video frame 751 and the reference video frame 752 as an example, the reference content (such as the first reference audio 710) can be used to guide the visual language large model M710 to understand the content of the reference video frame 751 and the reference video frame 752, so as to guide the visual language large model M710 to understand the semantic relationship between the reference video frame 751 and the reference video frame 752 through the reference content 710, and obtain the description text 760.
[0137] After obtaining the description text 760, the video generation large model M720 can be used to process the description text 760, the reference video frame 751 and the reference video frame 752, to obtain the second video segment 770 consistent with the description text 760 and visually coherent.
[0138] After obtaining the second video segment 770, the first video segment 741, the first video segment 742 and the second video segment 770 can be spliced according to the time sequence relationship to obtain the intermediate video 780, and the change process of the body region and the article display region of the target object in the intermediate video 780 matches the reference change process.
[0139] The body region of the target object refers to the image region of the main part of the body of the target object in the video, for example, including the torso, limbs, etc. The lip region of the target object refers to the image region of the mouth in the face of the target object, for example, a rectangular image region containing the lips and the muscle movement around the lips. The article display region refers to an image region for highlighting a product or a key article, for example, an image region where a mobile phone held by the host or a cosmetic region placed on the table.
[0140] In one embodiment, splicing the first video segment 741, the first video segment 742 and the second video segment 770 to obtain the intermediate video 780 can include: for the first video segment 741 and the first video segment 742, inserting the second video segment 770 between the two adjacent first video segments to obtain an intermediate video segment. According to the time sequence relationship, the plurality of intermediate video segments are spliced to obtain the intermediate video 780. The intermediate video segment refers to a longer and more coherent video segment formed by inserting a second video segment between two adjacent first video segments.
[0141] Splicing is the process of connecting multiple video segments in a time sequence on a timeline to form a continuous video. In one example, the video segments can be connected in chronological order. Alternatively, a smooth transition effect can be generated using optical flow estimation or frame interpolation algorithms before and after the splicing point (for example, the last 5 frames of the previous segment and the first 5 frames of the next segment), to avoid motion jumps at the splicing point and make the intermediate video segment smoother. Alternatively, different splicing strategies can be used for the article display region and the body region.
[0142] According to an embodiment of the present disclosure, by intelligently generating and inserting a customized second video segment for each pair of adjacent original video segments first, and then performing overall splicing, the smooth connection between video segments is achieved, which is coherent, natural and in line with semantics, and the fluency, professionalism and visual perception of the final synthesized video are improved.
[0143] The intermediate video 780 refers to a complete video obtained after splicing operation, which contains all the changes of actions and objects but has not been driven by lip shape. For the intermediate video 780, the lip region of the target object in the intermediate video 780 can be driven according to the reference content 710 by using the lip driving model M740, so that the change process of the lip region matches the reference change process, and the target video 790 is obtained.
[0144] According to an embodiment of the present disclosure, by decoupling the video generation task into two stages of "overall action and object splicing" and "local lip driving", the global coherence and physical rationality of the action and object interaction of the main character are ensured through splicing first, and then the precise audio-lip synchronization is realized by using the special lip driving model, which not only ensures the natural and smooth visual flow of the whole video, but also ensures the high matching of the mouth shape and the voice, improves the expressiveness, realism and professionalism of the automatically generated video, and realizes the efficient and high-quality digital human video synthesis.
[0145] The above is only an exemplary embodiment, but is not limited thereto, and other large model-based video generation methods known in the art can also be included, as long as the consistency between the change process of at least one region of the target object in the target video and the reference change process indicated by the reference content is improved, and the generation efficiency of the target video and the degree of coincidence with the reference content are improved.
[0146] Figure 8 A block diagram of a large model-based video generation apparatus according to an embodiment of the present disclosure is schematically shown.
[0147] As shown in Figure 8 The large model-based video generation apparatus 800 can include a time sequence alignment module 810 and a video generation module 820.
[0148] The time sequence alignment module 810 is configured to determine at least two reference video frames by time sequence alignment between the reference content and the first video segment, wherein the first video segment is extracted from the reference video according to the reference content, and the reference content indicates a reference change process related to a target object in the reference video.
[0149] The video generation module 820 is used to splice the first video segment and the second video segment to obtain the target video. The second video segment is obtained by processing the reference content and at least two reference video frames using a large video generation model. The change process of at least one region of the target object in the target video matches the change process of the reference.
[0150] According to embodiments of this disclosure, the large-model-based video generation apparatus 800 may further include a content understanding module and a first processing module.
[0151] The content understanding module is used to guide the visual language big model to understand the content of at least two reference video frames using reference content, and obtain descriptive text, which represents the intermediate changes between the reference video frames.
[0152] The first processing module is used to process the descriptive text and at least two reference video frames using a large video generation model to obtain a second video segment.
[0153] According to embodiments of this disclosure, the first processing module may include a first processing unit and a quality assessment unit.
[0154] The first processing unit is used to process descriptive text and at least two reference video frames using a large video generation model to obtain multiple candidate video segments.
[0155] The quality assessment unit is used to evaluate the quality of multiple candidate video segments using a large video assessment model to obtain a second video segment.
[0156] According to embodiments of this disclosure, the quality assessment unit may include a quality assessment subunit and a first determination subunit.
[0157] The quality assessment subunit is used to perform quality assessment on candidate video segments using a large video assessment model to obtain quality assessment values.
[0158] The first determining subunit is used to determine the candidate video segment as the second video segment in response to the quality assessment value meeting the preset quality conditions.
[0159] According to embodiments of this disclosure, the quality assessment value is obtained based on quality assessment items, which include at least one of the following: quality assessment items related to an object, quality assessment items related to an item, and quality assessment items related to a scene.
[0160] According to embodiments of this disclosure, the video generation apparatus 800 based on a large model may further include a display module, an update module, and a second processing module.
[0161] The display module is used to display at least two reference video frames and descriptive text in the interactive interface.
[0162] The updating module is configured to update the description text or the reference video frame to obtain updated description text or updated reference video frame in response to detecting an editing operation on at least one of the reference video frame and the description text via the interaction interface.
[0163] The second processing module is configured to process the updated description text or the updated reference video frame by using the video generation large model to obtain a second video clip.
[0164] According to an embodiment of the present disclosure, the reference video includes at least two candidate objects; and the target object is determined in at least one of the following manners: the target object is determined from the at least two candidate objects according to audio data corresponding to the reference video and a voiceprint feature of each of the at least two candidate objects; or the target object is determined from the at least two candidate objects according to a participation proportion of each of the at least two candidate objects in the reference content.
[0165] According to an embodiment of the present disclosure, for each of the at least two candidate objects other than the target object, the content understanding module includes a content understanding unit.
[0166] The content understanding unit is configured to guide the visual language large model to perform content understanding on the at least two reference video frames and constrain a change process of at least one region of the other candidate object according to the reference content, to obtain the description text.
[0167] According to an embodiment of the present disclosure, the reference change process includes at least one of the following: an action change process of the target object, an article change process with the target object, a foreground change process of a picture, and a background change process of the picture.
[0168] According to an embodiment of the present disclosure, the reference content includes at least one of reference text and first reference audio, and the first video clip is multiple; the time alignment module 810 includes a time alignment unit and a determination unit.
[0169] The time alignment unit is configured to time-align the first reference audio or second reference audio generated based on the reference text with the first video clip to determine a time sequence relationship of the multiple first video clips.
[0170] The determination unit is configured to determine at least two reference video frames in each two adjacent first video clips according to the time sequence relationship.
[0171] According to an embodiment of the present disclosure, for each two adjacent first video clips, the determination unit includes a second determination subunit and a third determination subunit.
[0172] The second determination subunit is configured to determine a tail video frame of a preceding video clip and a head video frame of a following video clip as the reference video frame.
[0173] The third determination sub-unit is configured to determine a first video frame with the largest data amount or the highest image quality in the preceding video segment and a second video frame with the largest data amount or the highest image quality in the following video segment as the reference video frame.
[0174] According to an embodiment of the present disclosure, the region includes at least one of a body region of the target object, a lip region of the target object, and an article display region; and the video generation module 820 can include a splicing sub-module and a driving sub-module.
[0175] The splicing sub-module is configured to splice the first video segment and the second video segment according to a time sequence relationship to obtain an intermediate video, wherein a change process of the body region and the article display region of the target object in the intermediate video matches the reference change process.
[0176] The driving sub-module is configured to drive the lip region of the target object in the intermediate video according to the reference content by using a lip driving model, so that a change process of the lip region matches the reference change process, to obtain the target video.
[0177] According to an embodiment of the present disclosure, the splicing sub-module can include an insertion unit and a splicing unit.
[0178] The insertion unit is configured to, for every two adjacent first video segments, insert a second video segment generated based on a reference video frame into the two adjacent first video segments to obtain an intermediate video segment.
[0179] The splicing unit is configured to splice the intermediate video segments according to a time sequence relationship to obtain the intermediate video.
[0180] According to an embodiment of the present disclosure, the large model-based video generation apparatus 800 can further include a scene semantic analysis module and an extraction module.
[0181] The scene semantic analysis module is configured to perform scene semantic analysis on the reference content by using a visual language large model to obtain scene semantic information used to describe the reference change process.
[0182] The extraction module is configured to extract a segment matching the scene semantic information from the reference video according to the scene semantic information to obtain the first video segment.
[0183] According to an embodiment of the present disclosure, the method is applied to a live streaming scenario, and the target object includes a digital human.
[0184] Figure 9 An illustrative structural block diagram of an agent of a large model according to an embodiment of the present disclosure is shown.
[0185] In embodiments of the present disclosure, inspired by the Von Neumann architecture in modern computer theory, as shown in FIG. 9, the AI agent 900 can include five core modules: an input module 910, a processing module 920, and an output module 930. Figure 9
[0186] The input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or external environments), and converting them into formats that the AI agent 900 can understand and process. The input module 910 is the first step in the interaction between the AI agent 900 and the outside world, enabling the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to it.
[0187] In examples, the input module 910 can input the reference content and the reference video described above.
[0188] In embodiments of the present disclosure, the processing module 920 can include a control module 921, a storage module 922, and a computation module 923. The processing module 920 is configured to determine a target task based on the input information received by the input module 910, determine a target large model based on the target task, execute a large model-based video generation method by calling the target large model, and obtain output information. In examples, the target large model includes at least one of a visual language large model, a video generation large model, and a video evaluation large model.
[0189] The control module 921 is the core support for the AI agent 900's ability to handle complex tasks. The control module 921 can execute the large model-based video generation method described above.
[0190] In examples, the control module 921 will constantly interact with the storage module 922, the computation module 923, and / or the output module 930 during operation. However, it is noted that in embodiments of the present disclosure, the control module 921 initiates communication with the storage module 922, the computation module 923, and / or the output module 930 as a single initiator, while there is no communication coupling between the storage module 922, the computation module 923, and / or the output module 930.
[0191] In examples, the performance of the control module 921 can be closely related to the large model on which the AI agent 900 is based. In order to fully utilize the capabilities of the large language model, the internal structure of the control module 921 can be designed to be highly configurable and extensible in order to cope with various types of tasks and demands in real-world scenarios.
[0192] The storage module 922 can be responsible for memorizing the visual language large model, the video generation large model, and the video evaluation large model. The foregoing visual language large model, video generation large model, and video evaluation large model can be included in the storage module 922.
[0193] In an example, after receiving the reference video and the reference content, the AI agent 900 can trigger a video generation process to obtain a target video and feed it back to the control module 921. Then, the control module 921 can pass the target video fed back to the output module 930.
[0194] The operation module 923 can be regarded as a pre-defined tool library. The tool for performing the foregoing time sequence alignment can be included in the operation module 923.
[0195] In an example, when the AI agent 900 needs to process data, the relevant tool can be called from the operation module 923 and fed back to the control module 921. Then, the control module 921 can process the relevant data by using the tool fed back. It can be understood that, although the large language model has excellent language understanding and generation capabilities, it, like a human, can solve only a limited number of tasks without the aid of any tool. When the AI agent 900 is endowed with the ability to call tools, it can implement a time sequence alignment task by using the tool for performing time sequence alignment.
[0196] The output module 930 can output the target video described above.
[0197] The AI agent 900 according to the embodiments of the present disclosure can simply and effectively improve the intelligent degree and improve the flexibility and versatility.
[0198] Figure 10 A block diagram schematically showing an electronic device suitable for implementing the large model-based video generation method according to the embodiments of the present disclosure is shown. The electronic device is intended to represent a variety of forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent a variety of forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0199] As Figure 10As shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0200] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0201] The computing unit 1001 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the large model based video generation method. For example, in some embodiments, the large model based video generation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the large model based video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the large model based video generation method by any other appropriate means, e.g., by means of firmware.
[0202] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0203] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0204] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0205] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0206] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0207] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0208] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, so long as the desired results of the technology of the present disclosure are achieved.
[0209] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above.
Claims
1. A video generation method based on a large model, comprising: At least two reference video frames are determined by temporally aligning reference content with a first video clip, wherein the first video clip is extracted from a reference video based on the reference content, the reference content indicating a reference change process related to a target object in the reference video; and The first video segment and the second video segment are spliced together to obtain the target video. The second video segment is obtained by processing the reference content and at least two reference video frames using a large video generation model. The change process of at least one region of the target object in the target video matches the reference change process.
2. The method according to claim 1, further comprising: Using the reference content, a large visual language model is guided to perform content understanding on at least two of the reference video frames to obtain descriptive text, wherein the descriptive text characterizes the intermediate changes between the reference video frames; and The second video segment is obtained by processing the descriptive text and at least two reference video frames using the video generation large model.
3. The method according to claim 2, wherein, The process of using the video generation large model to process the descriptive text and at least two reference video frames to obtain the second video segment includes: The video generation model is used to process the descriptive text and at least two reference video frames to obtain multiple candidate video segments; and The quality of the multiple candidate video segments is evaluated using a large video evaluation model to obtain the second video segment.
4. The method according to claim 3, wherein, The process of using a large video evaluation model to assess the quality of the multiple candidate video segments to obtain the second video segment includes: The candidate video segments are evaluated using the aforementioned video evaluation model to obtain quality evaluation values; and In response to the quality assessment value meeting the preset quality conditions, the candidate video segment is determined as the second video segment.
5. The method according to claim 4, wherein, The quality assessment value is obtained based on quality assessment items, which include at least one of the following: quality assessment items related to the object, quality assessment items related to the item, and quality assessment items related to the scene.
6. The method according to claim 2, further comprising: Display at least two of the reference video frames and the description text on the interactive interface; In response to detecting an editing operation on at least one of the reference video frame and the description text via the interactive interface, update the description text or the reference video frame to obtain the updated description text or the updated reference video frame; as well as The updated descriptive text or the updated reference video frame is processed using the video generation large model to obtain the second video segment.
7. The method according to claim 2, wherein, The reference video includes at least two candidate objects; The target object is determined using at least one of the following methods: The target object is determined from at least two candidate objects based on the audio data corresponding to the reference video and the voiceprint features of each of the at least two candidate objects. or The target object is determined from at least two candidate objects based on the participation percentage of each of the at least two candidate objects in the reference content.
8. The method according to claim 7, wherein, For at least two of the candidate objects other than the target object, The step of using the reference content to guide the visual language big data model to perform content understanding on at least two of the reference video frames to obtain descriptive text includes: Based on the reference content, the visual language big model is guided to perform content understanding on at least two of the reference video frames and to constrain the change process of at least one region of the other candidate objects to obtain the descriptive text.
9. The method according to claim 2, wherein, The reference change process includes at least one of the following: the action change process of the target object, the object change process of the target object, the foreground change process of the image, and the background change process of the image.
10. The method according to any one of claims 1 to 9, wherein, The reference content includes at least one of reference text and first reference audio, and the first video segment is multiple; The step of determining at least two reference video frames by temporally aligning the reference content with the first video segment includes: The first reference speech or the second reference audio generated based on the reference text is time-aligned with the first video segment to determine the temporal relationship of multiple first video segments; as well as Based on the temporal relationship, at least two reference video frames are determined in every two adjacent first video segments.
11. The method according to claim 10, wherein, For every two adjacent first video segments, determining the at least two reference video frames in each pair of adjacent first video segments includes at least one of the following methods: The reference video frames are determined from the last video frame of the preceding video segment and the first video frame of the following video segment; or The first video frame with the largest data volume or the highest image quality in the preceding video segment and the second video frame with the largest data volume or the highest image quality in the following video segment are determined as the reference video frames.
12. The method according to claim 10, wherein, The area includes at least one of the main body area of the target object, the lip area of the target object, and the item display area; The step of splicing the obtained second video segment and the first video segment to generate the target video includes: Based on the temporal relationship, the first video segment and the second video segment are spliced together to obtain an intermediate video, wherein the change process of the main area of the target object and the item display area in the intermediate video matches the reference change process; and Using a lip-driven model, the lip region of the target object in the intermediate video is driven according to the reference content, so that the change process of the lip region matches the change process of the reference, thereby obtaining the target video.
13. The method according to claim 12, wherein, The step of splicing the first video segment and the second video segment according to the temporal relationship to obtain an intermediate video includes: For every two adjacent first video segments, a second video segment generated based on the reference video frame is inserted into the two adjacent first video segments to obtain an intermediate video segment; and Based on the aforementioned temporal relationship, the intermediate video segments are spliced together to obtain the intermediate video.
14. The method according to claim 1, further comprising: The reference content is subjected to scene semantic analysis using a large visual language model to obtain scene semantic information describing the change process of the reference. as well as Based on the scene semantic information, a segment matching the scene semantic information is extracted from the reference video to obtain the first video segment.
15. The method according to claim 1, wherein, The method is applied to live streaming scenarios, and the target object includes digital humans.
16. A video generation apparatus based on a large model, comprising: A timing alignment module is used to determine at least two reference video frames by timing-aligning reference content with a first video segment, wherein the first video segment is extracted from a reference video based on the reference content, the reference content indicating a reference change process related to a target object in the reference video; and A video generation module is used to splice the first video segment and the second video segment to obtain a target video, wherein the second video segment is obtained by processing the reference content and at least two reference video frames using a large video generation model, and the change process of at least one region of the target object in the target video matches the reference change process.
17. An intelligent agent based on a large model, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a target large model based on the target task, and execute the method described in claims 1 to 15 by calling the target large model to obtain output information. The target large model includes at least one of a visual language large model, a video generation large model, and a video evaluation large model. as well as An output module is used to output the output information obtained by the processing module.
18. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 15.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 15.
20. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 15.
Citation Information
Cited By
Industrial defect data generation method based on visual language large model
CN121686146A
An industrial defect data generation method based on a visual language large model
CN121686146B