A method and apparatus for generating digital video based on key frames
Through the video generation method based on keyframes, the video frame deviation comparison is used to segment the video, key action information and frame groups are determined, randomly sorted and intermediate frames are generated, which solves the problem of serious traces of digital video synthesis and achieves a more natural and smooth video generation effect.
Patent Information
- Application Number
- CN202510238050.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the existing digital video generation methods, video synthesis traces are serious, and optimization is needed to improve the smoothness of the video and reduce the splicing traces.
By obtaining the deviation comparison data between video frames of the video to be processed, the video is divided into video material segments, the key action information is determined and the keyframe group is matched, the video material segments are randomly sorted and the target keyframes are inserted, and the preset video generation algorithm is used to generate intermediate frames to connect the material segments, and the target video is generated.
It improves the smoothness of connection between video material segments, weakens the clipping traces of videos, and improves the naturalness and quality of generated videos.
Smart Images

Figure CN119767106B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium and program product for generating digital videos based on key frames. Background Art
[0002] Digital customer service is a virtual human using artificial intelligence technology, aiming to imitate real customer service representatives and interact with users in various ways such as text, voice or image. They can understand users' questions, provide answers, and handle affairs, thereby improving customer service efficiency and user experience. Digital human technology combines multiple fields such as computer graphics, artificial intelligence, and animation technology, aiming to create virtual characters that are almost indistinguishable from real humans. Digital human videos refer to video content containing digital human characters, which may be completely virtual or generated through motion capture and expression capture technologies of human actors. Digital human videos usually have very high visual quality and realism, and can present vivid character images and action performances. The application fields of digital human video technology are very extensive, including but not limited to movie and TV drama production, game development, virtual reality and augmented reality, advertising and marketing. Generally speaking, digital human videos refer to video content containing digital human characters generated by digital human technology, with high realism and visual effects, and are widely used in fields such as movies, games, and virtual reality.
[0003] In related technologies, digital customer service virtual humans usually have the following characteristics: intelligent interaction: using technologies such as natural language processing (NLP) and machine learning to understand users' intentions and make corresponding responses; multi-channel support: can provide services on multiple channels such as websites, APPs, and social media; 7x24 hours online: not restricted by time and can respond to users' needs at any time; personalized service: can provide customized services according to user portraits; cost-effectiveness: reduce labor costs and improve service efficiency. All in all, digital customer service virtual humans are an important tool for enterprises to optimize customer service and improve operation efficiency. In the synthesis of digital videos based on digital humans, the steps usually include: background production: creating or selecting the background of the video, which can be a static picture, a dynamic video or a 3D scene; synthesis: synthesizing the digital human model with the background. Video editing software (such as Adobe After Effects, Premiere Pro, etc.) can be used for synthesis; editing: performing operations such as video clipping, color correction, and adding special effects on the video to make it more attractive.
[0004] However, the current processing methods for digital video generation have the following technical problems:
[0005] The video synthesis traces obtained through manual processing are serious and need to be optimized. Summary of the Invention
[0006] Based on this, in view of the above technical problems, it is necessary to provide a key-frame-based digital video generation method, device, computer device, computer-readable storage medium, and computer program product that can improve the smoothness of digital videos and weaken the splicing traces of videos.
[0007] In a first aspect, the present application provides a key-frame-based digital video generation method. The method includes:
[0008] Obtain a video to be processed, determine a segmentation point according to the deviation comparison data between video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation point;
[0009] Respond to the application requirements of the target video, and determine a number of key action information, where the key action information is used to characterize specific digital human actions;
[0010] Based on the key action information, determine a key frame group that matches each key action information in the video material segment, where the key frame group includes a number of key frames;
[0011] Based on the application requirements, randomly sort the video material segments, and insert target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, where the target key frames are randomly selected from the corresponding key frame groups;
[0012] Generate intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video.
[0013] In one embodiment, the determining, based on the key action information, a key frame group that matches each key action information in the video material segment, where the key frame group includes a number of key frames includes:
[0014] Obtain a reference action sequence corresponding to the key action information, where the reference action sequence includes a number of reference frames, and the reference frames include a simplified digital human joint model making a specific action;
[0015] Set a number of sampling points in the simplified digital human joint model, and determine the key frame group based on the deviation information of the sampling points between the reference frames and the video frames.
[0016] In one embodiment, the obtaining a reference action sequence corresponding to the key action information, where the reference action sequence includes a number of reference frames, and the reference frames include a simplified digital human joint model making a specific action includes:
[0017] Determine the key action classification based on the key action information, where the key action classification includes limb action types and / or facial action types;
[0018] Determine the model architecture of the digital human joint simplified model based on the key action classification, where the model architecture includes a torso joint model and a facial capture model, and the digital human joint simplified model includes a single architecture or a combination of multiple architectures.
[0019] In one embodiment, generating the intermediate frames between the video material segment and the target key frame based on a preset video generation algorithm to obtain the target video includes:
[0020] Determine a group of transition video frames at the head and tail of the video material segment to be connected based on a preset sampling window, where the number of frames in the group of transition video frames is greater than 1, and the duration of the group of transition video frames is not less than a preset ratio of the total duration of the video material segment to which it belongs;
[0021] Generate a group of intermediate frames based on the group of transition video frames and the target key frame.
[0022] In one embodiment, generating the intermediate frames between the video material segment and the target key frame based on a preset video generation algorithm to obtain the target video includes:
[0023] The group of intermediate frames includes a number of the intermediate frames to be inserted. Insert the intermediate frames into the group of transition video frames and the target key frame so that the video material segment is connected to the target key frame, and the distribution pattern of the intermediate frames is sparse on both sides and dense in the middle.
[0024] In one embodiment, after generating the intermediate frames between the video material segment and the target key frame based on a preset video generation algorithm to obtain the target video, it further includes:
[0025] Generate control accessories associated with the target video, and adjust the target video based on the control accessories. The control accessories include easing functions and / or spline curves, and the spline curve includes a sequence of control points having a mapping relationship with the video frame group in the target video.
[0026] In a second aspect, the present application also provides a digital video generation device based on key frames. The device includes:
[0027] A video segmentation module, configured to obtain a video to be processed, determine a segmentation point according to the deviation comparison data between the video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation point;
[0028] A demand data module, configured to determine a plurality of key action information in response to the application demand of a target video, where the key action information is used to characterize specific digital human actions;
[0029] A key frame module, configured to determine a key frame group matching each of the key action information in the video material segment based on the key action information, where the key frame group includes a plurality of key frames;
[0030] A material sequence module, configured to randomly sort the video material segment based on the application demand, and insert target key frames into the sequence of the sorted video material segment to obtain a material sequence to be combined, where the target key frame is randomly selected from the corresponding key frame group;
[0031] A video generation module, configured to generate intermediate frames between the video material segment and the target key frame based on a preset video generation algorithm to obtain the target video.
[0032] In one embodiment, the key frame module includes:
[0033] A simplified model module, configured to obtain a reference action sequence corresponding to the key action information, where the reference action sequence includes a plurality of reference frames, and the reference frames include a simplified model of a digital human joint making a specific action;
[0034] A sampling module, configured to set a plurality of sampling points in the simplified model of the digital human joint, and determine the key frame group based on the deviation information of the sampling points between the reference frame and the video frame.
[0035] In one embodiment, the simplified model module includes:
[0036] An action classification module, configured to determine a key action classification based on the key action information, where the key action classification includes a limb action type and / or a facial action type;
[0037] A model architecture module, configured to determine the model architecture of the simplified model of the digital human joint based on the key action classification, where the model architecture includes a torso joint model and a facial capture model, and the simplified model of the digital human joint includes a single architecture or a combination of multiple architectures.
[0038] In one embodiment, the video generation module includes:
[0039] A transition frame group module, configured to determine a transition video frame group at the head and tail of the video material segment to be connected based on a preset sampling window, where the number of frames in the transition video frame group is greater than 1, and the duration of the transition video frame group is not less than a preset ratio of the total duration of the video material segment to which it belongs;
[0040] An intermediate frame group module, configured to generate an intermediate frame group based on the transition video frame group and the target key frame.
[0041] In one embodiment, the video generation module includes:
[0042] An interpolation frame module, where the intermediate frame group includes several intermediate frames to be inserted, and the intermediate frames are inserted into the transition video frame group and the target key frame to connect the video material segment with the target key frame. The distribution form of the intermediate frames is sparse on both sides and dense in the middle.
[0043] In one embodiment, after the video generation module, there is further included:
[0044] A control accessory module, configured to generate control accessories associated with the target video, and adjust the target video based on the control accessories. The control accessories include easing functions and / or spline curves, and the spline curve includes a control point sequence having a mapping relationship with the video frame group in the target video.
[0045] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in a method for generating a digital video based on key frames as described in any one of the embodiments in the first aspect are implemented.
[0046] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in a method for generating a digital video based on key frames as described in any one of the embodiments in the first aspect are implemented.
[0047] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in a method for generating a digital video based on key frames as described in any one of the embodiments in the first aspect are implemented.
[0048] The above method, device, computer device, storage medium, and computer program product for generating a digital video based on key frames, deduced through the technical features in the claims, can achieve the following beneficial effects for the technical problems in the corresponding background technology:
[0049] A method for generating a digital video based on key frames provided by this application includes: obtaining a video to be processed, determining a segmentation point according to the deviation comparison data between video frames of the video to be processed, and segmenting the video to be processed into video material segments based on the segmentation point; in response to the application requirements of the target video, determining a number of key action information, where the key action information is used to represent specific digital human actions; determining a key frame group matching each key action information in the video material segments based on the key action information, and the key frame group includes a number of key frames; randomly sorting the video material segments based on the application requirements, and inserting target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, where the target key frames are randomly selected from the corresponding key frame groups; generating intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video. In implementation, after obtaining the video to be processed, at least one cutting point is obtained through deviation comparison data, so as to cut the long video to be processed into multiple video material segments, which can make the similarity at the joints between the cut material video segments relatively high. Finally, by generating transition materials between the material video segments and realizing the splicing of the material video segments through the intermediate frame group, it helps to further improve the smoothness of the connection between the material video segments, weaken the editing traces of the target video, and improve the quality of the produced digital video. On the other hand, since key frames representing specific actions are set, if the same key frames are repeatedly repeated, the repetition degree of specific actions will be relatively high and the artificial traces will be obvious. This application sets a key frame group including a number of key frames, and improves the randomness of key frame setting through random selection. Finally, it can weaken the repetition of key frames in the target video, improve the naturalness of the generated target video, and weaken the splicing and synthesis traces. Description of the Drawings
[0050] Figure 1 It is the first process schematic diagram of a method for generating a digital video based on key frames in an embodiment;
[0051] Figure 2 It is a schematic diagram of sampling points in an embodiment;
[0052] Figure 3 It is the structural block diagram of a video generation device in an embodiment;
[0053] Figure 4 It is the internal structure diagram of a computer device in an embodiment. Detailed Embodiment
[0054] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not used to limit this application.
[0055] In related technologies, digital customer service virtual humans usually have the following characteristics: intelligent interaction: using technologies such as natural language processing (NLP) and machine learning to understand user intentions and make corresponding responses; multi-channel support: can provide services on multiple channels such as websites, APPs, and social media; 7x24-hour online: not restricted by time and can respond to user needs at any time; personalized service: can provide customized services according to user portraits; cost-effectiveness: reduce labor costs and improve service efficiency. All in all, digital customer service virtual humans are an important tool for enterprises to optimize customer service and improve operation efficiency. In the synthesis of digital videos based on digital humans, the steps usually include: background production: creating or selecting the background of the video, which can be a static picture, a dynamic video, or a 3D scene; synthesis: synthesizing the digital human model with the background. Video editing software (such as Adobe After Effects, Premiere Pro, etc.) can be used for synthesis; editing: performing operations such as video clipping, color correction, and adding special effects on the video to make it more attractive.
[0056] However, the current processing methods for digital video generation have the following technical problems:
[0057] The video synthesis traces obtained through manual processing are serious and need to be optimized.
[0058] Based on this, this application provides a key-frame-based digital video generation method. In one embodiment, as Figure 1 shown, a key-frame-based digital video generation method is provided. In this embodiment, this method is illustrated by taking its application to a terminal as an example. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0059] Step 102: Obtain the video to be processed, determine the segmentation points according to the deviation comparison data between the video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation points.
[0060] Among them, a video frame is the basic unit in video data processing, representing a static image of a video frame at a specific time point. Video frames are arranged in sequence in a video file or video stream to present a continuous dynamic picture.
[0061] Among them, the cutting point can refer to the selected segmentation position for segmenting the video to be processed. The segmentation point can be determined based on the video time or the video frame. In the process of fine segmentation, the segmentation point is usually determined based on the video frame. For example, if the 30th frame of the video to be processed is used as the segmentation point for segmentation, the first 30 frames are segmented into the front video segment, and the 31st frame and onwards are segmented into the back video segment.
[0062] Among them, the deviation comparison result can refer to the parameterized quantization result obtained by comparing the deviation degree between two video frames.
[0063] Step 104: In response to the application requirements of the target video, determine a number of key action information, where the key action information is used to characterize specific digital human actions;
[0064] Step 106: Based on the key action information, determine a key frame group that matches each key action information in the video material segment, where the key frame group includes a number of key frames;
[0065] Step 108: Based on the application requirements, randomly sort the video material segment, and insert target key frames into the sequence of the sorted video material segment to obtain a material sequence to be combined, where the target key frame is randomly selected from the corresponding key frame group;
[0066] Step 1010: Generate intermediate frames between the video material segment and the target key frame based on a preset video generation algorithm to obtain the target video.
[0067] In the above digital video generation method based on key frames, through reasonable derivation in combination with the technical features in the embodiments, the following beneficial effects can be achieved to solve the technical problems proposed in the background technology:
[0068] A method for generating a digital video based on key frames provided by the present application includes: obtaining a video to be processed, determining a segmentation point according to the deviation comparison data between video frames of the video to be processed, and segmenting the video to be processed into video material segments based on the segmentation point; in response to the application requirements of the target video, determining a plurality of key action information, where the key action information is used to represent specific digital human actions; determining a key frame group matching each key action information in the video material segments based on the key action information, where the key frame group includes a plurality of key frames; randomly sorting the video material segments based on the application requirements, and inserting target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, where the target key frames are randomly selected from the corresponding key frame groups; generating intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video. In implementation, after obtaining the video to be processed, at least one cutting point is obtained through deviation comparison data, so as to cut the long video to be processed into multiple video material segments, which can make the similarity at the connection points between the cut material video segments relatively high. Finally, by generating an intermediate frame group between the material video segments, it helps to further improve the smoothness of the connection between the material video segments, weaken the editing traces of the target video, and improve the quality of the produced digital video. On the other hand, since key frames representing specific actions are set, if the same key frames are repeatedly repeated, the repetition degree of specific actions will be relatively high and the artificial traces will be obvious. The present application sets key frame groups including a plurality of key frames, and improves the randomness of key frame setting through random selection, and finally can weaken the repetition of key frames of the target video, improve the naturalness of the generated target video, and weaken the splicing and synthesis traces.
[0069] In one embodiment, as Figure 2 shown, step 108 includes:
[0070] Step 202: Obtain a reference action sequence corresponding to the key action information, where the reference action sequence includes a plurality of reference frames, and the reference frames include a simplified digital human joint model making a specific action;
[0071] Step 204: Set a plurality of sampling points in the simplified digital human joint model, and determine the key frame group based on the deviation information of the sampling points between the reference frames and the video frames.
[0072] In this embodiment, by setting a simplified joint model, it helps to simplify the description difficulty of key actions, and at the same time improve the recognition efficiency and accuracy of key frames.
[0073] In one embodiment, step 202 includes:
[0074] Step 402: Determine the key action classification based on the key action information, where the key action classification includes limb action types and / or facial action types;
[0075] Step 404: Determine the model architecture of the simplified digital human joint model based on the key action classification. The model architecture includes a torso joint model and a facial capture model, and the simplified digital human joint model includes a single architecture or a combination of multiple architectures.
[0076] In one embodiment, the step 1010 includes:
[0077] Step 502: Determine the transitional video frame group at the head and tail of the video material segment to be connected based on a preset sampling window. The number of frames in the transitional video frame group is greater than 1, and the duration of the transitional video frame group is not less than a preset ratio of the total duration of the video material segment to which it belongs.
[0078] Step 504: Generate an intermediate frame group based on the transitional video frame group and the target key frame.
[0079] In this embodiment, by including multiple frames at the head and tail parts of the video material segment in the processing scope of the connection generation algorithm, it helps to enhance the connection effect of the intermediate frames and improve the smoothness of the obtained synthesized video.
[0080] In one embodiment, the step 1010 includes:
[0081] Step 602: The intermediate frame group includes several intermediate frames to be inserted. Insert the intermediate frames into the transitional video frame group and the target key frame to connect the video material segment with the target key frame. The distribution pattern of the intermediate frames is sparse on both sides and dense in the middle.
[0082] In one embodiment, after the step 1010, it further includes:
[0083] Step 702: Generate control accessories associated with the target video, and adjust the target video based on the control accessories. The control accessories include easing functions and / or spline curves, and the spline curves include a sequence of control points having a mapping relationship with the video frame group in the target video.
[0084] In this embodiment, by setting control accessories, it helps to adjust the target video according to usage requirements and improves the flexibility of video generation.
[0085] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0086] Based on the same inventive concept, an embodiment of the present application further provides a key-frame-based digital video generation device for implementing the key-frame-based digital video generation method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the key-frame-based digital video generation device provided below can refer to the limitations on the key-frame-based digital video generation method in the above text, and will not be repeated here.
[0087] In one embodiment, as Figure 3 shown, a key-frame-based digital video generation device is provided, including: a video segmentation module, a demand data module, a key frame module, a material sequence module, and a video generation module, where:
[0088] The video segmentation module is used to obtain the video to be processed, determine the segmentation points according to the deviation comparison data between the video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation points;
[0089] The demand data module is used to determine a number of key action information in response to the application requirements of the target video, and the key action information is used to represent specific digital human actions;
[0090] The key frame module is used to determine a key frame group matching each key action information in the video material segments based on the key action information, and the key frame group includes a number of key frames;
[0091] The material sequence module is used to randomly sort the video material segments based on the application requirements, and insert target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, and the target key frames are randomly selected from the corresponding key frame groups;
[0092] A video generation module, configured to generate intermediate frames between the video material segment and the target key frames based on a preset video generation algorithm, so as to obtain the target video.
[0093] In one embodiment, the key frame module includes:
[0094] A simplified model module, configured to obtain a reference action sequence corresponding to the key action information, where the reference action sequence includes a plurality of reference frames, and the reference frames include a simplified model of a digital human joint making a specific action;
[0095] A sampling module, configured to set a plurality of sampling points in the simplified model of the digital human joint, and determine the key frame group based on the deviation information of the sampling points between the reference frames and the video frames.
[0096] In one embodiment, the simplified model module includes:
[0097] An action classification module, configured to determine a key action classification based on the key action information, where the key action classification includes a limb action type and / or a facial action type;
[0098] A model architecture module, configured to determine the model architecture of the simplified model of the digital human joint based on the key action classification, where the model architecture includes a torso joint model and a facial capture model, and the simplified model of the digital human joint includes a single architecture or a combination of multiple architectures.
[0099] In one embodiment, the video generation module includes:
[0100] A transition frame group module, configured to determine a transition video frame group at the head and tail of the video material segment to be connected based on a preset sampling window, where the number of frames of the transition video frame group is greater than 1, and the duration of the transition video frame group is not less than a preset ratio of the total duration of the video material segment to which it belongs;
[0101] An intermediate frame group module, configured to generate an intermediate frame group based on the transition video frame group and the target key frames.
[0102] In one embodiment, the video generation module includes:
[0103] An interpolation frame module, where the intermediate frame group includes a plurality of intermediate frames to be inserted, and the intermediate frames are inserted into the transition video frame group and the target key frames, so that the video material segment is connected to the target key frames, and the distribution form of the intermediate frames is sparse on both sides and dense in the middle.
[0104] In one embodiment, after the video generation module, it further includes:
[0105] A control accessory module is used to generate a control accessory associated with the target video, and adjust the target video based on the control accessory. The control accessory includes an easing function and / or a spline curve, and the spline curve includes a sequence of control points having a mapping relationship with a group of video frames in the target video.
[0106] Each module in the above digital video generation device based on key frames can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0107] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a method for generating a digital video based on key frames. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0108] Those skilled in the art can understand that Figure 4 the structure shown in
[0109] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.
[0110] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0111] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0114] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0115] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A key-frame-based digital video generation method, characterized in that The method includes: Obtain a video to be processed, determine a segmentation point according to the deviation comparison data between video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation point; In response to the application requirements of the target video, determine a number of key action information, where the key action information is used to characterize specific digital human actions; Based on the key action information, determine a key frame group matching each key action information in the video material segments, where the key frame group includes a number of key frames; Based on the application requirements, randomly sort the video material segments, and insert target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, where the target key frame is randomly selected from the corresponding key frame group; Generate intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video.
2. The method for generating a digital video based on key frames according to claim 1, wherein, The determining, based on the key action information, a key frame group matching each key action information in the video material segments, where the key frame group includes a number of key frames includes: Obtain a reference action sequence corresponding to the key action information, where the reference action sequence includes a number of reference frames, and the reference frames include a simplified digital human joint model making a specific action; Set a number of sampling points in the simplified digital human joint model, and determine the key frame group based on the deviation information of the sampling points between the reference frames and the video frames.
3. A method for generating a digital video based on key frames according to claim 1, characterized in that, The obtaining a reference action sequence corresponding to the key action information, where the reference action sequence includes a number of reference frames, and the reference frames include a simplified digital human joint model making a specific action includes: Determine a key action classification based on the key action information, where the key action classification includes limb action types and / or facial action types; Determine the model architecture of the simplified digital human joint model based on the key action classification, where the model architecture includes a torso joint model and a facial capture model, and the simplified digital human joint model includes a single architecture or a combination of multiple architectures.
4. A method for generating a digital video based on key frames according to claim 1, characterized in that, The generating intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video includes: Based on a preset sampling window, determine a transition video frame group at the head and tail of the video material segments to be connected, where the number of frames of the transition video frame group is greater than 1, and the duration of the transition video frame group is not less than a preset ratio of the total duration of the video material segment to which it belongs; Generate an intermediate frame group based on the transition video frame group and the target key frames.
5. A method for generating a digital video based on key frames according to claim 4, characterized in that The generating intermediate frames between the video material segments and the target key frames based on a preset video generation algorithm to obtain the target video includes: The intermediate frame group includes a number of the intermediate frames to be inserted. Insert the intermediate frames into the transition video frame group and the target key frames to connect the video material segments and the target key frames, and the distribution form of the intermediate frames is sparse on both sides and dense in the middle.
6. A method for generating a digital video based on key frames according to claim 1, characterized in that, After generating the intermediate frames between the video material segment and the target key frames based on a preset video generation algorithm to obtain the target video, the following steps are further included: Generate a control accessory associated with the target video, and adjust the target video based on the control accessory. The control accessory includes an easing function and / or a spline curve, and the spline curve includes a control point sequence having a mapping relationship with a video frame group in the target video.
7. A digital video generation device based on key frames, characterized in that, The device includes: A video segmentation module, configured to obtain a video to be processed, determine segmentation points according to deviation comparison data between video frames of the video to be processed, and segment the video to be processed into video material segments based on the segmentation points; A requirement data module, configured to determine a plurality of key action information in response to the application requirements of the target video, where the key action information is used to represent specific digital human actions; A key frame module, configured to determine a key frame group matching each key action information in the video material segment based on the key action information, where the key frame group includes a plurality of key frames; A material sequence module, configured to randomly sort the video material segments based on the application requirements, and insert target key frames into the sequence of the sorted video material segments to obtain a material sequence to be combined, where the target key frames are randomly selected from the corresponding key frame groups; A video generation module, configured to generate intermediate frames between the video material segment and the target key frames based on a preset video generation algorithm to obtain the target video.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Video generation method and device, electronic equipment and storage medium
CN115942039A
Video processing method and device, terminal equipment and storage medium
CN116419010A