A method, apparatus, device and medium for generating intermediate frames

CN115761065BActive Publication Date: 2026-09-01HISENSE VISUAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211430663.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2026-09-01
Estimated Expiration
2042-11-15

AI Technical Summary

Benefits of technology

[0018]本公开实施例提供的技术方案与现有技术相比具有如下优点:首先基于输入的语音信息,确定待生成中间帧的时间信息,并根据时间信息获取与待生成中间帧关联的待处理视频帧,其中,输入的语音信息用于驱动虚拟数字人进行动作,然后将待处理视频帧输入至光流估计网络模型中,得到对应的光流估计结果和融合图,最后基于光流估计结果和融合图,生成对应的中间帧,通过上述过程能够生成中间帧,通过中间帧有利于确保虚拟数字人在状态转换过程中自然过渡,使得虚拟数字人能在语音驱动下连贯地完成相应动作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761065B_ABST
    Figure CN115761065B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, device, and medium for generating intermediate frames, particularly in the fields of computer vision and image processing. The method includes: determining the temporal information of the intermediate frame to be generated based on input speech information, and acquiring a video frame to be processed associated with the intermediate frame based on the temporal information, wherein the input speech information is used to drive a virtual digital human to perform actions; inputting the video frame to be processed into an optical flow estimation network model to obtain corresponding optical flow estimation results and a fusion map; and generating the corresponding intermediate frame based on the optical flow estimation results and the fusion map. The embodiments of this disclosure can generate intermediate frames through the above process. Intermediate frames help ensure a natural transition during state transitions in the virtual digital human, enabling the virtual digital human to coherently complete corresponding actions under speech-driven guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer vision and image processing technology, and in particular to an intermediate frame generation method, apparatus, device and medium. Background Technology

[0002] In the development of virtual digital humans, the virtual digital human action library mainly involves customizing standardized action paradigms based on business scenario needs, such as greeting, gesturing, or nodding. In virtual digital human projects, real-time interaction with the virtual digital human via voice is required, enabling the virtual digital human to switch from a standard state to a state corresponding to the desired action during the interaction. To ensure a smooth transition during state transitions, intermediate frames are needed to assist in stitching together the virtual digital human's action segments; therefore, the generation of intermediate frames is particularly important. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems or at least partially solve the above-mentioned technical problems, this disclosure provides an intermediate frame generation method, apparatus, device and medium. The generated intermediate frames help to ensure that the virtual digital human transitions naturally during state transitions, so that the virtual digital human can complete the corresponding actions smoothly under voice drive.

[0004] To achieve the above objectives, the technical solutions provided by the embodiments of this disclosure are as follows:

[0005] In a first aspect, this disclosure provides an intermediate frame generation method, the method comprising:

[0006] Based on the input voice information, the time information of the intermediate frame to be generated is determined, and the video frame to be processed associated with the intermediate frame to be generated is obtained according to the time information, wherein the input voice information is used to drive the virtual digital human to perform actions.

[0007] The video frames to be processed are input into the optical flow estimation network model to obtain the corresponding optical flow estimation results and fusion map;

[0008] Based on the optical flow estimation results and the fusion map, the corresponding intermediate frame is generated.

[0009] Secondly, this disclosure provides an intermediate frame generation apparatus, the apparatus comprising:

[0010] The acquisition module is used to determine the time information of the intermediate frame to be generated based on the input voice information, and to acquire the video frame to be processed associated with the intermediate frame to be generated according to the time information, wherein the input voice information is used to drive the virtual digital human to perform actions.

[0011] The determination module is used to input the video frame to be processed into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map;

[0012] The generation module is used to generate corresponding intermediate frames based on the optical flow estimation results and the fusion map.

[0013] Thirdly, this disclosure also provides an electronic device, including:

[0014] One or more processors;

[0015] Storage device for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the intermediate frame generation methods described in the embodiments of this disclosure.

[0017] Fourthly, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the intermediate frame generation methods described in the embodiments of this disclosure.

[0018] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: First, based on the input voice information, the time information of the intermediate frame to be generated is determined, and the video frame to be processed associated with the intermediate frame to be generated is obtained according to the time information. The input voice information is used to drive the virtual digital human to perform actions. Then, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. Finally, based on the optical flow estimation result and fusion map, the corresponding intermediate frame is generated. The above process can generate intermediate frames, which helps to ensure that the virtual digital human transitions naturally during state transitions, so that the virtual digital human can complete the corresponding actions smoothly under voice drive. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1A A flowchart illustrating an intermediate frame generation method provided in this embodiment of the disclosure;

[0022] Figure 1B A schematic diagram illustrating the principle of an intermediate frame generation method provided in this embodiment of the disclosure;

[0023] Figure 2A A flowchart illustrating another intermediate frame generation method provided in this embodiment of the disclosure;

[0024] Figure 2B A schematic diagram illustrating the principle of another intermediate frame generation method provided in this embodiment of the disclosure;

[0025] Figure 2C A schematic diagram illustrating an intermediate frame insertion process provided in an embodiment of this disclosure;

[0026] Figure 3A A flowchart illustrating yet another intermediate frame generation method provided in this disclosure embodiment;

[0027] Figure 3B This is a schematic diagram of the structure of an optical flow estimation network model provided in an embodiment of the present disclosure;

[0028] Figure 3C This is a schematic diagram of the structure of the computational unit in the optical flow estimation network model provided in the embodiments of this disclosure;

[0029] Figure 3D This is a schematic diagram of another computing unit provided in an embodiment of the present disclosure;

[0030] Figure 3E A schematic diagram illustrating the computational principle of a certain computing unit provided in an embodiment of this disclosure;

[0031] Figure 4A A schematic diagram illustrating the principle of determining image fusion results according to an embodiment of this disclosure;

[0032] Figure 4B This is a schematic diagram of the structure of a semantic segmentation network model provided in an embodiment of the present disclosure;

[0033] Figure 4C A schematic diagram illustrating the principle of obtaining image fusion results through a semantic segmentation network model, as provided in this embodiment of the disclosure;

[0034] Figure 5 This is a schematic diagram of the structure of an intermediate frame generation apparatus provided in an embodiment of the present disclosure;

[0035] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0036] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0037] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0038] It should be noted that the brief descriptions of terms in this disclosure are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0039] It should be noted that in this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. For example, a product or apparatus that comprises a list of components is not necessarily limited to all components expressly listed, but may include other components not expressly listed or inherent to such product or apparatus.

[0040] With the maturation of related technologies, virtual digital humans have been applied in many scenarios, providing a variety of services such as business knowledge introduction, information inquiry, and intelligent customer service, thereby enhancing the user experience.

[0041] During the development of virtual digital humans, the creation of the virtual digital human action library requires ensuring smooth transitions between action segments. The action library design module mainly customizes standardized action paradigms based on the needs of different business scenarios, such as greeting, pointing, or nodding. In virtual digital human projects, real-time interaction with the virtual digital human is typically achieved through voice, prompting the virtual digital human to perform corresponding actions. Therefore, during the interaction, the virtual digital human needs to smoothly transition from a standard state to a state performing the corresponding action. This transition must be natural, allowing the digital human to continuously perform actions driven by voice, avoiding stuttering or prolonged pauses.

[0042] In the process of driving virtual digital humans, in addition to driving the virtual digital human's mouth shape, it is also necessary to drive the virtual digital human's body movements. Usually, an action library is created in advance, and standardized action paradigms are customized for different business scenarios. The static form of the virtual digital human itself and its form when performing actions can be achieved by simulating pre-recorded videos of human models. However, how to reasonably arrange and splice these videos is a problem that needs to be solved.

[0043] The existing video stitching approach involves triggering associated action videos during the playback of static videos of a virtual digital human, based on the fulfillment of interactive conditions, causing the virtual digital human to perform responsive actions. However, because static videos contain multiple frames, the uncertainty of the action trigger timing prevents continuous recording of static and action videos. Therefore, static and action videos are recorded separately. During this separate recording process, the form of the human model corresponding to the virtual digital human often changes, resulting in differences between a frame in the static video to be triggered and the starting frame in the action video, or between the last frame in the action video and the next frame in the static video. Consequently, when driving the virtual digital human, if an action response command arrives, noticeable stuttering occurs during the transition between static and action forms, as well as between different action form segments, impacting the overall continuity of the virtual digital human's form.

[0044] Based on the above issues, intermediate frames play an extremely important role in the video editing process of virtual digital humans, as they can assist in splicing together action segments of virtual digital humans.

[0045] Video frame interpolation involves adding several intermediate frames between every two frames of the original video, shortening the display time between each frame. This solves problems such as stuttering and jitter that occur when virtual digital humans switch between static and dynamic states, as well as between different dynamic state segments.

[0046] Video frame interpolation algorithms are generally classified into three types: phase-based methods, adaptive convolutional kernel-based methods, and optical flow-based methods. Phase-based methods rely on optimized algorithms to predict the spatial phase decomposition of intermediate frames using an encoder-decoder network, and then reconstruct them. This method performs poorly in scenes with large motion or complex occlusion. Adaptive convolutional kernel-based methods first perceive scene motion, and then synthesize new perspectives—the intermediate frames to be inserted—based on this motion. This method unifies motion perception and intermediate frame generation into a series of convolutional operations, using adaptive convolution and deformable convolution to perceive large motion and directly synthesizing the texture of the intermediate frames. Optical flow-based methods typically include an optical flow estimation (motion estimation) stage and an intermediate frame synthesis stage. This method offers good frame interpolation results and has a wide range of applications.

[0047] Therefore, this disclosure provides an intermediate frame generation method, which uses an optical flow estimation network model to predict the video frame to be processed, and generates a corresponding intermediate frame based on the predicted optical flow estimation result and fusion map. The intermediate frame generated by the above process helps to ensure that the virtual digital human transitions naturally during state transitions, so that the virtual digital human can complete the corresponding actions smoothly under voice drive.

[0048] To illustrate the intermediate frame generation scheme in more detail, the following will use an illustrative approach. Figure 1A To explain, it is understandable that Figure 1A The steps involved may include more or fewer steps in actual implementation, and the order of these steps may also be different, depending on whether the intermediate frame generation method provided in the embodiments of this application can be implemented.

[0049] Figure 1A This is a flowchart illustrating an intermediate frame generation method provided in an embodiment of the present disclosure. Figure 1B This is a schematic diagram illustrating the principle of an intermediate frame generation method provided in this embodiment. This embodiment is applicable to generating intermediate frames during the splicing of static video frames and motion video. The method of this embodiment can be executed by an intermediate frame generation device, which can be implemented in hardware / software and configured in an electronic device.

[0050] like Figure 1A As shown, the method specifically includes the following steps:

[0051] S110: Based on the input voice information, determine the time information of the intermediate frame to be generated, and obtain the video frame to be processed associated with the intermediate frame to be generated according to the time information.

[0052] The input voice information is used to drive the virtual digital human to perform actions. The voice information can be user-inputted or pre-recorded; this embodiment does not limit this. The voice information includes keyword information, which can be understood as information that drives the virtual digital human to perform actions, such as "Hello," "Thank you," etc. The intermediate frames to be generated can be understood as transition frames that need to be generated to ensure the continuity between the static video frames and the motion video frames of the virtual digital human. The video frames to be processed can be understood as the video frames at the moment before and after the intermediate frames to be generated during the video stitching process.

[0053] Based on the input voice information, by identifying the keywords contained in the voice information, the virtual digital human can be driven to perform actions corresponding to the keywords. Therefore, after identifying the keywords from the input voice information, the time information corresponding to the keywords is the time information for generating intermediate frames. Based on this time information, through static video frames and the action video corresponding to the keywords, the video frames to be processed associated with the intermediate frames to be generated can be obtained.

[0054] S120, input the video frame to be processed into the optical flow estimation network model to obtain the corresponding optical flow estimation results and fusion map.

[0055] In this context, the fusion map can be understood as an additional channel equivalent to optical flow, representing the weighted fusion weights of the video frames after warping transformation. Optical flow, or the movement of light, in computer vision represents the movement of objects in an image. This movement can be caused by camera movement or object movement, specifically referring to the amount of movement of pixels representing the same object in one frame of a video image to the next frame, represented by a two-dimensional vector. The optical flow estimation network in the optical flow estimation network model can be any of the FlowNet series networks, PWC-net networks, MaskFlowNet networks, or LiteFlowNet networks, etc., and this embodiment is not limited to any particular type.

[0056] After obtaining the video frames to be processed, inputting them into the optical flow estimation network model will yield the corresponding optical flow estimation results and fusion map.

[0057] In some embodiments, optionally, the optical flow estimation network model can employ a knowledge distillation strategy during the training phase, training the student model using supervision information from the teacher model and ground truth images. The teacher model is trained with an optical flow estimation network equipped with additional computational units to improve the student model's performance, and is jointly trained using a privileged distillation loss function, which can be expressed as follows:

[0058]

[0059] in, This represents the optical flow results from time t to time i under the Student Model. Let L represent the optical flow result from time t to time i under the Teacher Model, where L represents the loss value and i∈{0,1}.

[0060] In some embodiments, optionally, the loss function of the optical flow estimation network model during the training phase can also be the Connectionist Temporal Classification (CTC) loss function, the multi-class cross-entropy loss function, and the mean square loss function, etc. The specific loss function can be determined according to actual usage requirements, or it can be set by the user. This disclosure does not limit this.

[0061] S130, based on the optical flow estimation results and the fusion map, generates the corresponding intermediate frame.

[0062] After obtaining the optical flow estimation results and the fusion map, the corresponding intermediate frames can be obtained through the corresponding calculations.

[0063] The formula for calculating intermediate frames can be as follows:

[0064]

[0065] in, F represents the intermediate frame at time t. t→0 F represents the optical flow estimation result from time 0 to time t. t→1 denoted as the optical flow estimation result from time t to time 1, and M represents the fusion map.

[0066] It should be noted that: typically, the video frames to be processed input into the optical flow estimation network model are two images, therefore, the optical flow estimation results obtained are also two.

[0067] The intermediate frame generation method provided in this embodiment first determines the time information of the intermediate frame to be generated based on the input voice information, and obtains the video frame to be processed associated with the intermediate frame to be generated according to the time information. The input voice information is used to drive the virtual digital human to perform actions. Then, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. Finally, the corresponding intermediate frame is generated based on the optical flow estimation result and fusion map. The above process can generate intermediate frames, which helps to ensure that the virtual digital human transitions naturally during state transitions, so that the virtual digital human can complete the corresponding actions smoothly under voice drive.

[0068] Figure 2A This is a flowchart illustrating another intermediate frame generation method provided in an embodiment of this disclosure. Figure 2B This is a schematic diagram illustrating the principle of another intermediate frame generation method provided in this embodiment. This embodiment is an optimization based on the above embodiment. Optionally, this embodiment provides a detailed explanation of the process of acquiring the video frames to be processed and obtaining the stitched video. Figure 2A As shown, the method specifically includes the following:

[0069] S210, based on the input voice information, determines the timing information of the intermediate frame to be generated.

[0070] S220, a video frame is determined from the static library based on the time information, and a video frame is determined from the target action video corresponding to the voice information in the action library to obtain the video frame to be processed, wherein the video frame to be processed includes the first video frame and the second video frame.

[0071] The target action video can be understood as the video corresponding to the action that corresponds to the keyword information contained in the voice information.

[0072] Understandably, the insertion position of the intermediate frame is between a static video frame and an action video frame; that is, between a static video frame and the first frame of the target action video, and / or between the last frame of the action video and the next static video frame after the static video frame. Therefore, after obtaining the time information of the intermediate frame to be generated, a keyframe is determined from the static library based on this time information, and a video frame (usually the first frame of the target action video) is determined from the action library corresponding to the voice information. This yields the video frame to be processed. Alternatively, a video frame (usually the last frame of the target action video) is determined from the action library corresponding to the voice information, and the next video frame after the keyframe is retrieved from the static library. This yields the video frame to be processed. In other words, the first video frame can be a keyframe determined from the static library, and the second video frame can be the first frame of the target action video; the first video frame can also be the last frame of the target action video, and the second video frame can also be the next video frame after the keyframe in the static library.

[0073] S230, input the video frame to be processed into the optical flow estimation network model to obtain the corresponding optical flow estimation results and fusion map.

[0074] S240 generates the corresponding intermediate frame based on the optical flow estimation results and the fusion map.

[0075] S250, insert the intermediate frame between the first video frame and the second video frame to obtain the corresponding spliced ​​video.

[0076] After generating intermediate frames, they are inserted between the first and second video frames. In other words, the intermediate frames serve as a connector between the first and second video frames, thus obtaining the corresponding spliced ​​video.

[0077] In this embodiment, firstly, based on the input voice information, the time information of the intermediate frame to be generated is determined. Then, based on the time information, a video frame is determined from the static library, and a video frame is determined from the target action video corresponding to the voice information in the action library, resulting in a video frame to be processed. The video frame to be processed includes a first video frame and a second video frame. Then, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. Based on the optical flow estimation result and fusion map, the corresponding intermediate frame is generated. Finally, the intermediate frame is inserted between the first video frame and the second video frame to obtain the corresponding spliced ​​video. Through the above process, the intermediate frame is first generated by the optical flow estimation network model, and then the first video frame and the second video frame are connected by the intermediate frame. This ensures that the virtual digital human transitions naturally during state transitions, enabling the virtual digital human to complete the corresponding actions smoothly under voice drive. This provides a responsive real-time response for voice-driven virtual digital humans and improves the service quality of virtual digital humans.

[0078] In some embodiments, during the development of virtual digital humans, corresponding motion material libraries can be pre-created for different virtual digital humans. The design principle of motion material is to reduce the repetition frequency of individual dynamic materials and reduce the response time of calling individual dynamic actions, while ensuring the synchronization of the digital human's limb movements and voice. Simultaneously, in the process of voice-driven virtual digital humans, it is necessary to arrange the images to be driven (i.e., stitching) according to the voice information. Based on this, in this embodiment, different material libraries are pre-created based on different settings of the human model (e.g., static, dynamic, etc.), mainly divided into static libraries, motion libraries, and transition libraries. In the static library, the model's image is in a default standard state, remaining static and standing. When there is no voice-driven action, the images in this library are added to the arrangement. In the motion library, the model's image is in a dynamic state, including corresponding actions required by business needs and voice-driven actions. When voice-driven actions occur, the images in this library are added to the arrangement. The transition library consists of pre-generated intermediate frames based on different actions between the static library and the motion library. When voice-driven actions occur, the images in this library are added to the arrangement, thereby ensuring smooth image transitions when the virtual digital human switches states.

[0079] For example, Figure 2C This is a schematic diagram illustrating an intermediate frame insertion process provided in an embodiment of this disclosure. Figure 2C As shown, the static library stores multiple static images. The model in each image is in a default standard state, standing still (not shown in the figure). When keyword information in the voice message is detected, the virtual digital human needs to perform a target action. The target action video corresponding to the target action is stored in the action library. A single target action video may contain multiple frames. To drive the virtual digital human from a static state to perform the target action and to ensure the continuity of this process, a keyframe and its following frame are retrieved from the static library. An action is inserted between the keyframe and its following frame. At least one intermediate frame is inserted between the keyframe and the first frame of the target action video. At least one intermediate frame is inserted between the last frame of the target action video and the frame following the keyframe. Then, the keyframe, intermediate frames, target action video, intermediate frames, and the frame following the keyframe are stitched together to obtain the final stitched video. The intermediate frames can be pre-generated and stored in the stitching library. The number of inserted intermediate frames depends on the specific situation and is not limited in this embodiment.

[0080] It is understandable that, in some embodiments, in order to improve the efficiency of the video arrangement and splicing process, these images can be named in advance according to a pre-set naming rule based on the correspondence between different keyframes, different target action videos, and different intermediate frames. In this way, when driving the virtual digital human to perform corresponding actions, the keyframes, intermediate frames, and target action videos (which are related) can be directly determined according to the naming rule before video arrangement and splicing can be performed, thereby improving efficiency and enabling a faster response.

[0081] Figure 3A This is a flowchart illustrating another intermediate frame generation method provided in this embodiment. This embodiment is an optimization based on the above embodiments. Optionally, this embodiment provides a detailed explanation of the specific structure of the optical flow estimation network model. Figure 3A As shown, the method specifically includes the following:

[0082] S310 determines the time information of the intermediate frame to be generated based on the input voice information, and obtains the video frame to be processed associated with the intermediate frame to be generated according to the time information.

[0083] S320, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. The optical flow estimation network model includes multiple computing units, and adjacent computing units are connected through a residual network. Each computing unit includes a first twist layer, a second twist layer, a splicing layer, at least one convolutional layer and a deconvolutional layer.

[0084] The video frames to be processed include a first video frame and a second video frame. A first distortion layer performs a distortion transformation on the first video frame and the output of the previous computational unit to obtain a first transformation result; a second distortion layer performs a distortion transformation on the second video frame and the output of the previous computational unit to obtain a second transformation result; a stitching layer stitches together the first video frame, the second video frame, the first transformation result, the second transformation result, the output of the previous computational unit, and the interval time to obtain a first vector; at least one convolutional layer extracts features from the first vector to obtain a second vector; and a deconvolutional layer restores features from the second vector to obtain the output of the current computational unit. The interval time is the value between the time corresponding to the first video frame and the time corresponding to the second video frame.

[0085] Specifically, the optical flow estimation network model comprises multiple computational units, connected via residual networks. For one of these units, the video frame to be processed and the output of the previous unit are twisted using a first twisting layer and a second twisting layer, resulting in two transformations: a first transformation and a second transformation. Next, the video frame to be processed, the time interval, the output of the previous unit, the first transformation, and the second transformation are concatenated using a concatenation layer to obtain a first vector. This first vector is then input into at least one convolutional layer for feature extraction, yielding a second vector. This second vector is then input into a deconvolutional layer to obtain the output of the current computational unit, which is then fed into the next unit. After computation by multiple units, the output of the last unit is the optical flow estimation result and the fused image. The optical flow estimation result includes the optical flow estimation results for the first and second video frames.

[0086] It should be noted that the number of convolutional layers is not limited in this embodiment.

[0087] S330 generates the corresponding intermediate frame based on the optical flow estimation results and the fusion map.

[0088] In this embodiment, the time information of the intermediate frame to be generated is first determined based on the input voice information, and the video frame to be processed associated with the intermediate frame to be generated is obtained according to the time information. Then, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. The optical flow estimation network model includes multiple computing units, and adjacent computing units are connected through a residual network. Each computing unit includes a first twist layer, a second twist layer, a splicing layer, at least one convolutional layer and a deconvolutional layer. Finally, the corresponding intermediate frame is generated based on the optical flow estimation result and fusion map. Predicting the optical flow estimation result and fusion map through the optical flow estimation network model with the above structure is beneficial to improving the quality of the generated intermediate frame, enabling the virtual digital human to transition naturally during state transitions and complete corresponding actions coherently under voice drive, thereby improving the service quality of the virtual digital human.

[0089] Figure 3B This is a schematic diagram of the structure of an optical flow estimation network model provided in an embodiment of this disclosure. Figure 3B As shown, the optical flow estimation network model includes computational unit 1, residual network, computational unit 2, ...

[0090] It should be noted that, Figure 3B The number of computational units included in the optical flow estimation network model is merely illustrative and is not intended to limit its scope.

[0091] Figure 3CThis is a schematic diagram of the computational unit in the optical flow estimation network model provided in this embodiment of the disclosure. Figure 3C As shown, in each computing unit, the first twisted layer and the second twisted layer are connected in parallel, and after they are connected in parallel, they are connected to the splicing layer. The splicing layer is connected to at least one convolutional layer, and at least one convolutional layer is connected to a deconvolutional layer. The structure of the computing unit and the function of each layer have been described in detail in the above embodiments. To avoid repetition, they will not be described again here.

[0092] In some embodiments, optionally, each computing unit further includes a shrinking layer and a magnifying layer, the shrinking layer being located between the splicing layer and the at least one convolutional layer, and the magnifying layer being located after the deconvolutional layer;

[0093] The scaling layer is used to scale the first vector according to a scaling factor;

[0094] The amplification layer is used to amplify the output result of the current computing unit according to the amplification factor.

[0095] In this embodiment, if there are multiple convolutional layers, the scaling layer is located between the stitching layer and the first convolutional layer; if there is only one convolutional layer, the scaling layer is located between the stitching layer and that convolutional layer. The values ​​of the scaling factor and the magnification factor can be preset or determined according to specific circumstances; this embodiment does not impose any limitations on this.

[0096] To better handle the significant motion changes between the first and second video frames, each computation unit employs a coarse-to-fine approach to gradually increase the resolution. This means that "coarse" optical flow is calculated first at a low resolution, and then "fine" optical flow is calculated at a high resolution. Specifically, a scaling layer located between the stitching layer and the first convolutional layer scales the first vector according to a scaling factor, and a scaling layer located after the deconvolutional layer amplifies the output of the current computation unit according to a scaling factor.

[0097] In this embodiment, the above structure enables the optical flow estimation network model to have a wider range of applications, which is beneficial for outputting more accurate optical flow estimation results and fusion maps.

[0098] Figure 3D This is a schematic diagram of another computing unit provided in an embodiment of this disclosure. (See attached diagram.) Figure 3D As shown, in Figure 3C Based on the above, a shrinking layer and a magnifying layer have been added. The functions and connections of the two have been described in detail in the above embodiments, and will not be repeated here to avoid repetition.

[0099] For example, Figure 3EThis is a schematic diagram illustrating the computational principle of a certain computing unit provided in an embodiment of this disclosure. For example... Figure 3E As shown, assuming the first video frame is the image at time 0, denoted as I0, and the second video frame is the image at time 1, denoted as I1, with an interval t of any time between 0 and 1, for one of the multiple computational units, I0, I1, and the output of the previous computational unit (denoted as i-1) are used. and M i-1 Perform a distortion transformation operation, specifically: change I0, and M i-1 The input is fed into the first twisting layer to obtain the first transformation result. I1, and M i-1 The input is fed into the second twisting layer to obtain the second transformation result. Next, I0, I1, and M i-1 , t, first transformation result and the result of the second transformation The stitching operation is performed through the stitching layer, followed by the effects of the shrinking layer, convolutional layer, deconvolutional layer, and magnification layer, ultimately yielding the output of the current computational unit. and M i ).

[0100] It should be noted that: Figure 3E The number of convolutional layers is for illustrative purposes only and is not a limitation.

[0101] In some embodiments, optionally, after generating the corresponding intermediate frame based on the optical flow estimation result and the fusion map, the method may further include:

[0102] Based on the first video frame, the second video frame, the optical flow estimation result, the fusion map, and the target transformation result corresponding to the last computational unit in the optical flow estimation network model, the image fusion residual is obtained through the semantic segmentation network model.

[0103] Based on the intermediate frame and the image fusion residual, the corresponding image fusion result is obtained.

[0104] In this context, the target transformation result corresponding to the last computational unit can be understood as the result obtained after distortion transformation through the first and second distortion layers in the last computational unit. The semantic segmentation network model can use RefineNet or other networks; this embodiment does not impose any limitations.

[0105] Specifically, after obtaining the intermediate frames, the first video frame, the second video frame, the optical flow estimation result, the fusion map, and the target transformation result corresponding to the last computational unit in the optical flow estimation network model are input into the semantic segmentation network model to obtain the image fusion residual. Fusing the intermediate frames and the image fusion residual yields the corresponding image fusion result.

[0106] In this embodiment, since the generated intermediate frames may have problems such as blurring, artifacts and deformation, resulting in low quality and unnatural transitions, the image fusion process described above can optimize the intermediate frames and generate high-quality intermediate frames, which is more conducive to the subsequent virtual digital human state transition process.

[0107] Figure 4A This is a schematic diagram illustrating the principle of determining image fusion results according to an embodiment of the present disclosure. Figure 4A The process of determining the image fusion result has been described in detail in the above embodiments, and will not be repeated here to avoid repetition.

[0108] In some embodiments, the semantic segmentation network model may optionally include a context feature extraction layer, a third twisting layer, an encoding layer, a concatenation layer, and a decoding layer;

[0109] The context feature extraction layer is used to extract features from the first video frame and the second video frame to obtain feature vectors;

[0110] The third twisting layer is used to perform a twisting transformation on the feature vector and the optical flow estimation result to obtain a third transformation result;

[0111] The coding layer is used to encode the first video frame, the second video frame, the optical flow estimation result, the fusion map, the third transformation result, and the target transformation result to obtain a coding vector.

[0112] The splicing layer is used to splice the third transformation result and the encoding vector to obtain a spliced ​​vector;

[0113] The decoding layer is used to decode the spliced ​​vector to obtain the image fusion residual.

[0114] Specifically, the first and second video frames are input into the context feature extraction layer of the semantic segmentation network model for context feature extraction, which yields a feature vector. The feature vector and optical flow estimation result are input into the third distortion layer for distortion transformation, which yields the third transformation result. The first video frame, the second video frame, the optical flow estimation result, the fusion map, the third transformation result, and the target transformation result are input into the encoding layer for encoding, which yields an encoded vector. The third transformation result and the encoded vector are concatenated through the concatenation layer, which yields the concatenated vector. Finally, the concatenated vector is decoded through the decoding layer to obtain the image fusion residual.

[0115] In this embodiment, by sampling at the feature resolution through the structure of the semantic segmentation network model described above, the occurrence of problems such as blurring, artifacts and deformation can be reduced, which is beneficial for optimizing intermediate frames.

[0116] Figure 4B This is a schematic diagram of the structure of a semantic segmentation network model provided in an embodiment of this disclosure. Figure 4B As shown, the context feature extraction layer is connected to the third twist layer, the third twist layer is connected to the encoding layer, the encoding layer is connected to the concatenation layer, and the concatenation layer is connected to the decoding layer. The structure of the semantic segmentation network model and the role of each layer in the structure have been described in detail in the above embodiments. To avoid repetition, they will not be repeated here.

[0117] In some embodiments, optionally, the encoding layer includes at least one convolutional unit, and the decoding layer includes at least one deconvolutional unit, wherein the number of convolutional units is the same as the number of deconvolutional units, and the kernel parameters of the convolutional units are the same as the kernel parameters of the deconvolutional units.

[0118] Specifically, since the convolutional units in the coding layer and the deconvolutional units in the decoding layer are opposite operations, in order to ensure consistency, the number of convolutional units must be the same as the number of deconvolutional units, and the kernel parameters in the convolutional units must be the same as the kernel parameters in the deconvolutional units.

[0119] In this embodiment, the above features ensure the consistency of the encoding-decoding structure.

[0120] For example, the number of convolutional units can be 4, the kernel parameters in the convolutional unit can be 3*3, and each convolutional unit can include at least one convolutional layer. This embodiment does not limit this.

[0121] In some embodiments, optionally, the context feature extraction layer may employ an encoder structure, which may include at least one convolutional unit, and each convolutional unit may include at least one convolutional layer.

[0122] Figure 4C This is a schematic diagram illustrating the principle of obtaining image fusion results through a semantic segmentation network model, as provided in an embodiment of this disclosure. Figure 4C As shown, assume the first video frame is the image at time 0, denoted as I0, and the second video frame is the image at time 1, denoted as I1, F t→0 F represents the optical flow estimation result from time 0 to time t. t→1 This represents the optical flow estimation results from time t to time 1, where M represents the fused image. This represents the intermediate frame at time t. I0 and I1 are input into the context feature extraction part of the semantic segmentation network model for context feature extraction, resulting in a feature vector. The feature vector and F... t→0 and F t→1 The third distortion layer is used to perform a distortion transformation, resulting in the third transformation C. t→0 and C t→1 I0, I1, F t→0 F t→1 C t→0 C t→1 Target transformation results ( and The inputs M and C are fed into the coding layer for encoding to obtain the encoded vector; C is then fed into the coding layer for encoding to obtain the encoded vector. t→0 C t→1 The encoded vector is concatenated through a concatenation layer to obtain a concatenated vector; finally, the concatenated vector is decoded through a decoding layer to obtain the image fusion residual Δ, which is then used to fuse the intermediate frames. By fusing with Δ, a more accurate image fusion result is obtained.

[0123] It should be noted that: Figure 4C The number of convolutional and deconvolutional units is for illustrative purposes only and is not intended to limit the number of units.

[0124] Figure 5 This is a schematic diagram of an intermediate frame generation apparatus provided in an embodiment of this disclosure. This apparatus is configured in an electronic device and can implement the intermediate frame generation method described in any embodiment of this application. Figure 5 As shown, the device specifically includes the following:

[0125] The acquisition module 501 is used to determine the time information of the intermediate frame to be generated based on the input voice information, and to acquire the video frame to be processed associated with the intermediate frame to be generated according to the time information, wherein the input voice information is used to drive the virtual digital human to perform actions.

[0126] The determination module 502 is used to input the video frame to be processed into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map;

[0127] The generation module 503 is used to generate corresponding intermediate frames based on the optical flow estimation results and the fusion map.

[0128] As an optional implementation of this disclosure, the acquisition module 501 is specifically used for:

[0129] Based on the input voice information, determine the timing information of the intermediate frames to be generated;

[0130] Based on the time information, a video frame is determined from the static library, and a video frame is determined from the target action video corresponding to the voice information in the action library to obtain the video frame to be processed, wherein the video frame to be processed includes a first video frame and a second video frame.

[0131] The device further includes:

[0132] The stitching module is used to generate corresponding intermediate frames based on the optical flow estimation results and the fusion map, and then insert the intermediate frames between the first video frame and the second video frame to obtain the corresponding stitched video.

[0133] As an optional implementation of this disclosure, the optical flow estimation network model includes multiple computational units, adjacent computational units are connected through a residual network, each computational unit includes a first twist layer, a second twist layer, a splicing layer, at least one convolutional layer and a deconvolutional layer, and the video frame to be processed includes a first video frame and a second video frame.

[0134] The first distortion layer is used to perform distortion transformation on the first video frame and the output result of the previous computing unit to obtain a first transformation result;

[0135] The second distortion layer is used to perform distortion transformation on the second video frame and the output result of the previous calculation unit to obtain a second transformation result;

[0136] The splicing layer is used to splice the first video frame, the second video frame, the first transformation result, the second transformation result, the output result of the previous calculation unit, and the interval time to obtain a first vector;

[0137] The at least one convolutional layer is used to extract features from the first vector to obtain a second vector;

[0138] The deconvolutional layer is used to perform feature restoration on the second vector to obtain the output result of the current computing unit.

[0139] As an optional implementation of this disclosure, each computing unit further includes a shrinking layer and a magnifying layer, wherein the shrinking layer is located between the splicing layer and the at least one convolutional layer, and the magnifying layer is located after the deconvolutional layer;

[0140] The scaling layer is used to scale the first vector according to a scaling factor;

[0141] The amplification layer is used to amplify the output result of the current computing unit according to the amplification factor.

[0142] As an optional implementation of this disclosure, the apparatus further includes:

[0143] The residual determination module is used to obtain the image fusion residual based on the first video frame, the second video frame, the optical flow estimation result, the fusion map, and the target transformation result corresponding to the last calculation unit in the optical flow estimation network model, through the semantic segmentation network model.

[0144] The fusion module is used to obtain the corresponding image fusion result based on the intermediate frame and the image fusion residual.

[0145] As an optional implementation of this disclosure, the semantic segmentation network model includes a context feature extraction layer, a third distortion layer, an encoding layer, a concatenation layer, and a decoding layer;

[0146] The context feature extraction layer is used to extract features from the first video frame and the second video frame to obtain feature vectors;

[0147] The third twisting layer is used to perform a twisting transformation on the feature vector and the optical flow estimation result to obtain a third transformation result;

[0148] The coding layer is used to encode the first video frame, the second video frame, the optical flow estimation result, the fusion map, the third transformation result, and the target transformation result to obtain a coding vector.

[0149] The splicing layer is used to splice the third transformation result and the encoding vector to obtain a spliced ​​vector;

[0150] The decoding layer is used to decode the spliced ​​vector to obtain the image fusion residual.

[0151] As an optional implementation of this disclosure, the encoding layer includes at least one convolutional unit, and the decoding layer includes at least one deconvolutional unit. The number of convolutional units is the same as the number of deconvolutional units, and the kernel parameters of the convolutional units are the same as the kernel parameters of the deconvolutional units.

[0152] The intermediate frame generation apparatus provided in this embodiment first determines the time information of the intermediate frame to be generated based on the input voice information, and obtains the video frame to be processed associated with the intermediate frame to be generated according to the time information. The input voice information is used to drive the virtual digital human to perform actions. Then, the video frame to be processed is input into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map. Finally, the corresponding intermediate frame is generated based on the optical flow estimation result and fusion map. The above process can generate intermediate frames, which helps to ensure that the virtual digital human transitions naturally during state transitions, so that the virtual digital human can complete the corresponding actions smoothly under voice drive.

[0153] The intermediate frame generation apparatus provided in this disclosure can execute the intermediate frame generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0154] This disclosure provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any of the intermediate frame generation methods described in this disclosure.

[0155] The electronic device may be a personal computer (PC), a server, or a mainframe computer, etc., and this disclosure does not specifically limit it.

[0156] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 6 As shown, the electronic device includes a processor 610 and a storage device 620; the number of processors 610 in the electronic device can be one or more. Figure 6 Taking a processor 610 as an example; the processor 610 and the storage device 620 in the electronic device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0157] Storage device 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the intermediate frame generation method in the embodiments of this disclosure. Processor 610 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in storage device 620, thereby implementing the intermediate frame generation method provided in the embodiments of this disclosure.

[0158] Storage device 620 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on terminal usage. Furthermore, storage device 620 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, storage device 620 may further include memory remotely located relative to processor 610, which can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0159] The electronic device provided in this embodiment can be used to execute the intermediate frame generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0160] This disclosure provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described intermediate frame generation method and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0161] The computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0162] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments above is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of the embodiments suitable for specific application considerations.

Claims

1. A method for generating intermediate frames, characterized in that, The method includes: Based on the input voice information, the time information of the intermediate frame to be generated is determined, and the video frame to be processed associated with the intermediate frame to be generated is obtained according to the time information, wherein the input voice information is used to drive the virtual digital human to perform actions. The video frames to be processed are input into the optical flow estimation network model to obtain the corresponding optical flow estimation results and fusion map; Based on the optical flow estimation results and the fusion map, a corresponding intermediate frame is generated; The optical flow estimation network model includes multiple computational units, which are connected to each other through a residual network. Each computational unit includes a first twist layer, a second twist layer, a splicing layer, at least one convolutional layer, and a deconvolutional layer. The video frames to be processed include a first video frame and a second video frame. The first distortion layer is used to perform distortion transformation on the first video frame and the output result of the previous computing unit to obtain a first transformation result; The second distortion layer is used to perform distortion transformation on the second video frame and the output result of the previous calculation unit to obtain a second transformation result; The splicing layer is used to splice the first video frame, the second video frame, the first transformation result, the second transformation result, the output result of the previous calculation unit, and the interval time to obtain a first vector; The at least one convolutional layer is used to extract features from the first vector to obtain a second vector; The deconvolutional layer is used to perform feature restoration on the second vector to obtain the output result of the current computing unit; After generating the corresponding intermediate frame based on the optical flow estimation result and the fusion map, the method further includes: Based on the first video frame, the second video frame, the optical flow estimation result, the fusion map, and the target transformation result corresponding to the last computational unit in the optical flow estimation network model, the image fusion residual is obtained through the semantic segmentation network model. Based on the intermediate frame and the image fusion residual, the corresponding image fusion result is obtained; The semantic segmentation network model includes a context feature extraction layer, a third twisting layer, an encoding layer, a concatenation layer, and a decoding layer. The context feature extraction layer is used to extract features from the first video frame and the second video frame to obtain feature vectors; The third twisting layer is used to perform a twisting transformation on the feature vector and the optical flow estimation result to obtain a third transformation result; The coding layer is used to encode the first video frame, the second video frame, the optical flow estimation result, the fusion map, the third transformation result, and the target transformation result to obtain a coding vector. The splicing layer is used to splice the third transformation result and the encoding vector to obtain a spliced ​​vector; The decoding layer is used to decode the spliced ​​vector to obtain the image fusion residual.

2. The method according to claim 1, characterized in that, The step of obtaining the video frame to be processed associated with the intermediate frame to be generated based on the time information includes: Based on the time information, a video frame is determined from the static library, and a video frame is determined from the target action video corresponding to the voice information in the action library to obtain the video frame to be processed, wherein the video frame to be processed includes a first video frame and a second video frame. Accordingly, after generating the corresponding intermediate frame based on the optical flow estimation result and the fusion map, the method further includes: The intermediate frame is inserted between the first video frame and the second video frame to obtain the corresponding spliced ​​video.

3. The method according to claim 1, characterized in that, Each computing unit further includes a shrinking layer and a magnifying layer, wherein the shrinking layer is located between the splicing layer and the at least one convolutional layer, and the magnifying layer is located after the deconvolutional layer; The scaling layer is used to scale the first vector according to a scaling factor; The amplification layer is used to amplify the output result of the current computing unit according to the amplification factor.

4. The method according to claim 1, characterized in that, The encoding layer includes at least one convolutional unit, and the decoding layer includes at least one deconvolutional unit. The number of convolutional units is the same as the number of deconvolutional units, and the kernel parameters of the convolutional units are the same as the kernel parameters of the deconvolutional units.

5. An intermediate frame generation apparatus, characterized in that, The device includes: The acquisition module is used to determine the time information of the intermediate frame to be generated based on the input voice information, and to acquire the video frame to be processed associated with the intermediate frame to be generated according to the time information, wherein the input voice information is used to drive the virtual digital human to perform actions. The determination module is used to input the video frame to be processed into the optical flow estimation network model to obtain the corresponding optical flow estimation result and fusion map; The generation module is used to generate corresponding intermediate frames based on the optical flow estimation results and the fusion map; The optical flow estimation network model includes multiple computational units, which are connected to each other through a residual network. Each computational unit includes a first twist layer, a second twist layer, a splicing layer, at least one convolutional layer, and a deconvolutional layer. The video frames to be processed include a first video frame and a second video frame. The first distortion layer is used to perform distortion transformation on the first video frame and the output result of the previous computing unit to obtain a first transformation result; The second distortion layer is used to perform distortion transformation on the second video frame and the output result of the previous calculation unit to obtain a second transformation result; The splicing layer is used to splice the first video frame, the second video frame, the first transformation result, the second transformation result, the output result of the previous calculation unit, and the interval time to obtain a first vector; The at least one convolutional layer is used to extract features from the first vector to obtain a second vector; The deconvolutional layer is used to perform feature restoration on the second vector to obtain the output result of the current computing unit; After generating the corresponding intermediate frame based on the optical flow estimation result and the fusion map, the process further includes: Based on the first video frame, the second video frame, the optical flow estimation result, the fusion map, and the target transformation result corresponding to the last computational unit in the optical flow estimation network model, the image fusion residual is obtained through the semantic segmentation network model. Based on the intermediate frame and the image fusion residual, the corresponding image fusion result is obtained; The semantic segmentation network model includes a context feature extraction layer, a third twisting layer, an encoding layer, a concatenation layer, and a decoding layer. The context feature extraction layer is used to extract features from the first video frame and the second video frame to obtain feature vectors; The third twisting layer is used to perform a twisting transformation on the feature vector and the optical flow estimation result to obtain a third transformation result; The coding layer is used to encode the first video frame, the second video frame, the optical flow estimation result, the fusion map, the third transformation result, and the target transformation result to obtain a coding vector. The splicing layer is used to splice the third transformation result and the encoding vector to obtain a spliced ​​vector; The decoding layer is used to decode the spliced ​​vector to obtain the image fusion residual.

6. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Session processing method, device and equipment based on virtual object

    CN114691922A