Video playing method and related device

By predicting the time to generate the second video and determining a reasonable reserved duration, the problems of stuttering and unnatural video playback were solved, achieving a natural transition and a high-quality user experience.

CN119697403BActive Publication Date: 2025-11-07MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411808847.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-11-07
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Existing video playback methods are prone to stuttering, frame skipping, and unnatural playback when responding to user input, resulting in a poor user experience.

Method used

By predicting the time required to generate the second video corresponding to the input information, a reasonable reserve time is determined, and the switching video frame is determined from multiple video frames based on the playback status of the first video, the playback of the first and second videos is dynamically controlled to ensure a natural transition.

Benefits of technology

It achieves a smooth transition in video playback, avoiding stuttering and frame skipping, and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697403B_ABST
    Figure CN119697403B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video playing method and related equipment, the method comprising: in the process of playing a first video, in response to obtaining input information, determining a predicted time of generating a second video corresponding to the input information, and determining a switching video frame from a plurality of video frames included in the first video based on a playing condition of the first video and a reserved time length, the reserved time length being determined based on the predicted time, generating the second video based on the switching video frame, and determining an actual time of generating the second video, and playing the first video and the second video according to the actual time and the predicted time. The embodiments of the present application can reasonably determine the switching video frame, so as to ensure the video playing effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a video playing method and related device. BACKGROUND

[0002] At present, when playing a video, some information is usually inputted to change the playing of the video, thus, how to play a video based on the input information of a user becomes a hot research issue in the process of playing a video. SUMMARY

[0003] Embodiments of the present application provide a video playing method and related device, which can reasonably determine a switching video frame to ensure the video playing effect.

[0004] In a first aspect, embodiments of the present application provide a video playing method, which comprises:

[0005] In the process of playing a first video, in response to obtaining input information, a predicted time for generating a second video corresponding to the input information is determined, and a switching video frame is determined from a plurality of video frames included in the first video based on the playing situation of the first video and a reserved time length; the reserved time length is determined based on the predicted time;

[0006] A second video is generated based on the switching video frame, and an actual time for generating the second video is determined;

[0007] The first video and the second video are played according to the actual time and the predicted time.

[0008] In a second aspect, embodiments of the present application provide a video playing device, which comprises a determination unit, a generation unit and a playing control unit, wherein,

[0009] The determination unit is configured to, in the process of playing a first video, in response to obtaining input information, determine a predicted time for generating a second video corresponding to the input information, and determine a switching video frame from a plurality of video frames included in the first video based on the playing situation of the first video and a reserved time length; the reserved time length is determined based on the predicted time;

[0010] The generation unit is configured to generate a second video based on the switching video frame, and determine an actual time for generating the second video;

[0011] The playing control unit is configured to play the first video and the second video according to the actual time and the predicted time.

[0012] In a third aspect, an electronic device is provided, which includes a processor, a memory, a communication interface, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the processor. The programs include instructions for performing the steps in the first aspect of the embodiments.

[0013] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program for electronic data exchange. The computer program causes a computer to perform some or all of the steps described in the first aspect of the embodiments.

[0014] In a fifth aspect, a computer program product is provided, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments. The computer program product can be a software installation package.

[0015] By implementing the embodiments of the present application, the following beneficial effects are achieved:

[0016] The video playing method and related device described in the present application, in the process of playing the first video, in response to obtaining the input information, determine the predicted time of generating the second video corresponding to the input information, and determine the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video and the reserved time length, generate the second video based on the switching video frame, and determine the actual time of generating the second video, play the first video and the second video according to the actual time and the predicted time. In this way, when it is necessary to change the playing of the first video according to the input information, the predicted time required for generating the second video corresponding to the input information is first predicted, then a reasonable reserved time length corresponding to the predicted time is determined, and the switching video frame in the first video can be accurately and reasonably determined in combination with the reserved time length and the playing situation of the current first video. Then, based on the difference between the actual time and the predicted time, the playing of the first video and the second video is dynamically controlled. Since the second video is generated based on the switching video frame, the video playing method of the present application can ensure that the first video is naturally transitioned to the second video through the switching video frame, that is, the transition of the first video and the second video when playing can be ensured to be more natural, and thus the video playing effect can be ensured. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0018] Figure 1 is a flow diagram of a video playing method provided by an embodiment of the present application;

[0019] Figure 2 is a demonstration diagram of an 18-person body key point model provided by an embodiment of the present application;

[0020] Figure 3 is another demonstration diagram of an 18-person body key point model provided by an embodiment of the present application;

[0021] Figure 4 is a structure diagram of a time prediction model provided by an embodiment of the present application;

[0022] Figure 5 is another structure diagram of a time prediction model provided by an embodiment of the present application;

[0023] Figure 6 is a structure diagram of an electronic device provided by an embodiment of the present application;

[0024] Figure 7 is a function unit composition block diagram of a video playing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0026] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0027] Reference to“an embodiment” or“the embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” or“in at least one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a particular embodiment which is“preferred” over other embodiments. As used herein, the term“exemplary” or“illustrative” means serving as an example, instance or illustration. Any implementation described herein as“exemplary” or“illustrative” is not necessarily to be construed as preferred or advantageous over other implementations.

[0028] The electronic device described in the embodiments of the present application can include a smart phone (such as an Android phone, an iOS phone, a Windows Phone phone, etc.), a tablet computer, a palm computer, a driving recorder, a notebook computer, a mobile internet device (MID), or a wearable device (such as a smart watch, a Bluetooth headset), etc. The above are only examples and are not exhaustive. The electronic device can include, but is not limited to, the above electronic devices. The electronic device can also include a server, which can be an outsourcing server, a cloud server, an edge server, a conversational robot, a server cluster, etc. The present application is not limited in this regard.

[0029] The related terms involved in the present application will be introduced first.

[0030] The first video can specifically refer to a group of non-real-time generated frames, i.e., originally photographed video frames, which are used for loop playing.

[0031] The second video can be understood as a video generated based on input information in response to input information during the playing of the first video. For example, when the user asks a question, a new second video is generated according to the response text of the conversational robot.

[0032] The reserved time length is the time length between the first time when the first video frame in the first video being played is obtained and the second time when the first video is played to the switching video frame. That is, the first video frame in the first video is played until the switching video frame within the reserved time length.

[0033] The switching video frame can be understood as a video frame of the first video. Specifically, the first video frame in the first video being played when the input information is obtained is determined as the second video frame, and the second video frame corresponding to the reserved time length is determined from the second video frame, and the determined second video frame is determined as the switching video frame. After playing the switching video frame, the second video can be played.

[0034] A transition frame, which is a transition frame generated by the interpolation algorithm based on the switching video frame in the first video, is the first frame of the second video and is used to splice the first video and the second video, so that there is a smooth transition between the first video and the second video.

[0035] Full-link response time: the response time of the full link is composed of the automatic speech recognition (ASR) time, the dialogue robot time, the text to speech (TTS) time, the digital human driving and rendering algorithm time, and the network transmission time. Specifically, for example, the full-link response time = ASR time + dialogue robot time + TTS time + digital human driving and rendering algorithm time + network transmission time.

[0036] Stutter, that is, the playback stops at a certain frame and does not move, at which time the screen is static.

[0037] Backward, that is, the playback starts to reverse, that is, the playback is in the opposite direction of the original playback direction.

[0038] Frame skipping, which can be understood as a video composed of a group of continuous pictures, the difference between each adjacent two frames is small, if several frames are missing in the middle, the playback will skip several frames, and there will be obvious picture discontinuity.

[0039] In the related art, taking a 2D digital human as an example, a 2D digital human as a kind of real person level effect digital human has been widely used, for example: enterprise digital employees, product explainers, travel explainers, public services, etc. There are some digital humans “on duty”. The essence of a 2D digital human is to re-edit a video, for example, only the region of half a face is generated by an algorithm, and other non-generated regions are taken from the original video, and then the generated part and the original video are fused, and a 2D digital human can be obtained.

[0040] In practical applications, a 2D real-time digital human system responds in the form of image and voice in a very short time after the user asks a question. When there is no question, the digital human is in standby state, at which time the first video is played in a loop. Once the user asks a question, the second video will be generated according to the response text of the dialogue robot, and the second video will be played after the switching video frame. Since the digital human driving and rendering algorithm needs a period of time to generate the second video, the first video needs to be played to wait for the algorithm to complete the generation of the second video.

[0041] Specifically, the following two playback modes are mainly used:

[0042] The first mode: a short video (first video) loop is used, and when the second video has been generated, it also needs to wait for the short video loop to be played out before switching to play the second video, that is, the first frame of the second video is determined based on a fixed frame in the first video. Although the short video loop is fast in switching, the loop is very serious and looks unnatural.

[0043] The second mode: a long video (first video) loop is used, and at this time the first frame of the second video is determined based on a frame in the first video, but the frame is not fixed. Since the reserved time length is manually specified, the video may appear to be turning back, which is very unnatural and has poor user experience.

[0044] In view of the defects of the related technologies, the current target is: 1) shortening the full-link response time; 2) the second video is continuous and natural, and does not appear to be stuttering, frame skipping, or frequently turning back. Since the digital human driving and rendering algorithm needs a period of time to generate the second video, the first video needs to be continuously played within the reserved time length to wait for the second video to be generated. If the reserved time length is too long, the response time will be increased, and if the reserved time length is too short, the playing will appear to be stuttering and frame skipping. Therefore, the embodiment of the present application provides a video playing method, which comprises: in the process of playing a first video, in response to obtaining input information, determining a predicted time for generating a second video corresponding to the input information, and determining a switching video frame from a plurality of video frames included in the first video based on a playing situation of the first video and a reserved time length; the reserved time length is determined based on the predicted time; generating a second video based on the switching video frame, and determining an actual time for generating the second video; and playing the first video and the second video according to the actual time and the predicted time.

[0045] In the embodiment of the present application, in the case where it is necessary to change the playing of the first video according to the input information, the predicted time required for generating the second video corresponding to the input information is first predicted, and then a reasonable reserved time length corresponding to the predicted time is determined. The reserved time length and the current playing situation of the first video can be combined to accurately and reasonably determine the switching video frame in the first video. Then, based on the difference between the actual time and the predicted time, the playing of the first video and the second video is dynamically controlled. Since the second video is generated based on the switching video frame, the video playing method of the present application can ensure that the first video is naturally transitioned to the second video through the switching video frame, that is, the transition of the first video and the second video when playing can be ensured to be more natural, and thus the video playing effect can be ensured.

[0046] The embodiments of the present application will be described in detail below.

[0047] Please refer to Figure 1 , Figure 1is a flowchart of a video playing method provided by an embodiment of the present application, as shown in the figure, the video playing method comprises:

[0048] 101. In the process of playing the first video, in response to obtaining the input information, determining the predicted time of generating the second video corresponding to the input information, and determining the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video and the reserved time length.

[0049] The reserved time length is determined based on the predicted time. The input information can be understood as input information obtained by the electronic device, and the input form of the input information can include at least one of the following: text content, image content, video content, audio content, sensor collection, etc., which is not limited here. For example, the input information can include voice, text, gestures, lip language, etc.

[0050] The playing situation of the first video can include at least one of the following: the playing order of the first video, the playing position of the first video, the playing rate of the first video, the playing resolution of the first video, etc., which is not limited here. The playing order can be understood as the video playing order, which can include forward playing order or reverse playing order, the forward playing order means playing in the normal video playing order, and the reverse playing order means playing in the playing order opposite to the normal video playing order. The playing position of the first video can be understood as which frame is currently played.

[0051] In a specific implementation, in the process of playing the first video, the user can be waited for input information, and after detecting the input information, the input information can be obtained in response, and the switching video frame can be determined from the plurality of video frames included in the first video based on the playing situation of the first video and the reserved time length. The reserved time length corresponds to the switching video frame, since the reserved time length corresponds to the input information, the accurate predicted time is determined based on the input information, and the reasonable reserved time length corresponding to the predicted time is determined, so that the reasonable switching video frame is determined based on the reserved time length, avoiding the phenomenon of stuttering and frame skipping, etc., to ensure the digital human playing effect.

[0052] In an embodiment of the present application, the original video can be converted into a loop video for a digital human, specifically, the original video can be converted into a loop video by using an interpolation algorithm, and the loop video can be used as the first video.

[0053] The original video can be an original video for a digital human, and the digital human can include a 2D digital human or a 3D digital human. The digital human not only refers to a virtual character image, but also can refer to a cartoon character, an animal shape digital human, an object shape digital human, etc., which is not limited here. Of course, the original video can not include a digital human, but can be processed into a loop video including a digital human.

[0054] In a specific implementation, since the digital human video of different lengths needs to be generated, a long enough loop video is needed. In addition, the video shot is ultimately of a limited length, and therefore a loop video that can be played in a loop is needed.

[0055] In a specific implementation, for the original video of the digital human, since the original video is shot to ensure that the character will return to the initial action within a period of time, the condition for forming a loop video is met, and therefore only the frame returning to the initial action in a video segment needs to be screened out, and a certain time length is selected, so as to obtain an original video segment to be converted into a loop video. The following steps A1-A3 can be referred to:

[0056] A1, the pose key point model can be used to perform pose key point labeling on all frames to obtain key point coordinates (x i ,y i )∈P f , where P f represents the key point set of each frame of a certain video segment f, i=1, 2,..., n is the i-th key point of each frame, and n is the number of key points of each frame.

[0057] In a specific implementation, since there is usually only one character in the shot, the action amplitude of the character is not large, and the overall front face is kept facing the shot. Therefore, a 2D pose key point algorithm can be used. The 2D pose key point detection algorithm is a common technique in the art. For example, as shown in Figure 2 , it provides an 18 human body key point model.

[0058] The human body key point model is a subtask of human pose estimation. Taking the COCO key point model as an example, the 18-point model trained in the COCO dataset. The key points and serial numbers defined in the COCO dataset are as follows: nose-0, neck-1, right shoulder-2, right elbow-3, right wrist-4, left shoulder-5, left elbow-6, left wrist-7, right hip-8, right knee-9, right ankle-10, left hip-11, left knee-12, left ankle-13, right eye-14, left eye-15, ear-16, left ear-17, background-18 (not shown in the figure), which can be referred to for details Figure 3For example, the key point detection model that can be applicable is a VGGNet model, the first 10 layers of the VGGNet model are used to create feature maps for the input image, resulting in a plurality of features, and then the features can be input into two parallel branches of a convolutional layer. The first branch predicts a set of confidence maps (18), each of which represents a specific component of the human pose skeleton graph, and the key point detection task corresponds to the first branch. The second branch predicts a set of Part Affinity Fields (PAF, 38), which represent the degree of association between components.

[0059] Specifically, taking the opencv function as an example, the input of the key point detection model can be an original image without any processing, and the image can be pre-processed by the opencv self-provided function, including mean subtraction, scale, cropping, channel exchange, etc., to return a 4-channel blob (which can be simply understood as an N-dimensional array for the input of the neural network), and the N-dimensional array is input into the VGGNet network to obtain a two-dimensional confidence map, that is, the body part position (such as elbow joint, knee, etc.) is estimated. The confidence map is to show the possibility of the appearance of the human body part in the degree of gray and white.

[0060] A2, find the frame of returning to the initial action in the video. A certain frame can be selected as a reference initial action frame, and all frames of returning to the initial action in the video are found based on the reference initial action frame.

[0061] Specifically, step A2 can include the following steps B1-B3:

[0062] B1, select a certain frame f0 as a reference initial action frame.

[0063] B2, calculate the distance of each frame from f0 according to the following formula, that is, the sum of key point distances, which is as follows:

[0064]

[0065] Where i represents the ith frame in the video, j = 1, 2,..., n is the jth key point of each frame, n is the number of key points of each frame, is the key point set of the f0th frame, P i is the key point set of the ith frame.

[0066] B3, give a very small positive number ε. If D i < ε, mark the ith frame as a frame of returning to the initial action. If no frame is found in one traversal, increase ε to 1.05ε as the new ε, and traverse again. In this way, until ≥1 frame of returning to the initial action is found.

[0067] A3, select a certain length of video and generate a transition frame.

[0068] In a specific implementation, a video of more than 1 minute can be selected as a cycle period according to experience, so that the cycle of the video is difficult to be perceived. A video of more than 1 minute is selected from the first frame of the video, and a certain frame found in step A2) that returns to the initial action is taken as the last frame f e . A transition frame needs to be created from the f e frame to the first frame. By calling an interpolation algorithm, k frames of transition frames can be inserted, where k is an odd number. The size of k depends on the transition time. According to the frames per second (FPS) of the current video, the playing time of each frame can be determined, and the value of k can be calculated according to the required transition time. The length of the transition time can be determined subjectively. Generally, k = 3.

[0069] Among them, video interpolation aims to improve the frame rate and smoothness of the video, making the video look more "smooth". The general process of transition frame generation (interpolation algorithm) is as follows: based on deep learning, the original video adjacent two frames are usually taken as the input of the neural network, combined with optical flow neural network, occlusion estimation and other technologies, to predict the intermediate frame between the two frames. Deep learning method can extract image semantic information, and often performs better in occlusion estimation and other aspects, can calculate the dense optical flow between two frames, and use the optical flow information to warp the front and back two frames to the intermediate time, thereby synthesizing the intermediate frame. Specifically, for example: TOFlow based on optical flow interpolation algorithm, in the video interpolation task, the input frame number is 3 frames, thus, 1 cycle video is obtained.

[0070] In some possible examples, the playing situation includes that a first video frame in the first video is being played when the input information is acquired; and the plurality of video frames included in the first video are arranged and played in a first order;

[0071] The step 101 of determining the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video can include the following steps:

[0072] A21, determining a video frame located after the first video frame in the first video as a second video frame;

[0073] A22, determining a second video frame corresponding to the reserved time length from the second video frame, and taking the determined second video frame as the switching video frame.

[0074] The first order can include an order corresponding to a forward playing order.

[0075] In specific implementations, the playing condition can include that a first video frame in the first video is being played when the input information is obtained, so that a video frame in the first video that is located after the first video frame can be determined as a second video frame, and a second video frame corresponding to the reserved time length can be determined from the second video frame, and the determined second video frame can be used as the switching video frame. Thus, based on the first playing time of the first video frame in the first video that is being played when the input information is obtained, the second playing time corresponding to the switching video frame can be determined based on the first playing time and the reserved time length, and the corresponding second video frame can be obtained as the switching video frame based on the second playing time. In this way, the corresponding switching video frame can be determined based on the playing condition of the first video and the reserved time length.

[0076] The reserved time length can be related to the input information, and of course, the reserved time length can also be related to the predicted time, that is, the reserved time length is determined based on the predicted time. Since the corresponding second video is generated based on the input information, the corresponding reserved time length can be determined based on the generation time of the second video. The predicted time can be understood as the time when the second video is generated.

[0077] In some possible embodiments, the step 101 of determining the predicted time for generating the second video corresponding to the input information can include the following steps:

[0078] B11, obtaining an environment parameter;

[0079] B12, obtaining attribute information of a response text corresponding to the input information;

[0080] B13, inputting the environment parameter and the attribute information into a time prediction model to obtain the predicted time.

[0081] In the embodiments of the present application, the environment parameter can include at least one of the following: service concurrency, digital human driving and rendering algorithm generation efficiency, hardware related parameters when the digital human driving and rendering algorithm is tested, network parameters, and the like, which are not limited herein. The hardware related parameters when the digital human driving and rendering algorithm is tested can include at least one of the following: central processing unit (CPU) related parameters, graphics processing unit (GPU) related parameters, neural network processing unit (NPU) related parameters, memory related parameters, computing power, and the like, which are not limited herein. The network parameters can include at least one of the following: network bandwidth, network transmission rate, and the like, which are not limited herein.

[0082] The CPU-related parameters can include at least one of the following: a main frequency, a core number, a thread number, a floating point operations per second (FLOPS), and the like, without limitation, the GPU-related parameters can include at least one of the following: a main frequency, a core number, a thread number, a FLOPS, and the like, without limitation, the NPU-related parameters can include at least one of the following: a main frequency, a core number, a thread number, a FLOPS, and the like, without limitation, and the memory-related parameters can include at least one of the following: a memory size, a memory processing rate, a memory delay, and the like, without limitation.

[0083] In the embodiments of the present application, the attribute information of the response text corresponding to the input information can include at least one of the following: an audio length corresponding to the response text, a language type of the response text, an expression manner of the response text, and the like, without limitation. The language type of the response text can include a local or national language or a special code (Morse code), and the language type of the response text can include at least one of the following: Chinese, English, Japanese, and local dialects, without limitation. The expression manner of the response text can include at least one of the following: a gesture, a dance, a lip language, an audio, and the like, without limitation.

[0084] In specific implementations, the environment parameters can be obtained, and then the attribute information of the response text corresponding to the input information is obtained, and then the environment parameters and the attribute information are input into a time prediction model to obtain a predicted time. Since the environment parameters and the attribute information of the response text corresponding to the input information are used for time prediction, on the one hand, the prediction result is in-depth consistent with the actual network environment and device environment (software environment, hardware environment, etc.), and on the other hand, the prediction result is also in-depth consistent with the inherent characteristics of the response text corresponding to the input information, so that the predicted time is more consistent with the actual situation, and thus the reserved time length is more consistent with the actual situation, thereby reasonably determining the switching video frame to ensure the digital human playing effect.

[0085] In some possible embodiments, the number of predicted times is a plurality; and the method can further include the following steps:

[0086] C1, determining an average value and a standard deviation of the plurality of predicted times;

[0087] C2, determining a reserved time length according to the average value and the standard deviation.

[0088] In the embodiments of the present application, the number of predicted times can be multiple, and specifically, multiple predictions can be performed through the time prediction model to obtain multiple predicted times. The mean value and standard deviation are obtained by using the multiple predicted times for mean value operation and standard deviation operation, and the reserved duration can be obtained according to the mean value and the standard deviation. Since the model is constantly learning and the environment is constantly changing, the mean value of multiple predicted times is used for calculation, which can reduce the prediction error. The standard deviation reflects the prediction fluctuation to a certain extent, and the reserved duration is adjusted to a certain extent by using the standard deviation, so that the reserved duration is more in line with the actual situation, thereby ensuring the digital human playing effect.

[0089] In a specific implementation, wherein σ is the standard deviation of t gen , and s is an adjustment factor, which is a positive number, for example, the adjustment factor is 3.

[0090] For example, for a new computing environment, the environment parameters can be processed and input into the time prediction model to predict the required predicted time t gen . Finally, the reserved duration is wherein σ is the standard deviation of t gen . The switching video frame can be determined according to t cache . Since the input environment parameters come from the current computing environment and are relatively fixed, the prediction can be performed offline, without occupying the time for generating the second video in real time.

[0091] In a specific implementation, the reserved duration is determined based on the predicted time. Since the prediction time depth is related to the environment and the attribute depth of the response text itself, the reserved duration also conforms to the environment and the attribute of the response text corresponding to the input information. Based on this, the switching video frame corresponding to the reserved duration can be obtained from the first video, so that the switching video frame can be reasonably determined to ensure the digital human playing effect.

[0092] In some possible embodiments, the following steps can also be included:

[0093] D1, obtaining sample data, the sample data including sample environment parameters and / or sample attribute information;

[0094] D2, preprocessing the sample data to obtain target sample data;

[0095] D3, training the target sample data on a preset time prediction model to reach a preset condition to obtain the time prediction model.

[0096] ​In the embodiments of the present application, the sample environment parameters can include at least one of the following: service concurrency, digital human driving and rendering algorithm generation efficiency, hardware related parameters during testing of the digital human driving and rendering algorithm, network parameters, and the like, which are not limited herein. The hardware related parameters during testing of the digital human driving and rendering algorithm can include at least one of the following: CPU related parameters, GPU related parameters, NPU related parameters, memory related parameters, computing power, and the like, which are not limited herein. The network parameters can include at least one of the following: network bandwidth, network transmission rate, network delay, and the like, which are not limited herein.

[0097] It can be understood that the hardware related parameters not only include the front-end hardware related parameters, but also include the back-end hardware related parameters. The front-end can include a terminal, and the back-end can include servers.

[0098] The CPU related parameters can include at least one of the following: frequency, core number, thread number, FLOPS, and the like, which are not limited herein. The GPU related parameters can include at least one of the following: frequency, core number, thread number, FLOPS, and the like, which are not limited herein. The NPU related parameters can include at least one of the following: frequency, core number, thread number, FLOPS, and the like, which are not limited herein. The memory related parameters can include at least one of the following: memory size, memory processing rate, memory delay, and the like, which are not limited herein.

[0099] The sample attribute information can include at least one of the following: audio length corresponding to the response text, language type of the response text, expression manner of the response text, and the like, which are not limited herein. The language type of the response text can include local or national language or special code (Morse code). The language type of the response text can include at least one of the following: Chinese, English, Japanese, local dialect, and the like, which are not limited herein. The expression manner of the response text can include at least one of the following: gesture, dance, lip language, audio, and the like, which are not limited herein.

[0100] The sample data can include the sample environment parameters, or the sample data can include the sample attribute information, or the sample data can include both the sample environment parameters and the sample attribute information.

[0101] The preprocessing can include at least one of the following: normalization, one-hot vectorization, value padding, and the like, which are not limited herein. The purpose of the preprocessing is to constrain each input parameter within a certain range, for example, between 0 and 1.

[0102] The preset condition can be pre-set or system default, for example, the preset condition can be reaching a set training number, for example, the preset condition can be model convergence, for example, the preset condition can be that the accuracy of the model reaches a set accuracy, wherein the set training number and the set accuracy can be pre-set or system default.

[0103] The preset time prediction model can be pre-set or system default. The input (target sample data) of the time prediction model is denoted as x, and the output (predicted time) is denoted as y. By setting different inputs x, different training data can be obtained, such as setting different concurrent request quantities, setting TTS audio of different lengths, setting different CPU core numbers, thread numbers, setting different memory limits, using different GPU cards, and making different speed limits on the network, different prediction times t gen can be obtained. gen That is, y.

[0104] For model training, the sample data is pre-processed, such as normalization, one-hot vectorization, value padding, and the like, to process the input x and the output y to make them numerical, in the same scale and without null values. A machine learning model or a deep learning model is used for training (fitting an x->y function) to obtain a time prediction model (such as a regression model), as shown in FIG. 3, the time prediction model can be regarded as an x->y function. The time prediction model can include a machine learning model, which can use machine learning algorithms such as Gradient Boosting Decision Tree (GBDT), Boosting tree, or Deep Neural Networks (DNN) with stronger learning ability. Further, as shown in FIG. 4, the sample data can be pre-processed to obtain x, and then x is input into the time prediction model to obtain y. Figure 4 Figure 5

[0105] In a specific implementation, sample data (sample environment parameters and / or sample attribute information) can be obtained, and the sample data is pre-processed to obtain target sample data, each data in the target sample data is constrained within a certain range, and the target sample data is trained on the preset time prediction model to reach a preset condition to obtain the time prediction model.

[0106] In a specific implementation, the full-link response time of the second video can be predicted by the time prediction model to obtain a predicted time.

[0107] ​​In the embodiments of the present application, the time prediction model can realize function adaptation, and then the full-link response time of the second video corresponding to the response text is generated through function adaptive prediction.

[0108] Wherein, the input x and output y of the function can be defined, the input of the time prediction model is x, and the output is y. The prediction model is established by using the method of machine learning or deep learning. First, the elements affecting the generation time need to be found as input x. The main factors affecting the time are: audio length, service concurrency, digital human driving and rendering algorithm generation efficiency, CPU / GPU / memory related parameters (such as frequency, core number, thread number, FLOPS, memory size, etc.) of the digital human driving and rendering algorithm in the test, CPU / GPU / memory related parameters (such as frequency, core number, thread number, FLOPS, memory size, etc.) of the current server, and current network parameters (such as bandwidth, whether it is a backbone network, etc.). Y can be understood as how long it takes to start playing the second video stably and continuously without pause.

[0109] 102, generating the second video based on the switching video frame and determining the actual time of generating the second video.

[0110] Wherein, in the embodiments of the present application, the second video can include a group of video frames or two groups of video frames. The actual time of the second video can be understood as the time of actually generating the second video.

[0111] In specific implementation, whether the second video is generated can be detected in real time, and then the actual time of generating the second video can be obtained based on this method.

[0112] In specific implementation, a time prediction model can be used to predict the full-link response time of the second video corresponding to the input information based on the switching video frame, and at least one predicted time is obtained.

[0113] In the embodiments of the present application, the second video corresponding to the input information can be generated based on the switching video frame, and the actual time of generating the second video is determined. Since the second video is generated based on the switching video frame, the association between the first video and the second video is established by switching the video frame, which helps to realize the natural transition between the first video and the second video based on the actual time and the predicted time based on the switching video frame.

[0114] 103, playing the first video and the second video according to the actual time and the predicted time.

[0115] In the embodiments of the present application, the first video and the second video can be dynamically played according to the order between the actual time and the predicted time. Since the association between the first video and the second video is established by switching the video frames, the natural transition between the first video and the second video can be realized based on the switching of the video frames.

[0116] In specific implementations, one of the at least one predicted time can be selected and compared with the actual time, or the average of the at least one predicted time can be determined and compared with the actual time.

[0117] In some possible embodiments, the step 103 of playing the first video and the second video according to the actual time and the predicted time can be implemented in the following manner:

[0118] A31, when the actual time is earlier than or equal to the predicted time, continue playing the first video;

[0119] A32, when the playing of the switching video frame in the first video is completed, play the second video.

[0120] In specific implementations, since the actual time refers to the actual time of generating the second video, and the predicted time refers to the predicted time of generating the second video, when the actual time is earlier than the predicted time, it means that the generation of the second video has been completed before the predicted time arrives, but the playing has not reached the switching video frame. In addition, when the actual time is earlier than the predicted time, it means that the second video has been generated before the predicted time arrives. In order to ensure the natural transition between the first video and the second video based on the switching of the video frames, the first video needs to be continued to be played until the switching video frame, that is, after the switching video frame is played, the second video is played. Since the second video is generated based on the switching of the video frames, the natural transition between the first video and the second video is realized, and the playing effect of the digital person is ensured.

[0121] Of course, when the actual time is equal to the predicted time, it means that the playing has reached the switching video frame, and the second video has been generated. Therefore, after the switching video frame is played, the second video can be directly played.

[0122] In some possible embodiments, the plurality of video frames included in the first video are arranged in a first order; and the step 103 of playing the first video and the second video according to the actual time and the predicted time can be implemented in the following manner:

[0123] B31, when the actual time is later than the predicted time, continue playing the first video;

[0124] B32, after the switching video frame in the first video is played, if the second video has not been generated, continue to play the video frames after the switching video frame in the first video in the first order until the second video has been generated;

[0125] B33, play from a third video frame to the switching video frame in a second order; the third video frame refers to the video frame being played in the first video when the second video is generated;

[0126] B34, after the switching video frame is played, play the second video in the first order.

[0127] Wherein, the first order and the second order are both frame arrangement orders in the video. The first order and the second order can be pre-set or system default, and the first order and the second order are opposite to each other.

[0128] Wherein, the third video frame can be understood as the video frame being played when the second video is generated.

[0129] In the specific implementation, since the actual time refers to the actual time of generating the second video, and the predicted time refers to the predicted time of generating the second video, when the actual time is later than the predicted time, the first video will be continued to be played first, and after the switching video frame in the first video is played, the predicted time will be reached, but the second video has not been generated. At this time, in order to ensure the digital human playing effect, the video frames after the switching video frame in the first video will be continued to be played in the first order until the second video has been generated. Since the second video is generated based on the switching video frame, in order to realize the natural transition of the first video to the second video through the switching video frame, it is necessary to return to the switching video frame, and then play the second video after the switching video frame is played. Therefore, when the second video is generated, the third video frame will be played from the switching video frame in the second order; the third video frame refers to the video frame being played in the first video when the second video is generated. Therefore, the first video can be naturally transitioned to the second video through the switching video frame, so as to avoid the phenomenon of frame skipping and lag, and ensure the digital human playing effect.

[0130] To illustrate, taking the first video played in forward order as an example, if the actual time is later than the predicted time, the predicted time will arrive, but the second video will not be generated yet. In this case, the frames after the switching video frame can continue to be played, that is, the video frames in the first video that are after the switching video frame are played until the second video is generated. When the second video is generated, it starts from the current playback frame when the second video is generated and plays back in reverse order until it returns to the switching video frame. That is, after playing the switching video frame, it starts back in forward order again. By adjusting the reserved time, the phenomenon of the second video being generated later than the predicted time can be reduced, while ensuring that the overall time consumption does not increase significantly.

[0131] To illustrate further, regarding the handling strategy for actual times being earlier or later than predicted times, during playback, the playback begins with the first frame of the first video, following frame order. After user feedback, playback continues according to the reserved duration t. cache Confirm switching video frame f t That is, switching video frames after playback is complete. t Then, the second video began to play.

[0132] To further illustrate, if the second video is generated earlier than the predicted time, it means that the time to switch video frames has not yet been played. t However, the second video has already been generated, since the second video is based on the switched video frame f. t If generated, it will continue to play until f. t Then, play the second video F. gen If the second video is generated later than the predicted time, a situation will occur where the switching video frame has already passed. t However, if the second video has not yet been generated, you need to wait for it to be generated before you can continue playing. t The following video frames {f t+1 ,f t+2 ,...,f t+k After the second video is generated, it is played backwards from the currently playing video frame at the moment the second video was generated, i.e., {f}. t+k ,f t+k-1 ,...,f t}, where k is greater than f t The next frame played. Return to switching video frames. t Then turn back again and play the second video F in forward playback order. gen Adjusting the reserved time can reduce the phenomenon of generating the second video later than the prediction time, while ensuring that the end-to-end response time does not increase significantly.

[0133] In some possible embodiments, the second video includes a first group of video frames and a second group of video frames; the first group of video frames includes m video frames corresponding to the first order and the input information, which are generated based on the switch video frame; the second group of video frames includes n video frames corresponding to the second order and the input information, which are generated based on the switch video frame; m and n are positive integers.

[0134] The step 103 of playing the first video and the second video according to the actual time and the predicted time can be implemented in the following manner:

[0135] C31, when the actual time is earlier than or equal to the predicted time, continuing to play the first video in the first order until the switch video frame in the first video is played completely, and playing the first group of video frames in the first order;

[0136] C32, when the actual time is later than the predicted time, playing the first video in the first order;

[0137] C33, if the second group of video frames has not been generated after the switch video frame in the first video is played completely, continuing to play the video frames in the first video after the switch video frame in the first order until the second group of video frames is generated;

[0138] C34, playing from a fourth video frame to the switch video frame in the second order, and playing the second group of video frames in the second order after the switch video frame is played completely; the fourth video frame refers to the video frame being played in the first video when the second group of video frames is generated.

[0139] The first order and the second order can be pre-set or system defaults, and the first order and the second order are opposite to each other. The fourth video frame can be understood as the video frame being played when the second group of video frames is generated.

[0140] The second video can include a first group of video frames and a second group of video frames, the first group of video frames can include m video frames corresponding to the first order and the input information, which are generated based on the switch video frame, and the second group of video frames can include n video frames corresponding to the second order and the input information, which are generated based on the switch video frame; m and n are positive integers, and m and n can be equal or not equal. Since there is a difference in the video generation direction between the first group of video frames and the second group of video frames, the number of video frames corresponding to the input information can also be different.

[0141] In the embodiments of the present application, the first group of video frames and the second group of video frames can be generated in parallel, that is, one process is used to generate the first group of video frames corresponding to the switching video frame generation input information, and another process is used to generate the second group of video frames corresponding to the switching video frame generation input information, or, that is, one thread is used to generate the first group of video frames corresponding to the switching video frame generation input information, and another thread is used to generate the second group of video frames corresponding to the switching video frame generation input information.

[0142] In specific implementations, in the case where the second video includes the first group of video frames and the second group of video frames, the actual time can be understood as the actual time when the first group of video frames and the second group of video frames are both generated, and the predicted time can be understood as the predicted time when the first group of video frames and the second group of video frames are both generated.

[0143] Among them, due to the difference in the generation mode of the first group of video frames and the second group of video frames, the actual time of the first group of video frames and the second group of video frames may also differ, the first group of video frames can correspond to one actual time, and the second group of video frames can also correspond to one actual time, although the difference between the two actual times is not large, but there may be some difference, then the actual time corresponding to the second video can be the latest time among the actual times corresponding to the two groups of video frames.

[0144] In specific implementations, since the actual time refers to the actual time of generating the second video, and the predicted time refers to the predicted time of generating the second video, the first video is continued to be played in the first order first, and when the actual time is earlier than or equal to the predicted time, it means that the generation of the second video has been completed before the predicted time arrives, but the switching video frame has not been played to. In order to ensure the natural transition between the first video and the second video based on the switching video frame, it is necessary to continue to play the first video until the switching video frame, and under the condition that the first video is continued to be played in the first order until the switching video frame in the first video is played, the first group of video frames is played in the first order, since the first group of video frames is generated based on the switching video frame, and thus, the natural transition between the first video and the first group of video frames is realized, and the digital human playing effect is ensured.

[0145] In a specific implementation, since the actual time refers to the actual time of generating the second video, and the predicted time refers to the predicted time of generating the second video, the first video is played in the first order first. When the actual time is later than the predicted time, it indicates that after the switching video frame in the first video is played, the predicted time is reached, but the second set of video frames has not been generated. At this time, in order to ensure the digital human playing effect, after the switching video frame is played, the first video is played in the first order. If the second set of video frames has not been generated after the switching video frame in the first video is played, the video frames in the first video after the switching video frame are continued to be played in the first order until the second set of video frames is generated, the fourth video frame is played from the fourth video frame to the switching video frame in the second order, and after the switching video frame is played, the second set of video frames is played in the second order. The fourth video frame refers to the video frame being played in the first video when the second set of video frames is generated. Since the second set of video frames is generated based on the switching video frame, and the second set of video frames corresponds to the second order which is opposite to the first order, the number of backtracking can be reduced (once less), and the playing effect is more natural as a whole. In addition, the first video can be naturally transitioned to the second set of video frames through the switching video frame, so as to avoid the phenomenon of frame skipping and lag, and ensure the digital human playing effect.

[0146] For example, in a specific implementation, when the second video includes two sets of video frames, the next frame f t and the previous frame f t+1 of the switching video frame f t-1 may be generated at the same time to generate two sets of video frames respectively in the backward and forward directions. At this time, if the second video returns earlier than the predicted time, the second video is played after the switching video frame f t in the forward playing order. If the second video is generated later than the predicted time, the second video is continued to be played after the switching video frame f t in the reverse playing order. In this way, the number of backtracking can be reduced once, and the playing effect is more natural as a whole.

[0147] In a specific implementation, since there will be two times of backtracking once the actual time is later than the predicted time, which will cause the video to appear temporarily unnatural, and the user may perceive the repeated playing of the forward playing order and the reverse playing order. Therefore, when the computing resources allow, the next frame f t and the previous frame f t+1 of the switching video frame f t-1 may be generated at the same time to generate two sets of video frames respectively in the backward and forward directions. The two sets of video frames are respectively, i.e., the first set of video frames and the second set of video frames, specifically, the first set of video frames {f t+1 ,f t+2 ,...,f t+l} and the first set of video frames {f t-1 ,f t-2 ,...,ft-l}, wherein, l is the number of frames needed to generate the video. That is, when the actual time is earlier than or equal to the predicted time, the first video is continuously played until the switching video frame, and then the first group of video frames is played; otherwise, when the actual time is later than the predicted time, the first video is continuously played until the switching video frame, and then the second group of video frames is played in the reverse playing order. Specifically, when the actual time of the second video is earlier than the predicted time, the second video is played to the switching video frame f t , and then the first group of video frames {f t+1 ,f t+2 ,...,f t+l} is played in the forward playing order; and when the actual time of the second video is later than the predicted time, the second video is played to the switching video frame f t , and then the second group of video frames {f t-1 ,f t-2 ,...,f t-l} is continuously played in the reverse playing order. In this way, one backtracking is saved, and the overall playing effect is more natural.

[0148] Further, in the embodiments of the present application, after the second video is played, the first video is returned to. If the playing order when the first video is returned to is the forward playing order, the first video is continuously played in the forward playing order; otherwise, if the playing order when the first video is returned to is the reverse playing order, the first video is continuously played in the reverse playing order. Similarly, the user is waited for to ask a question under the first video played in the reverse playing order. In this way, in the application scenario of the real-time 2D digital person playing, a real-time 2D digital person playing system can be obtained.

[0149] The video playing method described in the present application, in the process of playing the first video, in response to obtaining the input information, determines the predicted time of generating the second video corresponding to the input information, and determines the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video and the reserved time length, the reserved time length is determined based on the predicted time, the second video is generated based on the switching video frame, and the actual time of generating the second video is determined, the first video and the second video are played according to the actual time and the predicted time. In this way, in the case of needing to change the playing of the first video according to the input information, the predicted time required for generating the second video corresponding to the input information is first predicted, and then a reasonable reserved time length corresponding to the predicted time is determined. Combined with the reserved time length and the current playing situation of the first video, the switching video frame in the first video can be accurately and reasonably determined. Then, based on the difference between the actual time and the predicted time, the playing of the first video and the second video is dynamically controlled. Since the second video is generated based on the switching video frame, the video playing method of the present application can ensure that the first video is naturally transitioned to the second video through the switching video frame, that is, the transition of the first video and the second video when playing can be ensured to be more natural, and further, the video playing effect can be ensured.

[0150] In some possible examples, the playing situation includes that a first video frame in the first video is being played when the input information is acquired; and the plurality of video frames included in the first video are arranged and played in a first order. Figure 6 Figure 6 is a structural schematic diagram of an electronic device provided in an embodiment of the present application, as shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs, the one or more programs are stored in the memory and configured to be executed by the processor, in the embodiment of the present application, the program includes instructions for executing the following steps:

[0151] In the process of playing the first video, in response to acquiring the input information, a predicted time for generating a second video corresponding to the input information is determined, and a switching video frame is determined from a plurality of video frames included in the first video based on a playing situation of the first video and a reserved time length; the reserved time length is determined based on the predicted time;

[0152] The second video is generated based on the switching video frame, and an actual time for generating the second video is determined;

[0153] The first video and the second video are played according to the actual time and the predicted time.

[0154] In some possible examples, the playing situation includes that a first video frame in the first video is being played when the input information is acquired; and the plurality of video frames included in the first video are arranged and played in a first order;

[0155] In the aspect of determining the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video, the program includes instructions for executing the following steps:

[0156] A video frame located after the first video frame in the first video is determined as a second video frame;

[0157] A second video frame corresponding to the reserved time length is determined from the second video frame, and the determined second video frame is taken as the switching video frame.

[0158] In some possible examples, in the aspect of playing the first video and the second video according to the actual time and the predicted time, the program includes instructions for executing the following steps:

[0159] When the actual time is earlier than or equal to the predicted time, the first video is continued to be played;

[0160] When the switching video frame in the first video is played completely, the second video is played.

[0161] ​In some possible examples, the first video includes a plurality of video frames arranged in a first order; and the program for playing the first video and the second video according to the actual time and the predicted time includes instructions for performing the following steps:

[0162] playing the first video continuously when the actual time is later than the predicted time;

[0163] playing the video frames after the switch video frame in the first video in the first order until the second video is generated after the switch video frame in the first video is played completely;

[0164] playing from a third video frame to the switch video frame in a second order; the third video frame is a video frame being played in the first video when the second video is generated;

[0165] playing the second video in the first order after the switch video frame is played completely.

[0166] In some possible examples, the second video includes a first group of video frames and a second group of video frames; the first group of video frames includes m video frames generated based on the switch video frame, corresponding to the first order and the input information; the second group of video frames includes n video frames generated based on the switch video frame, corresponding to the second order and the input information; m and n are positive integers;

[0167] In some possible examples, the program for playing the first video and the second video according to the actual time and the predicted time includes instructions for performing the following steps:

[0168] playing the first group of video frames in the first order when the actual time is earlier than or equal to the predicted time and the switch video frame in the first video is played completely in the first order;

[0169] playing the first video in the first order when the actual time is later than the predicted time;

[0170] playing the video frames after the switch video frame in the first video in the first order until the second group of video frames is generated after the switch video frame in the first video is played completely;

[0171] play the fourth video frame to the switching video frame in a second order, and play the second group of video frames in the second order after the switching video frame is played completely; the fourth video frame refers to a video frame being played in the first video when the second group of video frames is generated.

[0172] In some possible examples, the number of prediction times is multiple; the program further includes instructions for performing the following steps:

[0173] determining an average value and a standard deviation of the multiple prediction times;

[0174] determining a reserved time length according to the average value and the standard deviation.

[0175] In some possible examples, in the determining of the prediction time of generating the second video corresponding to the input information, the program includes instructions for performing the following steps:

[0176] obtaining an environmental parameter;

[0177] obtaining attribute information of a response text corresponding to the input information;

[0178] inputting the environmental parameter and the attribute information into a time prediction model to obtain the prediction time.

[0179] The electronic device described in the present application, in the process of playing the first video, in response to obtaining the input information, determines the prediction time of generating the second video corresponding to the input information, and determines the switching video frame from the multiple video frames included in the first video based on the playing situation of the first video and the reserved time length, the reserved time length being determined based on the prediction time, generates the second video based on the switching video frame, and determines the actual time of generating the second video, and plays the first video and the second video according to the actual time and the prediction time. In this way, in the case of needing to change the playing of the first video according to the input information, the prediction time required for generating the second video corresponding to the input information is first predicted, then a reasonable reserved time length corresponding to the prediction time is determined, and the reserved time length and the playing situation of the current first video can be combined to accurately and reasonably determine the switching video frame in the first video, and then the difference between the actual time and the prediction time is used to dynamically control the playing of the first video and the second video. Since the second video is generated based on the switching video frame, the video playing method of the present application can ensure that the first video is naturally transitioned to the second video through the switching video frame, that is, the transition of the first video and the second video when playing can be ensured to be more natural, and thus the video playing effect can be ensured.

[0180] Figure 7is a functional unit constituent block diagram of a video playing device 700 involved in embodiments of the present application. The video playing device 700 comprises a determining unit 701, a generating unit 702 and a playing control unit 703, wherein,

[0181] The determining unit 701 is configured to, in a process of playing a first video, in response to obtaining input information, determine a predicted time of generating a second video corresponding to the input information, and determine a switching video frame from a plurality of video frames included in the first video based on a playing situation of the first video and a reserved time length, wherein the reserved time length is determined based on the predicted time;

[0182] The generating unit 702 is configured to generate the second video based on the switching video frame, and determine an actual time of generating the second video;

[0183] The playing control unit 703 is configured to play the first video and the second video according to the actual time and the predicted time.

[0184] In some possible examples, the playing situation comprises that a first video frame in the first video is being played when the input information is obtained; and the plurality of video frames included in the first video are arranged and played in a first order;

[0185] In the aspect of determining the switching video frame from the plurality of video frames included in the first video based on the playing situation of the first video, the determining unit 701 is specifically configured to:

[0186] determine a video frame located after the first video frame in the first video as a second video frame;

[0187] determine a second video frame corresponding to the reserved time length from the second video frame, and determine the determined second video frame as the switching video frame.

[0188] In some possible examples, in the aspect of playing the first video and the second video according to the actual time and the predicted time, the playing control unit 703 is specifically configured to:

[0189] when the actual time is earlier than or equal to the predicted time, continue to play the first video;

[0190] when the switching video frame in the first video is played completely, play the second video.

[0191] In some possible examples, the plurality of video frames included in the first video are arranged in the first order; and in the aspect of playing the first video and the second video according to the actual time and the predicted time, the playing control unit 703 is specifically configured to:

[0192] continue playing the first video when the actual time is later than the predicted time;

[0193] continue playing the video frames in the first video after the switch video frame in the first video is played, and until the second video is generated, according to the first order;

[0194] play from a third video frame to the switch video frame according to a second order; the third video frame is a video frame being played in the first video when the second video is generated;

[0195] play the second video according to the first order after the switch video frame is played.

[0196] In some possible examples, the second video includes a first group of video frames and a second group of video frames; the first group of video frames includes m video frames generated based on the switch video frame, corresponding to the first order and the input information; the second group of video frames includes n video frames generated based on the switch video frame, corresponding to the second order and the input information; m and n are positive integers;

[0197] In the playing of the first video and the second video according to the actual time and the predicted time, the playing control unit 703 is specifically configured to:

[0198] continue playing the first video according to the first order when the actual time is earlier than or equal to the predicted time, until the switch video frame in the first video is played, and play the first group of video frames according to the first order;

[0199] continue playing the first video according to the first order when the actual time is later than the predicted time;

[0200] continue playing the video frames in the first video after the switch video frame in the first video is played, and until the second group of video frames is generated, according to the first order;

[0201] play from a fourth video frame to the switch video frame according to a second order, and play the second group of video frames according to the second order after the switch video frame is played; the fourth video frame is a video frame being played in the first video when the second group of video frames is generated.

[0202] In some possible examples, the predicted time is a plurality of times; and the video playing apparatus is further specifically configured to:

[0203] determining an average value and a standard deviation of the plurality of predicted times;

[0204] determining a reserved duration according to the average value and the standard deviation.

[0205] In some possible examples, in the determining, the determining unit 701 is specifically configured to:

[0206] obtain an environmental parameter;

[0207] obtain attribute information of a response text corresponding to the input information;

[0208] input the environmental parameter and the attribute information into a time prediction model to obtain the predicted time.

[0209] The video playing apparatus described in the present application, in the process of playing a first video, in response to obtaining input information, determines a predicted time of generating a second video corresponding to the input information, and determines a switching video frame from a plurality of video frames included in the first video based on a playing situation of the first video and a reserved duration, the reserved duration being determined based on the predicted time, generates the second video based on the switching video frame, and determines an actual time of generating the second video, and plays the first video and the second video according to the actual time and the predicted time. In this way, in the case where it is necessary to change the playing of the first video according to the input information, the predicted time required for generating the second video corresponding to the input information is first predicted, then a reasonable reserved duration corresponding to the predicted time is determined, and the switching video frame in the first video can be accurately and reasonably determined in combination of the reserved duration and the playing situation of the current first video, and the playing of the first video and the second video is dynamically controlled based on the difference between the actual time and the predicted time. Since the second video is generated based on the switching video frame, the video playing method of the present application can ensure that the first video is naturally transitioned to the second video through the switching video frame, that is, the transition of the first video and the second video when playing can be ensured to be relatively natural, and thus the video playing effect can be ensured.

[0210] It can be understood that the functions of each program module of the video playing apparatus of the present embodiment can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can be referred to the related descriptions of the above method embodiments, which will not be described herein again.

[0211] The present embodiment further provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute part or all steps of any method described in the above method embodiments, and the above computer includes an electronic device.

[0212] The embodiment of the present application further provides a computer program product, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The computer program product can be a software package, and the computer comprises an electronic device.

[0213] It should be noted that, for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0214] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0215] In several embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the above-described system embodiments are merely illustrative. For example, the division of the above units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical or other forms.

[0216] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0217] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0218] If the above integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned method of each embodiment of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0219] A person of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods of the embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0220] The embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present application, the specific implementation manner and application range will be changed, and the above description of the embodiments is not used to limit the present application.

Claims

1. A video playing method, characterized in that, The method comprises the following steps: During playing the first video, in response to obtaining the input information, a predicted time for generating a second video corresponding to the input information is determined, and a switching video frame is determined from a plurality of video frames included in the first video based on a playing condition of the first video and a reserved time length; The second video is generated based on the switching video frame, and an actual time for generating the second video is determined; The first video and the second video are played according to the actual time and the predicted time.

2. The method of claim 1, wherein, The playing condition comprises that a first video frame in the first video is being played when the input information is obtained; the plurality of video frames included in the first video are arranged and played in a first order; The switching video frame is determined from the plurality of video frames included in the first video based on the playing condition of the first video, which comprises the following steps: A video frame located after the first video frame in the first video is determined as a second video frame; A second video frame corresponding to the reserved time length is determined from the second video frame, and the determined second video frame is taken as the switching video frame.

3. The method according to claim 1 or 2, characterized in that, The first video and the second video are played according to the actual time and the predicted time, which comprises the following steps: When the actual time is earlier than or equal to the predicted time, the first video is continuously played; When the switching video frame in the first video is played completely, the second video is played.

4. The method according to claim 1 or 2, characterized in that, The plurality of video frames included in the first video are arranged in a first order; the first video and the second video are played according to the actual time and the predicted time, which comprises the following steps: When the actual time is later than the predicted time, the first video is continuously played; After the switching video frame in the first video is played completely, if the second video has not been generated, the video frames located after the switching video frame in the first video are continuously played in the first order until the second video has been generated; From a third video frame to the switching video frame, the third video frame is a video frame being played in the first video when the second video is generated, the second video is played in a second order; After the switching video frame is played completely, the second video is played in the first order.

5. The method according to claim 1 or 2, characterized in that, The second video comprises a first group of video frames and a second group of video frames; the first group of video frames comprises m video frames generated based on the switching video frame, corresponding to a first order and the input information; the second group of video frames comprises n video frames generated based on the switching video frame, corresponding to a second order and the input information; m and n are positive integers; The first video and the second video are played according to the actual time and the predicted time, which comprises the following steps: When the actual time is earlier than or equal to the predicted time, the first video is continuously played in the first order until the switching video frame in the first video is played completely, and then the first group of video frames is played in the first order; When the actual time is later than the predicted time, the first video is played in the first order; if the second set of video frames has not been generated after the switching video frame is played in the first video, continue playing the video frames in the first video after the switching video frame in the first order until the second set of video frames is generated; playing from a fourth video frame to the switching video frame in a second order, and after the switching video frame is played, playing the second set of video frames in the second order; the fourth video frame refers to the video frame being played in the first video when the second set of video frames is generated.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes determining a predicted time for generating a second video corresponding to the input information, including: obtaining an environmental parameter; obtaining attribute information of a response text corresponding to the input information; inputting the environmental parameter and the attribute information into a time prediction model to obtain the predicted time.

7. A video playback device, comprising: The apparatus includes a determination unit, a generation unit, and a playing control unit, wherein the determination unit is configured to, in response to obtaining input information during playing of a first video, determine a predicted time for generating a second video corresponding to the input information, and determine a switching video frame from a plurality of video frames included in the first video based on a playing situation of the first video and a reserved time length; the reserved time length is determined based on the predicted time; the generation unit is configured to generate the second video based on the switching video frame, and determine an actual time for generating the second video; the playing control unit is configured to play the first video and the second video according to the actual time and the predicted time.

8. An electronic device, comprising: A computer program product including a processor and a memory configured to store one or more programs for execution by the processor, the programs including instructions for performing the steps of the method of any of claims 1-6.

9. A computer-readable storage medium, characterized in that, A computer program product for electronic data interchange, wherein the computer program product causes a computer to perform the method of any of claims 1-6.

10. A computer program product, characterised in that, A non-transitory computer-readable storage medium including computer-readable code, or carrying computer-readable code, which, when executed in a processor of an electronic device, causes the processor in the electronic device to perform the method of any of claims 1-6.

Citation Information

Patent Citations

  • Video tracking device and method, video playing device and method and electronic device

    CN115550689A

  • Video definition switching method and device, equipment and medium

    CN117376644A