Video playing method and video playing device

By generating transitional video clips during video switching, the problem of discontinuous video playback is solved, resulting in a smoother video playback experience.

CN121967798APending Publication Date: 2026-05-01IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2025-12-24
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, when switching from a pre-generated video to a real-time generated video or vice versa during video playback, there is a lack of smooth transition, resulting in disjointed visuals and affecting the user experience.

Method used

When a video switch is detected, the end video frame is determined from the currently unplayed video frames, and a transition video segment is generated with the first frame of the target video. After playing the end video frame, the system switches to the target video, using the transition video segment for a smooth transition.

Benefits of technology

By generating transitional video clips, visual differences are reduced, improving the smoothness of video playback and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967798A_ABST
    Figure CN121967798A_ABST
Patent Text Reader

Abstract

The invention discloses a video playing method and a video playing device, and the method comprises the steps: detecting that a first video containing a first digital entity needs to be switched to a second video containing a second digital entity in a process of playing the first video; determining an ending video frame from the currently unplayed video frames of the first video; generating a transition video clip based on the ending video frame and the first video frame of the second video; and after the video frame is played, firstly playing the transition video clip, and then switching to the second video for playing. Through the above mode, the video playing continuity can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Video playback method and video playback device Technical Field

[0001] This application relates to the field of computer technology, and in particular to a video playback method and a video playback device. Background Technology

[0002] With the rapid development of artificial intelligence technology, digital human technology has shown enormous potential in various fields such as education, entertainment, and customer service. Taking personalized education scenarios as an example, videos featuring digital humans generally fall into two categories: one is pre-produced offline videos, suitable for content introductions, operation demonstrations, or narrative explanations in fixed scenarios; the other is real-time generated videos, which can dynamically output content that precisely matches the interactive information during user interaction with the digital human. However, currently, the switching between these two types of videos often uses a direct jump mode, which easily causes broken screen transitions, resulting in insufficient continuity, obvious abrupt changes, and other problems that affect the smoothness of the teaching experience. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a video playback method and a video playback device that can improve the continuity of video playback.

[0004] To solve the above-mentioned technical problems, one technical solution adopted in this application is: to provide a video playback method, the method comprising: during the playback of a first video containing a first digital entity, detecting that a switch to a second video containing a second digital entity is required; determining an end video frame from the currently unplayed video frames of the first video; generating a transition video segment based on the end video frame and the first video frame of the second video; and after playing the end video frame, playing the transition video segment first, and then switching to the second video for playback.

[0005] To solve the above-mentioned technical problems, another technical solution adopted in this application is: providing a video playback device, the device including a detection module, a determination module, a generation module, and a playback module; the detection module is used to detect that a switch to a second video containing a second digital entity is required during the playback of a first video containing a first digital entity; the determination module is used to determine the end video frame from the currently unplayed video frames of the first video; the generation module is used to generate a transition video segment based on the end video frame and the first video frame of the second video; the playback module is used to play the transition video segment first after playing the end video frame, and then switch to the second video for playback.

[0006] The above solution, during the playback of the first video, if a switch to the second video is detected, determines the ending video frame from the currently unplayed frames of the first video, and uses this ending video frame and the first frame of the second video to generate a transition video segment. After playing the ending video frame, this transition video segment is played first, and then the playback of the second video is switched to. Compared to the method of directly jumping from the first video to the second video when a switch is detected, the playback method described in this application, which uses a generated transition video segment to connect the first and second videos, can smoothly transition the differences between the ending video frame and the first video frame, thereby improving the continuity of video playback. Attached Figure Description

[0007] Figure 1 is a flowchart illustrating an embodiment of the video playback method provided in this application; Figure 2 is a flowchart illustrating an embodiment of step S12 shown in Figure 1; Figure 3 is a flowchart illustrating an embodiment of step S21 shown in Figure 2; Figure 4 is a flowchart illustrating an embodiment of playing a second video provided in this application; Figure 5 is a schematic diagram of the framework of an embodiment of the video playback device provided in this application; Figure 6 is a schematic diagram of the framework of an embodiment of the electronic device provided in this application; Figure 7 is a schematic diagram of the framework of the computer-readable storage medium provided in this application. Detailed Implementation

[0008] To make the purpose, technical solution and effects of this application clearer and more explicit, the following describes this application in further detail with reference to the accompanying drawings and embodiments.

[0009] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0010] The video playback method provided in this application is applicable to any video playback scenario that requires switching between different videos. It can be, but is not limited to, a video playback scenario where the user interacts with a digital entity in real time, or a video playback scenario that requires switching between multiple videos.

[0011] For example, in the currently playing video A, a digital human is explaining the operating procedures of an instrument. During the playback of video A, the user suddenly interrupts and asks a question, wanting the digital human to explain what problems would occur if step A is performed incorrectly. In this case, it is necessary to respond to the user's question and play video B, which is a response video about what problems would occur if step A is performed incorrectly. Therefore, it is necessary to switch from the current video A to video B.

[0012] For example, if multiple pre-generated videos need to be played in a preset order, switching between different videos may be involved during the playback of these videos.

[0013] To facilitate understanding, the following is a brief explanation of the background of the video playback method provided in this application, using personalized education scenarios such as learning machines as an example: It should be noted that in personalized education scenarios such as learning machines, emotionally supportive digital humans can provide users with a more immersive and personalized learning experience. Currently, the implementation methods of digital humans mainly fall into two categories: one is based on offline generation technology, which generates high-quality, highly expressive video content through pre-calculation and rendering. This type of video typically involves content introductions, operation demonstrations, or narrative explanations in fixed scenarios; the other is based on real-time synthesis / generation technology, which generates videos in real time through voice-driven, lip-syncing, and other methods. This allows for dynamic output of content that precisely matches the interactive information during user interaction with the digital human.

[0014] In simple terms, the offline videos mentioned above are pre-generated and stored videos (pre-stored videos). Real-time generated videos can be videos generated in real time based on user interaction information, or they can be videos obtained by generating corresponding facial regions in real time based on user interaction information based on pre-stored videos, and replacing the facial regions in the original pre-stored videos with the generated facial regions.

[0015] Offline algorithms can generate visually appealing and expressive digital human videos, but because they are pre-generated, they lack the flexibility of real-time interaction and cannot dynamically adjust based on facial expressions, tone of voice, or questions asked. While real-time synthesized digital human videos can generate corresponding videos based on user interaction information, enabling instant interaction, their image quality, motion smoothness, and naturalness of expression often differ significantly from high-quality offline-generated videos due to limitations in algorithm complexity and real-time requirements. Traditional methods often abruptly switch between these pre-stored and real-time generated videos—for example, when switching from a preset teaching scenario (offline video) to free dialogue (online digital human), or vice versa—lacking a smooth visual transition. This results in abrupt scene changes that disrupt the continuity and immersion of the user experience.

[0016] To address the aforementioned problem of discontinuous video playback caused by direct video switching, this application proposes the following video playback method and video playback device. Specifically: Please refer to Figure 1, which is a flowchart illustrating an embodiment of the video playback method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment is not limited to the flow order shown in Figure 1. As shown in Figure 1, this embodiment includes: S11: During the playback of a first video containing a first digital entity, it is detected that a switch to a second video containing a second digital entity is required.

[0017] In this embodiment, the digital entity can be, but is not limited to, a digital entity with the image of a person, or a digital entity with the image of an animal, etc.; the first digital entity and the second digital entity can be digital entities with the same image, or digital entities with different images.

[0018] In one implementation scenario, at least one of the currently playing first and second videos is a pre-stored video in a database; the video library includes several pre-stored videos containing digital entities. For example, one of the currently playing first and second videos is a pre-stored video in the database; the other is a video generated in real time.

[0019] Real-time generated videos include, for example, videos containing digital humans generated in real time based on user interaction information (such as user-provided voice, images, or descriptive text); for example, a trained multimodal driving model can be used to generate relevant driving parameters (such as lip features, emotion parameters, and / or scene information) using user interaction information, and then use the relevant driving parameters to generate videos containing lip or facial motion information in real time.

[0020] For example, both the first and second videos are pre-stored videos in the database. Later, as needed, the corresponding pre-stored videos can be retrieved from the video library and played directly, or the pre-stored videos can be fine-tuned based on actual conditions before playback. For instance, based on user interaction information, corresponding facial region images can be dynamically generated, and the facial regions of the digital human in the pre-stored video can be replaced with the generated facial region images before playback. Alternatively, the video containing the digital human / entity generated in real-time can also be saved to the video library as a pre-stored video.

[0021] It should be noted that the real-time generated video contains not only images but also corresponding audio.

[0022] In some implementation scenarios, pre-stored videos can be generated in the following way: obtain digital entity reference data provided by the user, and use a video diffusion model to generate a video containing digital entities based on the digital entity reference data, which is then used as a pre-stored video; the digital entity reference data includes at least one of the following: at least one digital entity image reference image, a reference video containing digital entities or real entities, and descriptive text about digital entities.

[0023] Specifically, the digital entity reference data provided by the user can be one or more reference images of the digital entity (used to extract the entity's actions, facial expressions, emotions, etc.), or descriptive text about the digital entity, such as "a primary school student in a school uniform walking around in the classroom." It can also be a reference video containing the digital entity or a real entity; for example, a pre-recorded video. The digital entity reference data serves as a reference for video generation. A video diffusion model can be used to generate a video containing the digital entity, which can then be stored as a pre-saved video. In essence, a pre-saved video can be generated based on the user-provided digital entity reference data; therefore, videos containing corresponding digital entities can be customized according to user needs and stored as pre-saved videos.

[0024] Taking the generation of emotionally supportive digital human videos as an example, the pre-stored videos generated can include casual conversations, object displays, teaching demonstrations, singing and dancing, etc. These videos are mainly about people, and the people can perform any complex walking, body movements and facial expressions.

[0025] In other implementation scenarios, videos captured from real-world objects can also be used as pre-stored videos.

[0026] In other implementation scenarios, offline generation technology can be used to pre-calculate and render high-quality, highly expressive video content, which can then be stored as pre-saved videos. To better meet user interaction needs, high-quality pre-saved videos in various styles and scenarios can be generated in advance. For example, a video diffusion model can be used to generate the corresponding video, and then each frame in the generated video can be finely adjusted to obtain a high-quality pre-saved video.

[0027] It should be noted that before switching videos, you need to first determine the second video to switch to, and then you can switch.

[0028] In one implementation scenario, if no user interaction information is detected during the playback of the first video, the second video to be switched to can be determined directly based on the currently playing first video. The second video can be a video found in the video library that has a scene and / or content related to the first video; or it can be the next video after the currently playing first video, found according to the playback order of multiple videos set in advance.

[0029] In one implementation scenario, if user interaction information is detected during the playback of the first video, a second video related to that interaction information needs to be determined first. This second video can be a video generated in real-time based on the user interaction information, a pre-stored video found in a video library, or a pre-stored interactive video. Furthermore, the second video needs to simultaneously play corresponding audio; therefore, the audio to be played and whether switching to the second video is necessary need to be determined during the playback of the first video and before the playback of the second video.

[0030] Of course, if the corresponding second video is not determined, and it is determined that there is no need to switch to the second video at present, then only the audio to be played will be played, and no video switching will be performed.

[0031] In a preferred embodiment, during the playback of a first video containing a first digital entity, audio to be played is generated; then, a pre-stored video matching the audio scene to be played is searched in the video library; if a pre-stored video matching the audio scene to be played is found, the matching pre-stored video is used as the second video, and it is determined that the current device needs to switch to the second video.

[0032] Of course, if no pre-stored video matching the speech scene to be played is found, it can be determined that there is no need to switch to the second video at this time. Alternatively, a preset interactive video can be selected as the second video, and it can be determined that a switch to the second video is required. Or, a second video can be generated in real time based on the user's interaction information, and it can be determined that a switch to the second video is required. The preset interactive video is, for example, a pre-set, universal main interactive video that can be applied to multiple scene switching. For example, the preset interactive video is a video containing the movement of digital entity lips. Subsequently, the lip movements of the digital entity in the preset interactive video can be synchronized with the speech to be played based on the speech to be played.

[0033] In one embodiment, regardless of whether a pre-stored video matching the audio scene to be played is found, the preset interactive video can be used as the second video; for example, if a matching pre-stored video is found, the preset interactive video can be used as the first part of the second video and the matching pre-stored video can be used as the second part of the second video; if no matching pre-stored video is found, the preset interactive video is used as the complete second video.

[0034] In one specific embodiment, generating the speech to be played includes: generating text to be played based on user interaction information using a large model; and converting the text to be played into speech to be played.

[0035] The user's interaction information can be understood as the user's interaction request information. This interaction request information can be, but is not limited to, request text, voice or images, etc. Of course, it can also be a combination of at least two of text, voice and images.

[0036] In one embodiment, searching for pre-stored videos in a video library that match the scene of the data to be played includes the following steps: calculating the text similarity between the scene description text of each pre-stored video in the video library and the text to be played corresponding to the audio to be played; and identifying the pre-stored videos whose text similarity meets the scene similarity requirement as pre-stored videos that match the scene of the audio to be played. The similarity requirement is that the text similarity should be greater than or equal to a preset similarity threshold, and the specific preset similarity threshold can be determined based on practical experience.

[0037] For example, pre-trained models such as BERT are used to extract word vectors corresponding to the text to be played and the scene description text, respectively, and then the cosine similarity between the word vectors is calculated; if there is a scene description text with a pre-set similarity greater than or equal to a preset similarity threshold, then the pre-stored video corresponding to the scene description text is determined to be a pre-stored video that matches the speech scene to be played.

[0038] It should be noted that after generating each pre-saved video, each pre-saved video will be associated with its corresponding information. This includes associating the scene description text related to the pre-saved video with its corresponding preset video, as well as associating the scene type, action, and emotion type of the pre-saved video with its corresponding preset video. Scene types include, for example, "broadcast," "chat," "leisure," "singing," "knowledge teaching," "wandering alone," "anniversary reminder," "holiday greetings," "dress-up magician," and "special scene." Scene description text is a piece of natural language text used to describe the specific context or content of the corresponding pre-saved video, such as "chatting face-to-face with a digital human," "doing crafts alone in a room," and "wearing New Year's clothes and celebrating New Year's Eve with family." Actions describe the names of the physical actions performed by the digital entities in the pre-saved video. Emotion types include, for example, happy and sad.

[0039] S12: Determine the end video frame from the currently unplayed video frames of the first video.

[0040] Step S12 is used to find the switching frame from the first video, that is, the end video frame of the first video. In one embodiment, the currently playing video frame can be directly used as the end video frame of the first video; in another embodiment, the first video can continue to play, and a suitable end video frame can be determined from the currently unplayed video frames of the first video.

[0041] Understandably, the more similar two video frames are, the smaller the difference in visual changes during the transition, and the more seamless the image appears. Based on this, a video frame with a high similarity to the first frame of the second video to be switched to can be selected from the currently unplayed video frames and used as the ending video frame. Of course, to ensure real-time performance, only a preset number of unplayed video frames closest to the currently playing video frame can be selected as the ending video frame, choosing one with a high similarity to the first frame of the second video to be switched to.

[0042] Specifically, please refer to Figure 2, which is a flowchart illustrating an embodiment of step S12 shown in Figure 1. In this embodiment, determining the end video frame from the currently unplayed video frames of the first video includes: S21: calculating the similarity between a plurality of currently unplayed video frames in the first video and the first video frame of the second video, wherein the plurality of currently unplayed video frames are a preset number of unplayed video frames in the first video that are closest to the currently played video frame.

[0043] This embodiment is used to find a video frame with a high similarity to the first frame of the second video from a preset number of unplayed video frames that are closest to the currently played video frame, and use it as the end video frame.

[0044] In one embodiment, for each currently unplayed video frame, the image similarity between the currently unplayed video frame and the first frame of the second video can be calculated, as well as the pose similarity between the first digital entity in the currently unplayed video frame and the second digital entity in the first frame of the video frame. Then, the corresponding image similarity and pose similarity are fused to obtain the similarity between the currently unplayed video frame and the first frame of the video frame. In other embodiments, only the image similarity or pose similarity can be used as the similarity between the currently unplayed video frame and the first frame of the second video. For a detailed description of how the similarity is determined, please refer to the embodiment shown in Figure 3.

[0045] Image similarity is used to characterize the degree of similarity between two video frames at the image level, while pose similarity is used to characterize the degree of similarity between the poses of digital entities in two video frames. Pose is, for example, the limb movements of digital entities, the running state of the limbs, etc.

[0046] In other embodiments, the image similarity or pose similarity described above may be used as the similarity between the currently unplayed video frame and the first video frame of the second video.

[0047] S22: Select the currently unplayed video frame whose similarity meets the similarity requirement as the end video frame.

[0048] In this embodiment, the currently unplayed video frame corresponding to the highest similarity can be selected as the end video frame. Alternatively, a similarity threshold can be preset, and the video frame among the currently unplayed video frames that is greater than the similarity threshold and is closest to the currently played video frame can be selected as the end video frame.

[0049] S13: Generate a transition video segment based on the ending video frame and the first video frame of the second video.

[0050] In one embodiment, several transition frames can be obtained by directly interpolating the end video frame and the first video frame, wherein the several transition frames constitute a transition video segment.

[0051] In another embodiment, a matching transition video generation strategy can first be selected from several transition video generation strategies based on the similarity between the ending video frame and the first video frame; then, a transition video segment can be generated using the matching transition video generation strategy. Each transition video generation strategy corresponds to a different similarity level.

[0052] In this embodiment, the method of selecting the corresponding matching transition video generation strategy by utilizing the similarity between two video frames is beneficial for selecting the most suitable transition video generation strategy based on the actual similarity situation.

[0053] In one embodiment, the several transition video generation strategies include at least one of a first transition video generation strategy, a second transition video generation strategy, and a third transition video generation strategy. The first transition video generation strategy involves interpolating the ending video frame with the first video frame to obtain several transition frames, which together form a transition video segment. The second transition video generation strategy uses a video generation algorithm to generate a transition video segment based on the ending video frame and the first video frame. The third transition video generation strategy generates at least one special effects video as the transition video segment; the at least one special effect includes at least one of the following: gradient dissolve, wipe push-pull, particle effects, video widget masking, etc.

[0054] In one specific implementation, the several transition video generation strategies include: a first transition video generation strategy, a second transition video generation strategy, and a third transition video generation strategy.

[0055] In this embodiment, step S13 selects a matching transition video generation strategy from several transition video generation strategies based on the similarity between the end video frame and the first video frame, including the following cases: Case 1: If the similarity between the end video frame and the first video frame is greater than the first similarity threshold, the first transition video generation strategy is selected as the matching transition video generation strategy.

[0056] Scenario 2: If the similarity between the ending video frame and the first video frame is greater than the second similarity threshold and less than or equal to the first similarity threshold, the second transition video generation strategy is selected as the matching transition video generation strategy.

[0057] Case 3: If the similarity between the ending video frame and the first video frame is less than or equal to the second similarity threshold, the third transition video generation strategy is selected as the matching transition video generation strategy.

[0058] That is, in this embodiment, a suitable transition video generation strategy can be dynamically selected based on the similarity between the end video frame and the first video frame to generate transition video segments.

[0059] In this context, the first similarity score is greater than the second similarity score. For example, if the first similarity score is 0.9 and the second similarity score is 0.2, then if the similarity score between the ending video frame and the first video frame (denoted as the target similarity score) is greater than 0.9, it indicates that the changes in the image between the ending and first video frames are small, and the content is highly similar (e.g., the pose of digital entities and the background remain almost unchanged). In this case, interpolation is performed between the ending and first video frames to obtain several transition frames. For example, interpolation can be performed using the optical flow information of the images from the ending and first video frames to obtain several transition frames.

[0060] If the target similarity is greater than 0.2 but less than or equal to 0.9, it indicates that the core elements such as the digital entity and the background still maintain continuity. For example, the digital entity may have slight movements or facial expressions, and the background details may have been slightly adjusted. In this case, a video generation algorithm is used to generate a transitional video segment based on the end video frame and the first video frame. For example, a video diffusion model can be used to generate a transitional video segment based on the end video frame and the first video frame.

[0061] If the target similarity is less than or equal to 0.2, it means that there is a big change in the picture between the end video frame and the first video frame. For example, the digital entity image (such as clothing, hairstyle) and scene background are completely different, or a major situational change occurs. In this case, the corresponding special effect can be called from at least one special effect to switch the video and cover up the huge difference in the picture.

[0062] S14: After the last video frame is played, a transitional video clip is played first, and then the playback switches to the second video.

[0063] In this embodiment, each video frame can be decoded and played sequentially in the order of the end video frame, the transition video segment, and the second video.

[0064] Specifically, when playing the second video, at least some video frames of the second video can be used as target video frames. The lip features of the associated speech frames of the target video frames are generated, and the facial region image of the target video frames is generated based on the lip features. The facial region image of the target video frames replaces the facial region corresponding to the second digital entity in the target video frames. The associated speech frames of the video frames are the speech frames in the speech to be played that correspond to the time of the video frames. Then, each video frame of the second video is played in sequence, and the associated speech frames corresponding to each video frame are played synchronously.

[0065] In one embodiment, the target video frame is a video frame in the second video that contains a facial region. In this embodiment, for the target video frame, the lip features of the associated speech frame corresponding to the target video frame in the speech to be played are generated. Using these lip features, a facial region image of the second digital entity is generated. Then, the generated facial region image replaces the facial region of the corresponding second digital entity in the original target video frame. For non-target video frames that do not contain a facial region, since they do not contain a facial region, there is no need to perform facial region replacement processing based on the speech. The original video frame and the corresponding associated speech can be used directly for synchronized playback.

[0066] It should be noted that this embodiment can achieve synchronization of the lips and voice of the second digital entity in the second video. For specific implementation details, please refer to the relevant description of the embodiment shown in Figure 4 below.

[0067] The above solution, during the playback of the first video, if a switch to the second video is detected, determines the ending video frame from the currently unplayed frames of the first video, and uses this ending video frame and the first frame of the second video to generate a transition video segment. After playing the ending video frame, this transition video segment is played first, and then the playback of the second video is switched to. Compared to the method of directly jumping from the first video to the second video when a switch is detected, the playback method described in this application, which uses a generated transition video segment to connect the first and second videos, can smoothly transition the differences between the ending video frame and the first video frame, thereby improving the continuity of video playback.

[0068] Please refer to Figure 3, which is a flowchart illustrating an embodiment of step S21 shown in Figure 2. This embodiment includes: S31: For each currently unplayed video frame in a plurality of currently unplayed video frames in the first video, calculate the image similarity between the currently unplayed video frame and the first video frame of the second video, and calculate the pose similarity between the first digital entity in the currently unplayed video frame and the second digital entity in the first video frame.

[0069] In one embodiment, for each currently unplayed video frame, the first image feature of the currently unplayed video frame and the second image feature of the first video frame can be extracted first, and then the similarity between the first image feature and the second image feature can be calculated, and the similarity can be used as the image similarity.

[0070] The first and second image features can be extracted using a trained image recognition model. For example, they can be image features extracted from the layer preceding the classification layer of the image recognition model.

[0071] Image similarity can be the cosine similarity between the first image feature and the second image feature, or it can be the Euclidean distance between the first image feature and the second image feature, etc.

[0072] In one embodiment, for each currently unplayed video frame, the first pose feature of the first digital entity in the currently unplayed video frame and the second pose feature of the second digital entity in the first video frame can be detected by the pose detection model; then the similarity between the first pose feature and the second pose feature is calculated as the pose similarity.

[0073] S32: Combine the image similarity and pose similarity between the current unplayed video frame and the first video frame to obtain the similarity between the current unplayed video frame and the first video frame.

[0074] In one embodiment, the image similarity and pose similarity between the current unplayed video frame and the first video frame, along with the sum of a first value, are obtained. The ratio of this sum to a second value is used as the similarity between the current unplayed video frame and the first video frame. Both the second and first values ​​are positive integers, and the second value is greater than the first value. Furthermore, the first and second values ​​are determined based on a range of similarity.

[0075] For example, refer to the following formula:

[0076] In the formula, S1 represents image similarity and S2 represents pose similarity. The first value is 2 and the second value is 4. This is because the range of each similarity is [-1,1], and this formula can normalize the final S to [0,1].

[0077] Please refer to Figure 4, which is a schematic flowchart of an embodiment of playing a second video provided by this application. In this embodiment, playing the second video includes: S41: taking at least a portion of the video frames of the second video as target video frames, generating lip features of the associated speech frames of the target video frames, generating a facial region image of the target video frames based on the lip features, and replacing the facial region corresponding to the second digital entity in the target video frames with the facial region image of the target video frames, wherein the associated speech frame of the video frame is the speech frame in the speech to be played that corresponds to the time of the video frame.

[0078] S42: Play each video frame of the second video in sequence, and simultaneously play the corresponding audio frames of each video frame.

[0079] It should be noted that, in order to make the second digital entity in the second video appear more natural, the lip movements and speech of the digital entity need to be accurately synchronized.

[0080] In one embodiment, to achieve precise synchronization between the lip movements of the second digital entity and the speech, face detection can be performed on each video frame in the second video to determine whether a face can be detected in each video frame, and the video frame in which the face of the second digital entity is detected is taken as the target video frame; and for each target video frame, the speech frame in the speech to be played that corresponds to the time of the video frame is taken as the associated speech frame of the target video frame, and the lip features (e.g., lip movement features) of the associated speech frame are generated, and then the facial region image (including at least the lip region) of the target video frame is generated based on the lip features, and the facial region image generated replaces the facial region corresponding to the second digital entity in the target video frame, thereby achieving precise matching between the target video frame and the corresponding speech frame.

[0081] Specifically, based on timeline alignment rules, audio frames corresponding to the time of each target video frame can be extracted from the frame sequence of the audio to be played, and these frames can be identified as the associated audio frames of the corresponding target video frames.

[0082] Furthermore, for video frames in the second video where the second digital entity's face is not detected, since there is no need to present lip movements, the original video frames can be used directly without performing operations such as facial region image generation and facial region image replacement.

[0083] In simple terms, face detection can be performed on each video frame of the second video. If the face of the second digital entity is detected, the facial region image of the second digital entity is generated using the lip features of its associated speech frame, and the facial region image is replaced to achieve precise driving of the second digital entity's face and lip movements. If the face of the second digital entity is not detected, the original video frame is used directly without performing operations such as facial region image generation and facial region image replacement. When playing each video frame of the second video in sequence, the associated speech frame corresponding to each video frame is played synchronously, ultimately achieving precise synchronization between the second digital entity's lip movements and speech.

[0084] In another embodiment, to further enhance the naturalness and vividness of the second digital entity's facial presentation, emotion prediction can be performed based on the text to be played corresponding to the speech to be played, obtaining emotion parameters (such as feature values ​​of dimensions like joy, calmness, and seriousness), and a facial region image containing both lip movements and emotional expressions can be generated based on the lip features and emotion parameters corresponding to the target video frame; the generated facial region image replaces the original facial region of the second digital entity in the target video frame, ultimately achieving coordinated real-time driving of the second digital entity's face, lip shape, and expression.

[0085] In one specific embodiment, the emotion parameter is obtained by predicting the emotion of the text to be played using a relevant model (e.g., an NLP model), and the facial region image of the target video frame is generated using a trained facial generation model. The facial generation model can be, but is not limited to, wav2lip, and can also be GAN, SyncNet, etc.

[0086] In one implementation scenario, at least one of the emotion parameters and lip features can be predicted using a facial generation model; furthermore, after obtaining the emotion parameters and lip features, a corresponding facial region image containing the mouth and expression area can be generated using the emotion parameters and lip features.

[0087] Please refer to Figure 5, which is a schematic diagram of a framework of an embodiment of the video playback device provided in this application. In this embodiment, the video playback device 50 includes: a detection module 51, a determination module 52, a generation module 53, and a playback module 54. The detection module 51 is used to detect that a switch to a second video containing a second digital entity is needed during the playback of a first video containing a first digital entity; the determination module 52 is used to determine the end video frame from the currently unplayed video frames of the first video; the generation module 53 is used to generate a transition video segment based on the end video frame and the first video frame of the second video; the playback module 54 is used to play the transition video segment first after playing the end video frame, and then switch to the second video for playback.

[0088] In some embodiments, the determining module 52 determines the end video frame from the currently unplayed video frames of the first video, including: calculating the similarity between a plurality of currently unplayed video frames in the first video and the first video frame of the second video, wherein the plurality of currently unplayed video frames are a preset number of unplayed video frames in the first video that are closest to the currently played video frame; and selecting the currently unplayed video frame whose similarity meets the similarity requirement as the end video frame.

[0089] In some embodiments, calculating the similarity between a plurality of currently unplayed video frames in the first video and the first video frame of the second video includes: for each currently unplayed video frame in the plurality of currently unplayed video frames in the first video, calculating the image similarity between the currently unplayed video frame and the first video frame of the second video, and calculating the pose similarity between the first digital entity in the currently unplayed video frame and the second digital entity in the first video frame; fusing the image similarity and pose similarity between the currently unplayed video frame and the first video frame to obtain the similarity between the currently unplayed video frame and the first video frame.

[0090] In some embodiments, calculating the image similarity between the currently unplayed video frame and the first video frame of the second video includes: extracting a first image feature of the currently unplayed video frame and a second image feature of the first video frame, and calculating the similarity between the first image feature and the second image feature as the image similarity; and / or, calculating the pose similarity between a first digital entity in the currently unplayed video frame and a second digital entity in the first video frame includes: using a pose detection model to detect the first pose feature of the first digital entity in the currently unplayed video frame and the second pose feature of the first video frame, respectively. The second pose feature of the digital entity; calculating the similarity between the first pose feature and the second pose feature as the pose similarity; and / or, fusing the image similarity and pose similarity between the current unplayed video frame and the first video frame to obtain the similarity between the current unplayed video frame and the first video frame, including: obtaining the sum of the image similarity and pose similarity between the current unplayed video frame and the first video frame, and a first value, and using the ratio of the sum to a second value as the similarity between the current unplayed video frame and the first video frame, wherein the second value and the first value are positive integers, and the second value is greater than the first value.

[0091] In some embodiments, the generation module 53 generates a transition video segment based on the end video frame and the first video frame of the second video, including: selecting a matching transition video generation strategy from several transition video generation strategies based on the similarity between the end video frame and the first video frame, wherein each transition video generation strategy corresponds to a different similarity; and generating a transition video segment using the matching transition video generation strategy.

[0092] In some embodiments, the plurality of transition video generation strategies include at least one of a first transition video generation strategy, a second transition video generation strategy, and a third transition video generation strategy; the first transition video generation strategy involves interpolating the end video frame with the first video frame to obtain a plurality of transition frames, and the plurality of transition frames constitute a transition video segment; the second transition video generation strategy involves generating a transition video segment based on the end video frame and the first video frame using a video generation algorithm; the third transition video generation strategy involves generating at least one special effects video as a transition video segment; and / or, the plurality of transition video generation strategies include the first transition video generation strategy, the second transition video generation strategy, and the third transition video generation strategy; based on the structure The similarity between the end video frame and the first video frame is used to select a matching transition video generation strategy from several transition video generation strategies, including: selecting a first transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being greater than a first similarity threshold; selecting a second transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being greater than a second similarity threshold and less than or equal to the first similarity threshold; and selecting a third transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being less than or equal to the second similarity threshold.

[0093] In some embodiments, during the playback of a first video containing a first digital entity, the detection module 51 detects that a switch to a second video containing a second digital entity is required, including: generating audio to be played during the playback of the first video containing the first digital entity; searching a pre-stored video in a video library that matches the audio scene to be played, wherein the video library includes a plurality of pre-stored videos containing digital entities; in response to finding a pre-stored video that matches the audio scene to be played, using the matching pre-stored video as the second video, and determining that a switch to the second video is required; and / or, in response to not finding a pre-stored video that matches the audio scene to be played, determining that no such video is currently available. The user needs to switch to the second video, or select a preset interactive video as the second video and confirm that the user needs to switch to the second video. The playback steps of the second video include: taking at least some video frames of the second video as target video frames, generating lip features of the associated speech frames of the target video frames, generating facial region images of the target video frames based on the lip features, and replacing the facial region corresponding to the second digital entity in the target video frames with the facial region images of the target video frames. The associated speech frames of the video frames are the speech frames in the speech to be played that correspond to the time of the video frames. The user then plays each video frame of the second video in sequence and plays the associated speech frames corresponding to each video frame synchronously.

[0094] In some embodiments, generating the speech to be played includes: generating text to be played using a large model based on user interaction information; converting the text to be played into speech; and / or, searching a video library for a pre-stored video that matches the scene of the speech to be played, including: calculating the text similarity between the scene description text of each pre-stored video in the video library and the text to be played corresponding to the speech; determining the pre-stored video whose text similarity meets the scene similarity requirement as a pre-stored video that matches the scene of the speech to be played; and / or, the second video playback step further includes: performing emotion prediction based on the text to be played corresponding to the speech to be played to obtain emotion parameters; and generating a facial region image of the target video frame based on lip features, including: generating a facial region image of the target video frame based on the lip features and emotion parameters corresponding to the target video frame; and / or, using at least some video frames of the second video as target video frames, including: using each video frame of the second video containing a facial region as a target video frame.

[0095] In some embodiments, at least one of the first video and the second video is a pre-stored video in a video library, which includes a plurality of pre-stored videos containing data entities; the step of generating the pre-stored video includes: obtaining digital entity reference data provided by the user, and generating a video containing digital entities based on the digital entity reference data using a video diffusion model, as a pre-stored video; the digital entity reference data includes at least one of the following: at least one digital entity image reference image, a reference video containing digital entities or real entities, and descriptive text about digital entities; or, obtaining a video taken of a real entity, as a pre-stored video.

[0096] Please refer to Figure 6, which is a schematic diagram of the framework of an embodiment of the electronic device provided in this application. In this embodiment, the electronic device 60 includes a memory 61 and a processor 62 coupled to each other.

[0097] The memory 61 stores program instructions, and the processor 62 executes the program instructions stored in the memory 61 to implement the steps of any of the above-described method implementations. In a specific implementation scenario, the electronic device 60 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 60 may also include mobile devices such as laptops and tablets, which are not limited here.

[0098] Specifically, processor 62 controls itself and memory 61 to implement the steps of any of the above embodiments. Processor 62 may also be referred to as a CPU (Central Processing Unit). Processor 62 may be an integrated circuit chip with signal processing capabilities. Processor 62 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 62 may be implemented using integrated circuit chips.

[0099] Please refer to Figure 7, which is a schematic diagram of the framework of the computer-readable storage medium provided in this application. The computer-readable storage medium 70 of this application embodiment stores program instructions 71, which, when executed, implement the methods provided in any embodiment and any non-conflicting combination of the above methods. The program instructions 71 can form a program file and be stored in the computer-readable storage medium 70 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 70 includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.

[0100] The above solution, during the playback of the first video, if a switch to the second video is detected, determines the ending video frame from the currently unplayed frames of the first video, and uses this ending video frame and the first frame of the second video to generate a transition video segment. After playing the ending video frame, this transition video segment is played first, and then the playback of the second video is switched to. Compared to the method of directly jumping from the first video to the second video when a switch is detected, the playback method described in this application, which uses a generated transition video segment to connect the first and second videos, can smoothly transition the differences between the ending video frame and the first video frame, thereby improving the continuity of video playback.

[0101] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0102] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0103] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A video playback method, characterized in that, The method includes: during the playback of a first video containing a first digital entity, detecting that a switch to a second video containing a second digital entity is required; determining an end video frame from the currently unplayed video frames of the first video; generating a transition video segment based on the end video frame and the first video frame of the second video; and after playing the end video frame, playing the transition video segment first, and then switching to the second video for playback.

2. The method according to claim 1, characterized in that, The step of determining the end video frame from the currently unplayed video frames of the first video includes: calculating the similarity between a plurality of currently unplayed video frames in the first video and the first video frame of the second video, wherein the plurality of currently unplayed video frames are a preset number of unplayed video frames in the first video that are closest to the currently played video frame; and selecting the currently unplayed video frames whose similarity meets the similarity requirement as the end video frame.

3. The method according to claim 2, characterized in that, The step of calculating the similarity between several currently unplayed video frames in the first video and the first video frame of the second video includes: for each currently unplayed video frame in the first video, calculating the image similarity between the currently unplayed video frame and the first video frame of the second video, and calculating the pose similarity between a first digital entity in the currently unplayed video frame and a second digital entity in the first video frame; and fusing the image similarity and the pose similarity between the currently unplayed video frame and the first video frame to obtain the similarity between the currently unplayed video frame and the first video frame.

4. The method according to claim 3, characterized in that, The step of calculating the image similarity between the currently unplayed video frame and the first frame of the second video includes: extracting a first image feature of the currently unplayed video frame and a second image feature of the first frame, and calculating the similarity between the first image feature and the second image feature as the image similarity; and / or, the step of calculating the pose similarity between the first digital entity in the currently unplayed video frame and the second digital entity in the first frame includes: using a pose detection model to detect the first pose feature of the first digital entity in the currently unplayed video frame and the second pose feature of the second digital entity in the first frame. ; calculate the similarity between the first pose feature and the second pose feature as the pose similarity; and / or, the step of fusing the image similarity between the currently unplayed video frame and the first video frame and the pose similarity to obtain the similarity between the currently unplayed video frame and the first video frame includes: obtaining the sum of the image similarity between the currently unplayed video frame and the first video frame, the pose similarity, and a first value, and using the ratio of the sum to a second value as the similarity between the currently unplayed video frame and the first video frame, wherein the second value and the first value are positive integers, and the second value is greater than the first value.

5. The method according to claim 1, characterized in that, The step of generating a transition video segment based on the ending video frame and the first video frame of the second video includes: selecting a matching transition video generation strategy from a plurality of transition video generation strategies based on the similarity between the ending video frame and the first video frame, wherein each transition video generation strategy corresponds to a different similarity; and generating the transition video segment using the matching transition video generation strategy.

6. The method according to claim 5, characterized in that, The plurality of transition video generation strategies includes at least one of a first transition video generation strategy, a second transition video generation strategy, and a third transition video generation strategy; the first transition video generation strategy involves interpolating the end video frame with the first video frame to obtain a plurality of transition frames, which together form the transition video segment; the second transition video generation strategy involves generating the transition video segment based on the end video frame and the first video frame using a video generation algorithm; the third transition video generation strategy involves generating at least one special effects video as the transition video segment; and / or, the plurality of transition video generation strategies includes a first transition video generation strategy, a second transition video generation strategy, and a third transition video generation strategy; the transition video segment based on the end video frame... The similarity between the end video frame and the first video frame is used to select a matching transition video generation strategy from several transition video generation strategies, including: selecting the first transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being greater than a first similarity threshold; selecting the second transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being greater than a second similarity threshold and less than or equal to the first similarity threshold; and selecting the third transition video generation strategy as the matching transition video generation strategy in response to the similarity between the end video frame and the first video frame being less than or equal to the second similarity threshold.

7. The method according to claim 1, characterized in that, The step of detecting the need to switch to a second video containing a second digital entity during the playback of a first video containing a first digital entity includes: generating audio to be played during the playback of the first video containing the first digital entity; searching a pre-stored video in a video library that matches the audio scene to be played, wherein the video library includes a plurality of pre-stored videos containing digital entities; in response to finding a pre-stored video that matches the audio scene to be played, using the matching pre-stored video as the second video, and determining that the current switch to the second video is required; and / or, in response to not finding a pre-stored video that matches the audio scene to be played, determining that the current switch to the second video is not required, or selecting a preset interactive video as the second video, and determining that the current switch to the second video is required.

8. The method according to claim 7, characterized in that, The generation of the speech to be played includes: generating text to be played using a large model based on user interaction information; converting the text to be played into the speech to be played; and / or, searching for pre-stored videos in the video library that match the scene of the speech to be played includes: calculating the text similarity between the scene description text of each pre-stored video in the video library and the text to be played corresponding to the speech to be played; determining the pre-stored videos whose text similarity meets the scene similarity requirement as pre-stored videos that match the scene of the speech to be played; and / or, the playback step of the second video further includes: performing emotion prediction based on the text to be played corresponding to the speech to be played to obtain emotion parameters; and, the generation of the facial region image of the target video frame based on the lip features includes: generating the facial region image of the target video frame based on the lip features and the emotion parameters corresponding to the target video frame; and / or, using at least some video frames of the second video as target video frames includes: using each video frame of the second video that contains a facial region as the target video frame.

9. The method according to claim 1, characterized in that, At least one of the first video and the second video is a pre-stored video in a video library, which includes several pre-stored videos containing data entities. The step of generating the pre-stored video includes: obtaining digital entity reference data provided by the user, and generating a video containing digital entities based on the digital entity reference data using a video diffusion model, as the pre-stored video. The digital entity reference data includes at least one of the following: at least one digital entity image reference image, a reference video containing digital entities or real entities, and descriptive text about digital entities; or, obtaining a video of a real entity as the pre-stored video.

10. A video playback device, characterized in that, The device includes: a detection module, configured to detect, during the playback of a first video containing a first digital entity, the need to switch to a second video containing a second digital entity; a determination module, configured to determine an end video frame from the currently unplayed video frames of the first video; a generation module, configured to generate a transition video segment based on the end video frame and the first video frame of the second video; and a playback module, configured to play the transition video segment first after playing the end video frame, and then switch to the second video for playback.