Video generation methods, apparatus, electronic devices and storage media

By identifying the ending frame of the target scene in a long video and generating a second image frame sequence, the problem of audio not finishing playing or images not finishing displaying in short videos is solved, thus improving the user's viewing experience.

CN116781992BActive Publication Date: 2026-08-04BEIJING IQIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING IQIYI TECH CO LTD
Filing Date
2023-06-27
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies often result in issues such as incomplete audio playback or incomplete image display when extracting short videos from long videos, leading to a decline in the user's viewing experience.

Method used

The final frame of the target scene is determined from the image frame sequence of the target video, a second image frame sequence is generated, and a video corresponding to the target scene is generated based on this sequence and audio. Necessary image frames are inserted to ensure the integrity of the audio and images.

Benefits of technology

It effectively reduces or avoids frame skipping in short videos, improving the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116781992B_ABST
    Figure CN116781992B_ABST
Patent Text Reader

Abstract

This application relates to a video generation method, apparatus, electronic device, and storage medium. The method includes: determining an image frame corresponding to the termination scene of a target scene from a first image frame sequence corresponding to a target video, thereby obtaining a fifth image frame; wherein the starting audio of the target video corresponds to the starting scene of the target scene; determining a second image frame sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between; and generating a video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video. Therefore, a video corresponding to a target scene can be generated based on the second image frame sequence and the audio corresponding to the target video, thus reducing or even avoiding frame skipping in the generated video corresponding to the target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a video generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] In existing technologies, when extracting a short segment from a long video to showcase its highlights, a high level of watchability is often required. An example application is extracting a short segment from a long video as a recap.

[0003] Currently, some long videos exhibit instances where the audio is incomplete during transitions. If the video is segmented based solely on the audio, the image from the next scene will be included in the short video clip, a problem referred to as a "frame clipping" at the end of the short video. If the segmentation is based solely on the image position corresponding to the transition point, the audio will still be incomplete even after the visuals have finished playing—meaning the dialogue is only partially completed. This degrades the user's viewing experience. If this problem remains unresolved, both audio-based and transition-point-based segmentation methods will negatively impact the user's viewing experience.

[0004] It is evident that reducing or even avoiding frame skipping in the generated target scene video is a technical issue worthy of attention. Summary of the Invention

[0005] In view of this, in order to solve some or all of the above-mentioned technical problems, embodiments of this application provide a video generation method, apparatus, electronic device and storage medium.

[0006] In a first aspect, embodiments of this application provide a video generation method, the method comprising: From the first image frame sequence corresponding to the target video, determine the image frame corresponding to the ending screen of the target scene to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; The sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between is determined as the second image frame sequence; Based on the second image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0007] In one possible implementation, generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video includes: The difference between the first quantity and the second quantity is determined to obtain the third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence; Based on the third quantity, the second image frame sequence, and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0008] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the end of the second image frame sequence, select an unselected image frame to obtain the sixth image frame; Based on the sixth image frame, perform the following replacement steps: The sixth image frame is copied a fourth number of times to obtain the fourth number of sixth image frames, wherein the fourth number is greater than 1; The fourth number of sixth image frames that are located at the end of the first image frame sequence and have not been replaced are replaced with the fourth number of sixth image frames; Determine whether the seventh image frame and the eighth image frame represent the same image frame, wherein the seventh image frame is the image frame preceding the sixth image frame in the second image frame sequence, and the eighth image frame is the last image frame in the first image frame sequence that has not been replaced; If the seventh and eighth image frames do not represent the same image frame, an unselected image frame is selected from the end of the second image frame sequence to obtain the sixth image frame; the replacement step is then performed based on the sixth image frame. When the seventh and eighth image frames represent the same image frame, the first image frame sequence obtained after replacement is determined as the third image frame sequence; based on the third image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated.

[0009] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the second image frame sequence, determine the third number of image frames to be copied; The third number of image frames are copied to obtain the third number of copied image frames; The third number of copied image frames are inserted into the second image frame sequence to obtain a fourth image frame sequence, so that the duration of the fourth image frame sequence is equal to the duration of the target video; Based on the fourth image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0010] In one possible implementation, determining the third number of image frames to be copied from the second image frame sequence includes: From the second image frame sequence, the third number of image frames are determined in reverse order to obtain the third number of image frames to be copied.

[0011] In one possible implementation, inserting the third number of copied image frames into the second image frame sequence includes: For each of the third number of copied image frames, the copied image frame is inserted into a target position, wherein the target position is the adjacent position of an image frame in the second image frame sequence that is the same as the copied image frame.

[0012] In one possible implementation, after generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video, the method further includes: Control the video corresponding to the target scene to play at a preset speed, wherein the preset speed represents a value greater than 1.

[0013] In one possible implementation, the audio corresponding to the starting frame of the target scene and the image frame corresponding to the starting frame of the target scene are located at the same position in the target video.

[0014] Secondly, embodiments of this application provide a video generation apparatus, the apparatus comprising: The first determining unit is configured to determine the image frame corresponding to the ending screen of the target scene from the first image frame sequence corresponding to the target video, thereby obtaining the fifth image frame; wherein the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene. The second determining unit is used to determine the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames between them as the second image frame sequence; The generation unit is used to generate a video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video.

[0015] In one possible implementation, generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video includes: The difference between the first quantity and the second quantity is determined to obtain the third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence; Based on the third quantity, the second image frame sequence, and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0016] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the end of the second image frame sequence, select an unselected image frame to obtain the sixth image frame; Based on the sixth image frame, perform the following replacement steps: The sixth image frame is copied a fourth number of times to obtain the fourth number of sixth image frames, wherein the fourth number is greater than 1; The fourth number of sixth image frames that are located at the end of the first image frame sequence and have not been replaced are replaced with the fourth number of sixth image frames; Determine whether the seventh image frame and the eighth image frame represent the same image frame, wherein the seventh image frame is the image frame preceding the sixth image frame in the second image frame sequence, and the eighth image frame is the last image frame in the first image frame sequence that has not been replaced; If the seventh and eighth image frames do not represent the same image frame, an unselected image frame is selected from the end of the second image frame sequence to obtain the sixth image frame; the replacement step is then performed based on the sixth image frame. When the seventh and eighth image frames represent the same image frame, the first image frame sequence obtained after replacement is determined as the third image frame sequence; based on the third image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated.

[0017] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the second image frame sequence, determine the third number of image frames to be copied; The third number of image frames are copied to obtain the third number of copied image frames; The third number of copied image frames are inserted into the second image frame sequence to obtain a fourth image frame sequence, so that the duration of the fourth image frame sequence is equal to the duration of the target video; Based on the fourth image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0018] In one possible implementation, determining the third number of image frames to be copied from the second image frame sequence includes: From the second image frame sequence, the third number of image frames are determined in reverse order to obtain the third number of image frames to be copied.

[0019] In one possible implementation, inserting the third number of copied image frames into the second image frame sequence includes: For each of the third number of copied image frames, the copied image frame is inserted into a target position, wherein the target position is the adjacent position of an image frame in the second image frame sequence that is the same as the copied image frame.

[0020] In one possible implementation, after generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video, the apparatus further includes: The control unit is used to control the video corresponding to the target scene to play at a preset speed, wherein the preset speed represents a value greater than 1.

[0021] In one possible implementation, the audio corresponding to the starting frame of the target scene and the image frame corresponding to the starting frame of the target scene are located at the same position in the target video.

[0022] Thirdly, embodiments of this application provide an electronic device, including: Memory, used to store computer programs; A processor is configured to execute a computer program stored in the memory, wherein, when the computer program is executed, it implements the method of any embodiment of the video generation method of the first aspect of this application.

[0023] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the method of any embodiment of the video generation method of the first aspect described above.

[0024] Fifthly, embodiments of this application provide a computer program comprising computer-readable code that, when executed on a device, causes a processor in the device to implement the method of any embodiment of the video generation method of the first aspect described above.

[0025] The video generation method provided in this application can determine the image frame corresponding to the ending scene of the target scene from the first image frame sequence corresponding to the target video, and obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting scene of the target scene; the ending audio of the target video corresponds to the ending scene of the target scene; then, the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between are determined as the second image frame sequence; then, based on the second image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated. Therefore, the video corresponding to the target scene can be generated based on the second image frame sequence and the audio corresponding to the target video, thus reducing or even avoiding the problem of frame skipping in the generated video corresponding to the target scene. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0029] Figure 1 A flowchart illustrating a video generation method provided in an embodiment of this application; Figure 2 A flowchart illustrating another video generation method provided in an embodiment of this application; Figure 3A A schematic diagram of a long video in a video generation method provided in an embodiment of this application; Figure 3B A schematic diagram of the target video in a video generation method provided in an embodiment of this application; Figure 3C This is a schematic diagram illustrating the generation of a video corresponding to a target scene in a video generation method provided in an embodiment of this application. Figure 4 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application; Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0030] Various exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this application.

[0031] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of this application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they indicate the logical order between them.

[0032] It should also be understood that in this embodiment, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.

[0033] It should also be understood that any component, data or structure mentioned in the embodiments of this application can generally be understood as one or more unless explicitly defined or given contrary guidance in the context.

[0034] Furthermore, the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.

[0035] It should also be understood that the description of the various embodiments in this application emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0036] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.

[0037] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0038] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0039] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. To facilitate understanding of the embodiments of this application, the application will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0040] To address the technical problem of reducing or even avoiding frame-interleaving in the generated video corresponding to the target scene, this application provides a video generation method that can reduce or even avoid frame-interleaving in the generated video corresponding to the target scene.

[0041] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of this application. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are imposed here.

[0042] like Figure 1 As shown, the method specifically includes: Step 101: Determine the image frame corresponding to the ending screen of the target scene from the first image frame sequence corresponding to the target video to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene.

[0043] In this embodiment, the first image frame sequence corresponding to the target video can be a sequence consisting of the image frames that make up the target video, arranged in the order in which they are played.

[0044] The fifth image frame can be the image frame representing the final screen of the target scene.

[0045] The starting audio of the target video can be the first frame of audio included in the target video.

[0046] The starting frame of the target video can be the first image frame among the various image frames included in the target video.

[0047] The ending audio of the target video can be the last frame of audio included in the target video.

[0048] The audio corresponding to the target scene can be audio from one or more complete scenes.

[0049] The visuals corresponding to the target scene (i.e., the first image frame sequence corresponding to the target video) depend on the audio corresponding to the target scene. Typically, the target video does not correspond to one or more complete scenes.

[0050] The playback duration of the audio corresponding to the target scene is equal to the playback duration of the first image frame sequence. In other words, the audio frames corresponding to the target scene correspond one-to-one with the image frames in the first image frame sequence.

[0051] Step 102: Determine the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between as the second image frame sequence.

[0052] In this embodiment, if the first image frame sequence corresponding to the target video is "image frame 1, image frame 2, image frame 3, image frame 4, image frame 5", and the fifth image frame is "image frame 3", then the second image frame sequence "image frame 1, image frame 2, image frame 3" can be obtained.

[0053] Step 103: Generate a video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video.

[0054] In this embodiment, the scene corresponding to the second image frame sequence can be the same as the scene corresponding to the audio of the target video. Both can correspond to one or more complete scenes.

[0055] The target image may include: audio corresponding to the target video and a first image frame sequence corresponding to the target video. Furthermore, the playback duration of the audio corresponding to the target video is equal to the playback duration of the first image frame sequence corresponding to the target video.

[0056] As an example, the audio corresponding to the target video can be composed of the following audio frames: "Audio Frame 1, Audio Frame 2, Audio Frame 3, Audio Frame 4, Audio Frame 5, Audio Frame 6". The first image frame sequence corresponding to the target video can include "Image Frame 1, Image Frame 2, Image Frame 3, Image Frame 4, Image Frame 5, Image Frame 6". If the fifth image frame is "Image Frame 3", then the second image frame sequence "Image Frame 1, Image Frame 2, Image Frame 3" can be obtained. In the above example, a video corresponding to the target scene can be generated based on the audio corresponding to the target video ("Audio Frame 1, Audio Frame 2, Audio Frame 3, Audio Frame 4, Audio Frame 5, Audio Frame 6") and the second image frame sequence ("Image Frame 1, Image Frame 2, Image Frame 3").

[0057] As an example, one or more image frames can be inserted into the second image frame sequence so that the playback duration of the image frame sequence after the insertion is equal to the playback duration of the audio corresponding to the target video, thereby obtaining the video corresponding to the target scene.

[0058] As another example, one or more audio frames in the audio corresponding to the target video can be compressed so that the playback duration of the compressed audio frame is equal to the playback duration of the second image frame sequence, thereby obtaining the video corresponding to the target scene.

[0059] In practice, when shortening a long video to create a shorter segment, audio may be interrupted during transitions. If the video is cut based on the audio, the next scene's image will be included in the shorter video clip; this is referred to as a "frame-clamping" issue at the end of the short video. The method described above, which generates the video corresponding to the target scene based on a second image frame sequence and the corresponding audio from the target video, can resolve this frame-clamping problem.

[0060] In some optional implementations of this embodiment, after performing step 103 above, the video corresponding to the target scene can also be controlled to play at a preset speed.

[0061] The preset speed multiplier represents a value greater than 1.

[0062] It is understood that, in the above optional implementation methods, the video obtained by inserting one or more image frames into the second image frame sequence can reduce the impact on the user's viewing experience during video playback if it is controlled to play at a preset speed.

[0063] In some optional implementations of this embodiment, the audio corresponding to the starting screen of the target scene and the image frame corresponding to the starting screen of the target scene are located at the same position in the target video.

[0064] It is understandable that, under normal circumstances, the frame-clamping problem described above only exists at the end of the captured short video, while it does not exist at the beginning of the short video, i.e., the beginning of the scene. Therefore, the above-mentioned optional implementation methods can avoid the presence of frame-clamping problems in the generated video corresponding to the target scene, thereby further improving the user's viewing experience.

[0065] The video generation method provided in this application can determine the image frame corresponding to the ending scene of the target scene from the first image frame sequence corresponding to the target video, and obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting scene of the target scene; the ending audio of the target video corresponds to the ending scene of the target scene; then, the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between are determined as the second image frame sequence; then, based on the second image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated. Therefore, the video corresponding to the target scene can be generated based on the second image frame sequence and the audio corresponding to the target video, thus reducing or even avoiding the problem of frame skipping in the generated video corresponding to the target scene.

[0066] Figure 2 This is a flowchart illustrating another video generation method provided in an embodiment of this application. Figure 2 As shown, the method specifically includes: Step 201: Determine the image frame corresponding to the ending screen of the target scene from the first image frame sequence corresponding to the target video to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene.

[0067] In this embodiment, step 201 and Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.

[0068] Step 202: The sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between is determined as the second image frame sequence.

[0069] In this embodiment, step 202 and Figure 1 Step 102 in the corresponding embodiment is basically the same, and will not be repeated here.

[0070] Step 203: Determine the difference between the first quantity and the second quantity to obtain the third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence.

[0071] In this embodiment, if the first image frame sequence is "image frame 1, image frame 2, image frame 3, image frame 4, image frame 5" and the second image frame sequence is "image frame 1, image frame 2, image frame 3", then the first quantity is 5, the second quantity is 3, and the third quantity is 2. Step 204: Based on the third quantity, the second image frame sequence, and the audio corresponding to the target video, generate the video corresponding to the target scene.

[0072] In this embodiment, the time difference between the playback duration of the first image frame sequence and the playback duration of the second image frame sequence can be determined based on a third quantity. Thus, the number of image frames to be inserted into the second image frame sequence can be further determined, and the corresponding number of image frames can be inserted into the second image frame sequence. Subsequently, the video corresponding to the target scene is generated by synthesizing the second image frame sequence obtained after insertion and the audio corresponding to the target video.

[0073] As an example, the image frame inserted into the second image frame sequence can be a preset image frame or an image frame selected from the second image frame sequence.

[0074] In some optional implementations of this embodiment, the video corresponding to the target scene can be generated based on the third quantity, the second image frame sequence, and the audio corresponding to the target video in the following manner: The first step is to select an unselected image frame from the end of the second image frame sequence to obtain the sixth image frame.

[0075] The sixth image frame may be an image frame selected from the end of the second image frame sequence.

[0076] Here, when performing this step for the first time, selecting an image frame from the end of the second image frame sequence, the last image frame of the second image frame sequence can be selected. When performing this step for the second time, selecting an image frame from the end of the second image frame sequence, the second-to-last image frame of the second image frame sequence can be selected. And so on.

[0077] The sixth image frame can be updated after each selection. Accordingly, when the first step above is executed for the first time, the sixth image frame can be the last image frame in the second image frame sequence. When the first step above is executed for the second time, the sixth image frame can be the second to last image frame in the second image frame sequence. And so on.

[0078] The second step involves performing the following replacement steps based on the sixth image frame (including sub-steps one through three): Sub-step one: Copy the sixth image frame a fourth number of times to obtain the fourth number of sixth image frames.

[0079] Wherein, the fourth quantity is greater than 1. As an example, the fourth quantity could be 2.

[0080] Sub-step two: Replace the fourth number of unreplaced image frames located at the end of the first image frame sequence with the fourth number of sixth image frames.

[0081] Specifically, taking a fourth quantity of 2 as an example, we will illustrate this exemplarily. If it is the first replacement, then two sixth image frames can be used to replace the two image frames located at the end of the first image frame sequence. If it is the second replacement, then two sixth image frames can be used to replace the third-to-last and fourth-to-last image frames located at the end of the first image frame sequence. And so on.

[0082] Sub-step three: Determine whether the seventh image frame and the eighth image frame represent the same image frame.

[0083] The seventh image frame is the image frame preceding the sixth image frame in the second image frame sequence. The eighth image frame is the last image frame in the first image frame sequence that has not been replaced.

[0084] The third step involves selecting an unselected image frame from the end of the second image frame sequence to obtain the sixth image frame, and then performing the replacement step based on the sixth image frame.

[0085] Fourth step: In the case that the seventh and eighth image frames represent the same image frame, the first image frame sequence obtained after replacement is determined as the third image frame sequence, and the video corresponding to the target scene is generated based on the third image frame sequence and the audio corresponding to the target video.

[0086] The third image frame sequence can be the first image frame sequence obtained after replacement.

[0087] Here, a third image frame sequence and the audio corresponding to the target video can be synthesized to generate a video corresponding to the target scene.

[0088] It is understood that in the above optional implementation, the image frame located at the end of the first image frame sequence can be replaced. The image frame being replaced is usually not part of the target scene, or is part of the later part of the target scene. Therefore, by replacing the image frame, the impact on the user's viewing experience during video playback can be further reduced.

[0089] In some optional implementations of this embodiment, the video corresponding to the target scene can be generated based on the third quantity, the second image frame sequence, and the audio corresponding to the target video in the following manner: First, the third number of image frames to be copied are determined from the second image frame sequence.

[0090] As an example, the third number of image frames to be copied can be randomly determined from the second image frame sequence or according to a preset strategy. For example, a preset number of image frames in the second image frame sequence that require extended playback time, or a preset number of image frames included in a video segment that needs to be played repeatedly, can be determined as the third number of image frames to be copied.

[0091] Then, the third number of image frames are copied to obtain the third number of copied image frames.

[0092] Then, the third number of copied image frames are inserted into the second image frame sequence to obtain a fourth image frame sequence, so that the duration of the fourth image frame sequence is equal to the duration of the target video.

[0093] As an example, if a preset number of image frames in the second image frame sequence that need to have their playback duration extended are determined as the third number of image frames to be copied, then each copied image frame can be inserted into the adjacent position of the copied image frame.

[0094] Subsequently, based on the fourth image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0095] Here, the fourth image frame sequence and the audio corresponding to the target video can be synthesized to generate the video corresponding to the target scene.

[0096] It is understood that, in the above optional implementation, the third number of image frames to be copied can be determined from the second image frame sequence to obtain the fourth image frame sequence. Then, the fourth image frame sequence and the audio corresponding to the target video are used to generate the video corresponding to the target scene. In this way, the picture in the generated video corresponding to the target scene can be smoother, further reducing the impact on the user's viewing experience during video playback.

[0097] In some application scenarios of the above optional implementations, the third number of image frames to be copied can be determined from the second image frame sequence in the following manner: From the second image frame sequence, the third number of image frames are determined in reverse order to obtain the third number of image frames to be copied.

[0098] It is understandable that, in the target scene, the audio and image frames corresponding to the starting screen are often corresponding, while the audio and image frames corresponding to the ending screen often have frame gaps. Therefore, in the above application scenario, a third number of image frames can be selected in reverse order and copied. This can make the audio and image frames corresponding to the ending screen of the final generated target scene video more consistent, thereby reducing or even avoiding frame gaps in the generated target scene video.

[0099] In some application scenarios of the above optional implementation methods, the third number of copied image frames can be inserted into the second image frame sequence in the following manner: For each of the third number of copied image frames, insert the copied image frame into the target position.

[0100] The target position is the adjacent position of the same image frame as the copied image frame in the second image frame sequence.

[0101] It is understandable that in the above application scenario, the copied image frame can be inserted into the adjacent position of the same image frame. This can make the picture in the video corresponding to the target scene smoother and further reduce the impact on the user's viewing experience during video playback.

[0102] It should be noted that, in addition to the contents described above, this embodiment may also include... Figure 1 The corresponding technical features described in the corresponding embodiments, thereby achieving Figure 1 For details on the technical effects of the video generation method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0103] The video generation method provided in this application uses a third quantity to reflect the severity of frame clipping in the target scene. Therefore, based on the third quantity, the second image frame sequence, and the audio corresponding to the target video, a video corresponding to the target scene is generated, which can further reduce the impact on the user's viewing experience during video playback.

[0104] The embodiments of this application are described below by way of example. However, it should be noted that the embodiments of this application may have the features described below, but the following description does not constitute a limitation on the protection scope of the embodiments of this application.

[0105] In existing technologies, when a specific short segment is extracted from a long video and played as a highlight, the application scenario is often a recap. In this case, extremely high watchability of the short video is required.

[0106] For some long videos (such as TV dramas), there may be instances where the audio is not finished during transitions. If the video is segmented based on the audio, the content of the next scene (i.e., video frames after the fifth frame in the first frame sequence) will be extracted into the short video clip. This is called a "frame clipping" at the end of the short video, typically involving 2-5 frames. If the video is segmented based on the transition point, there may still be instances where the audio hasn't finished playing even after the visuals have finished playing (the dialogue is mostly complete), reducing the user's viewing experience. If this problem cannot be resolved, both audio-based and transition-point-based segmentation will negatively impact the user's viewing experience.

[0107] This method no longer uses Python's built-in ffmpeg toolkit to directly extract video as the final output. Instead, it first extracts a longer and more complete video (i.e., the first image frame sequence mentioned above) based on audio integrity. Then, it supplements other data for frames that cross shots, ensuring the integrity of the video content. Ultimately, this method can guarantee both audio integrity and prevent "frame gaps" in the video.

[0108] like Figure 3A , Figure 3A This is a schematic diagram of a long video in a video generation method provided in an embodiment of this application. Figure 3A In a long video, the image frame corresponding to the ending scene of the target scene (e.g., scene 1) is located between the second and third image frames of scene 2. If scenes 1 and 2 are used for cropping, the audio above will appear incomplete. This method uses copying of the last two frames from scene 1 to complete the first two frames of scene 2. This ensures both the integrity of the audio and prevents "frame gaps" in the video content.

[0109] The specific implementation steps are as follows: against Figure 3A The video frame clipping situation is as follows: If the clipping is performed at the end of Scene 1, the audio will be cut off. If the clipping is performed at the end of the audio, Scene 2 will be clipped in.

[0110] The specific algorithm is as follows: First, calculate the position of the video frame at the end of Scene 1 (i.e., the fifth image frame mentioned above). Then, calculate the position of the video frame at the end of the audio in Scene 2. The number of image frames between two positions is the number of sandwich frames.

[0111] Next, based on the complete audio clipping points, the long video is clipped using ffmpeg to obtain the corresponding short video (i.e., the target video mentioned above). Please refer to... Figure 3B , Figure 3B This is a schematic diagram of the target video in a video generation method provided in an embodiment of this application.

[0112] Then, full frame extraction is performed on the short video to obtain all the video frames of the short video.

[0113] Subsequently, audio information is extracted from the short video to obtain the audio corresponding to the target video.

[0114] Next, for the extracted video frames, by calculating the number of frames, we can obtain the image names and quantities in the extracted frames of scene 2 that need to be discarded.

[0115] Next, the positions of the discarded images are filled in using extracted frames from Scene 1. The filling method is as follows: Image frames in Scene 1 are copied from back to front, with each frame copied into two frames. Then, images are deleted from the target video's corresponding image frame sequence (i.e., the first image frame sequence mentioned above) from back to front. The copied image frames are then used to fill in the positions where no image frames were previously present. This process continues until the copied image positions coincide with the filled positions, meaning that the seventh and eighth image frames mentioned above represent the same image frame. Please continue to refer to... Figure 3C , Figure 3C This is a schematic diagram illustrating the generation of a video corresponding to a target scene in a video generation method provided in an embodiment of this application.

[0116] Finally, the new image frame sequence and the extracted audio data are used to re-synthesize the video using ffmpeg.

[0117] Here, since long videos typically have 2-5 frames "intercalated", and the extracted short videos are played at 1.2x speed (i.e. the preset speed mentioned above), this method can solve the frame intercalation problem without affecting the user's viewing experience.

[0118] It should be noted that, in addition to the contents described above, this embodiment may also include the technical features described in the above embodiments, thereby achieving the technical effects of the video generation method shown above. Please refer to the above description for details. For the sake of brevity, it will not be elaborated here.

[0119] The video generation method provided in this application solves the "frame clipping" problem that occurs when cutting long videos into short videos. This improves the user's viewing experience when watching short videos.

[0120] Figure 4 This is a schematic diagram of a video generation device provided in an embodiment of this application. Specifically, it includes: The first determining unit 401 is used to determine the image frame corresponding to the ending screen of the target scene from the first image frame sequence corresponding to the target video, and obtain the fifth image frame; wherein the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene. The second determining unit 402 is used to determine the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames between them as the second image frame sequence; The generation unit 403 is used to generate a video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video.

[0121] In one possible implementation, generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video includes: The difference between the first quantity and the second quantity is determined to obtain the third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence; Based on the third quantity, the second image frame sequence, and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0122] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the end of the second image frame sequence, select an unselected image frame to obtain the sixth image frame; Based on the sixth image frame, perform the following replacement steps: The sixth image frame is copied a fourth number of times to obtain the fourth number of sixth image frames, wherein the fourth number is greater than 1; The fourth number of sixth image frames that are located at the end of the first image frame sequence and have not been replaced are replaced with the fourth number of sixth image frames; Determine whether the seventh image frame and the eighth image frame represent the same image frame, wherein the seventh image frame is the image frame preceding the sixth image frame in the second image frame sequence, and the eighth image frame is the last image frame in the first image frame sequence that has not been replaced; If the seventh and eighth image frames do not represent the same image frame, an unselected image frame is selected from the end of the second image frame sequence to obtain the sixth image frame; the replacement step is then performed based on the sixth image frame. When the seventh and eighth image frames represent the same image frame, the first image frame sequence obtained after replacement is determined as the third image frame sequence; based on the third image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated.

[0123] In one possible implementation, generating the video corresponding to the target scene based on the third quantity, the second image frame sequence, and the audio corresponding to the target video includes: From the second image frame sequence, determine the third number of image frames to be copied; The third number of image frames are copied to obtain the third number of copied image frames; The third number of copied image frames are inserted into the second image frame sequence to obtain a fourth image frame sequence, so that the duration of the fourth image frame sequence is equal to the duration of the target video; Based on the fourth image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0124] In one possible implementation, determining the third number of image frames to be copied from the second image frame sequence includes: From the second image frame sequence, the third number of image frames are determined in reverse order to obtain the third number of image frames to be copied.

[0125] In one possible implementation, inserting the third number of copied image frames into the second image frame sequence includes: For each of the third number of copied image frames, the copied image frame is inserted into a target position, wherein the target position is the adjacent position of an image frame in the second image frame sequence that is the same as the copied image frame.

[0126] In one possible implementation, after generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video, the apparatus further includes: The control unit (not shown in the figure) is used to control the video corresponding to the target scene to play at a preset speed, wherein the preset speed represents a value greater than 1.

[0127] In one possible implementation, the audio corresponding to the starting frame of the target scene and the image frame corresponding to the starting frame of the target scene are located at the same position in the target video.

[0128] The video generation device provided in this embodiment can be as follows: Figure 4The video generation device shown can execute all the steps of the video generation methods described above, thereby achieving the technical effects of the video generation methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.

[0129] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 The illustrated electronic device 500 includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the electronic device 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to implement communication between these components. In addition to a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 5 The general designated all buses as Bus System 505.

[0130] The user interface 503 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0131] It is understood that the memory 502 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0132] In some implementations, memory 502 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 5021 and application program 5022.

[0133] The operating system 5021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 5022 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of this application embodiment can be included in application program 5022.

[0134] In this embodiment, by calling the program or instructions stored in memory 502, specifically the program or instructions stored in application program 5022, processor 501 executes the method steps provided in each method embodiment, including, for example: From the first image frame sequence corresponding to the target video, determine the image frame corresponding to the ending screen of the target scene to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene. The sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between is determined as the second image frame sequence; Based on the second image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0135] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 502. Processor 501 reads the information in memory 502 and, in conjunction with its hardware, completes the steps of the above method.

[0136] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described above in this application, or combinations thereof.

[0137] For software implementation, the techniques described herein can be implemented by units that perform the functions described above. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0138] The electronic device provided in this embodiment may be as follows: Figure 5 The electronic device shown can execute all the steps of the video generation methods described above, thereby achieving the technical effects of the video generation methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, further details are not provided here.

[0139] This application also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.

[0140] One or more programs in the storage medium can be executed by one or more processors to implement the video generation method described above that is executed on the electronic device side.

[0141] The processor described above is used to execute the video generation program stored in the memory to implement the following steps of the video generation method executed on the electronic device side: From the first image frame sequence corresponding to the target video, determine the image frame corresponding to the ending screen of the target scene to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene. The sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between is determined as the second image frame sequence; Based on the second image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

[0142] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0143] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0144] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0145] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method of video generation, the method comprising: The method includes: From the first image frame sequence corresponding to the target video, determine the image frame corresponding to the ending screen of the target scene to obtain the fifth image frame; wherein, the starting audio of the target video corresponds to the starting screen of the target scene; The sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames in between is determined as the second image frame sequence; The difference between the first quantity and the second quantity is determined to obtain the third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence; Select the third number of image frames from the second image frame sequence and copy them; Based on the third number of image frames, the second image frame sequence, and the audio corresponding to the target video, a video corresponding to the target scene is generated.

2. The method according to claim 1, characterized in that, The step of generating the video corresponding to the target scene based on the third number of image frames, the second image frame sequence, and the audio corresponding to the target video includes: From the end of the second image frame sequence, select an image frame that has not been selected before to obtain the sixth image frame; Based on the sixth image frame, perform the following replacement steps: The sixth image frame is copied a fourth number of times to obtain the fourth number of sixth image frames, wherein the fourth number is greater than 1; The fourth number of sixth image frames that are located at the end of the first image frame sequence and have not been replaced are replaced with the fourth number of sixth image frames; Determine whether the seventh image frame and the eighth image frame represent the same image frame, wherein the seventh image frame is the image frame preceding the sixth image frame in the second image frame sequence, and the eighth image frame is the last image frame in the first image frame sequence that has not been replaced; If the seventh and eighth image frames do not represent the same image frame, an unselected image frame is selected from the end of the second image frame sequence to obtain the sixth image frame; the replacement step is then performed based on the sixth image frame. When the seventh and eighth image frames represent the same image frame, the first image frame sequence obtained after replacement is determined as the third image frame sequence; based on the third image frame sequence and the audio corresponding to the target video, the video corresponding to the target scene is generated.

3. The method according to claim 1, characterized in that, The step of generating the video corresponding to the target scene based on the third number of image frames, the second image frame sequence, and the audio corresponding to the target video includes: From the second image frame sequence, determine the third number of image frames to be copied; The third number of image frames are copied to obtain the third number of copied image frames; The third number of copied image frames are inserted into the second image frame sequence to obtain a fourth image frame sequence, so that the duration of the fourth image frame sequence is equal to the duration of the target video; Based on the fourth image frame sequence and the audio corresponding to the target video, a video corresponding to the target scene is generated.

4. The method according to claim 3, characterized in that, The step of determining the third number of image frames to be copied from the second image frame sequence includes: From the second image frame sequence, the third number of image frames are determined in reverse order to obtain the third number of image frames to be copied.

5. The method according to claim 3, characterized in that, The step of inserting the third number of copied image frames into the second image frame sequence includes: For each of the third number of copied image frames, the copied image frame is inserted into a target position, wherein the target position is the adjacent position of an image frame in the second image frame sequence that is the same as the copied image frame.

6. The method according to any one of claims 1-5, characterized in that, After generating the video corresponding to the target scene based on the second image frame sequence and the audio corresponding to the target video, the method further includes: Control the video corresponding to the target scene to play at a preset speed, wherein the preset speed represents a value greater than 1.

7. The method according to any one of claims 1-5, characterized in that, The audio corresponding to the starting frame of the target scene and the image frame corresponding to the starting frame of the target scene are located at the same position in the target video.

8. A video generation apparatus, characterized in that, The device includes: The first determining unit is configured to determine the image frame corresponding to the ending screen of the target scene from the first image frame sequence corresponding to the target video, thereby obtaining the fifth image frame; wherein the starting audio of the target video corresponds to the starting screen of the target scene; and the ending audio of the target video corresponds to the ending screen of the target scene. The second determining unit is used to determine the sequence consisting of the first image frame in the first image frame sequence, the fifth image frame, and the image frames between them as the second image frame sequence; A generation unit is used to determine the difference between a first quantity and a second quantity to obtain a third quantity, wherein the first quantity represents the number of image frames included in the first image frame sequence, and the second quantity represents the number of image frames included in the second image frame sequence; selects the third quantity of image frames from the second image frame sequence for copying; and generates a video corresponding to the target scene based on the third quantity of image frames, the second image frame sequence, and the audio corresponding to the target video.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.