Video generation method, device, readable medium and electronic device

The target boundary frame is determined through preset boundary recognition models and feature maps, and the video lens is divided in a refined manner, which solves the problem of poor quality of video summary in the prior art and improves the user experience.

CN114117127BActive Publication Date: 2025-09-02DOUYIN VISION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111397234.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-09-02
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

The existing video digest generation methods cannot divide the video lenses in a refined manner, resulting in poor quality of video digests and poor user viewing experience.

Method used

The preset boundary recognition model outputs the feature map of the pending boundary frame and each video image, determines the target boundary frame and divides the video lens to generate a specified video.

Benefits of technology

More accurate video footage division has been achieved, improving the quality of video summary and user viewing experience, and avoiding the problems of not highlighting the key points and not obvious classic performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114117127B_ABST
    Figure CN114117127B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, readable medium, and electronic device. The video generation method can, after obtaining a pending boundary frame for dividing video shots through a preset boundary recognition model output, determine a target boundary frame based on the pending boundary frame and a feature map of each frame of video image; then determine multiple target video shots contained in the original video based on the target boundary frame; and finally generate a designated video based on the multiple target video shots and the feature map of each frame of video image in each target video shot. In this way, a more accurate target boundary frame for dividing video shots can be obtained, thereby ensuring a refined division of video shots and obtaining more accurate target video shots. A designated video is generated based on the more accurate target video shots. When the designated video is a video summary, the quality of the video summary can be effectively improved, and the user's viewing experience can also be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video data processing, and in particular, to a video generation method, device, readable medium, and electronic device. Background Art

[0002] A video summary is a brief summary of the video content. For example, a video summary of a game video can help users quickly understand the game content. Generally, the video summary generation process includes two parts: video shot division and shot selection. Accurate shot division helps highlight the classic content of the video, improve the quality of the video summary, and also help enhance the user's viewing experience.

[0003] However, current methods for generating video summaries are usually unable to finely divide video shots, which can easily lead to poor video summary quality due to inaccurate video shot division, and is not conducive to improving the user's viewing experience. Summary of the Invention

[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] The present disclosure provides a video generation method, device, readable medium, and electronic device.

[0006] In a first aspect, the present disclosure provides a video generation method, the method comprising:

[0007] Get the original video;

[0008] The original video is used as input to a preset boundary recognition model to output a pending boundary frame for dividing the video shot and a feature map of each frame of the video image in the original video;

[0009] Determine a target boundary frame according to the pending boundary frame and the feature map of each frame of video image;

[0010] Determine a plurality of target video shots contained in the original video according to the target boundary frame;

[0011] A designated video is generated according to the plurality of target video shots and a feature map of each frame of video image in each target video shot.

[0012] In a second aspect, the present disclosure provides a video generation device, the device comprising:

[0013] Acquisition module, used to obtain original video;

[0014] A first determination module is configured to use the original video as an input to a preset boundary recognition model to output a pending boundary frame for dividing a video shot and a feature map of each frame of the video image in the original video;

[0015] A second determining module is configured to determine a target boundary frame based on the pending boundary frame and the feature map of each frame of video image;

[0016] A third determining module is configured to determine a plurality of target video shots contained in the original video according to the target boundary frame;

[0017] A generating module is used to generate a specified video according to the plurality of target video shots and a feature map of each frame of video image in each target video shot.

[0018] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect above.

[0019] In a fourth aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect above.

[0020] The above technical solution, after obtaining the pending boundary frame for dividing the video shots through the output of the preset boundary recognition model, determines the target boundary frame based on the pending boundary frame and the feature map of each frame of the video image; then determines multiple target video shots contained in the original video based on the target boundary frame, and generates a designated video based on the multiple target video shots and the feature map of each frame of the video image in each of the target video shots. In this way, a more accurate target boundary frame for dividing the video shots can be obtained, thereby ensuring a refined division of the video shots and obtaining more accurate target video shots. The designated video is generated based on the more accurate target video shots. When the designated video is a video summary, it is helpful to avoid problems such as the generated video summary not highlighting the key points and the classic performance not being obvious, which can effectively improve the quality of the video summary and the user's viewing experience.

[0021] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:

[0023] Figure 1 is a flowchart of a video generation method shown in an exemplary embodiment of the present disclosure;

[0024] Figure 2 is based on Figure 1 The illustrated embodiment shows a flow chart of a method for generating a video;

[0025] Figure 3 is based on Figure 1 A flow chart of another video generation method shown in the illustrated embodiment;

[0026] Figure 4 is a block diagram of a video generating device according to an exemplary embodiment of the present disclosure;

[0027] Figure 5 It is a block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0029] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0030] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] Before introducing the specific implementation methods of the present disclosure in detail, the application scenarios of the present disclosure are first described as follows. The present disclosure can be applied to the generation process of a specified video, and in particular, can be applied to the scenario of generating a video summary for a game video. For example, the game video can be a video image obtained by recording the screen when the game screen is displayed on the terminal. The game video can be a video image including the entire game process, or it can be a video image including only one or several levels in a level-breaking game. Currently, in the process of generating video summaries for game videos in the related art, inaccurate video shot division often occurs, that is, the tail part of the previous video shot is mistakenly divided into one video shot with the current video shot, or the starting part of the current video shot is mistakenly attributed to the previous video shot. When generating a video summary, it is usually necessary to compose the video summary based on video shots. For example, some classic video shots are selected from the divided video shots to compose the video summary, or some classic video shots in the original video are highlighted to form the video summary. If the video shot division is inaccurate, it is easy to cause the classic video shots used to compose the video summary to be mixed with non-classic shot content, resulting in the generated video summary having problems such as lack of emphasis and unclear classic performance. This will be very detrimental to the improvement of the quality of the video summary and the improvement of the user's viewing experience.

[0035] In order to solve the above technical problems, the present disclosure provides a video generation method, device, readable medium and electronic device. The video generation method can, after obtaining a pending boundary frame for dividing video shots through a preset boundary recognition model output, determine a target boundary frame based on the pending boundary frame and a feature map of each frame of video image; then determine multiple target video shots contained in the original video based on the target boundary frame; finally, generate a specified video based on the multiple target video shots and the feature map of each frame of video image in each target video shot. In this way, a more accurate target boundary frame for dividing video shots can be obtained, thereby ensuring a refined division of video shots and obtaining more accurate target video shots. A specified video is generated based on the more accurate target video shot. When the specified video is a video summary, it is helpful to avoid problems such as a lack of emphasis and unclear classic performance in the generated video summary, thereby effectively improving the quality of the video summary and the user's viewing experience.

[0036] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0037] Figure 1 is a flowchart of a video generation method shown in an exemplary embodiment of the present disclosure; see Figure 1 , the method may include:

[0038] Step 101: Obtain the original video.

[0039] The original video may be a game video image including the entire game process, or may be a game video image of a certain level or several levels in a level game.

[0040] For example, in this step, when the game screen is displayed on the terminal, the game video image including the entire game process can be obtained through the screen recording function, and the game video image of the entire game process can be segmented to obtain the game video image corresponding to each level of the game. The game video image of each level of the game is used as the original video to generate a video summary of the game video image of each level.

[0041] In step 102, the original video is used as input to a preset boundary recognition model to output a pending boundary frame for dividing the video shots and a feature map of each frame of the video image in the original video.

[0042] In which, the preset boundary recognition model can be a bidirectional LSTM (Long Short Term Memory, a neural network with the ability to remember long-term and short-term information) model, and the training process of the preset boundary recognition model can include: obtaining preset second model training data, the second model training data consisting of multiple original video samples, and the annotation data of the boundary frame image of each video shot in each of the original video samples; training the preset second initial model through the second model training data to obtain the preset boundary recognition model, the second initial model can be a preset initial bidirectional LSTM model, and the preset boundary recognition model can include an input layer, an output layer and at least one hidden layer, the output end of the input layer is coupled to the input end of the hidden layer, and the output end of the hidden layer is coupled to the input end of the output layer.

[0043] In this step, the original video can be input to the hidden layer through the input layer; the feature map of each frame of the video image in the original video is output through the hidden layer, and the feature map of the frame video image extracted from the original video is input to the output layer through the hidden layer, so that the output layer outputs the pending boundary frame.

[0044] It should be noted that, when there are multiple hidden layers, the feature map output by any of the multiple hidden layers can be used as the feature map of each frame of the original video. The specific architecture of the bidirectional LSTM model can be found in the relevant description of the prior art. The specific architecture of the bidirectional LSTM model in the prior art is relatively easy to obtain and will not be repeated in this disclosure.

[0045] In addition, the extracted frame video image can be obtained by skipping frames of the original video according to a preset video image extraction step size. For example, one frame can be extracted every 5 frames to obtain the extracted frame video image. In this way, the efficiency of the preset boundary recognition model can be effectively improved, and the time for generating the pending boundary frame can be shortened, which is conducive to shortening the time for generating the video summary and improving the generation efficiency of the video summary.

[0046] Step 103: Determine a target boundary frame according to the pending boundary frame and the feature map of each frame of the video image.

[0047] The target boundary frame may be the last frame of each video shot, or the first frame of each video shot.

[0048] Step 104: Determine multiple target video shots included in the original video according to the target boundary frame.

[0049] One possible implementation of this step is as follows: when the target boundary frame is the last frame of each video shot, multiple frames of video images before (including) the current target boundary frame and after (excluding) the previous target boundary frame can be determined as a target video shot. The entire original video can be divided into video shots according to the above method, thereby obtaining multiple target video shots.

[0050] Another possible implementation is: when the target boundary frame is the first frame image of each video shot, multiple frames of video images before the current target boundary frame (excluding the current target boundary frame) and after the previous target boundary frame (including the previous target boundary frame) can be determined as a target video shot. The entire original video can be divided into video shots according to the above method, thereby obtaining multiple target video shots.

[0051] Step 105 : Generate a designated video based on the multiple target video shots and the feature map of each frame of the video image in each target video shot.

[0052] The designated video may be a video summary or a short video corresponding to the original video.

[0053] In this step, the video shot type of each target video shot can be first determined. The video shot type can include a first type of video shot and a second type of video shot. The specified video is generated according to the video shot type. Specifically, when the first type of video shot is a highlight video shot, the second type of video shot can be a non-highlight video shot. When the first type of video shot is a preset character video shot, the second type of video shot can be a non-preset character video shot, and the preset character video shot can be a video shot containing a preset character image.

[0054] For example, highlight video shots and non-highlight video shots can be determined from the multiple target video shots, and then a specified video can be generated based on only the highlight video shots, or a specified video containing the complete content of the original video can be generated based on the highlight video shots and the non-highlight video shots.

[0055] It should be noted that when generating a designated video based only on the highlight video shots, the determined multiple highlight video shots may be spliced ​​together according to the order of each highlight video shot in the original video screen to form the designated video.

[0056] When generating a designated video containing the complete content of the original video based on the highlight video shot and the non-highlight video shot, the first preset frame rate can be used as the frame rate of the highlight video shot, and the second preset frame rate can be used as the frame rate of the target video shot of the non-highlight video shot to obtain the designated video, wherein the first preset frame rate is lower than the second preset frame rate. In this way, the highlight video shot can be displayed at a slower playback speed, which is conducive to the user to better obtain the video content in the highlight video shot, and the video image corresponding to the non-highlight video shot can be displayed at a relatively fast playback speed, which can quickly display the non-critical content, thereby avoiding the problem of poor user experience caused by the display of the non-critical content taking too much time.

[0057] The above technical solution can obtain more accurate target boundary frames for dividing video shots, thereby ensuring refined division of video shots and obtaining more accurate target video shots. A specified video is generated based on the more accurate target video shots. When the specified video is a video summary, it is helpful to avoid problems such as lack of emphasis and unclear classic performance in the generated video summary, thereby effectively improving the quality of the video summary and the user's viewing experience.

[0058] Furthermore, the above Figure 1 The target boundary frame is determined according to the feature map of the pending boundary frame and each frame of the video image in step 103 by the following method: Figure 2 The steps shown are implemented, Figure 2 is based on Figure 1 The embodiment shown is a flowchart of a video generation method, see Figure 2 , step 103 may include:

[0059] Step 1031 : Acquire a target video segment with the to-be-determined boundary frame as a non-first frame image from the original video.

[0060] In this step, M video images before the undetermined boundary frame can be obtained from the original video, and M-1 frames after the undetermined boundary frame can be obtained from the original video to obtain a target video segment formed by 2M frames of video images. Where M can be a preset positive integer greater than 1.

[0061] Step 1032: determine a first distance between each frame of the video image in the target video segment and the start frame of the target video segment based on the feature map of each frame of the video image in the target video segment, and determine a second distance between each frame of the video image in the target video segment and the end frame of the target video segment based on the feature map of each frame of the video image in the target video segment.

[0062] In this step, the first distance and the second distance are both Euclidean distances. The feature map of each frame of video image in the current target video clip (i.e., 2M frame video image) can be screened out from the feature map of each frame of video image output by the preset boundary frame recognition model, and the first distance between each frame of video image in the 2M frame of video image and the starting frame of the 2M frame of video image is determined according to the feature map of each frame of video image in the 2M frame of video image, and the second distance between each frame of video image in the 2M frame of video image and the ending frame of the 2M frame of video image is determined.

[0063] The first distance and the second distance can be calculated using the following formula:

[0064]

[0065] In the above formula, X=(x1,x2…x n ), Y=(y1,y2…y n ), X is the i-th frame video image, (x1, x2…x n ) is the feature vector corresponding to the feature map of the i-th frame video image, Y is the end frame (or start frame), (y1, y2…y n ) is the feature vector corresponding to the feature map of the end frame (or start frame).

[0066] Step 1033: Obtain the sum of the first distance and the second distance corresponding to each frame of the video image in the target video segment.

[0067] Step 1034 : The target frame image with the largest sum value in the target video segment is used as the target boundary frame.

[0068] For example, if the target video clip includes 4 frames of video images, the distance between the 1st frame and the 4th frame is A, the distance between the 2nd frame and the 1st frame is B, the distance between the 2nd frame and the 4th frame is C, the distance between the 3rd frame and the 1st frame is D, the distance between the 3rd frame and the 4th frame is E, and the distance between the 4th frame and the 1st frame is A. If B+C>D+E>A, then the video image corresponding to the 2nd frame is determined as the target boundary frame.

[0069] The above technical solution can obtain the pending boundary frame for dividing video shots through the output of the preset boundary recognition model, and then determine the target boundary frame based on the pending boundary frame and the feature map of each frame of video image. This can obtain a more accurate target boundary frame for dividing video shots, thereby ensuring the accuracy of the video shot segmentation results and improving the quality of video summarization.

[0070] Furthermore, the above Figure 1 The step 105 described in the above method can generate a specified video based on the multiple target video shots and the feature map of each frame of the video image in each target video shot by the following method: Figure 3 The steps shown are implemented, Figure 3 is based on Figure 1 The embodiment shown is a flowchart of another method for generating a video. Figure 3 , step 105 may include:

[0071] Step 1051: Input the feature map of each frame of the target video shot into a preset classification model, so that the preset classification model outputs the video shot type of the target video shot.

[0072] Among them, the video lens type includes a first type of video lens and a second type of video lens. When the first type of video lens is a highlight video lens, the second type of video lens can be a non-highlight video lens; when the first type of video lens is a preset character video lens, the second type of video lens can be a non-preset character video lens, and the preset character video lens can be a video lens containing a preset character image.

[0073] It should be noted that the preset classification model can be trained in the following ways:

[0074] Acquire feature maps corresponding to multiple video shot samples, each of which includes labeling information of a first category of video shot or a second category of video shot; use the feature maps corresponding to the multiple video shot samples as first model training data, and train a preset first initial model using the first model training data to obtain the preset classification model.

[0075] Step 1052: Determine whether the original video includes multiple original video segments.

[0076] Each of the original video clips includes at least one target video shot.

[0077] In this step, if it is determined that the original video includes multiple original video segments, step 1053 is executed; if it is determined that the original video includes only one original video segment, step 1054 is executed.

[0078] Step 1053 : splicing the multiple target video shots corresponding to the original video according to the first order corresponding to the multiple original video segments and the second order of the target video shots in each of the original video segments.

[0079] For example, if the original video includes a game video clip of the first level, a game video clip of the second level, and a game video clip of the third level in a level game, wherein the game video clip of the first level includes 3 target video shots, namely, a target video shot from 0 to 2 seconds, a target video shot from 2 to 2.5 seconds, and a target video shot from 2.5 seconds to 3 seconds, the game video clip of the second level includes 2 target video shots, namely, a target video shot from 0 to 1 second, and a target video shot from 1 to 3 seconds, and the game video clip of the third level includes 4 target video shots, namely, a target video shot from 0 to 1 second, a target video shot from 1 to 2 seconds, a target video shot from 2 to 2.5 seconds, and a target video shot from 2.5 seconds to 3.5 seconds. When splicing multiple target video shots in the original video, the splicing is performed according to the first order (i.e., the first level, the second level, and the third level) and the second order (i.e., the time order of the target video shots) of each level, so that the 0 to 2 seconds target video shot of the first level is spliced ​​with the 2 to 2.5 seconds target video shot of the first level, the 2 to 2.5 seconds target video shot of the first level is spliced ​​with the 2.5 to 3 seconds target video shot of the first level, and the 2.5 to 3 seconds target video shot of the first level is spliced ​​with the 0 to 1 second target video shot of the second level. Video shot splicing: the 0 to 1 second target video shot of the second level is spliced ​​with the 1 to 3 second target video shot of the second level, the 1 to 3 second target video shot of the second level is spliced ​​with the 0 to 1 second target video shot of the third level, the 0 to 1 second target video shot of the third level is spliced ​​with the 1 to 2 second target video shot of the third level, the 1 to 2 second target video shot of the third level is spliced ​​with the 2 to 2.5 second target video shot of the third level, and the 2 to 2.5 second target video shot of the third level is spliced ​​with the 2.5 second to 3.5 second target video shot of the third level.

[0080] Step 1054 : Using the first preset frame rate as the frame rate of the first type of video shot, and using the second preset frame rate as the frame rate of the target video shot of the second type of video shot, to obtain the designated video.

[0081] The first preset frame rate is lower than the second preset frame rate.

[0082] For example, the second preset frame rate may be twice or more than twice the first preset frame rate.

[0083] The above technical solution can generate a designated video that effectively highlights the first type of video shot by playing the first type of video shot at a lower frame rate and the second type of video shot at a relatively higher frame rate. For example, when the first type of video shot is a highlight video shot and the second type of video shot may be a non-highlight video shot, the highlight video shot can be played at a lower frame rate and the non-highlight video shot can be played at a relatively higher frame rate, thereby generating a designated video that effectively highlights the highlight video shot. When the first type of video shot is a preset person video shot and the second type of video shot may be a non-preset person video shot, the preset person video shot can be played at a lower frame rate and the non-preset person video shot can be played at a relatively higher frame rate, thereby generating a designated video that effectively highlights the preset person video shot. In this way, not only can the designated video achieve the effect of highlighting the key points, but the designated video can also have complete video content, which is conducive to the user's comprehensive and focused understanding of the content involved in the original video, thereby effectively improving the user's viewing experience.

[0084] Figure 4 is a block diagram of a video generating device according to an exemplary embodiment of the present disclosure; Figure 4 , the apparatus may include:

[0085] Acquisition module 401, used to acquire original video;

[0086] A first determination module 402 is configured to use the original video as an input of a preset boundary recognition model to output a pending boundary frame for dividing a video shot and a feature map of each frame of the original video;

[0087] A second determining module 403 is configured to determine a target boundary frame based on the pending boundary frame and the feature map of each frame of the video image;

[0088] A third determining module 404 is configured to determine a plurality of target video shots included in the original video according to the target boundary frame;

[0089] The generating module 405 is configured to generate a designated video according to the plurality of target video shots and the feature map of each frame of the video image in each target video shot.

[0090] The above technical solution can obtain more accurate target boundary frames for dividing video shots, thereby ensuring refined division of video shots and obtaining more accurate target video shots. Generating a specified video based on the more accurate target video shots is conducive to avoiding problems such as lack of emphasis on key points and lack of obvious classic performance in the generated video summary, thereby effectively improving the quality of the video summary and the user's viewing experience.

[0091] Optionally, the preset boundary recognition model includes an input layer, an output layer, and at least one hidden layer, and the first determination module 402 is configured to:

[0092] Input the original video into the hidden layer through the input layer;

[0093] The feature map of each frame of video image in the original video is output through the hidden layer, and the feature map of the frame video image extracted from the original video is input to the output layer through the hidden layer, so that the output layer outputs the pending boundary frame.

[0094] Optionally, the second determining module 403 is configured to:

[0095] Obtaining a target video segment with the to-be-determined boundary frame as a non-first frame image from the original video;

[0096] determining a first distance between each frame of the target video segment and a start frame of the target video segment based on a feature map of each frame of the target video segment, and determining a second distance between each frame of the target video segment and an end frame of the target video segment based on the feature map of each frame of the target video segment;

[0097] Obtaining a sum of the first distance and the second distance corresponding to each frame of the video image in the target video segment;

[0098] The target frame image with the largest sum value in the target video segment is used as the target boundary frame.

[0099] Optionally, the generating module 405 is configured to:

[0100] Inputting a feature map of each frame of the target video shot into a preset classification model, so that the preset classification model outputs a video shot type of the target video shot;

[0101] Generate the specified video according to the video shot type.

[0102] Optionally, the video shot type includes a first type of video shot and a second type of video shot, and the generating module 405 is configured to:

[0103] The first preset frame rate is used as the frame rate of the first type of video shot, and the second preset frame rate is used as the frame rate of the target video shot of the second type of video shot to obtain the designated video, wherein the first preset frame rate is smaller than the second preset frame rate.

[0104] Optionally, the original video includes a plurality of original video segments, each of which includes at least one target video shot. The generating module 405 is further configured to:

[0105] The multiple target video shots corresponding to the original video are spliced ​​according to the first sequence corresponding to the multiple original video segments and the second sequence of the target video shots in each of the original video segments.

[0106] Optionally, the preset classification model is trained in the following manner:

[0107] Obtain feature maps corresponding to a plurality of video shot samples, each of the video shot samples including labeling information of a first type of video shot or a second type of video shot;

[0108] The feature maps corresponding to the multiple video shot samples are used as first model training data, and a preset first initial model is trained using the first model training data to obtain the preset classification model.

[0109] Optionally, the preset boundary recognition model is trained by the following training method:

[0110] Obtaining preset second model training data, where the second model training data consists of a plurality of original video samples and annotation data of a boundary frame image of each video shot in each of the original video samples;

[0111] The preset second initial model is trained using the second model training data to obtain the preset boundary recognition model, which includes at least one hidden layer.

[0112] The above technical solution can obtain more accurate target boundary frames for dividing video shots, thereby ensuring refined division of video shots and obtaining more accurate target video shots. Generating a specified video based on the more accurate target video shots is conducive to avoiding problems such as lack of emphasis on key points and lack of obvious classic performance in the generated video summary, thereby effectively improving the quality of the video summary and the user's viewing experience.

[0113] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0114] Reference below Figure 5, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0115] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0116] Typically, the following devices may be connected to the I / O interface 505: an input device 505 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0117] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0118] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0119] In some embodiments, the client can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can interconnect with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0120] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0121] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain an original video; use the original video as input to a preset boundary recognition model to output a pending boundary frame for dividing video shots, and a feature map of each video frame in the original video; determine a target boundary frame based on the pending boundary frame and the feature map of each video frame; determine multiple target video shots contained in the original video based on the target boundary frame; and generate a specified video based on the multiple target video shots and the feature map of each video frame in each target video shot.

[0122] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0124] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, the acquisition module may also be described as a "module for acquiring original video."

[0125] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0126] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] According to one or more embodiments of the present disclosure, Example 1 provides a video generation method, the method including:

[0128] Get the original video;

[0129] The original video is used as input to a preset boundary recognition model to output a pending boundary frame for dividing the video shot and a feature map of each frame of the video image in the original video;

[0130] Determine a target boundary frame according to the pending boundary frame and the feature map of each frame of video image;

[0131] Determine a plurality of target video shots contained in the original video according to the target boundary frame;

[0132] A specified video is generated according to a plurality of target video shots and a feature map of each frame of video image in each target video shot.

[0133] According to one or more embodiments of the present disclosure, Example 2 provides the method described in Example 1, wherein the preset boundary recognition model includes an input layer, an output layer, and at least one hidden layer. The original video is used as input to the preset boundary recognition model to output a pending boundary frame for dividing the video shot, and a feature map of each frame of the video image in the original video, including:

[0134] Input the original video into the hidden layer through the input layer;

[0135] The feature map of each frame of video image in the original video is output through the hidden layer, and the feature map of the frame video image extracted from the original video is input to the output layer through the hidden layer, so that the output layer outputs the pending boundary frame.

[0136] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein determining the target boundary frame based on the undetermined boundary frame and the feature map of each frame of video image includes:

[0137] Obtaining a target video segment with the to-be-determined boundary frame as a non-first frame image from the original video;

[0138] determining a first distance between each frame of the target video segment and a start frame of the target video segment based on a feature map of each frame of the target video segment, and determining a second distance between each frame of the target video segment and an end frame of the target video segment based on the feature map of each frame of the target video segment;

[0139] Obtaining a sum of the first distance and the second distance corresponding to each frame of the video image in the target video segment;

[0140] The target frame image with the largest sum value in the target video segment is used as the target boundary frame.

[0141] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein generating a specified video according to the multiple target video shots and a feature map of each frame of video image in each target video shot includes:

[0142] Inputting a feature map of each frame of the target video shot into a preset classification model, so that the preset classification model outputs a video shot type of the target video shot;

[0143] Generate the specified video according to the video shot type.

[0144] According to one or more embodiments of the present disclosure, Example 5 provides the method described in Example 4, wherein the video shot type includes a first type of video shot and a second type of video shot, and generating the specified video according to the video shot type includes:

[0145] The first preset frame rate is used as the frame rate of the first type of video shot, and the second preset frame rate is used as the frame rate of the target video shot of the second type of video shot to obtain the designated video, wherein the first preset frame rate is smaller than the second preset frame rate.

[0146] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein the original video includes multiple original video clips, each of the original video clips includes at least one target video shot, and the method further includes:

[0147] The multiple target video shots corresponding to the original video are spliced ​​according to the first sequence corresponding to the multiple original video segments and the second sequence of the target video shots in each of the original video segments.

[0148] According to one or more embodiments of the present disclosure, Example 7 provides the method described in Example 4, wherein the preset classification model is trained in the following manner:

[0149] Obtain feature maps corresponding to a plurality of video shot samples, each of the video shot samples including labeling information of a first type of video shot or a second type of video shot;

[0150] The feature maps corresponding to the multiple video shot samples are used as first model training data, and a preset first initial model is trained using the first model training data to obtain the preset classification model.

[0151] According to one or more embodiments of the present disclosure, Example 8 provides the method described in any one of Examples 1-7, wherein the preset boundary recognition model is trained by the following training method:

[0152] Obtaining preset second model training data, where the second model training data consists of a plurality of original video samples and annotation data of a boundary frame image of each video shot in each of the original video samples;

[0153] The preset second initial model is trained using the second model training data to obtain the preset boundary recognition model, which includes at least one hidden layer.

[0154] According to one or more embodiments of the present disclosure, Example 9 provides a video generating apparatus, the apparatus comprising:

[0155] Acquisition module, used to obtain original video;

[0156] A first determination module is configured to use the original video as an input of a preset boundary recognition model to output a pending boundary frame for dividing the video shot and a feature map of each frame of the video image in the original video;

[0157] A second determining module is configured to determine a target boundary frame based on the pending boundary frame and the feature map of each frame of video image;

[0158] A third determining module is configured to determine a plurality of target video shots contained in the original video according to the target boundary frame;

[0159] A generation module is used to generate a specified video based on multiple target video shots and a feature map of each frame of video image in each target video shot.

[0160] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon, which implements the steps of any one of the methods in Examples 1-8 above when executed by a processing device.

[0161] According to one or more embodiments of the present disclosure, Example 11 provides an electronic device, including:

[0162] a storage device having a computer program stored thereon;

[0163] A processing device is used to execute the computer program in the storage device to implement the steps of any one of the methods in Examples 1-8 above.

[0164] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0165] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0166] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.

Claims

1. A video generation method, characterized in that: The method comprises: Get the original video; The original video is used as input to a preset boundary recognition model to output a pending boundary frame for dividing the video shot and a feature map of each frame of the video image in the original video; Determine a target boundary frame according to the pending boundary frame and the feature map of each frame of video image, wherein the target boundary frame is the last frame image or the first frame image of each video shot; Determine a plurality of target video shots contained in the original video according to the target boundary frame; Generate a specified video according to a plurality of target video shots and a feature map of each frame of video image in each target video shot; The determining of the target boundary frame according to the pending boundary frame and the feature map of each frame of the video image includes: Acquire a target video segment with the to-be-determined boundary frame as a non-first frame image from the original video; determining a first distance between each frame of the target video image and a start frame of the target video segment according to a feature map of each frame of the target video segment, and determining a second distance between each frame of the target video image and an end frame of the target video segment according to the feature map of each frame of the target video segment; Obtaining a sum of the first distance and the second distance corresponding to each frame of the video image in the target video segment; The target frame image with the largest sum value in the target video segment is used as the target boundary frame.

2. The method according to claim 1, characterized in that The preset boundary recognition model includes an input layer, an output layer, and at least one hidden layer. The original video is used as the input of the preset boundary recognition model to output a pending boundary frame for dividing the video shot, and a feature map of each frame of the video image in the original video, including: Inputting the original video into the hidden layer through the input layer; The feature map of each frame of video image in the original video is outputted through the hidden layer, and the feature map of the frame video image extracted from the original video is inputted to the output layer through the hidden layer, so that the output layer outputs the pending boundary frame.

3. The method according to claim 1, characterized in that Generating a specified video according to the multiple target video shots and a feature map of each frame of video image in each target video shot includes: Inputting a feature map of each frame of video image in each target video shot into a preset classification model, so that the preset classification model outputs a video shot type of the target video shot; The designated video is generated according to the video shot type.

4. The method according to claim 3, characterized in that The video shot types include first-category video shots and second-category video shots, and generating the specified video according to the video shot types includes: A first preset frame rate is used as the frame rate of the first type of video shot, and a second preset frame rate is used as the frame rate of the second type of video shot to obtain the specified video, wherein the first preset frame rate is smaller than the second preset frame rate.

5. The method according to claim 1, wherein The original video includes a plurality of original video segments, each of which includes at least one target video shot. The method further includes: The multiple target video shots corresponding to the original video are spliced ​​according to the first order corresponding to the multiple original video segments and the second order of the target video shots in each of the original video segments.

6. The method according to claim 3, characterized in that The preset classification model is trained in the following way: Obtaining feature maps corresponding to a plurality of video shot samples, each of the video shot samples including labeling information of a first type of video shot or a second type of video shot; The feature maps corresponding to the multiple video shot samples are used as first model training data, and a preset first initial model is trained using the first model training data to obtain the preset classification model.

7. The method according to any one of claims 1 to 6, characterized in that The preset boundary recognition model is trained by the following training method: Acquire preset second model training data, where the second model training data consists of a plurality of original video samples and annotated data of a boundary frame image of each video shot in each of the original video samples; The preset second initial model is trained using the second model training data to obtain the preset boundary recognition model, where the preset boundary recognition model includes at least one hidden layer.

8. A video generating device, characterized in that: The device comprises: Acquisition module, used to obtain original video; A first determination module is configured to use the original video as an input to a preset boundary recognition model to output a pending boundary frame for dividing the video shot and a feature map of each frame of the video image in the original video; A second determining module is configured to determine a target boundary frame according to the pending boundary frame and the feature map of each frame of video image, wherein the target boundary frame is the last frame image or the first frame image of each video shot; A third determining module is configured to determine a plurality of target video shots contained in the original video according to the target boundary frame; A generating module, configured to generate a specified video based on a plurality of target video shots and a feature map of each frame of video image in each target video shot; The second determination module is used to obtain a target video segment with the to-be-determined boundary frame as a non-first frame image from the original video; determine a first distance between each frame of the target video segment and the start frame of the target video segment based on a feature map of each frame of the target video segment, and determine a second distance between each frame of the target video segment and the end frame of the target video segment based on the feature map of each frame of the target video segment; obtain the sum of the first distance and the second distance corresponding to each frame of the target video segment; and use the target frame image with the largest sum value in the target video segment as the target boundary frame.

9. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video abstract generating method, video abstract generating device, and computer readable storage medium

    CN108419145A

  • Video shot segmentation method, system and device and storage medium

    CN110766711A

  • Video processing method and device, electronic equipment, storage medium and program product

    CN113099132A