Abstract image generation device and its program

The summary video generation device improves the quality of summary videos by dividing videos into feature-based intervals, calculating integrated scores, and selecting key segments, effectively addressing the limitations of single-feature methods.

JP7691852B2Active Publication Date: 2025-06-12NIPPON HOSO KYOKAI
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021088951
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-27
Publication Date
2025-06-12
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Existing methods for generating summary videos often rely on single features, which limits the quality of the summary, and fail to accurately reflect multiple features across different video division units.

Method used

A summary video generation device that divides input videos into intervals based on different features such as image, speech, and acoustic features, calculates interval scores using pre-trained neural networks, and integrates these scores to select and concatenate unit video intervals for generating a summary video.

Benefits of technology

This approach allows for the generation of higher-quality summary videos by accurately combining multiple features and selecting video segments with high importance, resulting in a more comprehensive and engaging summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691852000002
    Figure 0007691852000002
  • Figure 0007691852000003
    Figure 0007691852000003
  • Figure 0007691852000004
    Figure 0007691852000004
Patent Text Reader

Abstract

To provide a summary video generation device capable of generating a summary video by combining multiple features.SOLUTION: A summary video generation device 1 includes: video dividing means 11 that divides an input video using different division methods for each feature and generates multiple video segment sequences; segment score calculation means 12 that calculates a segment score that represents a degree of the video in the video segment is a summary video on the basis of, the video feature of the video segment; unit segment dividing means 20 that divides an input video into a unit video segment; unit segment score calculation means 30 that calculates a unit segment score for each unit video segment by adding the segment score corresponding to the video segment depending on the multiple video segment series and the time proportion of overlapping video segments; segment selection means 41 that selects unit video segments on the basis of, the unit segment score; and video connecting means 42 that connects videos of the selected unit video segments.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a summary video generation device that generates a summary video by summarizing a video, and a program thereof.

Background Art

[0002] In recent years, due to the development of social media services and the like, there has been an increasing number of cases where a summary video is distributed on a network for the main purpose of promoting a broadcast program or a self-produced video. However, since the editing work of the summary video requires a great deal of labor, a technology for automatically generating a summary video is demanded.

[0003] Conventionally, as a technique for automatically generating a summary video, for example, the methods are proposed in Patent Documents 1 to 3. The method described in Patent Document 1 is a method of generating a summary video by extracting a video section with high importance from a video based on the image features of key frames of divided videos obtained by dividing the video. The method described in Patent Document 2 is a method of generating a summary video by analyzing a graph in which video sections are used as nodes and the similarity of video features between nodes is used as edges, and extracting videos of video sections with high importance from the video. The method described in Patent Document 3 first divides a video into a plurality of cut videos and calculates scores related to a plurality of elements. Then, this method calculates the total score of the cut videos based on the weight distribution of each element set by the user, and extracts the cut videos with high total scores to generate a summary video.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] The methods described in Patent Documents 1 and 2 generate a summary video using a single feature (only image features or only video features), without considering multiple features. Therefore, in order to improve the quality of the summary video, a method that considers multiple features has been desired. The method described in Patent Document 3 generates a summary video with a score that integrates multiple features (telop score, face recognition score, camera work score) for each cut video obtained by dividing the video. However, the features of a video are not necessarily divided in units of cut videos. For example, there may be a case where a certain feature exists only in a part of a cut video. Therefore, there has been a desire to more accurately reflect multiple features in the cut video.

[0006] In view of such demands of the prior art, the present invention has been made, and an object thereof is to provide a summary video generation device and a program thereof that can generate a higher-quality summary video than before by combining multiple features.

Means for Solving the Problems

[0007] In order to solve the above problems, a summary video generation device according to the present invention is a summary video generation device that generates a summary video from an input video, and includes a video division means, an interval score calculation means, a unit interval division means, a unit interval score calculation means, an interval selection means, and a video connection means.

[0008] In such a configuration, the summary video generation device divides the input video into video intervals by different division methods for each feature such as predetermined image features, speech features, and acoustic features by the video division means. As a result, the video division means generates a plurality of video interval series for each feature. Since these video interval series are divided by different division methods for each feature, the respective video intervals do not necessarily match.

[0009] Then, the summary video generation device calculates an interval score indicating the degree to which the video of the video interval is a summary video based on the video features of the divided video intervals by the interval score calculation means. For example, the interval score calculation means calculates the interval score from the video of the video interval using a pre-trained neural network.

[0010] Furthermore, the summary video generation device divides the input video into fixed-length unit video intervals that are units of the video for generating the summary video by the unit interval division means. Then, the summary video generation device calculates a unit interval score by adding the interval scores corresponding to the overlapping video intervals according to the time ratio of the video intervals overlapping with a plurality of video interval sequences for each unit video interval by the unit interval score calculation means. As a result, the unit interval score calculation means can integrate the scores calculated from a plurality of features in the unit video interval. Then, the summary video generation device selects unit video intervals within a predetermined time length in order from the ones with a higher degree of being a summary video based on the unit interval scores by the interval selection means. Then, the summary video generation device concatenates the videos of the unit video intervals selected by the interval selection means by the video concatenation means.

[0011] As a result, the summary video generation device can calculate the score (unit interval score) of the unit video interval that is a component of the summary video from the scores (interval scores) of the video intervals of a plurality of features with different division units, and can generate a summary video by combining a plurality of features. Note that the summary video generation device can be operated by a summary video generation program for causing a computer to function as each of the above-described means.

Advantages of the Invention

[0012] According to the present invention, a summary video can be generated by combining a plurality of features with different video division units. In addition, the present invention can accurately select a video segment with high importance by calculating a score combining a plurality of features, and can generate a summary video with higher quality than before.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0014] <Configuration of Summary Video Generation Device> First, with reference to FIG. 1, the configuration of a summary video generation device according to an embodiment of the present invention will be described.

[0015] The summary video generation device 1 generates a summary video that summarizes the content from a video. Note that the video is assumed to include audio. As shown in FIG. 1, the summary video generation device 1 includes a plurality of modal-based interval score calculation means 10, a unit interval division means 20, a unit interval score calculation means 30, and a video summary means 40.

[0016] The modal-based interval score calculation means 10 divides the video based on features, and calculates a score (interval score) indicating the degree of importance as a summary for each divided interval. Here, the modal-based interval score calculation means 10 includes a plurality of modal-based interval score calculation means 10 1 ,10 2 ,10 3 for each modality. In the present invention, a modality is a feature for dividing a video, and is, for example, an image feature, a speech feature, an acoustic feature, etc., a feature of the video itself, or a feature of the audio corresponding to the video.

[0017] For example, here, the modal-based interval score calculation means 10 1 divides the input video V according to the features of the frame images (image features) constituting the video V, and calculates the scores of the divided video intervals. Also, the modal-based interval score calculation means 10 2 divides the input video V according to the speech features extracted from the audio corresponding to the video V, and calculates the scores of the divided video intervals. Also, the modal-based interval score calculation means 10 3 divides the input video V according to the acoustic features extracted from the audio corresponding to the video V, and calculates the scores of the divided video intervals. The modal-based interval score calculation means 10 includes a video division means 11 and an interval score calculation means 12.

[0018] The video division means 11 divides the input video into video intervals by different division methods for each predetermined feature, and generates a plurality of video interval series for each feature. Here, the video division means 11 is configured as including a plurality of modal-based interval score calculation means 10 1 ,10 2 ,10 3 for each predetermined feature. For example, the video segmentation means 11 of the modal-specific interval score calculation means 10 1 detects the switching of the camera which is an image feature of the video V and the cut point which is an editing point, and divides the video V for each cut point. Note that, as a method for detecting the cut point from the video, a general method may be used. For example, the method disclosed in Japanese Patent Laid-Open No. 2008-33749 can be used.

[0019] Also, the video segmentation means 11 of the modal-specific interval score calculation means 10 2 performs speech recognition on the audio corresponding to the video V, and divides the video V at the time of the sentence break which is a speech feature in the speech recognition result. Note that, for speech recognition, a general method can be used.

[0020] Also, the video segmentation means 11 of the modal-specific interval score calculation means 10 3 detects a silent interval having a predetermined length or more which is an acoustic feature based on the acoustic level of the audio corresponding to the video V, and divides the video V at the silent interval (for example, the center of the silent interval).

[0021] The video segmentation means 11 outputs information (for example, the time code at the head of the video interval, the time length) for specifying the individual video intervals of the divided video V to the interval score calculation means 12. Also, here, the video segmentation means 11 outputs the interval length (time length) of the individual video intervals of the divided video V to the unit interval segmentation means 20.

[0022] The interval score calculation means 12 calculates an interval score indicating the degree to which the video of the video interval is a summary video based on the video features of the video intervals divided by the video segmentation means 11. Here, the interval score calculation means 12 is configured as different modal-specific interval score calculation means 10 1 , 10 2 , 10 3 for each predetermined feature, but it may be shared as one configuration.

[0023] The interval score calculation means 12 calculates a feature vector for each video interval, and uses a pre-trained learning model such as a neural network to calculate a score (interval score) indicating that the video of the video interval is a summary video. For example, the interval score calculation means 12 sequentially inputs the frame images constituting the video of the video interval into a pre-trained convolutional neural network (CNN) for image classification, and calculates the feature vector of the video of the video interval by averaging the outputs of the intermediate layers of the CNN over the number of frames. Then, the interval score calculation means 12 uses a learning model (video interval importance calculation model) of a neural network that outputs a score (importance) indicating that the video is a summary video, which is learned with the feature vectors of the videos used for the summary video as positive examples and the feature vectors of the videos not used for the summary video (non-summary videos) as negative examples, to calculate the score of the video interval. Note that the learning method of the video interval importance calculation model will be described later with reference to FIG. 8. The interval score calculation means 12 outputs the calculated interval score to the unit interval score calculation means 30 together with the information specifying each video interval.

[0024] That is, as shown in FIG. 2, the modal-specific interval score calculation means 10 divides the video V into modal n video intervals V n (1), V n (2), …, V n (N n ) by the video segmentation means 11 to generate a video interval series. Here, n indicates the type of modality (feature), and the modal n video interval indicates a video interval divided by a certain modality (feature) n. Also, N n indicates the number of divisions of the video V divided by the modality n. Here, the modal-specific interval score calculation means 10 1 , 10 2 , 10 3 correspond to modalities 1, 2, and 3, respectively. Then, as shown in FIG. 2, the modal-specific interval score calculation means 10 uses the interval score calculation means 12 to calculate the modal n video interval V n(1), V n (2), …, V n (N n ) For each, calculate the interval score, and set it as the modal n - interval score S n (1), S n (2), …, S n (N n ).

[0025] At this time, as shown in FIG. 3, the interval score calculation means 12 of the modal - specific interval score calculation means 10 calculates the feature vector f n (k) from the modal n video interval V n (k). Then, the interval score calculation means 12 calculates the modal n - interval score S n (k) from the feature vector f n (k) by performing the operation of the learned NN.

[0026] The unit interval division means 20 divides the input video V into fixed - length unit video intervals. The interval length (unit interval length) of the unit video interval is the length of the unit video that is a component of the generated summary video. Here, the unit interval division means 20 sets the average value of the interval lengths of all the video intervals divided by the video division means 11 of the modal - specific interval score calculation means 10 as the unit interval length. Note that the unit interval length does not necessarily have to use the average value of the interval lengths of all the video intervals divided by the video division means 11, and for example, a preset fixed value may be used. The unit interval division means 20 outputs information (for example, the start time code and time length of the unit video interval) for specifying each of the video intervals (unit video intervals) divided by the unit interval length to the unit interval score calculation means 30.

[0027] That is, as shown in FIG. 4, the unit interval division means 20 divides the video V into unit video intervals V(1), V(2), …, V(N)) for each unit interval length T. Here, N is the number of divisions (number of unit video intervals) of the video V. Note that the last unit video interval V(N) may be shorter than the unit interval length T.

[0028] The unit interval score calculation means 30 calculates a unit interval score by adding the interval scores corresponding to the overlapping video intervals according to the ratio of the time of the overlapping video intervals for each unit video interval and a plurality of video interval sequences. Specifically, the unit interval score calculation means 30 calculates the unit interval score S(k) of the k-th unit video interval V(k) (k = 1, …, N; N is the number of unit video intervals) according to the following formula (1).

[0029]

Equation

[0030] In formula (1), VOL(n, k) represents the set of modal n video intervals that overlap with the unit video interval V(k) in terms of time for each modality n. Also, N VOL (n, k) represents the number of modal n video intervals belonging to VOL(n, k).

[0031] For example, focusing on the unit video interval V(2) in the example shown in FIG. 5, the set of modal 1 video intervals VOL(1, 2) that overlap with V(2) in terms of time = {V 1 (1), V 1 (2)}, and the number of video intervals N VOL (1, 2) = 2. Similarly, the set of modal 2 video intervals VOL(2, 2) that overlap with V(2) in terms of time = {V 2 (1)}, and the number of video intervals N VOL (2, 2) = 1. Also, the set of modal 3 video intervals VOL(3, 2) that overlap with V(2) in terms of time = {V 3 (2), V 3 (3)}, and the number of video intervals N VOL (3, 2) = 2.

[0032] Also, in formula (1), R(i, k) which is multiplied by the modal n interval score S n (i) (i = 1, …, N M ; N M is the number of modalities) represents the temporal overlap rate between the modal n video interval V n (i) and the unit video interval V(k).

[0033] For example, when calculating the unit interval score S(2) in the unit video interval V(2) of FIG. 5, the unit interval score calculation means 30, according to Equation (1), in modality 1, for V that overlaps with V(2), 1 (1), V 1 (2), the ratio (temporal overlap ratio), S 1 (1), S 1 (2) are added. The same applies to other modalities. Thereby, the unit interval score calculation means 30 can calculate the unit interval score based on the interval scores calculated from a plurality of features. The unit interval score calculation means 30 outputs the calculated unit interval score to the video summarization means 40.

[0034] The video summarization means 40 extracts high-importance videos from the input video and generates a summary video based on the unit interval scores calculated by the unit interval score calculation means 30. The video summarization means 40 includes an interval selection means 41 and a video concatenation means 42.

[0035] The interval selection means 41 selects unit video intervals in order from the ones with a higher degree of being the summary video based on the unit interval scores calculated by the unit interval score calculation means 30. Here, the interval selection means 41 sorts the unit video intervals in order from the ones with higher importance and selects a predetermined number of unit video intervals with higher importance. Note that this number may also be set from the outside. Also, the interval selection means 41 may select unit video intervals up to a predetermined or user-set time length. The interval selection means 41 outputs information (for example, the start time code and time length of the video interval) for specifying the selected unit video intervals to the video concatenation means 42.

[0036] The video concatenation means 42 generates a summary video by concatenating the videos of the unit video intervals selected by the interval selection means 41. This video concatenation means 42 extracts the video specified in the unit video section from the input video V. Then, the video concatenation means 42 concatenates the videos of the extracted unit video sections in time series to generate a summary video SV.

[0037] That is, as shown in FIG. 6, the video summarization means 40, by the section selection means 41, selects the unit video sections V(k 1 ), V(k 2 ), …, V(k N′ ) in order from the highest importance within a predetermined time length in the unit section scores S(1), S(2), …, S(N). Then, the video summarization means 40, by the video concatenation means 42, concatenates the videos of the unit video sections V(k 1 ), V(k 2 ), …, V(k N′ ) in time series to generate a summary video SV.

[0038] As described above, the summary video generation device 1 can generate a summary video that combines different video division units according to modalities (features) and selects video sections with high importance for a plurality of features. Note that the summary video generation device 1 can be operated by a summary video generation program for causing a computer (not shown) to function as each of the above-described means.

[0039] <Operation of the summary video generation device> Next, with reference to FIG. 7 (refer to FIG. 1 as appropriate for the configuration), the operation of the summary video generation device according to the embodiment of the present invention will be described. Here, it is assumed that in the modality-specific section score calculation means 10 1 , 10 2 , 10 3 , steps S1 and S2 operate in parallel. However, the modality-specific section score calculation means 10 1 , 10 2 , 10 3 may also operate in order.

[0040] Steps S1, S1 1In this case, the modal-specific interval score calculation means 10 1 The video segmentation means 11 detects cut points from the video V according to the modality 1 (here, image features), and divides the video V into modality 1 video intervals. Steps S2, S2 1 In this case, the modal-specific interval score calculation means 10 1 The interval score calculation means 12 calculates an interval score indicating the degree of importance for each video interval divided in step S1 1

[0041] Also, in steps S1, S1 2 In this case, the modal-specific interval score calculation means 10 2 The video segmentation means 11 recognizes the audio corresponding to the video V according to the modality 2 (here, speech features), and divides the video V into modality 2 video intervals at the timing of the sentence break in the speech recognition result. Steps S2, S2 2 In this case, the modal-specific interval score calculation means 10 2 The interval score calculation means 12 calculates an interval score indicating the degree of importance for each video interval divided in step S1 2

[0042] Also, in steps S1, S1 3 In this case, the modal-specific interval score calculation means 10 3 The video segmentation means 11 detects a silent interval based on the acoustic level of the audio corresponding to the video V according to the modality 3 (here, acoustic features), and divides the video V into modality 3 video intervals at the silent interval. Steps S2, S2 3 In this case, the modal-specific interval score calculation means 10 3 The interval score calculation means 12 calculates an interval score indicating the degree of importance for each video interval divided in step S1 3

[0043] In step S3, the unit interval segmentation means 20 divides the input video V into unit video intervals with a unit interval length. Note that the unit interval length is in step S1 (S1 1 ​​​, S1 2 , S1 3 It is set as the average value of the interval lengths of all video intervals divided by (). Note that the unit interval length may be a preset fixed value.

[0044] In step S4, for each unit video interval divided in step S3, the unit interval score calculation means 30 calculates a unit interval score indicating the degree to which the interval is a summary video. Here, the unit interval score calculation means 30 calculates the unit interval score in the unit video interval by adding the interval scores for each modality divided in step S2 according to the ratio of the temporal overlap between the video interval for each modality divided in step S1 and the unit video interval according to the formula (1).

[0045] In step S5, the interval selection means 41 of the video summarization means 40 selects unit video intervals within a predetermined time length in order from the highest importance among the unit interval scores calculated in step S4.

[0046] In step S6, the video concatenation means 42 of the video summarization means 40 extracts the videos of the unit video intervals selected in step S5 from the input video V and concatenates them in time series to generate a summary video SV. By the above operations, the summary video generation device 1 can generate a summary video that selects video intervals with high importance for a plurality of features.

[0047] <Regarding the learning method of the video interval importance calculation model> Here, an example of the learning method of the video interval importance calculation model used by the interval score calculation means 12 to calculate the interval score of the video interval will be described. The learning of the video interval importance calculation model can be performed, for example, by the video interval importance calculation model learning device 2 shown in FIG. 8. As shown in FIG. 8, the video interval importance calculation model learning device 2 includes a feature vector generation means 50, a neural network learning means 60, and a video interval importance calculation model storage means 70.

[0048] The feature vector generation means 50 generates a feature vector from the learning video LV which is learning data. This feature vector generation means 50 uses the positive example interval video LV which is the video used for the summary video P and the negative example interval video LV which is the video not used for the summary video N as a pair of learning videos LV, and respectively as feature vectors, the positive example feature vector V P and the negative example feature vector V N are generated.

[0049] Note that the generation of the feature vector is the same as the method of calculating the feature vector by the interval score calculation means 12. That is, the feature vector generation means 50 inputs the frame images constituting the video of the video interval (positive example interval video LV P , negative example interval video LV N ) into a pre-trained convolutional neural network (CNN) for image classification in sequence, and calculates the feature vector (positive example feature vector V P , negative example feature vector V N ) by averaging the outputs of the intermediate layer of the CNN for the number of frames.

[0050] The learning video LV uses, for example, a self-made video and its edited summary video, a program video obtained from a broadcast wave and its summary video distributed via a communication line, etc., uses the summary video as the positive example interval video LV P and can generate a video obtained by deleting frame images similar to the summary video from the original video (self-made video, program video) as the negative example interval video LV N . Of course, if it is known which section of the original video the summary video uses, the negative example interval video LV N may be generated by deleting the section of the summary video from the original video.

[0051] Here, referring to FIG. 9, the learning video LV will be schematically described. Note that the rectangles shown in FIG. 9 indicate the frames of the video, but the frames are thinned out for simplicity of illustration. As shown in FIG. 9, the original video VORG From the summary video V SUM When generating, the extracted segment video LV P1 , LV P2 , … are used as the positive example segment videos LV P of the training video LV. Also, the original video V ORG From the summary video V SUM The segment videos LV P1 , LV P2 , … that have been deleted N1 , LV N2 , LV N3 , … are used as the negative example segment videos LV N of the training video LV.

[0052] Returning to FIG. 8, the description of the configuration of the video segment importance calculation model learning device 2 will be continued. The feature vector generation means 50 outputs the generated positive example feature vector V P and negative example feature vector V N that form a pair to the neural network learning means 60.

[0053] The neural network learning means 60 uses the feature vectors (positive example feature vector, negative example feature vector) generated by the feature vector generation means 50 to learn the internal parameters of the neural network as the parameters of the video segment importance calculation model. This neural network learning means 60 uses the video segment importance calculation model to learn the parameters of the video segment importance calculation model so that the value obtained by subtracting the importance calculated by inputting the negative example feature vector from the importance calculated by inputting the positive example feature vector is increased. The neural network learning means 60 includes a positive example NN operation means 61, a negative example NN operation means 62, and a parameter update means 63.

[0054] The positive example NN (neural network) operation means 61 inputs the positive example feature vector V P generated by the feature vector generation means 50 and calculates the video segment importance calculation model. The positive example NN calculation means 61 performs the calculation of the video segment importance calculation model using the values of the parameters of the video segment importance calculation model stored in the video segment importance calculation model storage means 70. Note that when there is an instruction for recalculation from the parameter update means 63, the positive example NN calculation means 61 inputs the same positive example feature vector V P again and performs the calculation. The positive example NN calculation means 61 outputs the calculation result to the parameter update means 63.

[0055] The negative example NN (neural network) calculation means 62 inputs the negative example feature vector V generated by the feature vector generation means 50 N and calculates the video segment importance calculation model. The negative example NN calculation means 62 performs the calculation of the video segment importance calculation model using the values of the parameters of the video segment importance calculation model stored in the video segment importance calculation model storage means 70. Note that when there is an instruction for recalculation from the parameter update means 63, the negative example NN calculation means 62 inputs the same negative example feature vector V N again and performs the calculation. The negative example NN calculation means 62 outputs the calculation result to the parameter update means 63.

[0056] The parameter update means 63 updates the parameters of the video segment importance calculation model based on the calculation results of the positive example NN calculation means 61 and the negative example NN calculation means 62. This parameter update means 63 updates the parameters so that the value obtained by subtracting the calculation result (importance) of the negative example NN calculation means 62 from the calculation result (importance) of the positive example NN calculation means 61 becomes larger. The parameter update means 63 stores the updated parameters in the video segment importance calculation model storage means 70.

[0057] The update of the parameters by this parameter update means 63 can be performed using the general error backpropagation method. After the parameter update, this parameter update means 63 gives an instruction for recalculation to the positive example NN calculation means 61 and the negative example NN calculation means 62. Then, when the parameter update means 63 reaches a predetermined number of times, or when the variation amount of the parameter update falls below a predetermined threshold value, it gives an instruction to the positive example NN calculation means 61 and the negative example NN calculation means 62 to perform calculations using new feature vectors.

[0058] As a result, the neural network learning means 60 can learn the parameters so that the output value when the positive example feature vector V P is input in the video segment importance calculation model becomes larger than the output value when the negative example feature vector V N is input. With the video segment importance calculation model learned in this way, when a feature vector of a certain segment video is input, the importance indicating whether the segment video is important as a summary video can be calculated based on the output value. The video segment importance calculation model storage means 70 stores the parameters of the video segment importance calculation model learned by the neural network learning means 60. As described above, the video segment importance calculation model learning device 2 can learn a video segment importance calculation model that calculates the importance (score) indicating whether a video is important as a summary video from the feature vector of the video.

[0059] As described above, the configuration and operation of the summary video generation device 1 according to the embodiment of the present invention, and the learning method of the video segment importance calculation model used to calculate the segment score have been described, but the present invention is not limited to this embodiment. For example, here, the modal-specific interval score calculation means 10 of the summary video generation device 1 is configured with three, but at least two or more are sufficient. Also, the video segmentation means 11 of the modal-specific interval score calculation means 10 is not limited to segmenting the video with the above-described features, and various features can be used. For example, the video segmentation means 11 may segment the video section according to changes such as the number of people appearing detected by face recognition, the amount of movement of camera work, etc.

[0060] Also, here, the interval score calculation means 12 calculates the interval score using the learned neural network learning model (video section importance calculation model), but it is not necessarily required to use a neural network. For example, various scores (such as telop score, face recognition score, camera work score, etc.) described in Patent Document 3 may be used.

Explanation of Reference Numerals

[0061] 1 Summary video generation device 10 Modal-specific interval score calculation means 11 Video segmentation means 12 Interval score calculation means 20 Unit interval segmentation means 30 Unit interval score calculation means 40 Video summarization means 41 Interval selection means 42 Video concatenation means 2 Video section importance calculation model 50 Feature vector generation means 60 Neural network learning means 61 Positive example NN calculation means 62 Negative example NN calculation means 63 Parameter update means 70 Video section importance calculation model storage means

Claims

1. A summary video generation device that generates a summary video from an input video, video segmentation means for dividing the input video into video segments by different segmentation methods for each predetermined feature and generating a plurality of video segment sequences for each feature; interval score calculation means for calculating an interval score indicating the degree to which the video of the video segment is the summary video based on the video features of the video segments divided by the video segmentation means; unit interval segmentation means for dividing the input video into unit video segments of a fixed length; unit interval score calculation means for adding the interval scores corresponding to the overlapping video segments according to the ratio of the time of the overlapping video segments with the plurality of video segment sequences for each unit video segment and calculating a unit interval score; interval selection means for selecting the unit video segments within a predetermined time length in order from the ones with a higher degree of being the summary video based on the unit interval scores; video concatenation means for concatenating the videos of the unit video segments selected by the interval selection means; A summary video generation device characterized by comprising the above.

2. One of the features is an image feature, and the video segmentation means is characterized in that it detects a cut point that becomes the image feature from the input video and divides the input video at the cut point. The summary video generation device according to Claim 1.

3. One of the features is a speech feature, and the video segmentation means is characterized in that it performs speech recognition on the audio corresponding to the input video and divides the input video at the time of the break of the sentence that becomes the speech feature. The summary video generation device according to Claim 1.

4. One of the features is an acoustic feature, and the video segmentation means is characterized in that it detects a silent interval that becomes the acoustic feature from the audio corresponding to the input video and divides the input video at the silent interval. The summary video generation device according to Claim 1.

5. The interval score calculation means is characterized in that it calculates the interval score based on a neural network that has learned in advance whether a video is a summary video or a non-summary video. The summary video generation device according to any one of Claims 1 to 4.

6. The unit interval segmentation means is characterized in that it sets the average value of the interval lengths of all the video segments divided by the video segmentation means as the interval length of the unit video segment. The summary video generation device according to any one of Claims 1 to 5.

7. An abstract video generation program for causing a computer to function as the abstract video generation device according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Abrasive device

    JP1980037285A

  • Apparatus for classifying and collecting mist

    JP1983098117A

  • Video detector and summarized video image production device

    JP2000175149A

  • Method, device and program for controlling electronic bulletin board and recording medium

    JP2005044165A

  • Information processing device, information processing system, information processing method and program

    JP2012039550A