Scene extraction system, scene extraction method and scene extraction program
The scene extraction system addresses the lack of personalization in existing technologies by using a division and similarity calculation approach to extract scenes based on user-specific importance levels for video, sound, and speech features, creating digest movies that cater to individual user preferences.
Patent Information
- Application Number
- JP2024048196
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2044-03-25
AI Technical Summary
Existing technologies fail to personalize scene extraction based on individual user preferences, as they do not account for varying user priorities among features such as video, sound, and speech when selecting scenes.
A scene extraction system that includes a storage unit, division unit, similarity calculation unit, favorite scene extraction unit, and similar scene extraction unit, which divides videos into scenes, stores importance levels for each user, calculates similarities based on video, sound, and speech features, and extracts preferred and similar scenes using user viewing history and importance levels.
Enables personalized scene extraction that matches user preferences by considering individual importance levels for video, sound, and speech features, allowing for the creation of digest movies that align with user interests and preferences.
Smart Images

Figure 2025147780000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a scene extraction system, a scene extraction method, and a scene extraction program. [Background technology]
[0002] Conventionally, there exists a technique for extracting a scene that matches a user's preferences based on a plurality of features in the scene.
[0003] For example, Patent Document 1 discloses a technique for extracting scenes that match a user's preferences based on features such as words, images, and sounds in the scenes. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 5338450 Summary of the Invention [Problem to be solved by the invention]
[0005] The features that users prioritize when watching videos vary from user to user. Therefore, in order to provide scenes that match the user's preferences, it is preferable to set for each user which features to prioritize when extracting scenes. However, while the technology of Patent Document 1 can extract scenes based on multiple features, it cannot extract scenes by setting for each user which features the user prioritizes.
[0006] The present invention has been made in view of the above-mentioned circumstances, and an object to be achieved is to provide a new technique for extracting scenes that match the preferences of a user. [Means for solving the problem]
[0007] In order to solve the above problems, the present invention provides a scene extraction system for extracting scenes from a video, comprising: the scene extraction system includes a storage unit, a division unit, a similarity calculation unit, a favorite scene extraction unit, and a similar scene extraction unit; The dividing unit divides the video into scenes, the storage unit stores, for each user in the divided split scenes, a video importance level for a video-related feature, a sound importance level for a sound-related feature, and an utterance importance level for an utterance-related feature, in association with a user ID; the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on a viewing history of the user; The similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the importance associated with the user ID.
[0008] The present invention also provides a scene extraction method executed by a scene extraction system that extracts scenes from a moving image, comprising: the scene extraction system includes a storage unit, a division unit, a similarity calculation unit, a favorite scene extraction unit, and a similar scene extraction unit; a step of the division unit dividing the video into scenes; the storage unit stores, for each user in the divided split scenes, a video importance level for a video feature, a sound importance level for a sound feature, and a speech importance level for a speech feature, in association with a user ID; a step in which the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; a step in which the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on the viewing history of the user; The similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scenes, the visual similarity, the audio similarity, the speech similarity, and the importance level linked to the user ID.
[0009] The present invention also provides a scene extraction program for extracting scenes from a video, comprising: a computer that functions as a storage unit, a division unit, a similarity calculation unit, a preferred scene extraction unit, and a similar scene extraction unit; The dividing unit divides the video into scenes, the storage unit stores, for each user in the divided split scenes, a video importance level for a video-related feature, a sound importance level for a sound-related feature, and an utterance importance level for an utterance-related feature, in association with a user ID; the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on a viewing history of the user; The similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the importance level associated with the user ID.
[0010] With this configuration, it is possible to extract scenes that are preferred by the user based on the features that each user values, and similar scenes that match the user's preferences can be extracted.
[0011] In a preferred embodiment of the present invention, the scene extraction system includes a change importance level creation unit, the change importance level creation unit creates a change importance level by changing all or a part of the importance level based on the importance level and a predetermined condition; The similar scene extraction unit extracts scenes with similar change importance levels from the split scenes based on the preferred scene, the video similarity level, the sound similarity level, the speech similarity level, and the change importance level.
[0012] By configuring in this way, it becomes possible to extract change importance similar scenes, which are similar scenes based on the change importance obtained by changing the importance linked to the user ID, and it is possible to find an importance that suits the user's preferences using the change importance and the importance linked to the user ID.
[0013] In a preferred embodiment of the present invention, the scene extraction system includes an importance updating unit, The importance level update unit updates the change importance level related to the change importance level similar scene as the user's importance level by linking it to a user ID based on the user's viewing history related to the change importance level similar scene.
[0014] With this configuration, it becomes possible to use the viewing history of similar scenes with various degrees of importance, and update the degree of importance to match the user's preferences.
[0015] In a preferred embodiment of the present invention, the scene extraction system includes a presentation unit, The presentation unit presents scenes with a similar change importance level that are lower in proportion than the similar scenes.
[0016] With this configuration, it is possible to simultaneously present similar scenes and scenes with small changes in importance to the user, and by comparing the viewing histories of each, it is possible to find an importance level that suits the user's preferences.
[0017] In a preferred embodiment of the present invention, the scene extraction system includes a digest movie creation unit, The digest movie creation unit creates a digest movie using a plurality of the similar scenes.
[0018] With this configuration, it becomes possible to create a digest movie using a plurality of similar scenes, and to provide additional content to the user.
[0019] In a preferred embodiment of the present invention, the scene extraction system includes a frame extraction unit and a frame similarity calculation unit, the frame extraction unit extracts the first and last frames of the plurality of similar scenes; the frame similarity calculation unit calculates frame similarity between the extracted first frame and last frame; The digest movie creation unit creates the digest movie based on the frame similarity.
[0020] This configuration makes it possible to create a digest video based on the similarity between the first and last frames of similar scenes, allowing the creation of a digest video in which the user does not notice the gaps between similar scenes.
[0021] In a preferred embodiment of the present invention, the scene extraction system includes a preference score calculation unit, the preference score calculation unit calculates a preference score based on the preference scene, the video similarity, the sound similarity, the speech similarity, and the importance level; The digest video creation unit sets the similar scene with the highest preference score as the beginning of the digest video, and creates the digest video by setting the similar scene with the highest frame similarity to the last frame of the first similar scene, among the first frames of the similar scenes excluding the first similar scene, as the next scene after the first similar scene.
[0022] By configuring in this way, it is possible to position the beginning of a digest movie with a similar scene that best suits the user's preferences, thereby creating a digest movie that attracts the user's interest. [Effects of the Invention]
[0023] According to the present invention, a new technique for extracting scenes that match the preferences of a user can be provided. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 1 is a block diagram showing the configuration of a scene extraction system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing the hardware configuration according to the present embodiment. [Figure 3] 3 shows an example of a data configuration stored in a storage unit in this embodiment. [Figure 4] 3 shows an example of a data configuration stored in a storage unit in this embodiment. [Figure 5] 10 is a flowchart of a scene extraction process according to the present embodiment. [Figure 6] 10 is a flowchart of importance update processing in the present embodiment. [Figure 7] 10 is a flowchart of a digest movie creation process according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0025] The scene extraction system of the present invention will be described below with reference to the drawings. The drawings show preferred embodiments. However, the present invention may be embodied in many different forms and is not limited to the embodiments set forth herein.
[0026] For example, although the configuration, operation, etc. of the scene extraction system are described in this embodiment, similar effects can be achieved by an executed method (steps), device, computer program, etc. The program in this embodiment may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server, or the program may be started on an external computer to implement its functions on a client terminal (so-called cloud computing).
[0027] In addition, in this embodiment, the term "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In this embodiment, "information" is represented by, for example, the physical value of a signal value representing voltage or current, the high or low value of a signal value as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculation can be performed on a circuit in the broad sense.
[0028] A circuit in the broad sense is a circuit realized by appropriately combining a circuit, circuitry, processor, memory, etc. That is, it includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), etc.
[0029] <System Overview> Fig. 1 is a block diagram showing the configuration of a scene extraction system according to this embodiment. As shown in Fig. 1, the scene extraction system 0 includes a scene extraction device 1 and a user terminal 2. The scene extraction device 1 is configured to be able to communicate with the user terminal 2 via a network NW.
[0030] The scene extraction device 1 extracts scenes that match a user's preferences from a video based on information about the video, the scene, the user, etc. The scene extraction device 1 may extract similar scenes from other split scenes in a video based on a user's preferred scene extracted from the same video. The scene extraction device 1 may also extract similar scenes from multiple split scenes in multiple videos based on the extracted preferred scene.
[0031] A general-purpose server computer, a personal computer, or the like can be used as the scene extraction device 1. It is also possible to configure the scene extraction device 1 using a plurality of computers.
[0032] A user watches videos and scenes via a user terminal 2. Furthermore, the user may input importance levels and the like via the user terminal 2 and transmit them to the scene extraction device 1. The user terminal 2 may be a terminal device such as a smartphone, a tablet terminal, or a personal computer.
[0033] In this embodiment, the network NW is an IP (Internet Protocol) network, but there is no limitation on the type of communication protocol, and there is also no limitation on the type and scale of the network.
[0034] <Hardware configuration> 2 is a hardware configuration diagram. As shown in FIG. 2(a), the information processing device 10 (scene extraction device 1) has a control unit 101, a storage unit 102, and a communication unit 103, which are used to perform the functions of each unit and each process.
[0035] The control unit 101 includes one or more processors such as a CPU (Central Processing Unit), and controls the overall operation and processing of the information processing device 10 by executing the scene extraction program of the present invention, an OS (Operating System), browser software, and other applications.
[0036] The storage unit 102 is a hard disk drive (HDD), a solid state drive (SSD), a read only memory (ROM), a random access memory (RAM), etc., and stores the scene extraction program according to the present invention and data used when the control unit 101 executes processing based on the program. The control unit 101 executes processing based on the scene extraction program stored in the storage unit 102, thereby realizing the functional configuration described below.
[0037] The communication unit 103 controls communication with the network NW, and performs input necessary for operating the information processing device 10 and output related to the operation results.
[0038] As shown in FIG. 2(b), the terminal device 9 (user terminal 2) has a control unit 91, a storage unit 92, a communication unit 93, an input unit 94, and an output unit 95, and is used to perform the functions of each unit and each process.
[0039] The control unit 91 of the terminal device 9 includes one or more processors such as a CPU, and controls the overall operation and processing of the terminal device 9. The storage unit 92 of the terminal device 9 is an HDD, SSD, ROM, RAM, or the like, and stores the above-mentioned applications and data used when the control unit 91 executes processing based on a program.
[0040] A communication unit 93 of the terminal device 9 controls communication with the network NW. An input unit 94 of the terminal device 9 is a mouse, keyboard, etc., and inputs operation requests from the user / provider to the control unit 91. An output unit 95 of the terminal device 9 is a display, etc., and displays the results of processing by the control unit 91, etc.
[0041] <Functional components> As shown in FIG. 2, the scene extraction device 1 includes a division unit 11, a feature generation unit 12, a similarity calculation unit 13, a preferred scene extraction unit 14, a preference score calculation unit 15, a similar scene extraction unit 16, a presentation unit 17, a change importance creation unit 18, an importance update unit 19, a frame extraction unit 1a, a frame similarity calculation unit 1b, and a digest video creation unit 1c.
[0042] The arrangement of these functional components is an example, and some of the functional components of the scene extraction device 1 may be arranged in one or more devices configured to be able to communicate with the scene extraction device 1 and the user terminal 2.
[0043] <Data structure> 3 and 4 show examples of data configurations stored in the storage unit in this embodiment. The storage unit of the scene extraction device 1 stores split scene information, video similarity information, audio similarity information, speech similarity information, viewing history information, importance information, similarity score information, preference score information, similar scene information, and similar scene similarity information.
[0044] The arrangement of each piece of data is also an example, and some or all of the data stored in the storage unit of the scene extraction device 1 may be stored in one or more devices configured to be able to communicate with the scene extraction device 1 and the user terminal 2.
[0045] <Creating split scenes> 5 is a flowchart of the scene extraction process in this embodiment. First, in step S501, the division unit 11 divides a video by scene to create divided scenes (for example, by chapter, by batter in baseball, or by shots and other scenes in a soccer game). The division unit 11 may divide the video into one or more scenes. In this embodiment, the division unit 11 divides all the videos it manages (stored in the storage unit, received via the user terminal) by scene.
[0046] The dividing unit 11 may divide the video at equal intervals of a fixed time (for example, 5 minutes), or may detect change points in the video footage and divide it automatically. Furthermore, the dividing unit 11 may divide the video footage into similar videos as one cluster by clustering the video footage in the time direction, or may divide the video footage by tags (for example, chapters) if the video footage has been assigned tags for division.
[0047] The dividing unit 11 stores in the storage unit split scene information relating to which moving image the split scene was split from and which section of the original moving image the split scene belongs to. The split scene information includes a moving image ID, a split scene ID, and a moving image section, as shown in FIG. 3(a). This allows the scene extraction device 1 to refer to, for example, that split scene ID "M1_C2" is a split scene of the moving image with moving image ID "M1" and is a scene from a section of the moving image from 10 to 20 minutes.
[0048] <Feature generation> In step S502, the feature generation unit 12 generates video features, sound features, and speech features for the split scenes created by the split unit 11. For example, a video or a scene can be decomposed into multiple elements, and features for each element can be generated. In this embodiment, a video or a scene is decomposed into three elements: video, sound, and speech content (meaning of the conversation), and the feature generation unit 12 generates each feature as a video feature, sound feature, and speech feature.
[0049] In this embodiment, video features are features related to things that appear in a video or scene and their movements (for example, the color, shape, size, movement, and positional relationship of an object). In this embodiment, sound features are features related to sounds in a video or scene (for example, the volume, speed, frequency, wavelength, etc.). In this embodiment, speech features are features related to the content of speech (audio) in a video or scene (meaningful content that can be understood by humans), and are generated, for example, by extracting human speech and performing natural language processing.
[0050] The feature generator 12 may generate features by vectorizing the video (moving and still) and sound (audio and non-audio) included in the split scene. As shown in Fig. 3(a), the feature generator 12 associates the generated features with the split scene ID and stores them in the storage unit.
[0051] For example, the feature generator 12 generates video feature amounts by breaking down the split scenes into frames (still images) and creating a group of image files.
[0052] Video features can be generated from a group of image files using a proprietary trained encoding model or a trained model available as open source software. For example, a Vision Transformer-based model can be used, which can generate 400-dimensional features.
[0053] Additionally, pre-trained models that describe image patterns in natural language are also publicly available, and features can also be generated by embedding natural language information extracted from images. A model that embeds natural language information can be independently trained, or a publicly available pre-trained model can be used. For example, one publicly available model is the encoding published by OpenAI as an API (Application Programming Interface), which can generate 1536-dimensional features.
[0054] For example, the feature generator 12 generates sound features by extracting only sound information from a split scene and creating a sound file. The extracted sound information may or may not include human voice.
[0055] Feature generation from audio files can be performed using a proprietary trained encoding model or a trained model available as open source software. For example, a Vision Transformer-based model can be used, which can generate 383232-dimensional features.
[0056] For example, the feature generation unit 12 extracts speech information related to human speech (voice) from the split scenes to generate speech features. The feature generation unit 12 may use the voice information extracted from the split scenes to perform voice recognition and convert it into natural language to create a speech language file.
[0057] The features from the spoken language file may be obtained using an independently trained encoding model or a pre-trained model published as open source software. For example, the encoding published as an API by OpenAI may be used, which can generate 1536-dimensional features.
[0058] <Calculation of similarity> In step S503, the similarity calculation unit 13 calculates the similarity between the split scenes. The similarity calculation unit 13 calculates the video similarity, audio similarity, and speech similarity between the split scenes based on the video feature amount, audio feature amount, and speech feature amount of the split scenes.
[0059] The similarity calculation unit 13 may calculate the similarities (visual similarity, audio similarity, speech similarity) of the respective feature amounts between all the split scenes and store them in a storage unit as shown in FIGS. 3(b) to 3(d). When there are N split scenes, the similarity calculation unit 13 performs N×N calculations to calculate the similarity of each feature amount. The similarity calculation unit 13 may calculate the similarity using cosine similarity, which calculates the similarity between vectors.
[0060] Furthermore, the similarity calculation unit 13 may or may not calculate the similarity between the same split scenes (for example, the similarity between split scene IDs "M1_C1" and "M1_C1"). If it is acceptable to extract scenes that the user has viewed once as similar scenes that are similar to the user's preferred scenes, the similarity calculation unit 13 may calculate the similarity between the same split scenes so that it is high. On the other hand, if it is not acceptable to extract scenes that the user has viewed once, the similarity calculation unit 13 may calculate the similarity between the same split scenes so that it is low, or may not calculate the similarity at all, as shown in FIGS. 3(b) to 3(d).
[0061] <Extraction of favorite scenes> In step S504, the favorite scene extraction unit 14 extracts favorite scenes that the user likes from the plurality of split scenes based on the user's viewing history. The favorite scene extraction unit 14 stores viewing history information related to the viewing history of split scenes for each user in a storage unit. The viewing history information includes a user ID, a split scene ID, and a cumulative viewing time, as shown in FIG. 3(e).
[0062] The favorite scene extraction unit 14 counts the viewing time for each split scene for each user and calculates the cumulative viewing time. This allows the scene extraction device 1 to refer to, for example, that the user ID "U1" has previously viewed the split scene with the split scene ID "M2_C1" for a total of 280 seconds.
[0063] The preference scene extraction unit 14 extracts preference scenes that are preferred by the user based on the cumulative viewing time. For example, if the cumulative viewing time of a split scene exceeds a certain time (threshold), the preference scene extraction unit 14 may extract the split scene as a preference scene. When extracting a preference scene, the preference scene extraction unit 14 may store the extracted preference scene in the storage unit as shown in FIG. 3(e) (by linking the split scene ID with a preference flag "1").
[0064] The favorite scene extraction unit 14 may also extract favorite scenes using the proportion of the cumulative viewing time to the playback time of the split scene. If the cumulative viewing time of a split scene is equal to or greater than a predetermined proportion (e.g., half) of the playback time of the scene, the favorite scene extraction unit 14 extracts the split scene as a favorite scene. In this embodiment, this proportion is set to half, but is not limited to this and the proportion can be set by an administrator or the like via a terminal.
[0065] For example, since the playback time of split scene ID "M1_C1" is 600 seconds, if the cumulative viewing time of user "U1" is 300 seconds, the favorite scene extraction unit 14 extracts split scene ID "M1_C1" as the favorite scene of user "U1." On the other hand, since the playback time of split scene ID "M2_C2" is 630 seconds, if the cumulative viewing time of user "U1" is 300 seconds, the favorite scene extraction unit 14 does not extract split scene ID "M2_C2" as the favorite scene of user "U1."
[0066] Alternatively, the number of favorite scenes may be limited. For example, the favorite scene extraction unit 14 may extract the top M (M=1, 2, . . .) split scenes in terms of cumulative viewing time as favorite scenes. Alternatively, the favorite scene extraction unit 14 may extract the top M split scenes in terms of the ratio of cumulative viewing time to the playback time of the split scenes as favorite scenes.
[0067] The favorite scene extraction unit 14 may further extract favorite scenes using a genre associated with the split scene. The storage unit may associate a genre with a moving image ID or a split scene ID and store the ID, so that the favorite scene extraction unit 14 may extract favorite scenes using the genre.
[0068] For example, by accepting a user's preferred genre in advance via a user terminal and storing the genre in a memory unit linked to the user ID, the preferred scene extraction unit 14 can extract preferred scenes based on the user's cumulative viewing time and the accepted genre.
[0069] <Calculation of preference score> In step S505, the preference score calculation unit 15 calculates a preference score based on the preferred scene, video similarity, sound similarity, speech similarity, and importance associated with the user ID. The storage unit stores the video importance, sound importance, and speech importance for each user, as shown in FIG. 4(f). The sum of the video importance, sound importance, and speech importance for each user may be 1. The sum of the video importance, sound importance, and speech importance may be different for each user. The preference score calculation unit 15 may receive the importance for each user via the user terminal 2.
[0070] When extracting similar scenes, which of video, sound, or speech is emphasized varies from user to user. The video importance level is information regarding how much importance the user places on the video element when extracting similar scenes. The sound importance level is information regarding how much importance the user places on the sound element when extracting similar scenes. The speech importance level is information regarding how much importance the user places on the speech element (natural language) when extracting similar scenes.
[0071] The preference score calculation unit 15 can calculate a preference score for extracting similar scenes suitable for each user by applying a degree of importance (weight) to each similarity calculated by the similarity calculation unit 13.
[0072] First, the preference score calculation unit 15 calculates similarity score information as shown in Fig. 4(g) based on the preferred scene, video similarity, sound similarity, and speech similarity. The brackets next to each score in Fig. 4(g) indicate the calculation used to calculate the respective score (this is written for convenience to show the correspondence with the similarities in Figs. 3(b) to (d)). For example, in the case of the video score of split scene ID "M1_C2," the video similarity of 0.12 between "M1_C2" and the preferred scene "M1_C1" and the video similarity of 0.22 between "M1_C2" and the preferred scene "M2_C1" are calculated (totaled) to obtain a video score of 0.34.
[0073] The preference score calculation unit 15 calculates a similarity score (video score, sound score, speech score) by adding up the similarity between each of the split scenes and the preference scene, based on the similarity between all of the preference scenes (preference flag is 1) and each of the split scenes.
[0074] For example, in FIG. 3(e), the preferred scenes for user ID "U1" are split scene IDs "M1_C1" and "M2_C1," so the similarity score is calculated by adding up the respective similarities. The video score for split scene ID "M1_C2" in FIG. 4(g) is 0.34, which is the sum of the video similarity of split scene IDs "M1_C1" and "M1_C2" (0.12) and the video similarity of split scene IDs "M2_C1" and "M1_C2" (0.22). In this way, a similarity score can be calculated that takes multiple preferred scenes into consideration, and similar scenes can be extracted.
[0075] In addition, the preference score calculation unit 15 may use the video similarity of 0.12 between split scene IDs "M1_C1" and "M1_C2" and the video similarity of 0.22 between split scene IDs "M2_C1" and "M1_C2" as the video score for split scene ID "M1_C2".
[0076] The preference score calculation unit 15 calculates a preference score based on the similarity score and the importance level. The preference score calculation unit 15 calculates a preference score for each split scene using, for example, formula (1), and stores the preference score in the storage unit in association with the user ID and split scene ID as shown in FIG. 4(h). The brackets in the preference score in FIG. 4(h) indicate the calculation using formula (1) (this is written for convenience to show the correspondence with the score in FIG. 4(g)). For example, the preference score for split scene ID "M2_C2" is 1.022, and the brackets indicate the calculation used to derive that score.
[0077]
number
[0078] <Extraction of similar scenes> In step S506, the similar scene extraction unit 16 extracts scenes similar to the user's favorite scenes from the split scenes based on the favorite scenes, video similarity, audio similarity, speech similarity, and importance associated with the user ID. The similar scene extraction unit 16 extracts scenes similar to the user's favorite scenes from the split scenes based on the preference scores calculated by the preference score calculation unit 15.
[0079] For example, the similar scene extraction unit 16 may extract split scenes having preference scores equal to or greater than a certain numerical value (threshold) as similar scenes. The similar scene extraction unit 16 may also extract the top L (L=1, 2, . . .) split scenes having preference scores as similar scenes.
[0080] <Recommendation of similar scenes> In step S507, the presentation unit 17 presents (recommends) similar scenes based on the similar scenes extracted by the similar scene extraction unit 16. The presentation unit 17 may further present similar scenes for each genre based on the genre associated with the video ID or the similar scene ID.
[0081] <Creating the change importance level> 6 is a flowchart of the importance level update process in this embodiment. Processes similar to those described above are denoted by the same reference numerals, and their description will be omitted. In step S601, the change importance level creation unit 18 creates a change importance level by changing all or part of the importance level based on the importance level and a predetermined condition.
[0082] The change importance level is an importance level obtained by changing the importance level associated with the user ID based on a predetermined condition. In order to extract a scene (a change importance level similar scene) different from the similar scene extracted based on the original importance level (the importance level associated with the user ID), the change importance level creation unit 18 creates an importance level (a change importance level) by changing the importance level associated with the user ID.
[0083] The change importance level creation unit 18 creates each change importance level (changed video importance level, changed sound importance level, changed speech importance level) using, for example, equations (2) to (4). The change importance level creation unit 18 creates a changed video importance level that is X times the video importance level (for example, X=1.1) using equations (2) to (4), and further creates a changed sound importance level and a changed speech importance level such that the sum of the changed video importance level, changed sound importance level, and changed speech importance level is 1.
[0084]
number
[0085] The change importance level creation unit 18 can create a changed sound importance level by multiplying the sound importance level by X and a changed speech importance level by multiplying the speech importance level by X, similarly to equations (2) to (4). X may be an integer or a non-integer fraction. Furthermore, the change importance level creation unit 18 may determine X randomly or may change X periodically.
[0086] <Extraction of scenes with similar change importance based on change importance> In step S602, the preference score calculation unit 15 may calculate a preference score based on the preferred scene, video similarity, sound similarity, speech similarity, and change importance level. The similar scene extraction unit 16 extracts change importance level similar scenes from the split scenes based on the preferred scene, video similarity, sound similarity, speech similarity, and change importance level.
[0087] <Recommendation of scenes with similar change importance> In step S603, the presentation unit 17 presents scenes with a similar change importance level at a certain rate.The presentation unit 17 presents scenes with a similar change importance level at a lower rate than the similar scenes.
[0088] For example, when presenting K (10, etc.) scenes, the presentation unit 17 presents 70% of similar scenes with the highest preference scores, 10% of similar scenes with a change emphasis level that has X times the video emphasis level and the highest preference scores, 10% of similar scenes with a change emphasis level that has X times the sound emphasis level and the highest preference scores, and 10% of similar scenes with a change emphasis level that has X times the speech emphasis level and the highest preference scores.
[0089] In this way, by presenting scenes including a small number of scenes with similar change importance levels among the scenes to be presented by the presentation unit 17, there is a possibility that the user will prefer to view the scenes with similar change importance levels. If the user prefers to view the scenes with similar change importance levels, which are presented in smaller numbers than the similar scenes, it can be determined that the change importance level is more preferred by the user than the importance level associated with the current user ID.
[0090] <Importance update> In step S604, the importance updating unit 19 updates the change importance related to the change importance similar scene as the user's importance by linking it to the user ID based on the user's viewing history for the change importance similar scene. If the importance updating unit 19 determines that the user prefers to watch a scene with a change importance similar to that scene with a video importance X times higher, the importance updating unit 19 updates the user's importance to the change importance (the changed video importance obtained by multiplying the video importance by X, and the sound importance and speech importance that have changed accordingly).
[0091] For example, if the split scene that the user watches for the longest time in one day is a scene with a similar change importance with the video importance multiplied by X, the importance update unit 19 determines that the user prefers to watch scenes with a similar change importance with the video importance multiplied by X, and updates the change importance as the user's importance by linking it to the user ID.
[0092] In this way, by periodically changing the change importance level and further updating the importance level based on the user's viewing history, it is possible to bring the importance level closer to the user's preferences.
[0093] <Frame extraction from similar scenes> 7 is a flowchart of the digest movie creation process in this embodiment. Processes similar to those described above are denoted by the same reference numerals, and their description will be omitted. In step S701, the frame extraction unit 1a extracts the first and last frames (still images) of the multiple similar scenes extracted by the similar scene extraction unit 16.
[0094] <Calculation of frame similarity> In step S702, the frame similarity calculation unit 1b calculates the frame similarity between the first frame and the last frame extracted by the frame extraction unit 1a. The frame similarity calculation unit 1b calculates the frame similarity between the first frame of the similar scene extracted by the frame extraction unit 1a and the last frame of the similar scene extracted by the frame extraction unit 1a, excluding the first similar scene.
[0095] The frame similarity calculation unit 1b generates feature amounts for the first and last frames of all similar scenes extracted by the similar scene extraction unit 16, and stores similar scene information such as that shown in Fig. 4(i) in a storage unit. It is conceivable that the frame similarity calculation unit 1b generates feature amounts in the same way as the feature generation unit 12 generates feature amounts.
[0096] The frame similarity calculation unit 1b further calculates similarities based on the feature amounts of the first frames of all similar scenes and the feature amounts of the last frames of all similar scenes, and stores similar scene similarity information such as that shown in FIG. 4(j) in the storage unit. The frame similarity calculation unit 1b may calculate the similarity of the feature amounts using cosine similarity, which calculates the similarity between vectors. In this embodiment, as shown in FIG. 4(j), the frame similarity calculation unit 1b does not calculate the similarity between the first frame and the last frame of the same similar scene.
[0097] <Creating a digest video> In step S703, the digest movie creation unit 1c creates a digest movie using the plurality of similar scenes extracted by the similar scene extraction unit 16. The digest movie creation unit 1c may create a digest movie by connecting the extracted similar scenes in descending order of preference score, for example.
[0098] The digest video creation unit 1c may create a digest video by selecting the extracted similar scene with the highest preference score calculated by the preference score calculation unit 15 as the beginning of the digest video, and selecting the similar scene with the highest frame similarity to the last frame of the first similar scene, among the first frames of the similar scenes excluding the first similar scene, as the next scene after the first similar scene.
[0099] The digest video creation unit 1c may further create a digest video by selecting the similar scene that has the greatest frame similarity with the last frame of the Jth (J=1, 2, ...) similar scene from the first frame of similar scenes excluding similar scenes used in creating previous digest videos, as the J+1th similar scene.
[0100] For example, suppose the similar scene extraction unit 16 extracts similar scene IDs "M4_C1," "M6_C2," "M10_C1," and "M11_C1" as similar scenes. Furthermore, if the similar scene ID "M10_C1" has the highest preference score, the digest movie creation unit 1c will select the similar scene ID "M10_C1" as the first scene.
[0101] According to FIG. 4(j), the similarity between the last frame of similar scene ID "M10_C1" and the first frame of similar scene ID "M4_C1" is the greatest, so the digest video creation unit 1c selects similar scene ID "M4_C1" as the next scene.
[0102] Furthermore, as shown in Figure 4(j), the similarity between the last frame of similar scene ID "M4_C1" and the first frame of similar scene ID "M10_C1" is the highest, but similar scene ID "M10_C1" has already been used to create the digest movie. Therefore, the digest movie creation unit 1c excludes this similar scene and selects similar scene ID "M11_C1," which has the highest frame similarity to the last frame of similar scene ID "M4_C1," as the next scene. The digest movie creation unit 1c repeats this process until all extracted similar scenes have been used to create the digest movie.
[0103] As described above, the configuration of the present invention can provide a new technique for extracting scenes that match the preferences of a user. [Explanation of symbols]
[0104] 0 Scene Extraction System 1 Scene extraction device 11 Division 12 Feature generation unit 13 Similarity calculation unit 14 Preferred Scene Extraction Unit 15 Preference score calculation section 16 Similar scene extraction unit 17 Presentation section 18 Change Importance Creation Department 19 Emphasis update section 1a Frame extraction section 1b Frame similarity calculation section 1c Digest Video Creation Department 2. User terminal NW Network
Claims
1. A scene extraction system for extracting scenes from a video, comprising: the scene extraction system includes a storage unit, a division unit, a similarity calculation unit, a favorite scene extraction unit, and a similar scene extraction unit; The dividing unit divides the video into scenes, the storage unit stores, for each user in the divided split scenes, a video importance level for a video-related feature, a sound importance level for a sound-related feature, and an utterance importance level for an utterance-related feature, in association with a user ID; the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on a viewing history of the user; the similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the importance level associated with the user ID; Scene extraction system.
2. the scene extraction system includes a change importance level creation unit, the change importance level creation unit creates a change importance level by changing all or a part of the importance level based on the importance level and a predetermined condition; the similar scene extraction unit extracts scenes with similar change importance levels from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the change importance level; The scene extraction system according to claim 1 .
3. The scene extraction system includes an importance update unit, the importance level update unit updates the change importance level related to the change importance level similar scene as the importance level of the user by linking the change importance level related to the change importance level similar scene with the user ID based on the user's viewing history of the change importance level similar scene; The scene extraction system according to claim 2 .
4. The scene extraction system includes a presentation unit, the presentation unit presents scenes with a similar change importance level that have a lower ratio than the similar scenes; The scene extraction system according to claim 3 .
5. the scene extraction system includes a digest movie creation unit; the digest movie creation unit creates a digest movie using a plurality of the similar scenes; The scene extraction system according to claim 1 .
6. the scene extraction system includes a frame extraction unit and a frame similarity calculation unit; the frame extraction unit extracts the first and last frames of the plurality of similar scenes; the frame similarity calculation unit calculates frame similarity between the extracted first frame and last frame; the digest movie creation unit creates the digest movie based on the frame similarity. The scene extraction system according to claim 5 .
7. The scene extraction system includes a preference score calculation unit, the preference score calculation unit calculates a preference score based on the preference scene, the video similarity, the sound similarity, the speech similarity, and the importance level; the digest movie creation unit creates the digest movie by designating the similar scene with the highest preference score as the first similar scene, and designating, among the first frames of similar scenes excluding the first similar scene, a similar scene with the highest frame similarity to the last frame of the first similar scene as the scene next to the first similar scene. The scene extraction system according to claim 6 .
8. A scene extraction method executed by a scene extraction system that extracts scenes from a video, comprising: the scene extraction system includes a storage unit, a division unit, a similarity calculation unit, a favorite scene extraction unit, and a similar scene extraction unit; a step of the division unit dividing the video into scenes; the storage unit stores, for each user in the divided split scenes, a video importance level for a video feature, a sound importance level for a sound feature, and an utterance importance level for an utterance feature, in association with a user ID; a step in which the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; a step in which the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on the viewing history of the user; and a step in which the similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the importance associated with the user ID. Scene extraction method.
9. A scene extraction program for extracting scenes from a video, a computer that functions as a storage unit, a division unit, a similarity calculation unit, a preferred scene extraction unit, and a similar scene extraction unit; The dividing unit divides the video into scenes, the storage unit stores, for each user in the divided split scenes, a video importance level for a video-related feature, a sound importance level for a sound-related feature, and an utterance importance level for an utterance-related feature, in association with a user ID; the similarity calculation unit calculates a video similarity, a sound similarity, and a speech similarity between the split scenes based on a video feature, a sound feature, and a speech feature of the split scenes; the preferred scene extraction unit extracts preferred scenes preferred by the user from a plurality of split scenes based on a viewing history of the user; the similar scene extraction unit extracts similar scenes from the split scenes based on the preferred scene, the video similarity, the sound similarity, the speech similarity, and the importance level associated with the user ID; Scene extraction program.
Citation Information
Patent Citations
Video attribute information output apparatus, video summarizing device, program, and method for outputting video attribute information
JP2008176538A
Broadcast program recording system, broadcast program recorder, and broadcast program recording method
JP2009088828A
Video reproducing device and video reproducing method
JP2010016618A
Device and method for providing information
JP2011107936A
Similar video output method, similar video output apparatus and similar video output program
JP2012222450A