Summary generating device and summary generating method
The summary generation device addresses the challenge of providing suitable summaries upon resuming playback by determining relevant summary sentences based on an index value calculated from the interruption position, ensuring contextually relevant and engaging summaries.
Patent Information
- Application Number
- JP2021031239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-02-26
AI Technical Summary
Conventional summary generation methods fail to provide suitable summaries when resuming playback of story-like content from an arbitrary interruption position, as they cannot effectively summarize sentences from previous and subsequent playback ranges.
A summary generation device and method that determine a first range from the interruption position, calculate an index value for sentences in a second range, and extract summary sentences based on this index value, ensuring the summary is relevant and easily understandable.
The solution generates summaries that are contextually relevant to the interruption position, enhancing user understanding and motivation to continue playback by providing a suitable summary of the content.
Smart Images

Figure 0007673910000001 
Figure 0007673910000002 
Figure 0007673910000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a summary generating device and a summary generating method. [Background technology]
[0002] When resuming playback of dramas, novels, manga, and other content with a story, the previous content may be forgotten. In this regard, content that consists of multiple episodes in chronological order may be played at the beginning of each episode, including a summary of the previous episode and a preview of the next episode.
[0003] In recent years, the playback of dramas and other story-based content has been changing from a passive form such as television to an active form such as online, where users are charged for the playback period through subscription services.
[0004] When playing content in this way, the conventional style of presenting summaries only at the beginning of an episode means that if playback is resumed from the interrupted point in the middle of an episode, the summary is not presented, and so the summary may not be appropriate for resuming playback.
[0005] In this regard, for example, Japanese Patent Laid-Open Publication No. 2007-336085 (hereinafter, Patent Document 1) discloses a method for generating a preview of the part following the position where the reproduction of the content was interrupted. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] JP 2007-336085 A Summary of the Invention
[0007] However, when playback is resumed from an arbitrary interruption point, depending on the interruption point, a summary that is easy to understand may not be obtained simply by extracting sentences from the previous playback range or from the subsequent range. Therefore, a summary generation device and a summary generation method that can generate an easy-to-understand summary according to the interruption point are desired.
[0008] Here, the summary generation device generates a summary of a target content and includes a processor configured to determine a first range from a pause position of the target content, calculate index values of sentences included in a second range determined from the playback pause position, which is a range from which constituent sentences of the summary are extracted, from the first range, and extract the constituent sentences of the summary from the second range based on the index values.
[0009] In addition, the summary generation method is a method for generating a summary of a target content, and includes determining a first range from the interruption position of the target content, calculating index values from the first range of sentences included in a second range, which is a range from which the constituent sentences of the summary are extracted, and extracting the constituent sentences of the summary from the second range based on the index values.
[0010] Further details will be described in the following embodiments. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a schematic diagram showing a configuration of a summary generating device according to the first embodiment and an example of a process executed by a summary generating method according to the first embodiment. [Diagram 2] FIG. 2 is a diagram for explaining the configuration of content for which a summary is to be generated by the summary generating device. [Diagram 3] FIG. 3 is a diagram for explaining a method for determining the range for generating a summary of content. [Figure 4] FIG. 4 is a diagram for explaining the process of extracting summary constituent sentences. [Diagram 5]FIG. 5 is a flowchart illustrating an example of a summary generating method according to the first embodiment. [Figure 6] FIG. 6 is a schematic diagram showing a configuration of a summary generating device according to the second embodiment and an example of a process executed by the summary generating method according to the second embodiment. [Figure 7] FIG. 7 is a flowchart illustrating an example of a summary generating method according to the second embodiment. [Figure 8] FIG. 8 is a diagram for explaining a specific example of a method for generating a summary of a target range using a partial summary provided in the content as a gold standard. [Figure 9] FIG. 9 is a diagram for explaining a specific example of a method for generating a summary of a target range using a partial summary provided in the content as a gold standard. [Figure 10] FIG. 10 is a diagram for explaining a specific example of a method for generating a summary of a target range using a partial summary provided in the content as a gold standard. [Figure 11] FIG. 11 is a schematic diagram showing a configuration of a summary generating device according to the third embodiment and an example of a process executed by the summary generating method according to the third embodiment. [Figure 12] FIG. 12 is a diagram showing an example of the distribution of appearances of constituent sentences of a partial summary generated for a certain piece of content in the content. [Figure 13] FIG. 13 is a diagram showing a specific example of the extraction data. [Figure 14] FIG. 14 is a diagram showing a specific example of the placement data. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] <1. Overview of summary generation device and summary generation method>
[0013] (1) A summary generation device according to one embodiment is a summary generation device that generates a summary of a target content, and includes a processor. The processor is configured to: determine a first range in the target content from a playback interruption point of the target content; calculate index values of sentences included in a second range determined from the playback interruption point, the second range being at least partially different from the first range, from which the constituent sentences of the summary in the target content are extracted; and extract the constituent sentences of the summary from the second range based on the index values.
[0014] The target content has a story and is played back in chronological order. For example, it corresponds to videos such as movies, dramas, and animations, literary works such as novels, and performances such as lectures and classes. The elements to be played back include text, music, images, and the like, but here we focus on text. The text may be dialogue, sentences, or nouns. Music or images that are played back simultaneously may be added to the text.
[0015] A summary is a short edited version of the story of a certain range of content (hereinafter referred to as the summary range). Summaries can be either primary summaries or secondary summaries. A primary summary is what is known as an outline, while a secondary summary is what is known as a preview. A primary summary refers to a summary that introduces the range before the point at which playback begins (hereinafter referred to as the start position). A secondary summary refers to a summary that introduces the range after the start position. A secondary summary may also include content beyond the start position.
[0016] The first range is a range used to calculate index values of sentences included in the second range, and is a range within the target content. The first range is, for example, an expected playback range. The expected playback range is a range after the playback interruption position, and is a predetermined range from the playback interruption position. The predetermined range may be a range set in advance. The predetermined range may be determined based on the attributes of the user, the attributes of the target content, the playback tendency of the user, etc. The playback tendency of the user is, for example, the playback behavior of the user. Alternatively, an average value of general playback amounts may be used.
[0017] The playback interruption position is the position where playback was interrupted. As an example, the playback interruption position coincides with the playback start position. In this case, the interruption position is the reference position for generating a summary.
[0018] The second range is a range to be summarized and is determined based on the playback interruption position, and may be different from, may be the same as, or may at least partially overlap with, the first range.
[0019] The summary target range is, for example, the range from the beginning of the story to the end of the expected playback range determined based on the playback interruption position. The expected playback range is determined by the user's viewing / reading behavior, the average amount of general viewing and reading, or the provider's regulations. By setting the summary target range appropriately, it is possible to generate a summary that is easier to understand.
[0020] The index value is a value calculated from the first range for the sentences included in the second range, and includes, for example, importance. By extracting the constituent sentences of the summary from the second range based on the index value calculated from the first range, the generated summary includes sentences that take into account the first range determined from the playback interruption position. This makes it possible to generate a summary that is suitable for resuming playback.
[0021] (2) Preferably, the index value includes a value obtained based on the words included in the first range. The value obtained based on the words included in the first range is, for example, importance. This allows the generated summary to include sentences that take into account the value obtained based on the words included in the first range.
[0022] (3) Preferably, the first range is a range after the playback interruption position and is a predicted playback range of the target content determined from the playback interruption position, thereby making it possible to generate a summary that takes into account the playback content to be resumed after the interruption.
[0023] (4) Preferably, the second range is determined depending on whether the summary is the first summary or the second summary, the first summary being a summary generated with the range not including the part after the playback interruption position as the second range, and the second summary being a summary generated with the range including the part after the playback interruption position as the second range. The first summary is a synopsis, and the second summary is a preview. As a result, the first summary includes sentences extracted from the range up to the playback interruption position. Therefore, a summary is generated that reminds the user of the contents up to the playback interruption position and increases the user's motivation to play. Also, the second summary includes sentences after the playback interruption position. Therefore, a summary is generated that increases the user's motivation to play.
[0024] (5) Preferably, extracting the constituent sentences of the summary includes generating a subset having the sequentially extracted sentences by sequentially extracting a plurality of sentences contained in the constituent sentences from the second range based on an index value, the index value including a similarity between the sentences contained in the second range and the subset. In this case, if the extraction based on the index value is to extract sentences with high similarity, the constituent sentences will be a set of sentences with uniformity in content, and the content of the summary sentence will be easy to clarify. On the other hand, if the extraction based on the index value is to extract sentences with low similarity, the constituent sentences will be a set of sentences with diversity, and the summary will be easy to have well-balanced content.
[0025] (6) Preferably, the subset includes sentences extracted from a summary prepared in advance for the target content. This allows the summary to be generated using the constituent sentences of the summary prepared in advance for the target content. This makes the process easier than extracting all of the constituent sentences based on their index values.
[0026] (7) Preferably, the processor is further configured to select reference content based on the target content, and extracting the constituent sentences of the summary from the second range includes extracting the constituent sentences of the summary from the second range with reference to extraction data associated with the reference content. The reference content is content different from the target content and is content for which a summary is prepared. The extraction data is data representing a tendency of the positions in the reference content of the constituent sentences of the summary prepared in the reference content. By extracting the constituent sentences of the summary from the second range using the extraction data, a summary of the target content can be easily generated, and a summary that is as easy to understand as the reference content can be generated.
[0027] (8) A summary generation method according to an embodiment is a method for generating a summary of target content, the method being a method for generating a summary of the target content in the summary generation device described in (1) to (7), thereby obtaining a summary generated by the summary generation device described in (1) to (7).
[0028] 2. Examples of summary generation methods and summary generation devices
[0029] [First embodiment]
[0030] The summary generating device 10 according to this embodiment generates a summary of content. The content handled in this embodiment has a story and is played back in chronological order. For example, this corresponds to videos such as movies, dramas, and animations, literary works such as novels, and performances such as lectures and classes. Elements to be played back include text, music, images, and the like, but here, attention is focused on text. The text may be dialogue, sentences, or nouns. Music and images that are played back simultaneously may be added to the text.
[0031] A summary is a short edited version of the story of a certain range of content (hereafter referred to as the summary range). Summaries can be either synopses (first type of summary) or previews (second type of summary). A synopsis is a summary that introduces the range before the point at which playback begins (hereafter referred to as the start position). A preview is a summary that introduces the range after the start position. A preview may include content beyond the start position.
[0032] 1, the summary generating device 10 is configured as a computer having a processor 11 and a memory 12. The processor 11 is, for example, a CPU. The memory 12 includes a flash memory, an EEPROM, a ROM, a RAM, etc. Alternatively, the memory 12 may be a primary storage device or a secondary storage device.
[0033] The memory 12 stores a generation program 121 executed by the processor 11. The processor 11 executes a summary generation process by executing the generation program 121. The summary generation process refers to a process for generating a summary of a target content for which a summary is to be generated (hereinafter, target content).
[0034] The memory 12 further stores one or more pieces of content information 122. The content information 122 is information related to the target content, and includes information related to a summary generation reference position. As an example, the summary generation reference position is the interruption position, in which case the content information 122 includes interruption position information 21. In the following description, it is assumed that the interruption position coincides with the start position.
[0035] The content information 122 may include one or more pieces of summary information 22. The summary information 22 is a summary prepared in advance for the target content, and will be described in detail later.
[0036] All or at least a part of the content information 122 may be stored in a device external to the summary generation device 10, such as the server 30 in Fig. 11. In this case, the summary generation device 10 accesses the external device as necessary and reads out and uses the content information 122. Alternatively, the summary generation device 10 may have the server 30.
[0037] 1 shows an example in which a summary generation device 10 according to an embodiment also functions as a content playback device 15. It is not essential that the summary generation device 10 also functions as the playback device 15. The summary generation device 10 may be mounted on the playback device 15, may obtain necessary information directly or indirectly from the playback device 15, or may be an independent device.
[0038] 1, the summary generating device 10 includes an operation unit 17 that receives user operations related to playback, etc. The summary generating device 10 also includes a playback device 15 that plays back content.
[0039] The processor 11 executes a playback process 116 for playing back the specified content on the playback device 15 in accordance with an operation signal input from the operation unit 17. The playback device 15 plays back the specified content in accordance with a control signal from the processor 11.
[0040] The playback process 116 includes a process of storing interruption position information 21, which indicates the position where the playback of the content was interrupted, in the memory 12 as content information 122 of the played back content. As a result, when the playback of the content is interrupted in the playback device 15, the interruption position information 21 for that content is stored in the memory 12.
[0041] The summary generation device 10 may have a display 14. The display 14 is an example of an output device that outputs the generated summary. The output device may provide another form of output, such as a speaker, instead of or in addition to the display 14. When the summary generation device 10 also functions as a content playback device, the display 14 is also an example of an output device for the played content.
[0042] The summary generation device 10 may have a communication device 13 capable of communicating with other devices via a network such as the Internet. As an example, the generated summary may be output to the other device by the communication device 13. In this case, the communication device 13 is also an example of an output device that outputs the generated summary. Furthermore, the content information 122 may be stored in the other device, and the summary generation device 10 may read the content information 122 from the other device by the communication device 13 accessing the other device.
[0043] The target content may be made up of one or more seasons, as shown in Fig. 2. Each season may make up a story, and one or more seasons may make up a larger story as a whole.
[0044] Each season may be divided into a plurality of episodes (stories), for example. An episode is a complete story, and it is assumed that the target content is played for each story. Specifically, season 1 includes a plurality of episodes EP11, EP12, ... EP1n. Season 2 includes a plurality of episodes EP21, EP22, ... EP2n. Each episode includes one or a plurality of sentences Q, which are audio or text.
[0045] A partial summary may be prepared for each episode in the target content. A partial summary refers to a summary whose summary target range is a previous episode. Specifically, episode EP12 includes a partial summary AB12 whose summary target range is episode EP11, and episode EP1n includes a partial summary AB1n whose summary target range is episodes EP11 to EP1(n-1).
[0046] Note that examples in which a partial summary is provided for the target content will be used in the second embodiment and onwards, and in the first embodiment, it is assumed that a partial summary is not provided for the target content, or that a partial summary that is provided is not used.
[0047] If the target content is played back in episode units as expected, the partial summary is played back prior to the playback of the episode, so that the user can check the plot up to the previous episode prior to the playback of the episode.
[0048] In detail, the constituent sentences CS1, CS2, CS3... of the partial summary are constituent units that make up the partial summary, and are extracted from sentence Q included in the summary target range. Specifically, the constituent sentences CS1, CS2, CS3... of the partial summary AB12 are a set (hereinafter also referred to as a subset) of one or more sentences q extracted from episode EP11.
[0049] With reference to FIG. 1, the summary generation process executed by the processor 11 of the summary generation device 10 includes a first determination process 111. The first determination process 111 includes determining an expected playback range (first range) in the target content from the interruption position. The expected playback range is a range after the interruption position, and is a predetermined range from the interruption position. The predetermined range may be a range set in advance. As another example, the predetermined range may be determined from the attributes of the user, the attributes of the target content, the playback tendency of the user, etc. The playback tendency of the user is, for example, the playback behavior of the user, etc. Alternatively, an average value of general playback amounts may be used.
[0050] The summary generation process includes a second determination process 112. The second determination process 112 includes determining a summary target range (second range) for the target content. The summary target range may be at least partially different from the expected playback range. At least partially different may be a completely different range, or the ranges may overlap.
[0051] The range to be summarized is determined based on at least the interruption position. Preferably, the range to be summarized is determined based on both the interruption position and whether the summary to be generated is a synopsis or a preview. The method for determining the range to be summarized will be described with reference to FIG. 3.
[0052] In Fig. 3, the arrow represents the target content, content C, and indicates that the content is played back in the direction of the arrow, that is, from left to right in chronological order. Position P0, which is the starting point of the arrow in Fig. 3, corresponds to the start position of a certain season of content C. In other words, Fig. 3 shows the playback of content C from the beginning of a certain season. The partial summaries ABn and ABn+1 shown in Fig. 3 do not have to be provided in content C. In the first embodiment, it is assumed that these partial summaries are not included in content C.
[0053] Position P1 is the start position of episode n, and position P4 is the end position of episode n. Position P2 is a position corresponding to the interruption position indicated in interruption position information 21 included in content information 122 of content C, and corresponds to the next start position. A range H3 from interruption position P2 to position P3 of content C is set as the predicted playback range determined by the first determination process 111.
[0054] In a second determination process 112, processor 11 determines the range to be summarized based on interruption position P2. In this example, processor 11 determines the range to be summarized based on both the interruption position and whether the summary to be generated is a synopsis or a preview.
[0055] When the summary to be generated is a preview, as one example, processor 11 sets the range H4 from the start position P0 of the season to a position P3 after the interruption position P2 as the summary target range. As another example of a preview, processor 11 may set the range H3 from position P2 to position P3, which coincides with the expected playback range, as the summary target range. In other words, when the summary is a preview, the summary target range may include the range before interruption position P2, or may only include the range after P2. In this way, the preview generated also includes sentences extracted from the expected playback range. This allows the user to imagine the contents of the expected playback range, increasing their desire to play.
[0056] However, if position P3 is too close to position P4, that is, if the summary target range is set to the end of the episode, the possibility that the ending of the episode will be included in the preview increases. In other words, there is a possibility that it will become a so-called spoiler. Therefore, it is preferable that the summary target range when the summary is a preview is set to a range up to a position before position P4, the end of the episode.
[0057] When the summary to be generated is a synopsis, for example, processor 11 sets the range H5 from the season start position P0 to the interruption position P2 as the summary target range. As a result, the generated synopsis also includes sentences extracted from the range up to the interruption position P2. This allows the user to remember the contents up to the interruption position P2, increasing the desire to play it back.
[0058] The summary generation process executed by the processor 11 includes an extraction process 114. The extraction process 114 includes extracting sentences as constituent sentences of the summary from the summary target range. In the extraction process 114, the processor 11 uses the extraction data 34 of the reference content selected by the selection process 111.
[0059] The summary generation process includes a calculation process 113. The calculation process 113 includes calculating an index value of each sentence included in the summary target range from the expected playback range. The index value includes a value obtained based on the words included in the expected playback range. The value obtained based on the words included in the expected playback range is, for example, importance. Specifically, the processor 11 calculates an index value of each sentence included in the summary target range using the words included in the expected playback range.
[0060] The importance here is a value that represents the importance of a phrase contained in each sentence included in the summary target range in the expected playback range. A specific method for calculating the importance is not limited. As an example, the importance may be the frequency of occurrence of a phrase. For example, the importance of a sentence may be calculated to be higher as the frequency of occurrence of a phrase contained in each sentence included in the summary target range in the expected playback range increases.
[0061] Instead of or in addition to the frequency of appearance, the importance may be calculated based on the playback status of the phrase. The playback status of the phrase refers to the excitement (volume, range, etc.) of the voice or background sound during playback in the expected playback range, changes in brightness and color of the video, specific objects in the video, etc., for the phrase contained in each sentence included in the summary target range.
[0062] As another example, the so-called page rank concept may be used to calculate the importance. That is, for a phrase contained in each sentence included in the summary target range, the more frequently a phrase is referenced in the expected playback range, the higher the importance of that sentence is calculated. A pre-stored function may be used for the calculation. As another example, the importance may be calculated using the degree of relevance to other phrases in the expected playback range, or a combination of these.
[0063] Preferably, the index value includes a similarity to the subset. The subset refers to, for example, a set of sentences extracted sequentially when a plurality of sentences to be constituent sentences are extracted sequentially from a summary target range by an extraction process 114 described later. The similarity is, for example, a similarity to the subset. In this case, the index value is, for example, an MMR (Maximal Marginal Relevance) score calculated using the importance and the similarity.
[0064] The processor 11 calculates the MMR score MMR(q) of each sentence q included in the summary target range using the importance I(q) of each sentence q and the similarity Sim(q,k) to the subset k, using the following formula (1). Note that the coefficient λ is a value between 0 and 1. As an example, the coefficient λ is set to 0.5. MMR(q)=λI(q)-(1-λ)Sim(q,k) …(1)
[0065] As shown in formula (1), the closer the coefficient λ is to 1, the more importance the MMR score will be, and the closer the coefficient λ is to 0, the more importance the MMR score will be. When importance is emphasized, the sentence that contains a phrase with high importance in the expected playback range will have a higher MMR score. On the other hand, when similarity is emphasized, the sentence that has low similarity to the subset will have a higher MMR score.
[0066] When extracting sentences with high MMR scores in extraction process 114 described later, sentences that include words of high importance in the expected playback range and are not similar to the subset are likely to be extracted. As a result, the constituent sentences can be related to the expected playback range and become a diverse set of sentences.
[0067] The summary generation process includes an extraction process 114. The extraction process 114 includes extracting sentences from the summary target range as constituent sentences of the summary based on the index value of each sentence included in the summary target range. It is assumed that the number of sentences to be used as constituent sentences is predefined. In this case, as one example, the processor 11 extracts as constituent sentences up to the predefined number of sentences from the sentences included in the summary target range in descending order of index value.
[0068] A specific example of the extraction process 114 will be described with reference to Fig. 4. In the example of Fig. 4, the MMR score is used as the index value. Sentences q1 to q4 are assumed to be included in the summary target range. In the calculation process 113, the importance of each of sentences q1 to q4 is calculated based on the words and phrases included in the expected playback range.
[0069] 4, at the start of extraction process 114 when no sentences have been extracted as constituent sentences, the number of subsets is 0, and the similarity of each of sentences q1 to q4 to the subsets is calculated as 0. Therefore, at this time, the MMR scores of each of sentences q1 to q4 match their importance. If the magnitude relationship of each MMR score is q2>q3>q1>q4, in extraction process 114, sentence q2 with the highest MMR score is extracted as a constituent sentence (step S1).
[0070] When sentence q2 is extracted, the similarity with the subset is calculated for each of sentences q1, q3, and q4 included in the extracted summary target range. The subset in this case is sentence q2. The calculated similarity is used to calculate the MMR score for each of sentences q1, q3, and q4 (step S2). If the magnitude relationship of each MMR score is q4>q3>q1, in the extraction process 114, sentence q4 with the highest MMR score is extracted as a constituent sentence (step S3).
[0071] In extraction step 114, processor 11 repeats the above process until a prescribed number of sentences are extracted. This allows sentences related to the expected playback range to be extracted as constituent sentences, while also taking into consideration the relationships between multiple sentences in the entire summary. When the MMR score is used as the index value, sentences with low similarity to the subset of previously extracted sentences are likely to be extracted, so there is a high possibility that a summary with well-balanced content will be generated.
[0072] The MMR score is an example of an index value using the importance and the similarity. As another example, an index value Iv obtained by adding the similarity to the importance as shown in the following formula (2) may be used. Iv(q)=λI(q)+(1-λ)Sim(q,k) …(2)
[0073] When extracting sentences with a high index value Iv in the extraction process 114 described later, sentences that include words of high importance in the expected playback range and are similar to the subset are likely to be extracted. As a result, the constituent sentences can be a set of sentences that are related to the expected playback range and have uniform content.
[0074] The summary generation process includes a generation process 115. The generation process 115 includes arranging the sentences extracted as constituent sentences in the extraction process 114. Here, the method of arrangement is not limited to a specific method. As an example, it may be a method of arranging according to the order of appearance in the summary target range. As another example, it may be a method of arranging according to the magnitude of the calculated index value.
[0075] At this time, the processor 11 may group a plurality of sentences using similarities or contrasts between the plurality of sentences and arrange them in groups. This allows a natural summary to be generated using a plurality of grouped sentences such as a dialogue.
[0076] The summary generation method according to this embodiment will be described with reference to Fig. 5. The process shown in the flowchart of Fig. 5 is a summary generation process according to the summary generation method according to this embodiment, and is realized by processor 11 executing generation program 121. The process of Fig. 5 is started when playback device 15 plays back the target content, or when a user operation instructing presentation of a summary is received.
[0077] 5, processor 11 reads the interruption position from content information 122 of the target content, and determines an expected playback range based on the interruption position (step S101). In step S101, as an example, processor 11 sets a preset range from the interruption position as the expected playback range.
[0078] Processor 11 also determines the range to be summarized based on both the interruption position and whether the summary to be generated is a synopsis or a preview (step S103). Either step S101 or step S103 may be processed first.
[0079] Processor 11 calculates the importance of each sentence included in the summarization target range determined in step S103 based on the words included in the expected playback range determined in step S101 (step S105). Processor 11 also calculates a similarity to the subset of sentences that have already been extracted as summary constituent sentences, for each sentence included in the summarization target range that has not been extracted as a summary constituent sentence (step S107).
[0080] Processor 11 substitutes the importance calculated in step S105 and the similarity calculated in step S107 into formula (1) to calculate an MMR score as an example of an index value for each sentence included in the summary target range that has not been extracted as a summary constituent sentence (step S109). Processor 11 then extracts the sentence with the highest MMR score as a constituent sentence (step S111).
[0081] If the number of sentences extracted as constituent sentences has not reached the specified number (NO in step S113), processor 11 repeats steps S107 to S111 described above. In this way, sentences to be constituent sentences are extracted one after another. Each time a sentence is extracted as a constituent sentence, the number of sentences that become a subset increases, and the similarity of unextracted sentences is recalculated accordingly. Therefore, the MMR score changes each time a sentence is extracted as a constituent sentence.
[0082] When the number of sentences extracted as constituent sentences reaches a prescribed number (YES in step S113), processor 11 generates a summary by arranging the extracted sentences (step S115).
[0083] [Second embodiment]
[0084] In the summary generation process, a summary generation device 10 according to the second embodiment generates a summary by using a partial summary prepared for the target content. In the second embodiment, a processor 11 applies the Silver Standard Summary Algorithm (SSSA) to the partial summary. For the application of SSSA, see Yamanishi Yoshinori, Nishihara Yoko, and Kaneda Daichi, "Generating Previews of Novels by Applying Partial Summary and the Silver Standard Summary Algorithm," [online], June 9, 2020, Japanese Society for Artificial Intelligence, [searched June 9, 2020], Internet<URL:https: / / doi.org / 10.11517 / pjsai.JSAI2020.0_3K5OS5b01> is disclosed in.
[0085] In the process using SSSA, the processor 11 uses a partial summary prepared for the target content as a gold standard. In this case, as shown in Fig. 6, in the second embodiment, the summary generation process further includes a third determination process 117. The third determination process 117 includes determining, based on the interruption position, a partial summary to be used for generating the summary from among a plurality of partial summaries prepared for the target content, as the gold standard.
[0086] In a third determination process 117, processor 11 determines, as the gold standard, a partial summary associated with a position within a gold standard determination range according to the interruption position and whether the summary to be generated is a synopsis or a preview. In this way, a partial summary corresponding to an episode whose start position is close to the interruption position is determined as the gold standard. A specific example is a partial summary corresponding to an episode next to the episode to which the interruption position belongs. When the summary to be generated is a preview, as an example, the gold standard determination range is the range from the interruption position onwards.
[0087] The third determination process 117 includes determining a silver standard summary from the gold standard depending on the interruption location and whether the summary to be generated is a synopsis or a preview. The silver standard summary refers to a sentence from the constituent sentences of the gold standard that is used to generate the summary. As an example, the processor 11 determines a sentence from the constituent sentences of the gold standard that appears in the target content within a silver standard summary determination range from the interruption location to the silver standard summary. When the summary to be generated is a preview, as an example, the silver standard summary determination range is the range from the interruption location onward.
[0088] The control method according to the second embodiment will be described with reference to FIG. 7. A specific example of the control method according to the second embodiment will be described with reference to FIG. 3 and FIGS. 8 to 10. FIGS. 8 to 10 show an example of generating a summary of "Night on the Galactic Railroad" (quoted from Miyazawa Kenji's "Night on the Galactic Railroad" Aozora Bunko) as target content C. The sentence numbers in the leftmost column are numbers assigned sequentially to all sentences from the beginning. Here, a case will be described in which a preview is generated in which the summary target range includes the range after the interruption point.
[0089] The flowchart in Fig. 7 differs from the flowchart in Fig. 5 showing a specific example of the summary generation method of the first embodiment in the processing of steps S201 and S203. That is, in the summary generation method of the second embodiment, processor 11 determines, from among partial summaries prepared for the target content, a partial summary to be used in generating the summary as the gold standard (step S201), and determines, from its constituent sentences, a sentence to be used in generating the summary as the silver standard summary (step S203).
[0090] 3, a partial summary ABn corresponding to episode n is placed at position P1 in content C. A partial summary ABn+1 corresponding to the next episode n+1 is placed at position P4. These partial summaries ABn, ABn+1 are shown in summary information 22 included in content information 122 of content C.
[0091] The summary range of partial summary ABn is range H1. That is, in this example, partial summary ABn is the summary of episode n, whose summary range is from position P0 to position P1. The summary range of partial summary ABn+1 is range H2. That is, in this example, partial summary ABn+1 is the summary of episode n+1, whose summary range is from position P0 to position P4.
[0092] When the summary to be generated is a preview, for example, the gold standard determination range is the range H6 from the interruption position onward. In this case, if the partial summary associated with a position in the range H6 from the interruption position P2 is determined as the gold standard in the third determination process 117, in the example of Figure 3, processor 11 determines the partial summary associated with position P4 where episode n ends as the gold standard.
[0093] When the summary to be generated is a preview, the range for determining the silver standard summary is, for example, the range H6 from the interruption point onward. In this case, processor 11 determines, as the silver standard summary, the sentences that appear in the target content within the range H6 from the interruption point onward among the constituent sentences of the gold standard. Note that, as an example here, the range for determining the partial summary to be the gold standard and the range for determining the silver standard summary among the constituent sentences of the gold standard are the same range H6, but these ranges may be different.
[0094] 8, it is assumed that "Night on the Galactic Railroad" (quoted from Miyazawa Kenji's "Night on the Galactic Railroad" Aozora Bunko) has been read up to sentence number 130. In this case, sentence number 130 is stored in content information 122 as interruption position P2.
[0095] When the end of episode n to which sentence number 130 belongs is sentence number 191, sentence number 192 becomes position P4, which is the start position of the next episode n+1. In this case, for example, as shown in Figure 8, a partial summary ABn+1 corresponding to episode n+1 is placed immediately before sentence number 192.
[0096] The range H2, which is the summary target range of partial summary ABn+1, is from sentence number 1 to sentence number 191. The position P4 corresponding to partial summary ABn+1 is included in the range H6 following the interruption position P2. Therefore, in the third determination process 117, the processor 11 determines that the partial summary ABn+1 is the gold standard.
[0097] Fig. 9 shows a specific example of constituent sentences of partial summary ABn+1. Referring to Fig. 9, partial summary ABn+1 has, as an example, eight sentences with sentence numbers 5, 21, 31, 55, 120, 163, 170, and 191 extracted from range H2 (sentence numbers 1 to 191) as constituent sentences.
[0098] In this case, group K1, which is composed of sentence numbers 5, 21, 31, 55, and 120, was extracted from before interruption position P2, and group K2, which is composed of sentence numbers 163, 170, and 191, was extracted from after interruption position P2. In other words, group K2 is a group of sentences included in range H6 after interruption position P2, and group K1 is a group of sentences that is not included. Therefore, in third determination step 117, processor 11 determines group K2 to be the silver standard summary.
[0099] In the summary generation method according to the second embodiment, processor 11 subsequently generates a summary in the same manner as in the summary generation method according to the first embodiment. That is, referring to Fig. 7, processor 11 sets the silver standard summary (group K2: sentence numbers 163, 170, 191) determined in step S203 as a subset, and calculates the similarity to the subset for each sentence included in the summary target range that has not been extracted as a summary constituent sentence (step S107).
[0100] As another example, processor 11 may treat one or more sentences that were not selected as the silver standard summary (e.g., group K1: sentence numbers 5, 21, 31, 55, 120) or all of the constituent sentences of the gold standard partial summary (e.g., group K1+K2: sentence numbers 5, 21, 31, 55, 120, 163, 170, 191) as a subset, and calculate the similarity of each sentence included in the summary target range that has not been extracted as a summary constituent sentence to this subset.
[0101] Processor 11 calculates an MMR score for each sentence included in the summary target range that has not been extracted as a summary constituent sentence using the importance and similarity (step S109), and then processor 11 extracts the sentence with the highest MMR score as a constituent sentence (step S111).
[0102] If the number of sentences extracted as constituent sentences has not reached the specified number (NO in step S113), processor 11 repeats the above steps S107 to S111.
[0103] As an example, assume that sentence numbers 131, 132, 133, 134, and 138, which are after sentence number 130, which is the end position of reading, are extracted from the range to be summarized. When generating a preview, as shown in Fig. 10, the processor 11 adds the newly extracted sentence numbers 131, 132, 133, 134, and 138 (group K3) to sentence numbers 163, 170, and 191 (group K2), which are the silver standard summary, to form the constituent sentences of the summary.
[0104] 10, the constituent sentences of the preview summary may include both sentences before and after the interruption position P2. In other words, the constituent sentences of the preview summary are not limited to only sentences after the interruption position P2. For example, if the constituent sentences include at least one sentence after the interruption position P2, the summary may be treated as a preview.
[0105] By using the gold standard in this way, fewer sentences are extracted than if all of the constituent sentences were extracted, making the process easier. In other words, the number of times steps S107 to S113 are repeated can be reduced. Also, by determining a silver standard summary from the gold standard and using it as a subset, a summary with more appropriate content can be generated.
[0106] [Third embodiment]
[0107] The summary generating device 10 according to the third embodiment uses reference content in the summary generating process. The reference content refers to content other than the target content, which is referred to when extracting constituent sentences of the summary of the target content, and for which a summary is prepared.
[0108] As an example, the summary generation device 10 acquires information about the content used as the reference content from another device. Fig. 11 is a diagram showing the summary generation device 10 according to the third embodiment. As shown in Fig. 11, the summary generation device 10 according to the third embodiment is capable of communicating with a server 30 as another device via a communication device 13.
[0109] The server 30 stores content data 31A, 31B,...31 associated with a plurality of contents, respectively. The content data 31A, 31B,...31 are used in the summary generation process executed by the processor 11 of the summary generation device 10. The content data 31 being associated with the contents means that the content data 31 does not have to include the content itself, but includes information such as a name or an identifier indicating the content.
[0110] 11, the server 30 is an external device to the summary generation device 10 and is accessed by the communication device 13 via the network 70. However, as another example, the server 30 may be a storage device mounted on the summary generation device 10.
[0111] Each piece of content data 31 includes attributes 32. The attributes 32 are information that indicates the characteristics of the story of the content, and are, for example, values obtained by converting the sentences of the entire story into word vectors. The attributes 32 may also be meta information such as genre, screenwriter, season number, characteristics of the person to be played, or may be converted into values obtained by converting the sentences of the entire story into word vectors.
[0112] Content data 31 that can be used as reference content includes partial summaries 33. Fig. 11 shows a case in which the content data 31 includes multiple partial summaries 33A, 33B, 33C, .... The partial summaries 33A, 33B, 33C, ... are prepared for each episode of the content, and refer to summaries whose summary range is a previous episode.
[0113] Each of the content data 31 includes extraction data 34. The extraction data 34 is data representing the tendency of the position of each of the constituent sentences of one or more partial summaries 33A, 33B, 33C, etc. in the content. The extraction data 34 will be specifically described with reference to FIG. 12.
[0114] Fig. 12 is a diagram showing the distribution of sentences that are used as constituent sentences of partial summaries in actual contents A and B. The horizontal axis of Fig. 12 indicates the episode number EP, and the vertical axis indicates the extraction position PP of the constituent sentence from the summary target range. In Fig. 12, the extraction position PP of each sentence that is used as a constituent sentence of the corresponding partial summary is plotted for each episode.
[0115] Contents A and B are globally popular serial dramas that have been broadcast for three or more seasons. Figure 12 shows the extraction positions PP of sentences that constitute the partial summaries of all episodes of the three seasons, using up to three of all seasons of contents A and B, from the range of the target of summarization. For the partial summaries, sentences within the first three minutes of each episode are used.
[0116] The extraction position PP(q) of sentence q extracted as a constituent sentence is obtained by the following procedure. That is, the value EPe(q) is obtained by normalizing the appearance position of sentence q in all sentences (number Ne) included in episode e from which sentence q was extracted, using the following formula (3). EPe(q)=Pe(q) / Ne …(3)
[0117] The value EPe(q) is then used to obtain the absolute position AP(q) defined by equation (4) below. AP(q) = EPe(q) + (e-1) … (4)
[0118] Then, we obtain the extraction position PP(q), which is the past relative position defined by the following formula (5). The past relative position represents the relative position of sentence q when the episode number e' of the target of partial summary, which has sentence q as a component, is set to 1. PP(q)=AP(q) / (e'-1) …(5)
[0119] As an example, if the absolute position AP of the constituent sentence of the partial summary corresponding to episode number 3 is 1.4591, then the extraction position PP = 0.7295 is obtained from formula (5). By expressing the extraction position PP as a relative past position, it becomes possible to take into account the relative positional relationship in the range from the start of the season to the position of the partial summary. This makes it possible to compare and consider the extraction position of a sentence across content.
[0120] The inventors investigated the distribution of sentences that were used as constituent sentences of partial summaries, similar to contents A and B, for multiple contents, including contents A and B, each of which was a globally popular TV drama series that had been broadcast for three or more seasons.
[0121] As a result, the inventors noticed that, common to multiple contents, for season 1, the extraction position PP tends to be concentrated at 0.0 and 1.0, as shown in Fig. 12. The smaller the extraction position PP value (closer to 0), the closer the extraction position is to the beginning of the summarization range, that is, the beginning of the season. The larger the value (closer to 1), the closer the extraction position is to the end of the summarization range, that is, the closer it is to the end of the episode immediately preceding the corresponding episode. Therefore, it was considered that, common to multiple contents, for season 1, sentences used as constituent sentences of partial summaries tend to be biased toward the beginning and end of the summarization range (Consideration 1).
[0122] The inventors also noticed that the plots of the extraction position PP exist continuously in a downward sloping manner to the right near the middle between 0.0 and 1.0, which is common to multiple contents. This is considered to be because each partial summary tends to use sentences at the same position in the corresponding season as constituent sentences (Consideration 2). In other words, it is considered that each partial summary tends to use sentences with a specific high level of importance, which is common to multiple contents.
[0123] The inventors also noticed that, as shown in Fig. 12, the distribution tendency of the extraction positions PP for each of seasons 1 to 3 may differ between contents A and B. This was because it was considered that the tendency to extract sentences that constitute partial summaries may differ for each content and / or each season (Consideration 3).
[0124] Based on Observations 1 and 2, the inventors divided the range of the summary into three parts, the beginning, middle, and end, and for each partial summary, the extraction ratio of each part of the entire sentences that constitute the partial summary is used as extraction data 34. As an example, the beginning s, middle c, and end e are the ranges of 20%, 60%, and 20% from the start of the season, respectively.
[0125] Focusing on episode 50 of content A, indicated by a dotted line in Figure 12, it can be seen that 50% of the sentences that have been made into constituent sentences of the partial summary corresponding to episode 50 are present in the beginning s, 30% in the middle c, and 20% in the end e of the summary target range. In other words, it can be seen that the partial summary corresponding to episode 50 of content A was generated by extracting 50% of the sentences from the beginning s, 30% of the sentences from the middle c, and 20% of the sentences from the end e as constituent sentences from the range from the start of season 3 to the end of episode 49.
[0126] The extraction data 34 represents the extraction ratio of sentences that constitute the partial summaries corresponding to each episode from each category. In other words, the extraction data 34 can be said to represent the tendency of the relative positions of sentences that constitute the partial summaries of the content in the content.
[0127] As an example, the extraction data 34 can be shown in a table format as shown in Fig. 13. That is, referring to Fig. 13, the extraction data 34 for content A indicates the extraction ratios of the beginning s, middle c, and end e of the summary target range of sentences that are constituent sentences of partial summaries for all episodes 1 to 69 up to season 3 of content A. For example, for episode 50, 50%, 30%, and 20% are specified for the beginning s, middle c, and end e, respectively.
[0128] Furthermore, based on Consideration 3, the inventors decided to use the content suitable for the target content as the reference content in the summary generation process. Therefore, as shown in Fig. 11, a plurality of content data 31A, 31B, ... 31 are stored in the server 30, and each of them includes extraction data 34.
[0129] 11, in the third embodiment, the summary generation process further includes a selection process 118. The selection process 118 includes selecting a reference content. The reference content is selected from content data 31A, 31B, ... 31 stored in the server 30, for example.
[0130] In the selection process 118, content data 31 having story characteristics related to the target content is selected based on attributes 32 included in the content data 31. In the selection process 118, as one example, content whose value obtained by converting the sentences of the entire story into word vectors is within a predetermined range from the value of the target content is extracted as reference content. As another example, content that matches or is similar to the target content in at least one of genre, screenwriter, season number, and playback target characteristics may be extracted as reference content.
[0131] In the third embodiment, in extraction process 114, processor 11 uses extraction data 34 associated with the reference content to extract sentences to be constituent sentences of the summary from the target content. In detail, processor 11 refers to the extraction ratios of the beginning s, middle c, and end e of the summary target range for the relevant episode, which are shown in extraction data 34 associated with the reference content. The relevant episode is an episode whose start position is near position P2, and in the example of Figure 6, it is, for example, episode n+1. In this example, it may also be episode n.
[0132] Assume that the extraction ratios of the beginning s, middle c, and end e of the summary target range for the summary ABn+1 of the corresponding episode n+1 are 50%, 30%, and 20%, respectively, as shown in Fig. 13. In this case, the processor 11 applies these ratios to the target content and extracts sentences based on the MMR score. In this example, if the number of sentences to be constituent sentences is 10, 5, 3, and 2 sentences are extracted from the beginning s, middle c, and end e of the summary target range, respectively, based on the MMR score.
[0133] This allows a summary to be generated for the target content by extracting sentences from the summary target range in a similar manner to partial summaries prepared for reference content that has related story characteristics, making it easy to generate summaries.
[0134] The summary generating device 10 according to the third embodiment may further use placement data in the summary generating process. The placement data is an evaluation value of the mismatch between the order of positions of the constituent sentences in the content for which a summary is prepared and the order of placement in the partial summary. The placement data 35 may be, for example, the Jaro-Winkler distance.
[0135] 11, each piece of content data 31 includes placement data 35. The placement data 35 for content A indicates, for example, up to season 3 of content A, for all episodes 1 to 69, the extraction ratios of sentences that constitute partial summaries for the beginning s, middle c, and end e of the summary target range. For example, for episode 50, 50%, 30%, and 20% are specified for the beginning s, middle c, and end e, respectively.
[0136] The inventors noticed that the order of sentences in the partial summary differs from the order of their positions in the content, that is, there are cases where there is a mismatch. Therefore, for multiple contents including content A, the Jaro-Winkler distance was calculated as an index value showing the mismatch between the order of sentences in the partial summary and the order of appearance in the content, and verified. The closer the Jaro-Winkler distance is to 1, the closer the order of sentences in the partial summary is to the order of sentences in the content, that is, the smaller the mismatch is.
[0137] The inventors investigated a large number of works and found that the average value of the Jaro-Winkler distance for each season was 0.65 to 0.85 for most works. This indicates that sentences are often arranged in a different order from the content in partial summaries.
[0138] The greater the discrepancy between the arrangement order in the partial summary and the order in which the content appears in the content, the more difficult it is to fully understand the content through the summary. Conversely, the smaller the discrepancy, the easier it is to understand the content. For this reason, it is thought that different methods are used depending on the attributes of the content.
[0139] That is, in the case of a content category such as suspense, where it is preferred that the content not be fully understood by the summary, it is considered that a balance between understanding of the content and the desire to play it back can be achieved by setting an appropriate mismatch depending on the range of the summary target, etc. Therefore, in the summary generating device 10 according to the third embodiment, the placement data 35 prepared for each content may be used to generate a summary of the target content.
[0140] The placement data 35 can be shown in a table format as shown in Fig. 14, for example. The placement data 35 in Fig. 14 shows, as an example, placement data (Jaro-Winkler distance) for seasons 1, 2, 3 and all seasons of contents A, B and C. The Jaro-Winkler distance is 1 when the order of the positions of the sentences in the content and the placement order in the partial summary are completely the same, and is 0 when there is no similarity at all. Note that, although the example of the placement data 35 in Fig. 14 shows the value of the placement data for each season, one value may be shown for each content.
[0141] In the summary generating device 10 according to the third embodiment, the placement data of the content related to the target content is used for placing the sentences extracted as constituent sentences. As an example, the placement data of the reference content is used. Alternatively, in the summary generating method according to the third embodiment, it may be possible to select whether the placement order of the extracted sentences in the summary is made to match the appearance order in the target content, or to use the placement data of the related content. The selection may be made, for example, by the summary generator.
[0142] Specifically, in the generation process 115, the processor 11 arranges the multiple sentences extracted as constituent sentences based on the arrangement data 35. As an example, the processor 11 rearranges the multiple extracted sentences so that the Jaro-Winkler distance is the same as the Jaro-Winkler distance shown in the arrangement data 35 of the reference content or is a value within a predetermined range. This makes it possible to make the content of the target content as easy to understand from the summary as that of the reference content.
[0143] <3. Notes> The present invention is not limited to the above-described embodiment, and various modifications are possible. For example, at least two of the first to third embodiments may be combined. [Explanation of symbols]
[0144] 10: Summary generator 11: Processor 12: Memory 13: Communication equipment 14: Display 15: Playback device 17:Operation section 21: Interruption location information 22: Summary information 30: Server 31: Content data 31A: Content data 31B: Content data 32: Attribute 33: Partial summary 33B: Partial summary 33C: Partial summary 34: Data for extraction 35: Placement data 70: Network 111: Selection process 112: Second decision process 113: Calculation process 114: Extraction process 115: Generation process 116: Regeneration processing 117: Third decision process 118: Selection process 121: Generator 122: Content information AB12: Partial summary AB1n: Partial summary ABn :Partial summary C: Content CS1: Constituent sentence CS2: Constituent sentence CS3: Composition EP10: Episode EP11: Episode EP12: Episode EP1n : Episode EP21: Episode EP22: Episode H1 : Range H2 : Range H3 : Scope H4 : Scope H5: Scope H6 : Scope K1: Group K2: Group K3: Group P0 :Start position P2: Interruption position PP: Extraction position q : statement q1: sentence q2: sentence q3: Statement q4 : Statement
Claims
1. 1. A summary generation device for generating a summary of target content, comprising: A processor is provided. The processor, determining a first range in the target content from a playback interruption position of the target content; calculating, from the first range, index values of sentences included in a second range, the second range being a range for extracting sentences constituting the summary in the target content, the second range being at least partially different from the first range; and extracting constituent sentences of the summary from the second range based on the index value. Summary generator.
2. The index value includes a value obtained based on a term included in the first range. The apparatus of claim 1 .
3. The first range is a range after the playback interruption position and is an expected playback range of the target content determined from the playback interruption position.
3. Apparatus for generating a summary according to claim 1 or 2.
4. the second range is determined depending on whether the abstract is a first abstract or a second abstract; the first summary is generated with a range that does not include the portion after the playback interruption position as the second range, The second summary is generated with the range including the portion after the playback interruption position as the second range. A summary generating device according to any one of claims 1 to 3.
5. extracting the constituent sentences of the summary includes sequentially extracting a plurality of sentences included in the constituent sentences from the second range based on the index values to generate a subset having the sequentially extracted sentences; The index value includes a similarity between the sentences included in the second range and the subset. A summary generating apparatus according to any one of claims 1 to 4.
6. The subset includes sentences extracted from a prepared summary of the target content.
6. Apparatus for generating a summary according to claim 5.
7. The processor further comprises: configured to select reference content based on the target content; Extracting sentences constituting the summary from the second range includes: and extracting the constituent sentences of the summary from the second range by referring to extraction data associated with the reference content. Summary generating apparatus according to any one of claims 1 to 6.
8. 1. A method for generating a summary of a target content, comprising: determining a first range in the target content from a playback interruption position of the target content; calculating, from the first range, index values of sentences included in a second range, the second range being a range for extracting constituent sentences of the summary in the target content and at least a part of which is different from the first range; extracting sentences constituting the summary from the second range based on the index value. Summary generation method.
Citation Information
Patent Citations
Video support system
JP1996292965A
System for browsing book summary information
JP2006155125A
Reproduction controller, reproduction control method, and program
JP2007208631A
Preview reproduction device
JP2007274594A
Trailer generator, trailer generating method, trailer generating server, trailer generating program, and recording medium
JP2007336085A