Dialogue video summarization device, dialogue video summarization method, and dialogue video summarization program

The system processes interactive video data to select important utterances and their reactions, generating a summary video that effectively conveys dialogue dynamics, addressing the limitations of conventional summarization techniques.

JP7692562B2Active Publication Date: 2025-06-16NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021183726
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-06-16
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Conventional techniques for summarizing interactive video data fail to effectively convey the listener's reaction to the speaker's utterance, making it difficult for users to grasp the dialogue dynamics.

Method used

A system comprising a dialogue video input unit, a dialogue act estimation unit, an importance estimation unit, and a summary video generation unit, which processes dialogue video data to select important utterances and their corresponding reactions, generating a summary video that highlights these key interactions.

Benefits of technology

Enables the creation of summaries that clearly depict the listener's response to the speaker's utterance, facilitating easier understanding of dialogue dynamics in interactive video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692562000001
    Figure 0007692562000001
  • Figure 0007692562000002
    Figure 0007692562000002
  • Figure 0007692562000003
    Figure 0007692562000003
Patent Text Reader

Abstract

To create a summary that a reaction from a listener against a speech of a speaker in a conversation video image data can be easily gripped.SOLUTION: A summery device 10 divides a conversation video image data in each speech to estimate whether or not each speech is a speech of a reaction in a conversation. Also, the summery device 10 estimates an importance in each speech. Then, the summery device 10 selects a speech with a high importance and a speech of a reaction of a speech with an importance that is a predetermined value or larger on the basis of the height of the importance in each speech of the conversation video image data and a fact that each speech is a speech of the reaction corresponded to the speech of the importance that is the predetermined value or larger, and generates a summary video image of the conversation video image data by using the selected speech.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive video summarization device, an interactive video summarization method, and an interactive video summarization program.

Background Art

[0002] Conventionally, there is a technique for structuring or summarizing discussions in interactive video data of multiple speakers. This technique is useful for understanding the content of the interactive video data. Here, when a user views a summary of the interactive video data, it is also important to understand not only the above-mentioned content but also the reaction of the listener to the speaker's utterance (for example, the degree of approval / judgment).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the conventional summary technique for interactive video data has a problem in that, although it is useful for understanding the content, it does not clearly convey the reaction of the listener to the speaker's utterance. Therefore, an object of the present invention is to solve the above-mentioned problem and create a summary that makes it easy to grasp the reaction of the listener to the speaker's utterance in the interactive video data.

Means for Solving the Problems

[0005] In order to solve the above-described problems, there is provided a dialogue video input unit that receives input of dialogue video data, a dialogue act estimation unit that divides the dialogue video data for each utterance and estimates whether each of the utterances is a response utterance in the dialogue, an importance estimation unit that estimates the importance for each utterance, and based on the height of the importance for each utterance and whether each of the utterances is a response utterance for an utterance having an importance equal to or higher than a predetermined value, from the utterances included in the dialogue video data, an important utterance and a response utterance for an utterance having an importance equal to or higher than the predetermined value are selected, and a summary video generation unit that generates a summary video of the dialogue video data using the selected utterances, and a summary video output unit that outputs the generated summary video.

Effect of the Invention

[0006] According to the present invention, it is possible to create a summary that makes it easy to grasp the listener's response to the speaker's utterance in the dialogue video data.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. The present invention is not limited to these embodiments.

[0009] [Overview] The overview of the dialogue video summary device (summary device) of the present embodiment will be described. The summary device estimates whether each utterance constituting the input dialogue video data is an utterance for advancing the dialogue or a reaction to the utterance. In addition, the summary device also estimates the importance of each utterance. Then, based on the results of the above estimations, the summary device generates a summary video that summarizes important utterances in the dialogue video data and the listener's reactions to the important utterances.

[0010] Thereby, the summary device can generate a summary video that makes it easy to grasp the listener's reaction to the speaker's utterance in the dialogue video data.

[0011] [Configuration Example] First, a configuration example of the summary device 10 will be described with reference to FIG. 1. The summary device 10 includes a summary perspective input unit 11, a dialogue video input unit 12, an information extraction unit 13, a dialogue division unit 14, a dialogue act estimation unit 15, an importance estimation unit 16, a summary evaluation value calculation unit 17, a summary video generation unit 18, and a summary video output unit 19.

[0012] [Summary Perspective Input Section] The summary perspective input section 11 receives an input of the perspective to be emphasized in the summary of the dialogue video data from the user of the summarization device 10. The perspectives here are, for example, the playback time of the summary video, content understanding, atmosphere understanding, etc. Also, the summary perspective input section 11 receives an input of a value indicating the degree of emphasis for each of the above perspectives (for example, in the range of values from 0.0 to 1.0. The higher the value, the higher the degree of emphasis).

[0013] For example, the summary perspective input section 11 receives an input of information indicating how much importance is attached to the shortness of the playback time of the summary video, how much importance is attached to the ease of understanding the content of the summary video, and how much importance is attached to the atmosphere of the dialogue in the summary video.

[0014] Here, it is assumed that the user cannot set the maximum value for all perspectives, and when the value of one perspective is set high, the values of other perspectives may be set relatively low. For example, the summary perspective input section 11 may receive an input of a value indicating the degree of emphasis for each perspective from the user through the user interface shown in FIG. 2.

[0015] [Dialogue Video Input Section] The dialogue video input section 12 receives an input of dialogue video data including acoustic information and video information showing the speaker.

[0016] [Information Extraction Section] The information extraction section 13 extracts, from the input dialogue video data, for each utterance, acoustic information, video information, language information of the utterance content separated by the speaker, non-verbal information such as the gestures and head movements of the speaker and the listener (for example, three-dimensional coordinates of body joints in time units), etc.

[0017] For example, when extracting the language information for each utterance, the information extraction section 13 recognizes the utterance section of the dialogue video data as shown in FIG. 3 (S1), separates the speakers for each utterance section (S2), and performs voice recognition for each speaker (S3).

[0018] [Dialogue segmentation part] Return to the description of FIG. 2. The dialogue segmentation unit 14 obtains the time of the entire dialogue video data (the time of the entire dialogue) from the dialogue video input unit 12, and divides the time into the beginning stage, the middle stage, and the end stage.

[0019] For example, the dialogue segmentation unit 14 may divide the time of the entire dialogue into three equal parts as described above, and regard each part as the beginning stage, the middle stage, and the end stage. Or, after dividing into three equal parts, the positions of division into the beginning stage, the middle stage, and the end stage may be adjusted based on the appearance time positions of keywords such as "Let's start" in the beginning stage and "It's about time~Shall we do it?" and "To summarize" in the end stage.

[0020] The dialogue segmentation unit 14 attaches the information of the above-mentioned beginning stage, middle stage, and end stage to the information (dialogue video information) extracted by the information extraction unit 13, and outputs it to the dialogue act estimation unit 15 and the importance estimation unit 16.

[0021] [Dialogue act estimation unit] The dialogue act estimation unit 15 estimates the intention (dialogue act) of the utterance unit of the dialogue video data. The dialogue act is, for example, whether the utterance is an utterance that advances the dialogue or an utterance as a reaction in the dialogue, etc. For example, classifications such as SWBD-DAMSL (see Reference 1) and ISO 24617-2 can be used. For example, the dialogue act estimation unit 15 estimates the dialogue act using the acoustic information, video information, language information, and non-verbal information output from the dialogue segmentation unit 14.

[0022] Reference 1: Daniel Jurafsky, Elizabeth Shriberg, and Debra Biasca. Switchboard SWBD-DAMSL Shallow-Discourse Function Annotation Coders Manual, Draft 13. Technical report, Institute of Cognitive Science Technical Report, 1997.

[0023] In addition, the dialogue act estimation unit 15 can use the learning and estimation method described in the following Document 2 for the transcription data in the estimation of dialogue acts.

[0024] Document 2: Japanese Patent Application Laid-Open No. 2020-173608

[0025] Here, ISO-24617-2 will be taken as an example for explanation. In ISO-24617-2, for each utterance, as a dialogue act, a dimension (Dimension) for classifying functions and the functions of its subcategories (DSF: Dimension Specific Function) are set. In the summary of the dialogue, in particular, the dimension of the Task for advancing the dialogue and the dimension of Auto-Feedback indicating the reaction to the utterance are targeted.

[0026] The dialogue act estimation unit 15 estimates the dimension and function for each utterance using, for example, the learned model described in Document 2, and outputs the estimation result as dialogue act information to the summary evaluation value calculation unit 17.

[0027] [Importance Estimation Unit] The importance estimation unit 16 estimates the importance of each utterance using the acoustic information, video information, language information, and non-verbal information output from the dialogue video input unit 12. For the estimation of the importance here, for example, an existing learned model (see, for example, the following Document 3 and Document 4) can be used. Then, the importance estimation unit 16 outputs the estimated importance for each utterance (for example, a value from 0.0 to 1.0).

[0028] Document 3: Nihei, et al., Estimation of Important Utterances for Discussion Summarization Based on Linguistic and Non-Linguistic Information, Transactions of the Institute of Electronics, Information and Communication Engineers A 2019, J102-A, pp.35-47.

[0029] Document 4: Nihei, et al., Verification of the Effectiveness of a Discussion Summary Browser Equipped with an Important Utterance Estimation Model Based on Multimodal Information, Journal of the Human Interface Society 22(2). DOI: https: / / doi.org / 10.11184 / his.22.2_137

[0030] [Summary evaluation value calculation unit] The summary evaluation value calculation unit 17 calculates an evaluation value (summary evaluation value) for evaluating utterances to be included in the summary video among the utterances of the dialogue video data.

[0031] For example, the summary evaluation value calculation unit 17 calculates a summary evaluation value (for example, a numerical value of 0 or more) based on the dialogue action information for each utterance output from the dialogue action estimation unit 15, the video information, language information, and importance information output from the importance estimation unit 16, and the information on the degree of emphasis on the playback time, content understanding, and atmosphere understanding output from the summary perspective input unit 11.

[0032] An example of the calculation process of the summary evaluation value will be described with reference to FIGS. 4 and 5. First, the summary evaluation value calculation unit 17 calculates the summary evaluation value for each utterance by the following formula (1) using the information output from the dialogue action estimation unit 15 and the importance estimation unit 16 (S11 in FIG. 5).

[0033] Summary evaluation value = Importance * (Weighting of utterance position) … Formula (1)

[0034] However, the summary evaluation value calculation unit 17 shall change the weighting to be multiplied for each utterance position (for example, the beginning stage, the middle stage, the end stage) (for example, beginning stage: 1.0, middle stage: 0.8, end stage: 1.0). Also, for the importance in Formula (1), the value shown in the importance information output from the importance estimation unit 16 is used.

[0035] By doing so, the summary evaluation value calculation unit 17 can calculate the summary evaluation value of the utterance in consideration of the importance of the utterance and the utterance position of the utterance in the dialogue video data (for example, the beginning stage, the middle stage, the end stage of the dialogue). As a result, the summary video generation unit 18 can preferentially incorporate, for example, utterances with high importance and utterances that appear at the beginning stage (for example, the start of the dialogue) and the end stage (for example, the conclusion of the dialogue) of the dialogue into the summary video. As a result, the summary video generation unit 18 can generate a summary video that is easy for the user to understand the content of the dialogue.

[0036] Next, the summary evaluation value calculation unit 17 updates the above summary evaluation value by the following formula (2) using the value of the content importance (degree of emphasizing content understanding) output from the summary perspective input unit 11 (S12 in FIG. 5).

[0037] Summary evaluation value = Summary evaluation value * (Content importance + 1) … Formula (2)

[0038] By doing so, the summary evaluation value calculation unit 17 can assign a high summary evaluation value to the utterances with high importance in the dialogue video data. Also, the summary evaluation value calculation unit 17 can adjust the weighting of the importance values according to the degree to which the user emphasizes the understanding of the dialogue content. As a result, the summary video generation unit 18 can generate a summary video that preferentially incorporates the utterances with high importance to the extent desired by the user.

[0039] Furthermore, the summary evaluation value calculation unit 17 sets the summary evaluation value of the utterance immediately following an important utterance to be high. For example, the summary evaluation value calculation unit 17 updates the summary evaluation value by the following formula (3) for the utterances (for example, the utterances started within 3 seconds from the end time of the utterances with an importance of 0.8 or more) that are uttered within a predetermined time from the end time of the utterances with an importance above a certain threshold among all the utterances in the dialogue video data. Note that the "atmosphere importance" in formula (3) is a value indicating the degree to which the summary perspective input unit 11 emphasizes the understanding of the atmosphere.

[0040] Summary evaluation value = Summary evaluation value * (Atmosphere importance / 2 + 1) … Formula (3)

[0041] For example, as shown in S13 of FIG. 5, the summary evaluation value calculation unit 17 updates the summary evaluation value of the utterances classified as Auto-Feedback by the dialogue act estimation unit 15 among the utterances starting within 3 seconds after the utterances with an importance of 0.8 or more according to the above formula (3).

[0042] By doing so, the summary evaluation value calculation unit 17 can assign a high summary evaluation value to the response utterance to the utterance with high importance. As a result, the summary video generation unit 18 can generate a summary video that more preferentially incorporates the response utterances to the utterances with high importance. As a result, for the user, it is possible to generate a summary video in which the reaction of the listener to the utterance is easy to understand.

[0043] Note that the summary evaluation value calculation unit 17 may update the summary evaluation value by a calculation formula using an attenuation function for all utterances uttered after an utterance with an importance equal to or higher than a predetermined threshold (utterance with high importance). For example, the summary evaluation value calculation unit 17 may update the summary evaluation value according to the following formula (4).

[0044] Summary evaluation value = Summary evaluation value * (atmosphere importance + 1) * F(t) … Formula (4)

[0045] However, F(t) in formula (4) is an attenuation function, and t is the elapsed time from the utterance with high importance.

[0046] For example, for an utterance A at time t when the importance is 0.8 or more, for an utterance B at time t after utterance A A the summary evaluation value calculation unit 17 multiplies it by the attenuation exponential function e-(t B -t B -t A ). Then, every time an utterance with an importance of 0.8 or more appears, the summary evaluation value calculation unit 17 performs the same process on the subsequent utterances.

[0047] By doing so, the summary evaluation value calculation unit 17 can assign a higher summary evaluation value to the utterance immediately after the utterance with high importance. That is, the summary evaluation value calculation unit 17 can assign a higher summary evaluation value to the utterance that is more likely to be a response to the utterance with high importance. As a result, the summary video generation unit 18 can preferentially incorporate into the summary video the utterances that are likely to be responses to the utterances with high importance. As a result, for the user, it is possible to generate a summary video in which the reaction of the listener to the utterance is easy to understand.

[0048] An example of the summary evaluation value calculated by the summary evaluation value calculation unit 17 is shown in FIG. 6. Note that the degree of emphasis on each perspective is: playback time: 0.2, content comprehension: 0.4, and atmosphere comprehension: 0.8.

[0049] For example, as shown in FIG. 6, for the utterance "How about yakitori?" spoken by speaker D from the start time of 110.00 to the end time of 120.00, the importance is 0.88, the dimension is Task, the DSF is Inform / Answer, and when the utterance time is 10.00, the summary evaluation value is 1.23. Also, for the utterance "yakitori" spoken by speaker B from the start time of 121.00 to the end time of 122.50, the importance is 0.50, the dimension is Task, the DSF is auto positive, and when the utterance time is 1.50, the summary evaluation value is 0.98.

[0050] [Summary video generation unit] The summary video generation unit 18 generates a video (summary video) obtained by thinning out a part of the dialogue video data output from the dialogue segmentation unit 14 based on the summary evaluation value calculated by the summary evaluation value calculation unit 17, the information on the playback time of the dialogue video data, and the information on the degree of importance of the playback time received by the summary perspective input unit 11.

[0051] The summary video generation unit 18 selects, from the utterances included in the dialogue video data, utterances with high importance and utterances that are responses to utterances with importance equal to or higher than a predetermined value based on the height of the importance for each utterance and whether each utterance is a response utterance to an utterance with importance equal to or higher than a predetermined value. Then, the summary video generation unit 18 generates a summary video of the dialogue video data using the selected utterances. The details of the summary video generation process performed by the summary video generation unit 18 will be described later.

[0052] [Summary video output unit] The summary video output unit 19 outputs the summary video generated by the summary video generation unit 18. For example, the summary video output unit 19 outputs the summary video generated by the summary video generation unit 18 to an external device.

[0053] According to the summarization device 10 described above, a user can generate a summary video that makes it easy to grasp the reaction of the listener to the speaker's utterance in the dialogue video data. Further, the summarization device 10 can generate a summary video that reflects the viewpoints that the user values.

[0054] [Example of processing procedure] Next, an example of the processing procedure executed by the summarization device 10 will be described with reference to FIG. 7. First, the summarization viewpoint input unit 11 of the summarization device 10 receives an input of the summarization viewpoint of the dialogue video data from the user (S21). For example, the summarization viewpoint input unit 11 receives an input of the importance degree of each summarization viewpoint (for example, the importance degree of playback time, the importance degree of content understanding, the importance degree of atmosphere understanding) from the user.

[0055] Also, the dialogue video input unit 12 receives an input of the dialogue video data to be summarized (S22). Then, the information extraction unit 13 extracts, from the dialogue video data input in S22, acoustic information, video information, language information of the speaker-separated utterance content, non-verbal information such as the gestures and head movements of the speaker and the listener, etc. (S23: Information extraction).

[0056] After S23, the dialogue segmentation unit 14 performs segmentation of the dialogue video data (S24). For example, the dialogue segmentation unit 14 obtains the time of the entire dialogue video data from the dialogue video input unit 12, and divides that time into the beginning stage, the middle stage, and the end stage. Then, the dialogue segmentation unit 14 assigns the information of the above-mentioned beginning stage, middle stage, and end stage to the information (dialogue video information) extracted by the information extraction unit 13 in S23 and outputs it.

[0057] After S24, the dialogue act estimation unit 15 estimates the dialogue act at the utterance unit using the dialogue video information (for example, acoustic information, video information, language information, non-verbal information) for each utterance output by the information extraction unit 13 (S25). For example, the dialogue act estimation unit 15 uses the dialogue video information for each utterance output from the dialogue segmentation unit 14 to estimate whether the utterance is an utterance that advances the dialogue or an utterance that is a reaction in the dialogue.

[0058] After S24, the importance estimation unit 16 estimates the importance of each utterance using the acoustic information, video information, language information, and non-verbal information output from the dialogue video input unit 12 (S26). Then, the importance estimation unit 16 outputs the video information, language information, and importance information indicating the estimated importance.

[0059] The summary evaluation value calculation unit 17 calculates the summary evaluation value for each utterance based on the dialogue act information for each utterance output from the dialogue act estimation unit 15 in S25, the video information, language information, and importance information output from the importance estimation unit 16 in S26, and the summary viewpoints (particularly, the importance of content understanding and the importance of atmosphere understanding) output from the summary viewpoint input unit 11 in S21 (S27).

[0060] After S27, the summary video generation unit 18 generates a summary video using the summary viewpoints (particularly, the importance of playback time) input in S21 and the summary evaluation value calculated in S27 (S28). First, the summary video generation unit 18 sets the playback time of the summary video to be short according to the importance of the playback time of the summary video. Then, within the playback time of the summary video, the summary video generation unit 18 preferentially selects the utterances with high summary evaluation values and generates a summary video using the selected utterances. After that, the summary video output unit 19 outputs the summary video generated in S28 (S29).

[0061] By executing the above processing, the summary device 10 can generate a summary video that enables the user to easily grasp the listener's reaction to the speaker's utterance in the dialogue data. In addition, the summary device 10 can generate a summary video that reflects the viewpoints that the user values.

[0062] [Details of the summary video generation process] Next, the summary video generation process in S28 of FIG. 7 will be described in detail. First, the process by which the summary video generation unit 18 selects an utterance to be used for the summary video from the dialogue video data will be described with reference to FIG. 8.

[0063] First, the summary video generation unit 18 determines the length of the summary video to be output (the playback time of the summary video) (S31).

[0064] For example, the summary video generation unit 18 acquires the total time of the entire dialogue (e.g., 20 minutes) from the dialogue segmentation unit 14. Then, the summary video generation unit 18 uses the importance degree of the playback time (e.g., 0.2) received by the summary perspective input unit 11 to set how much shorter or longer the summary video should be from the reference value of the length of the summary video (e.g., 50% of the playback time of the dialogue video data). For example, the summary video generation unit 18 calculates the summary video length by the following formula (5).

[0065] Summary video length = Total time of the entire dialogue * Reference value * (1 - Importance degree of playback time / 2) … Formula (5)

[0066] Next, the summary video generation unit 18 selects the sections to be left in the output summary video. In the selection of sections, it can be handled as a knapsack problem such that the total playback time obtained by adding up the playback times of each section falls within the above-mentioned summary video length, and the total of the summary evaluation values within the summary video length is maximized. Examples of the selection method include the greedy method, in which sections are selected in descending order of the summary evaluation value so that the total playback time falls within a predetermined time.

[0067] First, the summary video generation unit 18 creates an empty list (playback list) (S32). Next, the summary video generation unit 18 determines whether the output (output list) of the summary evaluation value calculation unit 17 is empty (S33). If the output list is empty (Yes in S33), the process ends.

[0068] On the other hand, if the output list is not empty (No in S33), the summary video generation unit 18 selects the utterance with the highest summary evaluation value from the output list (S34). Note that the utterance selected here is called the selected utterance.

[0069] After S34, the summary video generation unit 18 determines whether the time obtained by adding the time of the selected utterance and the total time of all the utterances in the playback list is shorter than the summary video length calculated in S31 (S35: Time of the selected utterance + Total playback time of the playback list < Summary video length?).

[0070] Here, if the sum of the time of the selected utterance and the total time of all the utterances in the playback list is shorter than the summary video length (Yes in S35), the process ends.

[0071] On the other hand, if the sum of the time of the selected utterance and the total time of all the utterances in the playback list is longer than the summary video length (No in S35), the summary video generation unit 18 adds the selected utterance to the playback list (S36) and removes the selected utterance from the output list (S37). Then, it returns to S33.

[0072] Then, the summary video generation unit 18 repeats the processes after S33 for all the utterances in the output list. After that, the summary video generation unit 18 finally sets the remaining playback list as a list (playback list) of sections to be left in the summary video.

[0073] By doing so, the summary video generation unit 18 can select utterances of the dialogue video data so as to fit within the set summary video length.

[0074] For example, the summary video generation unit 18 can select each utterance shown in FIG. 9 (the utterance of Speaker A, "Is there anything else?", the utterance of Speaker D, "How about yakitori?", the utterance of Speaker B, "Yakitori", and the utterance of Speaker A, "Sounds good") so as to fit within the set summary video length.

[0075] Next, with reference to FIG. 10, the process by which the summary video generation unit 18 eliminates the overlap of the utterance times in the playback list will be described.

[0076] Among the utterances in the playback list, there may be cases where there is an overlap between the start time and the end time. In this case, the summary video generation unit 18 concatenates the utterances and adds them to the merge list as a new utterance section. As a method for creating the merge list, for example, a method in which the summary video generation unit 18 sequentially compares the start time and the end time of each utterance can be mentioned.

[0077] First, the summary video generation unit 18 creates an empty list (merge list) (S41). Next, if the playback list is empty (Yes in S42), the summary video generation unit 18 ends the process.

[0078] On the other hand, if the playback list is not empty (No in S42), the summary video generation unit 18 selects the utterance with the earliest start time from the playback list (S43). The utterance selected here is called "selected utterance: previous".

[0079] After S43, if the playback list has only one line (Yes in S44), the summary video generation unit 18 adds the selected utterance: previous to the merge list (S55) and ends the process.

[0080] On the other hand, if the playback list has more than one line (No in S44), the summary video generation unit 18 selects the utterance with the second earliest start time from the playback list (S45). The utterance selected here is called "selected utterance: next".

[0081] After S45, if the start time of the selected utterance: next is after the end time of the selected utterance: previous (Yes in S46), the summary video generation unit 18 adds the selected utterance: previous to the merge list (S47). Then, after removing the selected utterance: previous from the playback list (S48), the summary video generation unit 18 returns to S42.

[0082] On the other hand, if the start time of the selected utterance: next is before or the same as the end time of the selected utterance: previous (No in S46), the process proceeds to S51. Here, if the end time of the selected utterance: next is after the end time of the selected utterance: previous (Yes in S51), the summary video generation unit 18 overwrites the end time of the selected utterance: previous with the end time of the selected utterance: next (S52). Then, the process proceeds to S53. On the other hand, if the end time of the selected utterance: next is the same as or earlier than the end time of the selected utterance: previous (No in S51), the process skips the S52 process and proceeds to S53.

[0083] Thereafter, the summary video generation unit 18 concatenates the selected utterance items other than the start time and end time later with the selected utterance: before (S53). The concatenation at this time may be a concatenation of character strings. Then, the summarization device 10 removes the selected utterance: later from the playback list (S54) and returns to S42.

[0084] For example, when the summary video generation unit 18 performs the above processing, for example, in FIG. 9, the information of each item of "yakitori" of speaker B and "that's fine" of speaker A whose speech times overlap is concatenated to obtain the information shown in FIG. 11. Then, the summary video generation unit 18 cuts and concatenates the acoustic information and video information output from the dialogue video input unit 12 based on the video section (start time, end time) of the utterance shown in FIG. 11 to generate one summary video.

[0085] By the summary video generation unit 18 performing the above processing, a summary video of the dialogue video data can be generated.

[0086] [System configuration, etc.] In addition, each component of each part shown in the figure is a functional concept, and it is not necessarily physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program executed by the CPU, or can be realized as hardware by wired logic.

[0087] In addition, among the processes described in the above-described embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified.

[0088] [Program] The above-described summarization device 10 can be implemented by installing a program (interactive video summarization program) as package software or online software on a desired computer. For example, by causing the information processing device to execute the above program, the information processing device can function as the summarization device 10. The information processing device mentioned here includes mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and further includes terminals such as PDAs (Personal Digital Assistant) within its scope.

[0089] FIG. 12 is a diagram showing an example of a computer that executes an interactive video summarization program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0090] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System), for example. The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100, for example. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0091] The hard disk drive 1090 stores, for example, an OS 1091, application programs 1092, program modules 1093, and program data 1094. That is, the programs defining each process executed by the above-described summarization device 10 are implemented as program modules 1093 in which computer-executable code is described. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, program modules 1093 for executing processes similar to the functional configurations in the summarization device 10 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0092] In addition, the data used in the processes of the above-described embodiments is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program modules 1093 and program data 1094 stored in the memory 1010 or the hard disk drive 1090 to the RAM 1012 and executes them as needed.

[0093] Note that the program modules 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, and may be stored, for example, in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program modules 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). Then, the program modules 1093 and program data 1094 may be read by the CPU 1020 from the other computer via the network interface 1070.

Explanation of Signs

[0094] 10 Summarization device (Interactive video summarization device) 11 Abstract viewpoint input section 12 Dialogue video input section 13 Information extraction section 14 Dialogue segmentation section 15 Dialogue act estimation section 16 Importance estimation section 17 Abstract evaluation value calculation section 18 Abstract video generation section 19 Abstract video output section

Claims

1. A dialogue video input unit that receives an input of dialogue video data, A dialogue act estimation unit that divides the dialogue video data for each utterance and estimates whether each utterance is a response utterance in the dialogue, An importance estimation unit that estimates the importance of each utterance, Based on the height of the importance of each utterance and whether each utterance is a response utterance to an utterance with an importance equal to or higher than a predetermined value, from the utterances included in the dialogue video data, select the utterances with high importance and the response utterances to the utterances with an importance equal to or higher than the predetermined value, and generate a summary video of the dialogue video data using the selected utterances. A summary video generation unit, A summary video output unit that outputs the generated summary video A dialogue video summarization device, characterized by comprising:

2. Further comprising a summary evaluation value calculation unit that calculates a summary evaluation value, which is an evaluation value for whether each utterance is used for generating the summary video, based on the height of the importance of the utterance and whether the utterance is a response utterance to an utterance with an importance equal to or higher than a predetermined value, The summary video generation unit Selects the utterances to be used for generating the summary video based on the magnitude of the calculated summary evaluation value of the utterance The dialogue video summarization device according to claim 1, characterized by the above.

3. Further comprising an importance input unit that receives an input from a user of the degree of emphasis on understanding the atmosphere of the dialogue in generating the summary video, The summary evaluation value calculation unit The larger the magnitude of the input degree of emphasis on understanding the atmosphere of the dialogue, the larger the value added to the summary evaluation value when the utterance is a response utterance The dialogue video summarization device according to claim 2, characterized by the above.

4. The importance input unit further Receives an input from a user of the degree of emphasis on understanding the content in generating the summary video, The abstract evaluation value calculation unit increases the value added to the abstract evaluation value according to the importance of the utterance as the importance of the input content understanding increases. The dialogue video summarization apparatus according to claim 3, characterized in that.

5. The importance input unit further receives an input of the importance of the playback time in the generation of the summary video from the user, The summary video generation unit further shortens the playback time of the generated summary video below a predetermined reference value as the importance of the input playback time increases. The dialogue video summarization apparatus according to claim 3, characterized in that.

6. further includes a dialogue splitting unit that splits the dialogue video data into a beginning stage, a middle stage, and an end stage, The summary evaluation value calculation unit further increases the value added to the summary evaluation value according to the importance of the utterance when the utterance is an utterance in the beginning stage or the end stage. The dialogue video summarization apparatus according to claim 4, characterized in that.

7. A dialogue video summarization method executed by a dialogue video summarization apparatus, comprising: a step of receiving an input of dialogue video data; a step of splitting the dialogue video data for each utterance and estimating whether each utterance is a reactive utterance in the dialogue; a step of estimating the importance of each utterance; based on the importance level of each utterance and whether each utterance is a reactive utterance to an utterance with an importance level equal to or higher than a predetermined value, selecting from the utterances included in the dialogue video data the utterances with high importance and the reactive utterances to the utterances with an importance level equal to or higher than the predetermined value, and generating a summary video of the dialogue video data using the selected utterances; a step of outputting the generated summary video and a dialogue video summarization method characterized by including.

8. A step of receiving an input of dialogue video data, a step of dividing the dialogue video data for each utterance and estimating whether each of the utterances is a response utterance in the dialogue, a step of estimating the importance of each utterance, a step of selecting, from the utterances included in the dialogue video data, an utterance with a high importance and a response utterance to an utterance with an importance equal to or higher than a predetermined value based on the height of the importance of each utterance and whether each of the utterances is a response utterance to an utterance with an importance equal to or higher than the predetermined value, and generating a summary video of the dialogue video data using the selected utterances, a step of outputting the generated summary video, and a dialogue video summarization program characterized by causing a computer to execute the above steps.

Citation Information

Patent Citations

  • Apparatus for classifying and collecting mist

    JP1983098117A

  • Video digesting device and video digesting program

    JP2012044390A

  • Argument structure update device, argument structure update method, and argument structure update program

    JP2018147196A

  • Speech summary generation apparatus, speech summary generation method, and program

    JP2020071676A

  • Dialogue action estimation device, dialogue action estimation method, dialogue action estimation model learning device and program

    JP2020173608A