Dialogue playback method, dialogue playback device, and program

By extracting and visualizing time-series data of feature amounts for participants, the method addresses the challenge of capturing mental movements in dialogues, enabling a deeper understanding of conversation dynamics and facilitating a third-party perspective.

JP7700853B2Active Publication Date: 2025-07-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023526778
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-10
Publication Date
2025-07-01
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture the mental movements of participants in dialogues, making it difficult to understand and analyze the conversation dynamics.

Method used

A method to extract time-series data of feature amounts for each participant, calculate weighted sums, and generate synchronized display data to visualize mental movements during dialogues, using a computer to facilitate the analysis of mental states through synchronized video playback.

Benefits of technology

Enables the visualization and understanding of participants' mental movements during dialogues, allowing for a deeper analysis of conversation dynamics and facilitating a third-party perspective on the conversation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700853000001
    Figure 0007700853000001
  • Figure 0007700853000002
    Figure 0007700853000002
  • Figure 0007700853000003
    Figure 0007700853000003
Patent Text Reader

Abstract

In the present invention, the emotions of a participant in a dialogue can be ascertained by causing a computer execute: an extraction process for extracting time-series data of characteristic quantities for each participant in the dialogue from time-series data related to the dialogue, where the values can change according to the emotions of the participants of the dialogue; a generation process for synchronizing with the time of the video of the dialogue and generating display data that visualizes the time-series data of the characteristic quantities together with the video; and a display process which displays the display data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive playback method, an interactive playback device, and a program.

Background Art

[0002] Conventionally, an activity called a living lab has been carried out in which various people including strangers gather and have conversations to create services. However, it is not easy for strangers to build a relationship in which they can have conversations in a free atmosphere and have better conversations. A process is required in which the speaker correctly conveys what they think, and the listener understands while relativizing it with their own thoughts. Participants in the conversation need to be familiar with these processes and have speaking and listening skills.

[0003] Therefore, there is a way of proceeding with a conversation in which, while looking back on the conversation content to assist thinking, it is connected to the next conversation.

[0004] As a method of looking back on the conversation content, there is also a technology that has already been disclosed. In Non-Patent Document 1, a technology for automatically transcribing the conversation content is disclosed.

[0005] Further, in Patent Document 1, a technology for summarizing a video by weighting various features extracted from the video, setting feature amounts by the viewer themselves, and extracting video sections is disclosed.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Non-Patent Documents

[0007]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] However, in the prior art, it is difficult to look back on the mental movements of the participants in the dialogue.

[0009] The present invention has been made in view of the above points, and an object thereof is to enable grasping of the mental movements of the participants in the dialogue.

Means for Solving the Problems

[0010] Therefore, in order to solve the above problems, time-series data of feature amounts for each participant in the dialogue are extracted from the time-series data related to the dialogue, the value of which can change according to the mental movements of the participants in the dialogue, For the same type of each participant the weighted sum of the feature amounts is As a feature quantity of all participants a computer executes an extraction procedure for calculating, a generation procedure for generating display data for visualizing the time-series data of the weighted sum together with the video of the dialogue in synchronization with the time of the video of the dialogue, and a display procedure for displaying the display data.

Effects of the Invention

[0011] It is possible to grasp the mental movements of the participants in the dialogue.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Embodiments for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a diagram showing a state where participants in a dialogue are having an online dialogue. Participants in the dialogue (hereinafter simply referred to as "participants") have an online dialogue, and the state (video (shared video and participant video in FIG. 3) and audio) is saved using a recording function or the like that the video conferencing system has. Alternatively, for the video, what is acquired by screen capture may be saved, and for the audio, it may be saved by an arbitrary recording function. The dialogue playback device 10 of the present embodiment executes processing for assisting the retrospective playback of the saved video. Note that in the online dialogue, the terminal used by each participant may be used as the dialogue playback device 10, or a device different from the terminal may be used as the dialogue playback device 10.

[0014] FIG. 2 is a diagram showing an example of the hardware configuration of the dialogue playback device 10 in the embodiment of the present invention. The dialogue playback device 10 in FIG. 2 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a CPU 104, an interface device 105, a display device 106, and an input device 107, etc., which are mutually connected by a bus B.

[0015] The program for realizing the processing in the dialogue playback device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the installation of the program does not necessarily have to be performed from the recording medium 101, and it may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program and also stores necessary files, data, etc.

[0016] When there is an instruction to start a program, the memory device 103 reads and stores the program from the auxiliary storage device 102. The CPU 104 realizes the functions related to the interactive playback device 10 according to the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to the network. The display device 106 displays a GUI (Graphical User Interface) or the like according to the program. The input device 107 is composed of a keyboard, a mouse, etc., and is used to input various operation instructions.

[0017] FIG. 3 is a diagram showing a functional configuration example of the interactive playback device 10 according to an embodiment of the present invention. In FIG. 3, the interactive playback device 10 includes a feature amount extraction unit 11, a video summary unit 12, a review video generation unit 13, and a display unit 14. Each of these units is realized by processing executed by the CPU 104 by one or more programs installed in the interactive playback device 10. The interactive playback device 10 also uses a data storage unit 15. The data storage unit 15 can be realized, for example, using the auxiliary storage device 102 or a storage device that can be connected to the interactive playback device 10 via a network.

[0018] Note that the interactive playback device 10 may not have the video summary unit 12. That is, the summary of the video data described later may not be performed.

[0019] Hereinafter, the processing executed by each unit will be described.

[0020] [Feature Amount Extraction Unit 11] The feature amount extraction unit 11 receives, for example, after the end of the dialogue or during the dialogue, the video data of the dialogue (the time-series data of the video output to the terminal of any one participant (the shared video and the participant video in FIG. 1)), the time-series data of the audio of the dialogue, and one or more other time-series input data. Hereinafter, the time-series data of the video will be referred to as "video data". The time-series data of the audio will be referred to as "audio data". The other time-series input data will be referred to as "time-series input data".

[0021] Time-series input data is data that has time information that can be associated (synchronized) with time information such as timestamps in video data or audio data. Although time-series input data is assumed to be mainly data acquired during a conversation, data having pre-prepared time information may be regarded as time-series input data.

[0022] Time-series input data refers to time-series data of numerical values that can change according to the mental state of participants, in addition to video and audio. Examples of time-series input data include, as equipment used, time-series pulse rates (numerical values indicating such) of each participant measured during a conversation by a wristwatch-type pulse meter, and time-series electromyograms (numerical values indicating such), electroencephalograms (EEGs) measured during a conversation by equipment capable of measuring an eye tracker, electromyogram (EMG), or electroencephalogram (EEG), and data (numerical values) arbitrarily input in time series for one or more predetermined items by a participant himself / herself using an input device such as a mouse or keyboard during a conversation, and data (numerical values) arbitrarily input in time series for one or more predetermined items by a third party (a person other than the participant) using an input device such as a mouse or keyboard after viewing the video of the conversation. Examples of the predetermined items include items representing negative mental states that are difficult to express during a conversation, such as "There were times when I couldn't understand the conversation content," and items representing positive mental states, such as "Interesting," "I can empathize." Any item that can represent the mood felt from the conversation may be used, and it should be data that can be output as a numerical value, such as by inputting using a Likert scale. Further, a single numerical value output by an arbitrary function that extracts values having the same time information from a plurality of items and uses the plurality of items as inputs may be regarded as time-series input data.

[0023] The feature extraction unit 11 synchronizes the timing (time) for all the received data (video data, audio data, time-series input data). As a method for synchronizing the video data and the audio data, for example, the method disclosed in Japanese Patent Application Laid-Open No. 2020-198510 may be used. The synchronization of the time-series data is, for example, when the video data is used as a master (hereinafter, the data selected as the master here is referred to as "master data"), the offsets of the timestamps of the audio data and the time-series input data are measured with respect to the video data, and the measured offsets are given to the timestamps of the audio data and the time-series input data (the timestamps are changed with the measured offsets).

[0024] Next, the feature extraction unit 11 extracts, for each participant, feature quantities that can change in value according to the mental movement of the participant from the synchronized data.

[0025] In the case of video data, for each participation in the dialogue, it is desirable to extract feature quantities that appear as the influence received from the dialogue and feature quantities that give influence to the dialogue. For example, feature quantities representing the facial expressions and body movements of each participant can be considered. For the extraction of feature quantities of facial expressions and body movements, any known method such as a method using an API such as OpenFace may be used.

[0026] Similarly, in the case of audio data, for each participant, it is desirable to extract feature quantities that appear as the influence received from the dialogue and feature quantities that give influence to the dialogue. For example, the magnitude and pattern of the sound pressure of the audio waveform of each participant can be considered. Feature quantities can be extracted for each participant by extracting sound pressure above / below a predetermined threshold with respect to the time-series sound pressure of each participant, or by extracting, for each participant, the location (sound pressure waveform) where a preset sound pressure waveform pattern appears in the time-series sound pressure of each participant.

[0027] Similarly, in the case of time-series input data, values that appear as the influence received from the dialogue may be extracted for each participant, or values of parts where there was input using an input device such as a mouse or keyboard and the input exceeded a certain threshold may be extracted as feature quantities for each participant. For example, in time-series input data, values above / below a predetermined threshold may be extracted as feature quantities for each participant, or in input time-series data, parts where a preset time-series change pattern appears may be extracted as feature quantities for each participant.

[0028] Further, the feature quantity extraction unit 11 may cause one or more of the above-mentioned feature quantities to be learned by some CNN model of machine learning, and extract one or more output values as feature quantities to be used in subsequent processing. The CNN model is not limited to a specific one. For example, the time-series change information of values such as Action Unit output by an API such as OpenFace and correct data regarding what kind of expression the value indicates may be learned by the corresponding CNN model. For example, by learning a combination of one or more Action Units indicating expressions such as joy, anger, sorrow, and happiness and the time-series change values of the combination with a CNN, for each participant, data in which the values of ActionUnit are output in time series using OpenFace is input to the learned CNN for the time-series feature quantities of the participant, and values of joy, anger, sorrow, and happiness may be output as output. The output values may also be used as feature quantities for subsequent processing.

[0029] The feature extraction unit 11 associates various feature amounts F(t) obtained for each participant with an ID (hereinafter referred to as "feature amount ID") that identifies the type of each feature amount F(t). Here, t is a variable representing the time during the conversation. Therefore, the feature amount F(t) indicates the feature amount at time t. Then, the feature extraction unit 11 associates, for example, a weight w (a weight that can be set for each participant and is stored in the DB in advance for each feature amount ID) with each feature amount F(t) based on the feature amount ID. The feature extraction unit 11 also stores various feature amounts F(t) associated with their respective feature amount IDs, the master data, and the time information of the master data in the data storage unit 15. However, the weight w may be set by the user at the time of feature extraction. When the weight w is set at the time of feature extraction, the set weight w is associated with the feature amount ID at the time of setting the weight w.

[0030] Next, the feature extraction unit 11 calculates the feature amount FA(t) of the entire participant for each ID. When the weight that can be set for each participant is w, FA(t) is calculated as follows. FA(t)=F1(t)*w1+F2(t)*w2+F3(t)*w3+F4(t)*w4 (in the case of 4 participants) Here, FN(t) (here, N = 1 to 4) indicates the feature amount of participant N at time t. wN indicates the weight for participant N.

[0031] The feature extraction unit 11 further stores various feature amounts FA(t) associated with their respective feature amounts D in the data storage unit 15.

[0032] In addition, as a feature quantity for calculating the feature quantity FA(t), the responses of each participant to the questionnaire regarding the impressions of the dialogue can also be utilized. The feature quantity extraction unit 11 may organize one or more responses including information that can be associated with the time information of the master data as time-series data having time information, and utilize the time-series data as a feature quantity. For example, when there is a response such as "It was fun until the topic of ○○ came up", in the time-series data, the value of the feature quantity related to the time-series data may be changed before and after the corresponding topic appears.

[0033] Also, the questionnaire responses may be reflected in the weight w. For example, when a personality such as not showing much emotion can be read from the questionnaire results, it is conceivable to set the weight w of that person to be arbitrarily larger than that of other participants.

[0034] [Video Summarization Unit 12] The video summarization unit 12 performs summarization of the video data. Summarization of the video data means extracting and combining video data of one or more partial intervals from the video data. Summarization of the video data may be performed using a known method as disclosed in Patent Document 1.

[0035] The video summarization unit 12 may utilize any feature quantity FA(t) extracted by the feature quantity extraction unit 11. For example, the video summarization unit 12 may perform summarization of the video data by combining partial intervals in the video data where any FA(t) is equal to or greater than a threshold value.

[0036] Also, when using any feature quantity F(t) extracted for each participant, the video summarization unit 12 performs the operations in paragraphs

[0033] to

[0036] for each participant in the digest video generation unit 14 described in Patent Document 1. That is, the video section after summarization is selected so that the total time is approximately equal to the video time input as a parameter using the feature quantity extracted for each participant. When using any feature quantity F(t) extracted for each participant, if there are 4 participants, there will be 4 patterns of the video sections after this selection. The video summarization unit 12 selects, from the partial videos included in any of these 4 patterns, the one with the largest linear sum of feature quantities, and when the difference between the total time and the video time becomes equal to or less than a predetermined value, proceeds to step S109 described in Patent Document 1 to generate summary video data. By doing so, it is possible to perform a summary that treats the feature quantities of each participant equally.

[0037] The video summarization unit 12 stores the generated summary video data in the data storage unit 15. The video summarization unit 12 also converts the time (timestamp) of the feature quantity F(t) or the feature quantity FA(t) at the time t on the time axis of the video data before summarization, which is included in the summary video data, into the time information on the time axis of the summary video data. That is, for the time series of the feature quantity F(t) or the feature quantity FA(t) generated by connecting the feature quantities F(t) or the feature quantities FA(t) in the time interval included in the summary video data among the feature quantities F(t) or the feature quantities FA(t), new times are set.

[0038] [Review video generation unit 13] The review video generation unit 13 receives the video data recorded during the conversation, or the summary video data when the video data has been summarized by the video summarization unit 12 (hereinafter, both cases are referred to as "target video data"), acquires the feature quantity F(t) or the feature quantity FA(t) associated with the target video data from the data storage unit 15, and generates video data (hereinafter, referred to as "review video data") in which the feature quantity is visualized (information indicating the feature quantity is superimposed) for the target video data.

[0039] The playback video generation unit 13 accepts settings from a user or the like regarding which one or more feature amounts (F(t) or FA(t)) are to be displayed and by what visualization method the feature amounts to be displayed are to be displayed (in what form (format) each feature amount is to be displayed). The visualization method can be set to various visualizations by utilizing libraries such as D3.js, for example.

[0040] The playback video generation unit 13 also accepts settings from a user or the like regarding where on the visualization video each feature amount to be displayed is to be displayed. For example, when the playback video generation unit 13 designates a certain feature amount F(t) as the display target, since the feature amount F(t) has been extracted for each participant, settings are accepted regarding the placement locations for all the participants. For example, when it is assumed that the feature amount F(t) is to be displayed on the display device 106, information on the x coordinate and y coordinate as the placement location, and vertical and horizontal pixel values as the size information of the area where the feature amount F(t) is to be displayed are set for each participant. The placement location is set in the same way when the feature amount FA(t) is the display target. However, in the case of the feature amount FA(t), settings for each participant are not necessary.

[0041] The visualization method and the setting of the placement location of the feature amount may be performed by preparing one or more sets of the visualization method setting and the setting of the placement location of the feature amount (x, y coordinates, and vertical and horizontal pixel values) and allowing selection from among them.

[0042] The feature amount to be the display target may be changed during the time series of the video data. When changing, it is also possible to set the change location in advance before generating the playback video data. Also, it is possible to change by receiving some input during video playback.

[0043] Based on the above settings, the playback video generation unit 13 generates playback video data in which the feature amount selected as the display target is visualized at the selected position by the selected visualization method for the target video data.

[0044] [Display unit 14] The display unit 14 displays on the display device 106 a screen (hereinafter referred to as the "replay screen") for playing back the replay video data generated by the replay video generation unit 13.

[0045] FIG. 4 is a diagram showing an example of the display of the replay screen. As shown in FIG. 4, the replay screen g1 is a screen including the controller c1 together with the replay area of the replay video data.

[0046] In the replay video data, for each of the participants A to D, a feature amount display area af for displaying the feature amounts of the respective participants in synchronization with the video is superimposed. In the example of FIG. 4, an example in which four types of feature amounts are the display targets is shown.

[0047] The controller c1 includes, for example, buttons for accepting playback, pause, etc. When the play button is pressed, the display unit 14 plays back the replay video data. When the pause button is pressed, the display unit 14 pauses the playback of the replay video data. When each participant's terminal is used as the interactive playback device 10, each participant can operate the controller c1. However, since the replay video data is played back in synchronization for all, when someone presses the pause button, the playback of all will pause at the same place. Each participant can press the pause button when there is something to say while viewing the replay video data, and can have a conversation with other participants.

[0048] As described above, according to the present embodiment, when playing back the video data of the conversation, the feature amounts indicating the mental movements of the participants in the conversation can be synchronously displayed. Therefore, it is possible to grasp the mental movements of the participants in the conversation. In addition, it is possible to view the conversation itself from the perspective of a third party. As a result, for example, it is possible to look back on the minute mental movements of each participant, the ideas that suddenly come to mind, and the parts of the conversation that a small number or one person considers important, so that the participants can view the conversation itself from a third-party perspective in an overview manner.

[0049] In addition, in the present embodiment, the feature amount extraction unit 11 is an example of an extraction unit. The playback video generation unit 13 is an example of a generation unit. The display unit 14 is an example of a display unit. The playback video data is an example of display data.

[0050] As described above in detail regarding the embodiments of the present invention, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

Explanation of Reference Numerals

[0051] 10 Interactive playback device 11 Feature amount extraction unit 12 Video summary unit 13 Playback video generation unit 14 Display unit 15 Data storage unit 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 CPU 105 Interface device 106 Display device 107 Input device B Bus

Claims

1. An extraction procedure for extracting time-series data of feature quantities for each participant in the conversation from time-series data related to the conversation, where the value can change according to the mental state of the participants in the conversation, and calculating the weighted sum of the same type of feature quantities of each participant as the feature quantity of the entire group of participants; A generation procedure for generating display data for visualizing the time-series data of the weighted sum together with the video of the conversation in synchronization with the time of the video of the conversation; A display procedure for displaying the display data; A conversation playback method characterized in that a computer executes the above.

2. A computer executes a summarization procedure for summarizing the video based on the feature quantities of each participant, and the generation procedure generates the display data for visualizing the time-series data of the weighted sum together with the summarized video in synchronization with the time of the summarized video. The conversation playback method according to Claim 1, characterized by the above.

3. An extraction unit that extracts time-series data of feature quantities for each participant in the conversation from time-series data related to the conversation, where the value can change according to the mental state of the participants in the conversation, and calculates the weighted sum of the same type of feature quantities of each participant as the feature quantity of the entire group of participants; A generation unit that generates display data for visualizing the time-series data of the weighted sum together with the video of the conversation in synchronization with the time of the video of the conversation; A display unit that displays the display data; A conversation playback device characterized by having the above.

4. It has a summarization unit that summarizes the video based on the feature quantities of each participant, and the generation unit generates the display data for visualizing the time-series data of the weighted sum together with the summarized video in synchronization with the time of the summarized video. The conversation playback device according to Claim 3, characterized by the above.

5. A program characterized by causing a computer to execute the conversation playback method according to Claim 1 or 2.

Citation Information

Patent Citations

  • Video digesting device and video digesting program

    JP2012044390A

  • Digest data generation device, digest data reproduction device, digest data generation system, digest data generation method, and program

    JP2019068300A

  • Mobile device for recording, reviewing, and analyzing video

    US20170111594A1

  • Emotion estimating device and emotion estimating method

    WO2016143759A1

  • Information processing system, information processing method, and recording medium

    WO2017022286A1