Information display device, information display method, and program
The information presentation device addresses the challenge of controlling psychological distance in online communication by using auditory feedback to adjust perceived space, enhancing communication comfort and effectiveness.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2022-06-15
- Publication Date
- 2026-05-26
AI Technical Summary
In online communication environments, such as web conferencing, it is difficult to effectively control the psychological distance between participants, leading to uncomfortable communication due to restricted nonverbal element transmission and uniform information presentation.
An information presentation device that acquires speech voice data from participants, determines the communication state, and outputs auditory information to adjust the perceived psychological distance through sound-induced self-motion sensations, using earphones and bone conduction earphones to present forward or backward-moving sounds based on predefined thresholds.
Facilitates comfortable communication by adjusting the perceived distance between participants, promoting smoother interactions by extending or contracting the perceived space through auditory feedback.
Smart Images

Figure 0007865383000003 
Figure 0007865383000004 
Figure 0007865383000005
Abstract
Description
[Technical Field]
[0001] One aspect of this invention relates, for example, to an information presentation device, an information presentation method, and a program for an online communication environment utilizing a network. [Background technology]
[0002] When people communicate with each other, they unconsciously gauge the physical or psychological distance between themselves and the other person. This sense of distance is sometimes called "maai" (the appropriate distance or timing), and people with strong communication skills are adept at gauging this timing. In recent years, voice communication in online environments, such as web conferencing, is becoming mainstream due to changing social trends. In such environments, the transmission of nonverbal elements is more restricted compared to face-to-face interaction, and communication tends to become uniform in its information presentation. In other words, it becomes difficult to communicate effectively. In face-to-face interactions, it is possible to understand the other person's state and control the sense of distance between people, but in online environments, it is difficult to control the sense of distance, making it difficult to maintain a comfortable distance for both oneself and the other person.
[0003] Incidentally, Non-Patent Document 1 reports the results of a psychological experiment showing that when there is a sense of forward movement, the near-body space expands forward (expansion of near-body space). Furthermore, Non-Patent Document 2 reports that presenting sounds from the front can induce a sense of moving forward, and presenting sounds from behind can induce a sense of moving backward (sound-induced self-motion sensation). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] JP Noel et.al,Full body action remapping of peripersonal space, Neuropsychologia(2012) [Non-Patent Document 2] Sakamoto et al., A Study on Self-Motion Sensation Induced by Auditory Information, VR Society National Conference, 2002. [Overview of the project] [Problems that the invention aims to solve]
[0005] It is known that by devising ways to present information, it is possible to expand a person's perceived near-body space and induce a sense of self-movement. Utilizing these phenomena could potentially lead to smoother online communication. This invention was made in view of the above circumstances, and its purpose is to provide a technology that can facilitate comfortable communication even in a remote environment. [Means for solving the problem]
[0006] An information presentation device according to one aspect of this invention comprises a user information acquisition unit, a determination unit, and an output unit. The user information acquisition unit acquires the speech voice data of the first participant and the speech voice data of the second participant in a two-way telecommunication environment between at least one first participant and one second participant, each equipped with an acoustic device. The determination unit determines the state of communication between the first participant and the second participant based on the speech voice data of the first participant and the speech voice data of the second participant. The output unit causes auditory information corresponding to the determined state to be output to the acoustic device of the first participant. [Effects of the Invention]
[0007] According to one aspect of this invention, it is possible to provide a technology that can facilitate comfortable communication even in a remote environment. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is a diagram illustrating the elemental technologies of the web conferencing system according to the embodiment. [Figure 2]FIG. 2 is a functional block diagram showing an example of the information presentation device 1 shown in FIG. 1. [Figure 3] FIG. 3 is a diagram for explaining the state data 12a shown in FIG. 2. [Figure 4] FIG. 4 is a diagram for explaining the threshold value 12b shown in FIG. 2. [Figure 5] FIG. 5 is a diagram for explaining the presentation content 12c shown in FIG. 2. [Figure 6] FIG. 6 is a flowchart showing an example of the processing procedure of the information presentation device 1 having the above configuration. [Figure 7] FIG. 7 is a flowchart showing an example of the processing procedure in step S10 of FIG. 6. [Figure 8] FIG. 8 is a flowchart showing an example of the processing procedure in step S11 of FIG. 6. [Figure 9] FIG. 9 is a diagram showing an example of the association between the determined communication state and the presentation content. [Figure 10] FIG. 10 is a diagram for explaining that the psychological distance can be controlled by auditory information. [Figure 11] FIG. 11 is a flowchart showing an example of the processing procedure in step S13 of FIG. 6. [Figure 12] FIG. 12 is a diagram for explaining a series of processing procedures in the information presentation device 1 of the embodiment. [Figure 13] FIG. 13 is a diagram for explaining the effects obtained by the embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the embodiments, in two-way telecommunication (online communication) using a network, a technique for constructing a distance at which participants (users) can easily interact with each other will be described.
[0010] Since a web conference requires at least two participants (let's call them User 1 and User 2), for simplicity, the following explanation will assume only User 1 and User 2. Of course, the same arguments apply to web conferences with three or more participants.
[0011] Figure 1 is a diagram illustrating the elemental technologies of a web conferencing system according to an embodiment. In Figure 1, the conference equipment 2 of the web conference participants communicates with the information presentation device 1 via a network 100, which is the so-called internet, for example via a VPN (Virtual Private Network). The information presentation device 1 acquires the user's spoken audio data acquired by the microphone 30 via the network 100 and determines the state of communication between the web conference participants. The information presentation device 1 transmits auditory information corresponding to the determination result to the acoustic device worn by the participant for output. The participant may wear only a regular earphone 41 as the acoustic device, or may wear both an earphone 41 and a bone conduction earphone 42. Here, the earphone 41 is an example of a first device that plays conversational audio, and the bone conduction earphone 42 is an example of a second device that plays acoustic information different from conversational audio.
[0012] <Structure> Figure 2 is a functional block diagram showing an example of the information presentation device 1 shown in Figure 1. Multiple participants' conference equipment 2 are connected to the network 100, and these communicate with each other using a common protocol with the information presentation device 1.
[0013] The information display device 1 is a computer comprising an interface unit 13, a processor 11, storage 12, and memory 14. The interface unit 13 establishes a communication link between the network 100 and each conference equipment 2, and exchanges various types of data.
[0014] The processor 11 is a computing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and it implements the processing functions of the embodiment according to the program 14a loaded from the storage 12 into the memory 14. The memory 14 is a semiconductor memory such as ROM (Read Only Memory) or RAM (Random Access Memory).
[0015] The storage 12 is a non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive), and stores basic software such as an OS (Operating System) and a program for realizing the processing according to the embodiment. In other words, the program can be installed on the information presentation device 1. Furthermore, storage 12 stores state data 12a, threshold 12b, and presentation content 12c.
[0016] The processor 11 includes, as processing functions according to one embodiment of this invention, a user information acquisition unit 111, a state data calculation unit 112, a state determination unit 113, a presentation content acquisition unit 114, and an output unit 115. The user information acquisition unit 111, the state data calculation unit 112, the state determination unit 113, the presentation content acquisition unit 114, and the output unit 115 are realized by the processor 11 executing a program 14a loaded into the memory 14.
[0017] In other words, program 14a includes instructions to make the processor 11 function as a user information acquisition unit 111, a state data calculation unit 112, a state determination unit 113, a presentation content acquisition unit 114, and an output unit 115.
[0018] The user information acquisition unit 111 acquires the speech voice data of the first user participating in the web conference and the speech voice data of the second user via the network 100. The state data calculation unit 112 calculates state data that reflects the state of telecommunication based on the acquired speech data.
[0019] Figure 3 is a diagram illustrating the state data 12a. The state data 12a includes, for example, the percentage of silent intervals in a conversation, the percentage of the duration of the user's (first user's) utterances, and the percentage of the duration of the other party's (second user's) utterances. The state data 12a is stored in storage 12, with each entry's numerical value associated with a record ID (IDentification) corresponding to a specific time. The record ID may increase over time, but if the number of entries reaches a default value, older entries may be deleted in order, and new rows may be added.
[0020] Returning to Figure 2, we will continue the explanation. The state determination unit 113 determines the state of communication between the first user and the second user based on the state data 12a calculated from the acquired speech voice data. Based on the state data 12a, the state determination unit 113 calculates an index indicating the psychological distance between the first user and the second user.
[0021] Indicators that indicate psychological distance include, for example, the average percentage of silent time, the average percentage of spontaneous speech, and the average percentage of time the other party speaks. These indicators can be considered to reflect the state of telecommunication. Their calculation will be explained later.
[0022] The state determination unit 113 further determines the state of communication between the first user and the second user based on a comparison of the calculated indicator with a predetermined threshold.
[0023] Figure 4 is a diagram illustrating threshold 12b. For example, threshold S for determining silence. th , threshold B for determining the bias state thThese and other parameters are predefined and stored in storage 12. Using these thresholds, along with the average silence time percentage, the average spontaneous speech time percentage, and the average interlocutor speech time percentage, the state of mutual communication among participants in a web conference can be determined. The calculation method will be described later.
[0024] Returning to Figure 2, we will continue the explanation. The content acquisition unit 114 acquires auditory information from the storage 12 according to the communication status determined by the status determination unit 113. The output unit 115 transmits the auditory information acquired by the content acquisition unit 114 to the audio device of the destination user for output. For example, the content acquisition unit 114 transmits auditory information to the earphones 41 and 42 of the second user according to the psychological distance between the first user and the second user. The auditory information is stored in the storage 12 as specific content (presentation content) to be presented to each participant.
[0025] Figure 5 is a diagram illustrating the presentation content 12c. Presentation content 12c is a table in which, for each entry corresponding to one of several communication states, the type of sound source, sound file, and specified playback time for improving that communication state are associated. Here, the type of sound source and sound file are examples of auditory information that can induce autokinetic sensation, and for example, noise sounds that do not interfere with conversation can be used.
[0026] Presentation content 12c is a table that manages auditory information that induces self-motor sensation, pre-associated with communication states. Presentation content 12c records the sound source to be presented according to the communication state, the actual audio file, and the specified playback time. For example, in the "bias improvement" state, auditory information is presented to speakers whose bias has been determined by a threshold. Note that multiple audio files may be prepared for a single state, in which case one will be randomly selected from the multiple audio files in presentation content 12c. As is already known, by playing forward-moving sounds (sounds that move from back to front), it is possible to create a sense of self-movement, such as the feeling that one is moving forward, or that the space in front of oneself is expanding. Conversely, backward-moving sounds (sounds that move from front to back) create the illusion that one is moving backward. In this embodiment, this effect is used to shorten or lengthen the psychological distance between users participating in a web conference.
[0027] <effect> Next, we will explain the operation of the above configuration. Figure 6 is a flowchart showing an example of the processing procedure for the information presentation device 1 with the above configuration. In Figure 6, the processor 11 acquires, for example, the speech voice data of each user at regular intervals and stores it in the storage 12 (step S10). Next, the processor 11 calculates state data from the acquired speech voice data and calculates an index that indicates the psychological distance between the first user and the second user. Based on this index, the processor 11 determines the state of communication between the first user and the second user (step S11).
[0028] Next, based on the determined state of communication, the processor 11 selectively retrieves the content to be presented (auditory information) from the storage 12 so that the above indicator falls within a certain range. For example, if the indicator indicates that the psychological distance is too far, the processor 11 retrieves the sound file of sound moving forward. If the indicator indicates that the psychological distance is too close, the processor 11 retrieves the sound file of sound moving backward. Then, the processor 11 transmits the retrieved sound information to the second user's acoustic device and presents the auditory information (step S13).
[0029] FIG. 7 is a flowchart showing an example of the processing procedure in step S10 of FIG. 6. In FIG. 7, the processor 11 records the user's speech at regular intervals (T) (step S21), analyzes the recorded data (step S22), and performs speaker separation (step S23). As is well known, to separate speakers from voice data, for example, data classification processing using AI (Artificial Intelligence) technology may be applied.
[0030] Next, the processor 11 calculates the silent interval time l s , the utterance interval time l us , and the utterance interval time l ur of the other party from the uttered voice data (step S24). The subscript s indicates (silent ), the subscript us indicates (utterance_sender), and the subscript ur indicates (utterance_receiver).
[0031] Next, the processor 11 calculates, for example, the silent interval time ratio R s , the utterance interval time ratio R us , and the utterance interval time ratio R ur of the other party according to, for example, Equation (1) (step S25).
[0032]
Equation
[0033] Then, the processor 11 records these calculated amounts in the storage 12 (step S26).
[0034] FIG. 8 is a flowchart showing an example of the processing procedure in step S11 of FIG. 6. In FIG. 8, the processor 11 reads the latest M rows of R us , R ur , R s (FIG. 3) recorded in the state data 12a of the storage 12 (step S31). Note that the reading range M may be set in advance.
[0035] Next, the processor 11 calculates the average silence time ratio R from the information it has read. s_ave , average spontaneous speech time ratio R us_ave , and the average percentage of time the other party speaks R ur_ave This is calculated using equation (2) (step S32).
[0036]
number
[0037] Next, the processor 11 retrieves the silent threshold S from the storage 12. th And the threshold B for biased state th Read out (step S33). Next, the processor 11 calculates the average silence time ratio R s_ave and S th Compare with (step S34), if R s_ave >S th If so, the communication state between oneself and the other party is determined to be [silence] (step S35). If the answer in step S34 is No, the processor 11 determines the average spontaneous speech time ratio R us_ave and B th Compare with (step S36), if R us_ave >B th Therefore, the state of communication between oneself and the other person is judged as [one's biased state] (Step S37).
[0038] If the answer in step S36 is No, then processor 11 calculates the average interlocutor utterance time ratio R ur_ave and B th Compare with (step S36), if R ur_ave >B th If so, the state of communication between oneself and the other party is determined to be [the other party's biased state] (step S39). If the answer in step S38 is No, the processor 11 concludes that the state of communication between oneself and the other party [does not require improvement] (step S40).
[0039] Figure 9 shows an example of the correspondence between the determined communication state and the presented content. In this embodiment, auditory information is selected to be presented to the user according to the items that need improvement during the speech action. That is, if silence is determined, auditory information that gives the sensation of moving forward (forward movement sound) is selected. Also, if there is a bias in the speaker, auditory information that gives the sensation of moving backward (backward movement sound) is selected (for the speaker who is speaking biasedly).
[0040] Figure 10 illustrates how psychological distance can be controlled by auditory information. As shown in Figure 10, in this embodiment, the psychological distance between the user and their communication partner is involuntarily controlled by the presentation of auditory information. Specifically, the technology described in Non-Patent Document 1 (extension of the near-body space in the direction of self-motion sensation) allows for control of the distance between oneself and the other person. Furthermore, the technology described in Non-Patent Document 2 (motion sensation through auditory information) allows for the induction of self-motion sensation by presenting noise to the auditory system, thereby controlling the near-body space. By combining these effects, psychological distance can be controlled.
[0041] In other words, as shown in Figure 9(a), extending the peri-personal space forward can reduce the perceived psychological distance from the other person. Conversely, as shown in Figure 9(b), shrinking the peri-personal space backward can increase the perceived psychological distance from the other person. Here, peri-personal space is a concept that refers to the space around the body within arm's reach.
[0042] When actual movement is involved, the surrounding space expands or contracts in the direction of the movement. In this embodiment, auditory effects induce a sense of self-motion solely through auditory information presentation, causing the user to feel as if they are moving forward or backward. By utilizing this, it is possible to involuntarily shorten or lengthen the perceived distance to another person, independently of the user's will.
[0043] Figure 11 is a flowchart showing an example of the processing procedure in step S13 of Figure 6. In Figure 11, the processor 11 communicates with the user's conference equipment 2 to obtain the number of audio devices worn by the user (step S51). If the user is wearing, for example, two devices such as earphones 41 and bone conduction earphones 42, it is determined that multiple devices are being worn (Yes in step S52). In this case, the processor 11 merges (superimposes) the presented information (auditory information) with the voice of another speaker and plays it back (step S53).
[0044] If the answer in step S52 is No, meaning the user is wearing either the earphone 41 or the bone conduction earphone 42, the processor 11 causes the device not playing the other speaker's voice to play the presented information (step S54). In other words, the output unit 115 sends the auditory information to the bone conduction earphone 42, which is the second device. For IP (Internet Protocol) based communication, the destination can be distinguished, for example, by port number.
[0045] Steps S52, S53, and S54 are repeated until playback for the specified playback time is complete, so that even short audio files can be played for a certain period of time. In other words, playback of the audio information is repeated until playback for the specified playback time is completed (Yes in step S55).
[0046] Figure 12 is a diagram illustrating a series of processing steps in the information presentation device 1 of the embodiment. The processor 11 of the embodiment determines the current state of communication based on the content of previous utterances (step S100: determination of communication state), and selects appropriate presentation information according to the determined state (step S200: selection of presentation information according to state). Then, the processor 11 presents the generated information to the auditory system to encourage improvement of the operation (step S300: auditory presentation of information to encourage improvement).
[0047] <Effects> As described above, in this embodiment, the information presentation device 1 acquires the user's speech actions during a web conference and determines whether improvement of the actions is necessary, taking into account past acquisition history. To this end, it records the speech actions and determines the duration of the speech. It also determines the occurrence of speech for each user, both for themselves and their conversation partner. These processes are performed at arbitrary sampling intervals and compared with a preset threshold to determine whether there are inappropriate states such as "silence" or "speaker bias."
[0048] Figure 13 is a diagram illustrating the effects obtained by the embodiment. As shown in Figure 13(a), auditory information controls the perceived distance between speakers, closing the distance when the psychological distance is far and increasing the distance when the distance is close. By providing auditory information as feedback according to the state of communication in this way, as shown in Figure 13(b), it is possible to guide each to an ideal state, maintaining a distance that facilitates conversation while promoting improved conversational satisfaction.
[0049] In other words, the embodiment focuses on the fact that the sense of distance in a conversation changes depending on auditory information, and presents auditory information to adjust the sense of distance appropriately. Furthermore, it determines whether the sense of distance is close or far depending on the conversation state and presents auditory information to adjust the sense of distance appropriately. As a result, according to this embodiment, it is possible to provide a technology that can promote comfortable communication even in a remote environment.
[0050] It should be noted that this invention is not limited to the embodiments described. For example, the embodiments envision all forms of communication via online communication tools that involve audio, such as web conferencing. However, the technology disclosed in the embodiments can also be implemented in face-to-face communication, for example, in the form of a smartphone application that can record and analyze the user's speech.
[0051] Furthermore, the number of speakers in a web conference may exceed two. In such cases, the duration of each speaker's utterance can be calculated, and speakers with disproportionate speaking time can be identified. In addition, when presenting auditory information, instead of simply playing it for a set period and ending the presentation, the system may continue to assess the user's state in real time and present the information as needed when the user's condition requires improvement. Furthermore, the number of utterances per unit of time and the duration of silence can also be used as indicators of psychological distance.
[0052] Furthermore, the present invention can be implemented by modifying its components without departing from its essence. Moreover, various inventions can be formed by appropriately combining the multiple components disclosed in the above embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of Symbols]
[0053] 1...Information presentation device 2…Conference facilities 11… Processor 12…Storage 12a...Status data 12b... threshold 12c…Presentation content 13… Interface section 14…Memory 14a...Program 30... Mike 41…Earphones 42...Bone conduction earphones 100…Network 111...User Information Acquisition Unit 112...Status data calculation unit 113... State determination unit 114…Presentation content acquisition unit 115... Output section.
Claims
1. A user information acquisition unit that acquires the speech voice data of the first participant and the speech voice data of the second participant in a two-way telecommunication environment between at least a first participant and a second participant, each equipped with an acoustic device, A determination unit that determines the state of communication between the first participant and the second participant based on the speech voice data of the first participant and the speech voice data of the second participant, based on an index indicating the psychological distance between the first participant and the second participant, An information presentation device comprising: an output unit that outputs auditory information capable of inducing a sense of self-movement according to the determined state to the acoustic device of the first participant.
2. A memory unit that stores auditory information capable of inducing the aforementioned sense of self-movement, The system further comprises an auditory information acquisition unit that acquires the auditory information from the storage unit, The auditory information acquisition unit acquires auditory information that has been pre-associated with the state from the storage unit, The information presentation device according to claim 1, wherein the output unit transmits the acquired auditory information to the acoustic device of the first participant.
3. The information presentation device according to claim 2, wherein the auditory information acquisition unit selectively acquires the auditory information from the storage unit such that the index falls within a certain range.
4. The memory unit stores forward-moving sounds and backward-moving sounds. The aforementioned auditory information acquisition unit, If the aforementioned indicator indicates that the psychological distance is too great, the forward-moving sound is acquired. The information presentation device according to claim 3, which acquires the rearward movement sound when the indicator indicates that the psychological distance is too close.
5. The information presentation device according to any one of claims 2 to 4, wherein the determination unit uses either the number of utterances per unit time or the duration of silence as an indicator of the sense of psychological distance.
6. The information presentation device according to claim 1, wherein the acoustic device includes either a first device for reproducing conversational audio or a second device for reproducing acoustic information different from the conversational audio.
7. In a computer comprising a processor and a memory unit, an information presentation method is performed by the processor, The processor includes a process for acquiring speech data from the first participant and speech data from the second participant in a two-way telecommunication environment between at least one first participant and one second participant, each equipped with an acoustic device. The process by which the processor determines the state of communication between the first participant and the second participant based on the speech data of the first participant and the speech data of the second participant, based on an index indicating the psychological distance between the first participant and the second participant. An information presentation method comprising the process of the processor outputting auditory information that can induce a sense of self-motion in accordance with the determined state to the acoustic device of the first participant.
8. In a program that includes instructions to be executed by the processor of a computer comprising a processor and a memory unit, Instructions to cause the processor to execute a process of acquiring the speech voice data of the first participant and the speech voice data of the second participant in a two-way telecommunication environment between at least one first participant and one second participant, each equipped with an acoustic device, The processor is given an instruction to execute a process of determining the state of communication between the first participant and the second participant based on the speech data of the first participant and the speech data of the second participant, based on an index indicating the psychological distance between the first participant and the second participant. A program comprising an instruction that causes the processor to execute a process of outputting auditory information that can induce a sense of self-motion in accordance with the determined state to the acoustic device of the first participant.