Information display device, information display method, and information display program

The information presentation device addresses the challenge of maintaining psychological distance in online communication by defining sound source positions based on relationship-based distances, improving comfort through adjusted audio and video presentation.

JP7845493B2Active Publication Date: 2026-04-14NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NIPPON TELEGRAPH & TELEPHONE CORP
Filing Date
2022-10-28
Publication Date
2026-04-14

Smart Images

  • Figure 0007845493000002
    Figure 0007845493000002
  • Figure 0007845493000003
    Figure 0007845493000003
  • Figure 0007845493000004
    Figure 0007845493000004
Patent Text Reader

Abstract

An information presentation device according to an embodiment of the present invention comprises a sound source position regulation unit and a speech presentation unit, and presents a plurality of speech information acquired via a network from each of one or more first participant terminals among a plurality of participant terminals participating in online communication, said speech information being presented via the network to a second participant terminal among the plurality of participant terminals. The sound source position regulation unit regulates a sound source position for each of one or more conversation partners using one or more first participant terminals on the basis of psychological distance information that expresses a psychological distance to each conversation partner as perceived by a subject using a second participant terminal, the psychological distance information being set for each conversation partner. On the basis of the sound source position for each conversation partner, the speech presentation unit generates sound field information in which speech information from the one or more first participant terminals has been subjected to sound image localization, and transmits the sound field information to the second participant terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of this invention relates to an information presentation device, an information presentation method, and an information presentation program. [Background technology]

[0002] Currently, online communication is primarily conducted via video calls using both video and audio. However, in business settings such as meetings, negotiations, and exhibitions, there are many cases where participants interact while viewing documents (slides), and in such cases, the conversation may proceed using only audio without displaying video.

[0003] In face-to-face communication, it is common for each participant to maintain a certain distance from each other, depending on their relationship with the other participants. This distance is known as personal space or the F-formation, and is an important element in achieving comfortable communication. For example, by moving away from an overbearing boss and closer to a cooperative colleague, discomfort during the conversation can be reduced to some extent.

[0004] In contrast, in online communication, the video and audio of all participants are consolidated onto a single screen and speaker.

[0005] Therefore, it is difficult for the listener to individually adjust how each conversation partner appears (size and position of face, etc.) and how they sound (volume, direction, etc.). As a result, they are more likely to be forced into conversations with people they find psychologically uncomfortable with, without being able to alleviate that discomfort.

[0006] Conversely, from the speaker's perspective, it's difficult to recognize the listener's viewing environment and understand their own appropriate visual and auditory perception. As a result, there's a risk of unintentionally being perceived as condescending by the listener and causing unnecessary discomfort.

[0007] Based on the above, in online communication, it is necessary to maintain an appropriate distance from each participant based on their relationship with one or more other participants (each conversation partner).

[0008] Therefore, efforts are being made in both research and practical services to express the distance between participants and their respective dialogue partners.

[0009] For example, as mentioned earlier, in business use, the priority of displaying video is lower, so we can focus particularly on techniques for expressing distance through sound.

[0010] Traditionally, for example, Apple's FaceTime® has implemented a feature that enhances realism by creating a spatial sound image using 3D audio technology, making it seem as if the voice is coming from the position where the person you are talking to is shown on the screen. However, this feature remains within the context of so-called reality reproduction, where the video and audio are consistent, and it is not clear whether the volume and direction of the reproduced audio are appropriate from the perspective of reducing discomfort for participants.

[0011] Furthermore, Non-Patent Document 1 proposes a spatial audio technology that intentionally separates the sound source positions of each participant to improve intelligibility. However, this simply separates each participant's sound source equally according to mechanical rules, and does not take into account the relationships between participants. In other words, no particular consideration has been given to reducing the discomfort that participants may feel towards other participants. [Prior art documents] [Non-patent literature]

[0012] [Non-Patent Document 1] M. Wong., R. Duraiswami, “SharedSpace: Spatial Audio and Video Layouts for Videoconferencing in a Virtual Room”, 2021 Immersive and 3D Audio: from Architecture to Automotive (I3DA), September 2021, DOI: 10.1109 / I3DA48870.2021.9610974 [Overview of the project] [Problems that the invention aims to solve]

[0013] This invention was made in view of the above circumstances and aims to provide an information presentation technology that can give the subject an appropriate sense of distance, based on the relationship between the subject (the participant) and each of the other participants who are their dialogue partners. [Means for solving the problem]

[0014] To solve the above problems, an information presentation device according to one aspect of the present invention presents a plurality of audio pieces of information acquired from one or more first participant terminals among a plurality of participant terminals participating in online communication via a network to a second participant terminal among a plurality of participant terminals via the network, and comprises a sound source position defining unit and an audio presentation unit. The sound source position defining unit defines the sound source position of each conversation partner based on psychological distance information that represents the psychological distance to each conversation partner from the perspective of a target person using a second participant terminal, which is set for each of the one or more conversation partners using one or more first participant terminals. The audio presentation unit generates sound field information with sound image localization based on the sound source position of each of the one or more conversation partners, and transmits it to the second participant terminal. [Effects of the Invention]

[0015] That is, according to one aspect of this invention, it is possible to provide an information presentation technology that can give an appropriate sense of distance to the target person based on the relationship between the target person and each conversation partner.

Brief Description of the Drawings

[0016] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an information presentation system according to the first embodiment of this invention. [Figure 2] FIG. 2 is a block diagram showing an example of the hardware configuration of a communication server as the first embodiment of the information presentation device of this invention. [Figure 3] FIG. 3 is a block diagram showing an example of the software configuration of the communication server. [Figure 4] FIG. 4 is a diagram showing an example of the stored content of the participant information database of the communication server. [Figure 5] FIG. 5 is a flowchart showing an example of the processing procedure and processing content of the preparation process executed by the control unit of the communication server. [Figure 6] FIG. 6 is a schematic diagram showing an example of the difference in positions with respect to each conversation partner. [Figure 7] FIG. 7 is a schematic diagram showing an example of the intimacy with respect to each conversation partner. [[ID=2,8]] [Figure 8] FIG. 8 is a diagram showing a sound source coordinate system that defines the sound source positions according to the difference in position and intimacy. [Figure 9] FIG. 9 is a schematic diagram showing the sound source positions of each conversation partner according to the difference in position. [Figure 10] FIG. 10 is a schematic diagram showing the sound source positions of each conversation partner according to the intimacy. [Figure 11] FIG. 11 is a schematic diagram showing the sound source positions of a plurality of conversation partners with the same difference in position and intimacy. [Figure 12] FIG. 12 is a flowchart showing an example of the processing procedure and processing content of the conversation process executed by the control unit of the communication server. [Figure 13]Figure 13 is a schematic diagram showing an example of the display screen on a participant's device. [Figure 14] Figure 14 is a block diagram showing an example of the software configuration of a communication server as a second embodiment of the information presentation device of the present invention. [Figure 15] Figure 15 is a flowchart showing an example of the processing procedure and processing content of the dialogue processing performed by the control unit of the communication server in the second embodiment. [Figure 16A] Figure 16A is a schematic diagram showing an example of the face area of ​​each conversation partner in the input video for each conversation partner. [Figure 16B] Figure 16B is a schematic diagram showing an example of the standardized video for each dialogue partner. [Modes for carrying out the invention]

[0017] Embodiments of this invention will be described below with reference to the drawings.

[0018] [First Embodiment] (Example configuration) (1) System Figure 1 shows an example of the configuration of an information presentation system in the first embodiment of this invention.

[0019] The information presentation system of this embodiment includes a communication server CS as the first embodiment of the information presentation device of this invention as its main component. The information presentation system enables the transmission of information data via a network NW between the communication server CS and multiple participant terminals PT used by multiple participants participating in online communication with multiple people. For each participant, the information presentation system treats the participant as the target and the other participants as the target's dialogue partners, and the communication server CS causes the participant terminal PT of the dialogue partner to present information acquired by the dialogue partner's participant terminal PT to the target's participant terminal PT. That is, the communication server CS treats each participant terminal PT as both the target's participant terminal PT and the dialogue partner's participant terminal PT.

[0020] The network NW is the internet. Of course, the network NW can be any network that is capable of transmitting the above-mentioned information data, such as a LAN (Local Area Network).

[0021] Online communication involving multiple participants encompasses all online communication that includes voice. Because it often involves participants with varying levels of psychological distance, its application is primarily intended for business settings such as meetings, negotiations, and exhibitions. Of course, it can also be applied to private conversations with family and friends.

[0022] (2) Equipment (2-1) Participant terminal PT Participant terminals (PTs) are not restricted as long as they can output audio and video and allow remote communication with others via a network such as the internet, including PCs (Personal Computers), smartphones, and glasses-type devices.

[0023] (2-2) Communication Server CS Figures 2 and 3 are block diagrams showing examples of the hardware and software configurations of the communication server CS.

[0024] The communication server CS consists of a server computer, for example, located on the web or in the cloud. The communication server CS may also be a PC, which is one of the multiple participant terminals PT.

[0025] The communication server CS comprises a control unit 1, to which a storage unit having a program storage unit 2 and a data storage unit 3, and a communication interface unit 4 are connected via a bus 5. In Figures 2 and 3, the interface is denoted as I / F.

[0026] The control unit 1 is a hardware processor such as a CPU (Central Processing Unit). For example, by using a multi-core and multi-threaded CPU, multiple information processing operations can be performed simultaneously. The control unit 1 may also be equipped with multiple hardware processors.

[0027] The communication interface unit 4, under the control of the control unit 1, transmits and receives information data with each participant terminal PT.

[0028] The program storage unit 2 is configured by combining, for example, a non-volatile memory that can be written to and read at any time, such as an HDD (Hard Disk Drive) or SSD (Solid State Drive), with a non-volatile memory such as ROM (Read Only Memory). In addition to middleware such as an OS (Operating System), the program storage unit 2 stores application programs necessary for inputting the above-mentioned information necessary for information presentation in the first embodiment and for transmitting registration requests thereof. Hereafter, the OS and each application program will be collectively referred to as a program.

[0029] The data storage unit 3 is, for example, a combination of a non-volatile memory that can be written to and read at any time, such as an HDD or SSD, and a volatile memory such as RAM (Random Access Memory), as a storage medium. The data storage unit 3 includes, in its storage area, a conference information database 31, a participant information database 32, and a sound field information database 33, which are the main storage units necessary for carrying out the first embodiment of this invention. In Figure 3, databases are abbreviated as DB.

[0030] The meeting information database 31 stores meeting information, which is information about each online communication involving multiple participants. This information is associated with a meeting ID to distinguish the online communication, and includes the date and time of the meeting and user information of the participants. User information includes login information such as user ID and password, name, etc. Meeting information can be set from the participant terminal PT used by the participant who is the organizer of the online communication.

[0031] The participant information database 32 stores participant information about each other participant (the dialogue partner) that each participant in each online communication has set from their own participant terminal PT. Participant information includes, for example, information indicating the difference in status and level of familiarity with the dialogue partner.

[0032] The sound field information database 33 stores sound field information for each participant, based on the sound field information for each of their conversation partners, using audio information acquired from each participant terminal PT of the online communication participants. Sound field information is information used to output audio information as a spatial sound image using stereophonic sound technology. Furthermore, the sound field information database 33 also stores video information of the display screen for each participant, adjusted for display position and display size based on the sound field information, using video information acquired from each participant terminal PT.

[0033] Furthermore, the control unit 1 includes, as processing function units necessary for implementing the first embodiment, a meeting information registration unit 11, a psychological distance stage setting unit 12, a psychological distance setting unit 13, a sound source position determination unit 14, an input information acquisition unit 15, a sound field position reflection unit 16, an audio output unit 17, and a video output unit 18. All of these processing function units are realized by causing the hardware processor of the control unit 1 to execute an application program stored in the program storage unit 2.

[0034] Furthermore, at least one or at least a portion of the processing functions within the processing function unit may be implemented using integrated circuits such as ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), FPGAs (field-programmable gate arrays), and GPUs (Graphics Processing Units), instead of being implemented by the application program and the hardware processor of the control unit 1.

[0035] The meeting information registration unit 11 communicates with the participant terminal PT of the participant who will be the organizer of the online communication via the network NW using the communication interface unit 4, receives meeting information from the participant terminal PT, assigns a unique meeting ID to the meeting information, and stores it in the meeting information database 31.

[0036] The psychological distance stage setting unit 12 communicates with each participant terminal PT of the online communication participants stored in the conference information database 31 via the network NW through the communication interface unit 4, and presents the content of the conference information to each participant terminal PT. Each participant terminal PT sets a psychological distance stage, which is the stage in terms of the difference in status and level of intimacy that they can take, according to the number of people they are talking to in the online communication as seen from their perspective. The psychological distance stage setting unit 12 receives the set psychological distance stages from the participant terminal PTs via the network NW through the communication interface unit 4 and transmits them to the psychological distance stage setting unit 13.

[0037] The psychological distance setting unit 13 communicates with each participant terminal PT of the online communication participant stored in the conference information database 31 via the network NW through the communication interface unit 4, and receives the setting of the psychological distance, indicated by psychological distance stages, for each other participant who is the person the participant is talking to, from the participant terminal PT. The psychological distance setting unit 13 stores the set psychological distance information in the participant information database 32.

[0038] In this embodiment, two elements are assumed to constitute psychological distance: differences in social standing and the degree of intimacy.

[0039] A difference in position refers to the objective role of each participant in any given dialogue and the resulting hierarchical relationship. For example, differences in position include superiors and subordinates within a company, professors and students in a university research lab, and customers and staff in customer support.

[0040] Intimacy refers to the degree of affection each participant feels for the other participants. For example, in a company setting, a close senior colleague (high intimacy) might be considered a close junior colleague (low intimacy).

[0041] The sound source location determination unit 14 determines the sound source location of other participants who are dialogue partners, relative to the target participant, based on the psychological distance information of each dialogue partner for each participant stored in the participant information database 32. The sound source location determination unit 14 stores the sound source location information for each determined dialogue partner in the participant information database 32. This method for determining sound source locations will be explained in detail in the operation description.

[0042] Figure 4 shows an example of the contents stored in the participant information database 32. The participant information database 32 holds psychological distance information for defining the sound source locations of other participants (a, b, c, ...) who are dialogue partners from the perspective of the target participant (n). Specifically, the participant information database 32 associates the target participant's user ID with the other participant ID, and the user IDs of each dialogue partner with the other participant ID, and stores psychological distance information for each dialogue partner set by the psychological distance setting unit 13. This information includes position information, which is a value representing the degree of difference in position, and intimacy information, which is a value representing the degree of intimacy. Furthermore, in addition to the psychological distance information, the participant information database 32 also stores the sound source coordinate values ​​of other participants that indicate the sound source location determined by the sound source location defining unit 14.

[0043] The meeting information registration unit 11, the psychological distance stage setting unit 12, the psychological distance setting unit 13, and the sound source location setting unit 14 basically operate at any point before the online communication takes place. However, during the online communication, the level of intimacy may change depending on the content of the conversation. Therefore, each processing function unit except for the meeting information registration unit 11, namely the psychological distance stage setting unit 12, the psychological distance setting unit 13, and the sound source location setting unit 14, may also operate during the online communication.

[0044] The participant information database 32 can also store a meeting ID to distinguish online communications, taking into account the possibility that the level of intimacy may change based on the content of the dialogue, so that the value of intimacy can be changed for each online communication.

[0045] The input information acquisition unit 15, the sound field position reflection unit 16, the audio output unit 17, and the video output unit 18 operate during the online communication session.

[0046] The input information acquisition unit 15 communicates with participant terminals PT of participants participating in online communication stored in the conference information database 31 via the network NW through the communication interface unit 4, and acquires audio and video information from each of these participant terminals PT. The input information acquisition unit 15 transmits the acquired audio and video information to the sound field position reflection unit 16.

[0047] The sound field position reflection unit 16, for each participant participating in online communication stored in the conference information database 31, generates sound field information for each dialogue partner based on the sound source coordinate values ​​of each dialogue partner stored in the participant information database 32, with the participant as the target. The sound field position reflection unit 16 then applies the voice information of each dialogue partner to the generated sound field information for each target. In other words, the sound field position reflection unit 16 generates sound field information with the voice information of each dialogue partner localized as a sound image. This sound field information with localized sound image is voice information for reproducing the voice of each dialogue partner in three-dimensional sound according to the sound field generated based on the psychological distance information of each target. The sound field position reflection unit 16 stores this generated sound field information with localized sound image for each target in the sound field information database 33.

[0048] Furthermore, the sound field position reflection unit 16 generates display video information, which is information about a display screen with adjusted display position and display size of the video information of each conversation partner, based on the generated sound field information for each target person, and stores it in the sound field information database 33.

[0049] The audio output unit 17 transmits, for each participant participating in the online communication stored in the conference information database 31, the sound field information corresponding to that participant, which is stored in the sound field information database 33, to the participant terminal PT of that participant via the network NW using the communication interface unit 4.

[0050] The video output unit 18 transmits, for each participant participating in the online communication stored in the conference information database 31, the display video information corresponding to that participant, stored in the sound field information database 33, to the participant terminal PT of that participant via the network NW using the communication interface unit 4.

[0051] (Example of operation) Next, we will explain an example of the operation of the communication server CS configured as described above. Basic operations such as login from participant terminals PT will not be explained. Furthermore, the operation of registering online communication meeting information to the meeting information database 31 by the meeting information registration unit 11 is a general operation, so a detailed explanation will be omitted here.

[0052] (1) Preparation process Then, at any point before the online communication session begins, in response to a preparation request from a participant terminal PT of a participant who intends to participate in the online communication session, the control unit 1 of the communication server CS executes a program stored in the program storage unit 2 to perform the preparation process shown in this flowchart. Similarly, at any point during the online communication session, if a preparation request is received from a participant terminal PT of a participant who is participating in the online communication session, the control unit 1 can also perform the preparation process shown below.

[0053] Figure 5 is a flowchart showing an example of the processing procedure and content of the preparation process executed by the control unit 1 of the communication server CS. For example, when the control unit 1 receives a preparation request from a participant terminal PT via the network NW through the communication interface unit 4, it starts this preparation process. The preparation process is basically a process between the requesting participant terminal PT and the control unit 1, and nothing is performed between it and other participant terminals PT.

[0054] When the preparation process begins, the control unit 1 operates as a psychological distance stage setting unit 12 and receives a specification of the online communication to be set from the requesting participant terminal PT (step S101). Specifically, the control unit 1 searches the online communication registered in the meeting information database 31 that is not currently completed and in which the participant's user ID is registered as a participant, communicates with the participant terminal PT via the network NW using the communication interface unit 4, presents the search results to the participant, and identifies the online communication to be set. Alternatively, the preparation request sent from the participant terminal PT may include information specifying the online communication to be set.

[0055] Next, the control unit 1 operates as a psychological distance stage setting unit 12 and performs the process of setting the psychological distance stage (step S102). Specifically, the control unit 1 obtains the psychological distance stage, which is the possible stage in terms of the difference in status and level of intimacy, from the participant terminal PT that made the request, based on the number of people the participant is talking to from their perspective, via the network NW using the communication interface unit 4.

[0056] Then, the control unit 1 operates as a psychological distance setting unit 13 and performs the process of setting the psychological distance for each of the other participants who are the conversation partners in the online communication to be set, as registered in the meeting information database 31 (step S103). Specifically, according to the psychological distance levels set in step S102, the control unit 1 presents the participant terminal PT that made the request with a selection of possible psychological distances for each conversation partner via the network NW through the communication interface unit 4 and accepts the selection of the psychological distance. Then, the control unit 1 stores the selected psychological distance for each conversation partner in the participant information database 32.

[0057] Figure 6 is a schematic diagram illustrating an example of the difference in status between different conversation partners. If, for example, three levels of status differences are set, then the difference in status can be assigned to each conversation partner, with the participant themselves as the reference point, and the three levels being superior, equal, and inferior. For example, if the online communication is a company meeting, then superiors and senior colleagues would be superior, colleagues would be equal, and subordinates and junior colleagues would be inferior. Note that while Figure 6 uses three levels of status differences, this can be increased to four or more levels if there are many conversation partners or a wide variety of statuses.

[0058] The control unit 1 stores the following values ​​in the participant information database 32 as position information, which represents the level of difference in position: "0" if the same rank is selected, "1" if a higher rank is selected, and "-1" if a lower rank is selected. In the example in Figure 6, the user ID of the superior, conversation partner Ca, is "a", the user ID of the colleague, conversation partner Cb, is "b", the user ID of the subordinate, conversation partner Cc, is "c", and the user ID of the participant, the participant terminal PT that made the request, is "n". In this case, as shown in Figure 4, the participant information database 32 will store "1" in the position difference field for the record of participant ID "n" and other participant ID "a", "0" in the position difference field for the record of participant ID "n" and other participant ID "b", and "-1" in the position difference field for the record of participant ID "n" and other participant ID "c".

[0059] Figure 7 is a schematic diagram illustrating an example of intimacy levels with each conversation partner. If, for example, seven levels of intimacy are set, then each conversation partner can be assigned one of these seven levels, ranging from "-3 (low)" to "3 (high)," with "0 (intermediate)" as the baseline. For example, if the online communication is a company meeting, one could select "3" for a close subordinate, "0" for a distant colleague, and "-2" for a boss with whom one does not get along. Note that the number of intimacy levels can also be increased or decreased depending on the number of conversation partners, etc.

[0060] The control unit 1 stores the value selected as the intimacy level as intimacy information in the participant information database 32. Here, as shown in the example in Figure 7, if "-2" is selected as the intimacy level of conversation partner Ca, "0" as the intimacy level of conversation partner Cb, and "3" as the intimacy level of conversation partner Cc, then the participant information database 32 will store, as shown in Figure 4, the intimacy information of "-2" in the intimacy field of the record for participant ID "n" and other participant ID "a", the intimacy information of "0" in the intimacy field of the record for participant ID "n" and other participant ID "b", and the intimacy information of "3" in the intimacy field of the record for participant ID "n" and other participant ID "c".

[0061] Returning to the explanation of Figure 5, the control unit 1 then operates as a sound source location determination unit 14 and performs the process of determining the sound source location for each participant (step S104). That is, for each online communication identified by the meeting ID stored in the participant information database 32, the control unit 1 determines the sound source location for each other participant who is the conversation partner of the person identified by the participant ID. Specifically, the control unit 1 determines the sound source location according to the position information and intimacy information stored in the participant information database 32, and stores the coordinates of the determined sound source location in the corresponding sound source coordinates item of the other participant in the participant information database 32.

[0062] Figure 8 shows a sound source coordinate system that defines the sound source position according to differences in status and intimacy. In the sound source coordinate system, the difference in status indicated by status information is assigned to the position (Y coordinate) in the vertical direction (Y axis direction) of the sound source, and the difference in status is represented by its vertical position. In addition, the intimacy indicated by intimacy information is assigned to the position (Z coordinate) in the depth direction (Z axis direction) of the sound source, and the intimacy is represented by its proximity position. There may be cases where the YZ coordinates overlap when there are dialogue partners with the same status and the same intimacy level, in which case the position (X coordinate) in the horizontal direction (X axis direction) of the sound source is made different.

[0063] Differences in status can be reproduced as the vertical position of the sound field output on the participant terminal PT display screen. Therefore, the control unit 1 determines the vertical Y coordinate of the sound source for each status, with the aim of leveling out the differences in status. Specifically, with the aim of reducing the sense of intimidation caused by differences in status, the height of the status and the height of the Y coordinate are made inversely proportional. That is, the control unit 1 determines the vertical position of the sound source so that the conversation partner indicating a higher status is positioned lower on the participant's display screen. As a result, the statements of participants with higher status are played back from a lower position, which can mitigate the sense of intimidation.

[0064] Figure 9 is a schematic diagram showing the sound source positions of each conversation partner according to their differences in status. In the example in Figure 6, as mentioned above, the status information is set as "1" for conversation partner Ca, who is the superior, "0" for conversation partner Cb, who is the colleague, and "-1" for conversation partner Cc, who is the subordinate. Therefore, as shown in Figure 9, the control unit 1 uses the Y coordinate "0" of the target person "n" as a reference, sets the Y coordinate of conversation partner Cb, who is the colleague, to "0", and the Y coordinate of conversation partner Ca, who is the superior, to "y na ", the Y coordinate of the conversation partner Cc, who is a subordinate, is "y nc (However, y nc >0>y na This is determined to be the case. This allows the participant terminal PT display screen of the target person "n" to hear the voice of the superior (who is in a higher position) from below, and the voice of the subordinate (who is in a lower position) from above.

[0065] Intimacy can be reproduced as the distance of the sound field output on the participant terminal PT's display screen. Therefore, the control unit 1 determines the distance (L) of the sound source according to the degree of intimacy, with the aim of reflecting intimacy. Specifically, based on the knowledge that "the relationship with the conversation partner affects the distance during the conversation," such as in the F formation, the intimacy level and the distance are made inversely proportional. Distance is basically achieved by changing the Z coordinate, which is a value in the depth direction (Z axis direction). In other words, the control unit 1 determines the depth direction position so that the closer the conversation partner is on the participant's display screen, the higher the intimacy information indicates.

[0066] Figure 10 is a schematic diagram showing the sound source positions of each conversation partner according to the intimacy level. In the example of Figure 7, as described above, intimacy information of "2" is set for the conversation partner Ca who is the superior, "4" for the conversation partner Cb who is a colleague, and "7" for the conversation partner Cc who is a subordinate. Therefore, as shown in Figure 10, the control unit 1 sets the distance from the target person "n" of the conversation partner Ca who is the superior to "l na ", the distance of the conversation partner Cb who is a colleague to "l nb ", and the distance of the conversation partner Cc who is a subordinate to "l nc ". The control unit 1 obtains each distance l, for example, as follows. As a result, on the display screen of the participant terminal PT of the target person "n", the voice of the superior with a low intimacy level can be heard from a distance, and the voice of the subordinate with a high intimacy level can be heard from nearby.

[0067]

Number

[0068] As described above, when there are conversation partners in the same position and intimacy level and the YZ coordinates overlap, the control unit 1 changes the position (X coordinate) in the horizontal direction of the sound source. Specifically, the control unit 1 arranges the X coordinates of the corresponding conversation partners evenly on the left and right.

[0069] Figure 11 is a schematic diagram showing the sound source positions of a plurality of conversation partners with the same difference in position and intimacy level. As shown in Figure 11, when there are conversation partners Cc, Cc', and Cc" with the same position and intimacy level, their coordinates should be the same (x nc , y nc , z nc ), but the control unit 1 changes the X coordinate to x nc , x nc ’, x nc ”. Note that when the value in the horizontal direction of the sound source is changed in this way, the distance l of the changed conversation partner changes. Therefore, when the value in the horizontal direction of the sound source is changed, the control unit 1 corrects the value in the depth direction so that the distance does not change. That is, the control unit 1 changes the Z coordinate to z nc , z nc ’, z nc"Let's assume that."

[0070] Furthermore, if the participant's terminal screen only displays audio and does not show the video of the other participant (the person they are talking to), then the X coordinates may remain the same.

[0071] Here, we will explain an example of the specific coordinate definition procedure performed by the control unit 1. In this example, the origin is set to the coordinates (0,0,0) of the subject.

[0072] i. The control unit 1 determines the y-coordinate of each dialogue partner's sound source by multiplying the difference in position value stored as position information in the participant information database 32 by an arbitrary variable. For example, participants are divided into three stages: higher, same, and lower. The vertical width of the participant terminal PT's display screen is set to "40", the bottom edge of the display screen to "-20", and the coordinates are changed by "10" for each stage, with the y-coordinate of the higher dialogue partner being "-10", the y-coordinate of the same dialogue partner being "0", and the y-coordinate of the lower dialogue partner being "10".

[0073] ii. The control unit 1 determines the distance l between the subject and the sound source of each conversation partner by multiplying the intimacy value stored as intimacy information in the participant information database 32 by an arbitrary variable. For example, it assigns the distance to seven levels from "-3" to "3", and sets the range of possible distances from "10" to "70", changing the distance by "10" for each level, setting the distance l to "10" when the intimacy is highest ("3") and "70" when it is lowest ("-3").

[0074] iii. The control unit 1 calculates the z coordinate such that the distance l is satisfied. For example, when y=30 and l=50, z 2 =50 2 -30 2 Therefore, z = 40 (when x = 0). In this case, instead of setting x = 0, you can also calculate the distance by varying the value within any ± range.

[0075] iv. If there are multiple sound sources with the same position and level of intimacy, the control unit 1 distributes the x-coordinates of the corresponding sound sources. For example, if there are three people and the width of the display screen is "80", the left edge of the display screen is set to "-40", and the x-coordinates of each are set to "-30", "0", and "30".

[0076] v. The control unit 1 modifies the z coordinate so that the distance l is satisfied. That is, it is calculated in the same way as in the case of x≠0 in iii. above.

[0077] Returning to the explanation of Figure 5, the control unit 1 determines whether or not to terminate this preparation process (step S105). For example, the control unit 1 terminates this preparation process when it receives a termination instruction from the participant terminal PT via the network NW through the communication interface unit 4. If it determines that it is not yet time to terminate, the control unit 1 proceeds to the process in step S101.

[0078] (2) Dialogue Processing Figure 12 is a flowchart showing an example of the processing procedure and content of the dialogue processing performed by the control unit 1. For each online communication, the control unit 1 executes the dialogue processing shown in this flowchart, targeting each participant, by executing the program stored in the program storage unit 2. The control unit 1 can perform the processing shown in this flowchart in parallel for multiple online communications held simultaneously.

[0079] When the communication interface unit 4 receives an online communication start command from the participant terminal PT of the target person via the network NW, the control unit 1 starts the dialogue processing for that online communication. Then, the control unit 1 operates as an input information acquisition unit 15 and determines whether or not it has acquired input information, that is, whether or not it has received audio and video information transmitted via the network NW from the participant terminal PT of the other participant who is the target person's dialogue partner (step S111). At this time, the control unit 1 can distinguish between the target person's participant terminal PT and the participant terminal PT of the dialogue partner of that participant, based on the online communication conference information registered in the conference information database 31. The control unit 1 repeats the process in step S111 until it has acquired input information.

[0080] Once input information is acquired, the control unit 1 operates as a sound field position reflection unit 16 and generates the sound field experienced by the subject, taking into account the positional relationship between the subject and each conversation partner (step S112). Specifically, the control unit 1 identifies the subject and conversation partners based on the online communication meeting information registered in the meeting information database 31, and generates sound field information for each conversation partner that the subject will experience, based on the sound source coordinates that take into account the positional relationship between the subject and each conversation partner stored in the participant information database 32. Then, the control unit 1 applies the acquired audio information to the sound field information of the conversation partner that is the source of the audio information acquired in step S111. That is, the control unit 1 generates sound field information with the acquired audio information of the conversation partner localized as a sound image. The control unit 1 stores this generated sound field information in the sound field information database 33. Furthermore, based on the generated sound field information, the control unit 1 adjusts the display position and size of the video information of the person it is talking to, generates display video information which is to be displayed on the display screen of the participant terminal PT, and stores it in the sound field information database 33.

[0081] Then, the control unit 1 operates as an audio output unit 17 to output audio (step S113) and also operates as a video output unit 18 to output video (step S114).

[0082] Specifically, the control unit 1 identifies the participant terminal PT of the target person based on the online communication meeting information registered in the meeting information database 31, and transmits the sound field information of each conversation partner, which corresponds to the target person and is stored in the sound field information database 33, to the identified participant terminal PT via the network NW using the communication interface unit 4. In addition, the control unit 1 transmits the display video information corresponding to the target person, which is stored in the sound field information database 33, to the identified participant terminal PT via the network NW using the communication interface unit 4. As a result, the participant terminal PT of the target person can simultaneously reproduce the voice of each conversation partner in 3D sound according to the sound field information of each conversation partner, and display and play the video of each conversation partner on the display screen.

[0083] Subsequently, the control unit 1 determines whether or not to terminate this dialogue process (step S115). For example, the control unit 1 terminates this dialogue process when it receives a termination instruction from the participant terminal PT of the target person via the network NW through the communication interface unit 4. If it determines that it is not yet time to terminate, the control unit 1 proceeds to the process in step S111.

[0084] Figure 13 is a schematic diagram showing an example of the display screen SC of the participant terminal PT of the target participant. The control unit 1 operates as a sound field position reflection unit 16 and generates display video information by drawing the video information CV of each conversation partner on the display screen SC based on the sound source coordinates of each defined conversation partner. The display video information includes image information of a depth design indicating depth as the background of the display screen SC, and the video information CV of the conversation partner is placed on this depth design. The depth design can be expressed, for example, by perspective lines PL according to perspective projection or by shades of color. It should be noted that it is not mandatory to provide a depth design as the background of the display screen SC, and it is of course not necessary to place a special image, such as a monochrome display. In addition, the size of the video information CV of the conversation partner is changed in proportion to the distance from the sound source indicated by the sound field information, and the closer the distance, the larger it is drawn. Figure 13 shows the closest distance l nc This is an example where the size of the video information CV of the conversation partner is drawn as large as possible. The coordinates of the sound source position can be the center of the video information CV, or it can be the area around the mouth of the conversation partner in the placed video information CV by extracting the face area from the video information CV using OpenCV or similar.

[0085] (Effects / Actions) As described above, in the first embodiment, the communication server CS functions as an information presentation device that presents multiple audio pieces of information acquired from one or more first participant terminals PT used by one or more participants who are conversation partners among the multiple participant terminals PT participating in online communication via the network NW to a second participant terminal PT used by a target participant among the multiple participant terminals PT via the network NW. The communication server CS includes a sound source position definition unit 14 that defines the sound source position of each conversation partner based on psychological distance information that represents the psychological distance to each conversation partner from the perspective of the target participant using the second participant terminal, which is set for each of the one or more conversation partners using the one or more first participant terminals, and a sound field position reflection unit 16, a sound field information database 33, and a sound output unit 17 that serve as an audio presentation unit that generates sound field information with sound image localization of the audio information from one or more first participant terminals based on the sound source position of each of the one or more conversation partners and transmits it to the second participant terminal. Therefore, according to the first embodiment, psychological distance information for each conversation partner is acquired, the sound source position is defined according to that psychological distance information, and the voice of each conversation partner is output according to the defined sound source position. Thus, it is possible to provide an information presentation technology that can give the subject an appropriate sense of distance based on the relationship between the subject and each conversation partner.

[0086] Furthermore, in the first embodiment, the psychological distance information includes position information indicating the position of the conversation partner from the perspective of the subject, and the sound source position defining unit 14 determines the vertical (Y-axis) position of the sound source such that the conversation partner with a higher position indicated by the position information is positioned lower on the display screen SC of the second participant terminal PT. Therefore, according to the first embodiment, it is possible to provide information presentation technology that can give the subject an appropriate sense of distance based on the position of the person they are talking to. In other words, the lower the position of the person, the more comfortable the conversation can be, as the audio is output in three-dimensional sound from the top of the display screen SC.

[0087] Furthermore, in the first embodiment, the psychological distance information includes intimacy information indicating the level of intimacy with the conversation partner from the perspective of the subject, and the sound source position defining unit 14 determines the sound source depth direction (Z-axis direction) position such that the closer the conversation partner is to the display screen SC of the second participant terminal PT, the higher the level of intimacy information indicates. Therefore, according to the first embodiment, it is possible to provide information presentation technology that can give the subject an appropriate sense of distance based on the degree of intimacy with the person they are talking to. In other words, the closer the person feels to the speaker, the closer their voice is output in 3D sound, thereby enabling a comfortable conversation. In particular, by focusing on two elements, differences in status and intimacy, and by averaging out the "differences in status" while correcting with "intimacy," it is possible to alleviate the feeling of pressure caused by the hierarchical relationship with the person they are talking to, and the tension of being surrounded by people they are not close to, thereby reducing discomfort during conversations for the subject.

[0088] Furthermore, in the first embodiment, the sound source position defining unit determines the horizontal (X-axis) position of the sound source such that dialogue partners with the same status and level of intimacy are in the same vertical sound source position on the display screen SC of the second participant terminal PT, but are in different left-right positions on the display screen SC. Therefore, according to the first embodiment, since conversation partners with the same status and level of intimacy can be presented side by side on the display screen SC, it is possible to provide an information presentation technology that can give the target person an appropriate sense of distance, even when there are many conversation partners.

[0089] Furthermore, in the first embodiment, for each of the one or more second participant terminals PT, the system further comprises a sound field position reflection unit 16, a sound field information database 33, and a video output unit 18, which are video presentation units that generate display video information to display video information from the first participant terminal PT at the vertical and horizontal positions of the sound source determined by the sound source position definition unit 14, with a size proportional to the depth position of the sound source determined by the sound source position definition unit 14, and transmit it to the second participant terminal PT, with the size being larger for depth positions that are closer. Therefore, according to the first embodiment, in addition to audio, by presenting video of each conversation partner based on the relationship between the subject and each conversation partner, it is possible to provide information presentation technology that can give the subject a more appropriate sense of distance.

[0090] [Second Embodiment] Next, a second embodiment will be described. Note that parts similar to those in the first embodiment will be given the same reference numerals as in the first embodiment, and their descriptions will be omitted.

[0091] (Example configuration) Figure 14 is a block diagram showing an example of the software configuration of a communication server CS as a second embodiment of the information presentation device of the present invention. In the second embodiment, the control unit 1 of the communication server CS includes, in addition to the same conference information registration unit 11, psychological distance stage setting unit 12, psychological distance setting unit 13, sound source position defining unit 14, input information acquisition unit 15, sound field position reflection unit 16, audio output unit 17, and video output unit 18 as in the first embodiment, an input information leveling unit 19 as a processing function unit necessary to implement the second embodiment.

[0092] The input information leveling unit 19 leveles the video and audio information, which are input information acquired from each participant terminal PT via the network NW by the input information acquisition unit 15, to generate leveled video information and leveled audio information, and supplies them to the sound field position reflection unit 16. This input information leveling method will be explained in detail in the operation description.

[0093] (Example of operation) Figure 15 is a flowchart showing an example of the processing procedure and processing content of the dialogue processing performed by the control unit 1 in the second embodiment. In the second embodiment, if the control unit 1 determines in step S111 that it has acquired input information from the participant terminal PT of another participant who is the dialogue partner of the subject, it operates as an input information leveling unit 19 and levels the acquired input information (step S116). Specifically, the control unit 1 corrects the acquired video information and audio information so that, for example, the size of the face in the video information of each dialogue partner, the volume of the voice in the audio information of each dialogue partner, etc., become equal. Then, the control unit 1 uses the leveled video information and leveled audio information obtained by these corrections as the information to be processed and executes the processing in step S112.

[0094] If there are variations in the video and / or audio information input to each participant terminal (PT) for each conversation partner, the target participant terminal (PT) will be unable to properly represent the sense of distance when outputting the video and audio of each conversation partner. For example, if the input volume of a conversation partner with low intimacy is relatively loud, even if the sound field is generated to move the sound source coordinates farther away based on intimacy information, the target participant terminal (PT) will still hear the voice of that conversation partner with low intimacy as loud. To prevent this, the way the video is seen (size and position of the face) and the way the audio is heard (volume) are standardized in advance.

[0095] To equalize the size and position of faces, the face area is extracted from the video information of each conversation partner using OpenCV or similar tools, and the faces of the other conversation partners are cropped and rendered, aligning them to the conversation partner with the largest face area in the video.

[0096] Figure 16A is a schematic diagram showing an example of the face area of ​​each conversation partner in the input video information of each conversation partner, and Figure 16B is a schematic diagram showing an example of the leveled video information of each conversation partner. In the example shown in Figure 16A, the input video information IVb of conversation partner Cb, a colleague, shows that the camera is far away and the size of the face area FA is small, and the input video information IVc of conversation partner Cc, a subordinate, shows that the position of the face area FA is shifted to the right. In such cases, as shown in Figure 16B, the control unit 1 uses the leveled video information LIa as is for conversation partner Ca, the superior, whose face is the largest and whose face area FA is in the center, without making any corrections. On the other hand, for the input video information IVb of conversation partner Cb, a colleague, the control unit 1 generates the leveled video information LIb by correcting it to match the size of the face in the input video information IVa of conversation partner Ca, the superior, whose face is the largest. Furthermore, with respect to the input video information IVc of the subordinate, Cc, the control unit 1 generates leveled video information LIc by performing a correction that trims the image so that the face is centered within a range where the face position can be corrected.

[0097] Furthermore, regarding volume leveling, the control unit 1 generates leveled speech information by performing a correction to match the volume of the person speaking with the quietest voice, similar to the size of a face. Alternatively, the control unit 1 generates leveled speech information by performing a correction to match the average volume of all people speaking with, by amplifying quiet voices and attenuating loud voices.

[0098] (Effects / Actions) As described above, the second embodiment includes an input information leveling unit 19 that levels the size and position of the face of the person the participant is talking to in the video information from one or more first participant terminals PT and supplies it to the video display unit, and / or levels the volume of the audio information from one or more first participant terminals PT and supplies it to the audio display unit. Therefore, according to the second embodiment, even if there is variation in the input information from each conversation partner, it is possible to provide an information presentation technology that can give the target person an appropriate sense of distance.

[0099] [Third Embodiment] The communication server CS, as a first or second embodiment of the information presentation device, may be configured to automatically acquire psychological distance in cooperation with other systems. Specifically, the psychological distance setting unit 13 of the control unit 1 of the communication server CS automatically inputs the "difference in status" and "level of intimacy" with each conversation partner in cooperation with other systems, without receiving settings from the participant terminal PT of the target participant.

[0100] For example, the psychological distance setting unit 13 can obtain job title information of each conversation partner from a system that manages employee information and set the difference in status. Alternatively, the psychological distance setting unit 13 can estimate and set the level of intimacy from the content of the conversations between the target person and each conversation partner on the chat tool.

[0101] One example of how this can be implemented is that the psychological distance setting unit 13 uses a score to determine the degree of intimacy from the conversation history, as disclosed in, for example, Reference 1 below.

[0102] (Reference 1) Yuto Hoshikawa, Kei Wakabayashi, and Tetsuji Sato, "Evaluation of a method for estimating intimacy using conversation content on Twitter," Proceedings of the 8th Forum on Data Engineering and Information Management, March 2016.

[0103] Thus, according to the third embodiment, by configuring the communication server CS to automatically acquire psychological distance in cooperation with other systems, it becomes possible to omit the task of setting the psychological distance of the target person.

[0104] [Fourth Embodiment] The communication server CS, as a first or second embodiment of the information presentation device, may dynamically change the sound source position during the conversation. That is, the sound source position defining unit 14 of the control unit 1 of the communication server CS dynamically changes the sound source position defined in the preparation process during the conversation.

[0105] If the level of intimacy with a particular conversation partner changes during the conversation, the coordinates of the sound source can be changed by updating that value. For example, if the conversation with a boss with whom the user previously had a strained relationship has improved, the level of intimacy with that boss may increase, and the sound source can be moved closer accordingly. As described in the first embodiment, the psychological distance stage setting unit 12, the psychological distance setting unit 13, and the sound source position setting unit 14 operate even during the conversation, enabling the user to manually set the sound source position.

[0106] In this fourth embodiment, in addition to this manual update, the sound source position defining unit 14 has a function to estimate the emotions of both the subject and the person they are talking to, and temporarily changes the coordinates of the sound source according to the degree of intimacy and emotions such as joy, anger, sadness, and happiness.

[0107] For example, if the subject is relaxed, you might bring the audio sources of all the people they're talking to closer together. Or, if a close junior colleague is expressing anger and shouting, you might temporarily move their audio source further away.

[0108] As an example of how this can be implemented, the sound source positioning unit 14 can perform emotion estimation using the audio alone or by utilizing facial expressions in the video, as disclosed in, for example, Reference 2 below.

[0109] (Reference 2) Kenji Nishida, Toru Yamada, Katsutoshi Itoyama, Kazuhiro Nakadai, "Investigation of Emotion Estimation Methods Using Facial Expressions and Speech," Abstracts of the 57th Annual Meeting of the Japanese Society for Artificial Intelligence AI Challenge Workshop, pp. 52-57, November 2020.

[0110] Thus, according to the fourth embodiment, by configuring the communication server CS to dynamically change the sound source position during a conversation, it is possible to provide an information presentation technology that can give an appropriate sense of distance according to the psychological distance of the target person at that time.

[0111] [Fifth Embodiment] In a second embodiment of the information presentation device, the communication server CS may be configured to personalize the leveling items. Specifically, the input information leveling unit 19 of the control unit 1 of the communication server CS modifies or adds items to the target items when performing leveling, according to the type of dialogue and the preferences of the target person.

[0112] In the second embodiment, since emphasis is placed on expressing distance, the size of the face and the volume of the voice were listed as basic items. However, in the fifth embodiment, the input information leveling unit 19 adds elements such as voice quality and speaking style to the leveling target and performs leveling.

[0113] For example, if there is a mix of high-pitched and low-pitched voices in the conversation, making it difficult to hear, the input information equalization unit 19 will bring the pitches of both voices closer together.

[0114] As an example of implementation, the input information leveling unit 19 extracts speech features and replaces them with similar synthesized speech that is closer to the average, as disclosed in, for example, Reference 3 below.

[0115] (Reference 3) D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, S. Khudanpur, "X-VECTORS: ROBUST DNN EMBEDDINGS FOR SPEAKER RECOGNITION", 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), PP. 5329-5333, April 2018, DOI: 10.1109 / ICASSP.2018.8461375.

[0116] [Other embodiments] In each of the embodiments described above, the sound source position determination unit 14 determines the sound source coordinates of other participants, indicating the sound source position of each dialogue partner, during the preparation process and stores them in the participant information database 32. However, it is not always necessary to determine the sound source coordinates of other participants in advance and store them in the participant information database 32. That is, the sound source position determination unit 14 may calculate the sound source coordinates of other participants each time during the dialogue based on the psychological distance information, i.e., the difference in position and the value of intimacy, stored in the participant information database 32, and transmit this information to the sound field position reflection unit 16.

[0117] Furthermore, this invention is applicable not only to online communication but also to some real-world (offline) use. For example, it can be applied to a scenario where each participant wears noise-canceling earphones and an intercom, and the location of the sound source is set to coordinates different from the actual location according to the level of intimacy with each conversation partner, and the audio is played back from there. Moreover, in such a scenario, a visual application is also conceivable where each participant wears MR (Mixed Reality) glasses with a camera facing forward, and after cutting out the actual video of each conversation partner, the video is repositioned at the coordinates of the sound source defined by this information presentation system.

[0118] Furthermore, although each embodiment shows the information presentation device being composed of a single communication server CS, it may also be composed of multiple servers. For example, the server that performs preparation processing may be separated from the server that performs dialogue processing, or the servers that perform dialogue processing may be divided according to the number of online communications held simultaneously and the number of participants.

[0119] Furthermore, it goes without saying that the flow of each process described with reference to the flowchart is not limited to the procedures described.

[0120] The program may be transferred while stored on an electronic device, or it may be transferred without being stored on an electronic device. In the latter case, the program may be transferred via a network, or it may be transferred while recorded on a recording medium. The recording medium is a non-temporary, tangible medium. The recording medium is a computer-readable medium. The recording medium can be any medium that is capable of storing a program and is readable by a computer, such as a CD-ROM or memory card, and its form is not restricted.

[0121] Although embodiments of this invention have been described in detail above, the above description is merely illustrative in all respects. It goes without saying that various improvements and modifications can be made without departing from the scope of this invention. In other words, when implementing this invention, specific configurations may be adopted as appropriate depending on the embodiment.

[0122] In short, this invention is not limited to the embodiments described above, and in the implementation stage, the components can be modified and implemented without departing from the gist of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the embodiments. For example, some components may be deleted from all the components shown in the embodiments. Moreover, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0123] 1…Control Unit 2…Program memory 3…Data storage unit 4…Communication Interface Section 5... Bus 11…Meeting Information Registration Department 12…Psychological Distance Stage Setting Section 13…Psychological distance setting section 14...Sound source position specification section 15... Input information acquisition unit 16...Sound field position reflection section 17…Audio output section 18…Video output section 19...Input Information Leveling Unit 31…Conference Information Database 32…Participant Information Database 33…Sound field information database Ca, Cb, Cc, Cc', Cc ” ...conversation partner CS...Communication Server CV…Video Information FA... Face area IVa, IVb, IVc… Input video information LIa, LIb, LIc… Normalized video information NW...Network PL...Perth Line PT...Participant terminal SC…display screen

Claims

1. An information presentation device that presents multiple audio pieces of information acquired from one or more first participant terminals among multiple participant terminals participating in online communication via a network to a second participant terminal among the multiple participant terminals via the network, A sound source position defining unit defines the sound source position for each of the conversation partners, based on psychological distance information that represents the psychological distance from the target person using the second participant terminal to each of the one or more conversation partners using the one or more first participant terminals, which is set for each of the one or more conversation partners using the one or more first participant terminals, A voice presentation unit generates sound field information that localizes the voice information from the one or more first participant terminals based on the sound source position of each of the one or more dialogue partners, and transmits it to the second participant terminal. An information display device equipped with the following features.

2. The aforementioned psychological distance information includes positional information that indicates the position of the person the subject is talking to, The sound source positioning unit determines the vertical position of the sound source such that the higher the position of the conversation partner indicated by the position information, the lower the position on the display screen of the second participant terminal. The information display device according to claim 1.

3. The aforementioned psychological distance information includes intimacy information indicating the level of intimacy with the person the subject is talking to, The sound source positioning unit determines the depth position of the sound source on the display screen of the second participant terminal such that the closer the interaction partner is to the one whose intimacy information indicates a higher level of intimacy. The information display device according to claim 1 or 2.

4. The aforementioned psychological distance information includes intimacy information indicating the level of intimacy with the person the subject is talking to, The sound source positioning unit determines the depth position of the sound source so that the closer the interaction partner is to the second participant terminal's display screen, the higher the level of intimacy indicated by the intimacy information. The sound source position defining unit determines the horizontal position of the sound source such that, for dialogue partners with the same status and level of intimacy, the vertical position of the sound source is the same on the display screen of the second participant terminal, but the horizontal position is different on the display screen. The information display device according to claim 2.

5. The system further comprises a video display unit that generates display video information for each of the one or more second participant terminals, which displays video information from the first participant terminal at the vertical and horizontal positions of the sound source determined by the sound source position defining unit, with a size proportional to the depth position of the sound source determined by the sound source position defining unit, and transmits this information to the second participant terminal. The aforementioned sizes are larger the closer they are to the depth. The information display device according to claim 4.

6. The system further comprises a leveling unit that levels the volume of the audio information from the one or more first participant terminals and supplies it to the audio presentation unit. The information display device according to claim 1.

7. An information presentation method performed by an information presentation device comprising a processor and memory, which presents multiple audio pieces of information acquired from one or more first participant terminals among a plurality of participant terminals participating in online communication via a network to a second participant terminal among the plurality of participant terminals via the network, The processor defines the sound source location for each of the one or more conversation partners using the one or more first participant terminals, based on psychological distance information representing the psychological distance from the target person using the second participant terminal to each of the conversation partners, and stores the defined sound source locations for each of the conversation partners in the memory. The processor generates sound field information that localizes the audio information from the one or more first participant terminals based on the sound source position of each of the one or more dialogue partners, and transmits it to the second participant terminal. A method of presenting information that includes this.

8. An information presentation program that causes a processor in the information presentation device to perform the processing that each part of the information presentation device described in claim 1 performs.

Citation Information

Patent Citations

  • Large room type virtual office system

    JP1997288645A

  • Voice output control device, voice output control method, program, and recording medium

    JP2014011509A

  • Remote conference system, server, photography device, audio output method, and program

    JP2022054192A

  • Generating content for a virtual reality system

    US20150058102A1