Information processing device, information processing method, and information processing system

The information processing device improves remote communication by identifying speaker participation and adjusting voice clarity, addressing issues of irrelevant voice emphasis and sound muffling, resulting in clearer and more natural conversations.

WO2025182639A1PCT designated stage Publication Date: 2025-09-04SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/005168
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2025-02-17
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Conventional remote communication systems struggle with emphasizing irrelevant voices and muffling surrounding sounds, leading to difficulty in focusing on conversations and creating an unnatural atmosphere.

Method used

An information processing device that includes a human sensing unit to determine speaker participation, a clarity adjustment unit to enhance or suppress voices based on participation, and a transmission unit to process and transmit audio, ensuring clear and natural communication.

Benefits of technology

Enhances clarity of relevant voices while reducing irrelevant sounds, maintaining a realistic and engaging conversation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025005168_04092025_PF_FP_ABST
    Figure JP2025005168_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an information processing device, an information processing method, and an information processing system which make it possible to promote a smoother dialogue. A human sensing unit determines, on the basis of images captured by a plurality of cameras of speakers present in a predetermined space, whether each speaker participates in a conversation. A clarity adjustment unit adjusts the clarity of speech voice of the speaker determined to participate in the conversation by the human sensing unit, the speech voice being included in voice in the predetermined space, which voice is collected by a microphone. The transmission unit transmits the voice in the predetermined space processed by the clarity adjustment unit.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and information processing system

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing system.

[0002] A remote communication system is known that enables remote communication between multiple people in remote areas. In such a remote communication system, a person in one remote area communicates with a person in another remote area using a microphone, speaker, camera, etc. of the remote communication system.

[0003] JP 2010-191544 A

[0004] The remote communication system uses a microphone placed in a remote area to pick up the voice of a speaker, and reproduces the picked up voice from a speaker placed in another remote area.

[0005] However, simply extracting and reproducing human voices has the problem that even voices unrelated to the conversation are emphasized, making it difficult to concentrate on the conversation with the other person.Furthermore, if surrounding sounds are completely muted in order to emphasize the conversation, it may be difficult to grasp the situation in the other person's space, and the atmosphere may become unnatural when there is no conversation.

[0006] The present disclosure has been made in consideration of the above-described circumstances, and provides an information processing device, an information processing method, and an information processing system that can promote smooth dialogue.

[0007] The information processing device disclosed herein includes a human sensing unit that determines whether each speaker is participating in a conversation based on images captured by a camera of multiple speakers present in a specified space, a clarity adjustment unit that adjusts the clarity of the speech of the speakers participating in the conversation determined by the human sensing unit to be included in the audio of the specified space picked up by a microphone, and a transmission unit that transmits the audio of the specified space processed by the clarity adjustment unit.

[0008] 1 is a block diagram of a remote communication device. FIG. 1 is a bird's-eye view of an environment in which a remote communication device is placed. FIG. 2 is a front view of an environment in which a remote communication device is placed. FIG. 3 is a diagram for explaining the processing of a person sensing unit. FIG. 4 is a diagram illustrating coordinate conversion from screen coordinates to normalized coordinates. FIG. 5 is a diagram illustrating individual sound separation processing. FIG. 6 is a diagram illustrating multiple opposed speaker balancing processing. FIG. 7 is a diagram illustrating processing using a normal distribution function for outfield speech sounds. FIG. 8 is a diagram illustrating an outline of data flow between a remote communication device on the user's side according to a first embodiment. FIG. 9 is a flowchart of transmission-side processing. FIG. 10 is a flowchart of person position determination processing. FIG. 11 is a flowchart of infield / outfield determination processing. FIG. 12 is a flowchart of individual sound separation processing. FIG. 13 is a flowchart of clarity adjustment processing for individual sounds. FIG. 14 is a flowchart of multiple opposed speaker balancing processing. FIG. 15 is a flowchart of background sound adjustment processing. FIG. 16 is a flowchart of infield speech sound adjustment processing. FIG. 17 is a flowchart of outfield speech sound adjustment processing. FIG. 18 is a flowchart of localization center speaker unit derivation processing. A block diagram of a remote communication device according to a second embodiment. FIG. 19 is a diagram illustrating an outline of data flow between a remote communication device on the user's side according to the second embodiment. A block diagram of a remote communication device according to a third embodiment. FIG. 10 is a diagram showing an overview of data flow between a remote communication device on the user's side according to a third embodiment. FIG. 11 is a block diagram of a remote communication device according to a fifth embodiment. FIG. 12 is a diagram for explaining audio signal processing by a remote communication device according to the fifth embodiment. FIG. 13 is a diagram for explaining audio reproduction by a remote communication device according to a sixth embodiment. FIG. 14 is a diagram for explaining audio reproduction by a remote communication device according to a seventh embodiment. FIG. 15 is a hardware configuration diagram showing an example of a computer that realizes the arithmetic unit of a remote communication device that is an information processing device according to the first to seventh embodiments.

[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted. The description will be given in the following order.

[0010] 1. Remote communication device according to the first embodiment 1.1. Definition of terms 1.1.1. Remote and local 1.1.2. Individual sound 1.1.3. Infield and outfield 1.1.4. Multiple facing speakers 1.2. Transmission unit 1.2.1. Human sensing unit 1.2.2. Audio separation unit 1.2.3. Acoustic signal processing unit 1.2.4. Output integration unit 1.2.5. Transmission unit 1.3. Receiving unit 1.3.1. Receiving unit 1.3.2. Audio output control unit 2. Data flow between remote communication devices according to the first embodiment 3. Remote communication processing 3.1. Transmission-side processing 3.1.1. Human position determination processing 3.1.2. Infield / outfield determination processing 3.1.3. Individual sound separation processing 3.1.4. Clarity adjustment processing 3.2. 1. Multiple Facing Speaker Balancing Processing (Receiving Side Processing) 3.2.1. Background Sound Adjustment 3.2.2. Infield Speech Audio Adjustment 3.2.3. Outfield Speech Audio Adjustment 3.2.4. Localization Center Speaker Unit Derival 4. Effects 5. Remote Communication Device According to Second Embodiment 5.1. Acoustic Signal Processing Unit 5.1.1. Background Sound Separation Unit 5.1.2. Clarity Adjustment Unit 5.2. Output Integration Unit 5.3. Audio Output Control Unit 6. Data Flow Between Remote Communication Devices According to Second Embodiment 7. Effects 8. Remote Communication Device According to Third Embodiment 8.1. Acoustic Signal Processing Unit 8.1.1. Audio Synthesis Unit 9. Data Flow Between Remote Communication Devices According to Third Embodiment 10. Effects 11. Remote Communication Device According to Fourth Embodiment 11.1. Human Sensing Unit 11.2. Acoustic Signal Processing Unit 11.3. Other Voice Recognition 11.4. Effects 12. Remote Communication Device According to Fifth Embodiment 12.1. Human Sensing Unit 12.2. Acoustic Signal Processing Unit 12.3. Effects 13. Remote Communication Device According to Sixth Embodiment 13.1. Audio Output Control Unit 13.2. Effects 14. Remote Communication Device According to Seventh Embodiment 14.1. Human Sensing Unit 14.2. Acoustic Signal Processing Unit 14.3. Audio Output Control Unit 14.4. Effects 15. Hardware Configuration

[0011] 1. A remote communication device according to a first embodiment In a conventional remote communication system, when multiple people speak at the same time, the sounds are mixed together during playback, making it difficult to hear, or the voice of the person you want to hear is arbitrarily suppressed, making it seem unnatural. Furthermore, in conventional remote communication systems, sounds other than the speech are often cut out.

[0012] 1 is a block diagram of a remote communication device. The remote communication device 1 according to the embodiment processes background sounds and spoken voices separately, and further divides the spoken voices into two types depending on the person being spoken to. The remote communication device 1 then processes each of the divided voices using different parameters, thereby adjusting the audibility of the voices while maintaining a sense of connection with the remote location. Here, the sense of connection corresponds to a sense of realism that makes it feel as if you are having a face-to-face conversation with the person being spoken to in the same space.

[0013] The remote communication device 1 has a transmitting unit 10 and a receiving unit 20. The remote communication device 1 shown in Figure 1 is a device that performs two-way communication with a remote communication device 1 of a partner, and the remote communication device 1 of the partner also has a transmitting unit 10 and a receiving unit 20.

[0014] However, when two-way transmission is not performed, the remote communication device 1 on the transmitting side only needs to have at least a transmission unit 10, and the remote communication device 1 on the receiving side only needs to have at least a receiving unit 20. In other words, the information processing system that realizes remote communication according to the embodiment has the remote communication device 1 as a transmitting device, and the remote communication device 1 as a receiving device.

[0015] The remote communication device 1 on the sending side and the remote communication device 1 on the receiving side are connected via a network 7 .

[0016] Fig. 2A is a perspective view of the environment in which the remote communication device is placed, and Fig. 2B is a front view of the environment in which the remote communication device is placed. The dotted lines in Figs. 2A and 2B indicate the wiring of signal lines.

[0017] The remote communication device 1 is connected to an audio interface 5. The audio interface 5 is connected to a plurality of microphones 3 and a plurality of opposed speakers 4. The audio interface 5 outputs sounds picked up by the plurality of microphones 3 to the remote communication device 1. The audio interface 5 also outputs sounds output from the remote communication device 1 to the plurality of opposed speakers 4.

[0018] The camera 2 captures an image of a specific space in which it is placed and outputs the captured image to the remote communication device 1. The microphone 3 is an omnidirectional microphone. A plurality of microphones 3 are arranged in a predetermined configuration to form a microphone array 30. The camera 2 and the microphone array 30 are preferably aligned in approximately the same position. In particular, they are preferably aligned in a horizontal line from the position where people are lined up in FIG. 2A toward the camera 2 and the microphone array 30.

[0019] A display 6 is also connected to the remote communication device 1. The display 6 displays the video received by the remote communication device 1. As shown in Fig. 2B , a speaker array 41 and a speaker array 42 of multiple opposed speakers 4 are arranged above and below the display 6, respectively.

[0020] In the following explanation, it is assumed that people who will be conversing using the remote communication device 1 are lined up in a row parallel to the display 6 at equal distances from the display 6. The distance from the display 6 to the people is, for example, approximately 2.0 m. In the following explanation, the surface of the display 6 that actually displays the image will be referred to as the "screen." Furthermore, the direction connecting the head and feet of the photographed person, which is the up-down direction of the image when projected onto the screen, will be referred to as the "vertical direction." Furthermore, the left-right direction when the photographed person is facing forward, which is the left-right direction when the image is projected onto the screen, will be referred to as the "horizontal direction." In particular, the left side of the image when projected onto the screen will be referred to as the "left," and the right side will be referred to as the "right."

[0021] <1.1. Definitions of Terms> Next, definitions of terms used in the embodiments will be described.

[0022] <1.1.1. Remote and Local> In the embodiment, the remote communication device 1 connects two physically separated locations and performs remote communication between people in each space. Of the two locations where remote communication is performed, the location where the transmitting remote communication device 1 is located is referred to as the "remote" location, and the location where the transmitting remote communication device 1 is located is referred to as the "local" location. In other words, the following description will focus on the case where video and audio from the remote location are sent to the local location. Furthermore, the remote person who is the conversation partner of the local location is referred to as the speaker, and the conversation partner to whom the speaker wants to send a message is referred to as the listener.

[0023] <1.1.2. Individual Sound> Individual sound refers to the individual voice of each speaker. In many-to-many communication, where many remote people converse with many local people, it is expected that multiple people will be speaking simultaneously at one of the locations (remote). In such a situation, simply recording with an omnidirectional microphone will result in the recording of multiple voices. Therefore, the process of extracting each speaker's voice individually from the recorded voices of multiple people is called "individual sound separation."

[0024] <1.1.3. Infield and Outfield> The infield refers to the position of a remote speaker who is currently paying attention to the local speaker, such as a local speaker who is currently having a conversation with the local speaker. For example, when a conversation is taking place between a local person and a remote person, the remote person is treated as the "infield" from the local person's perspective. The infield can also refer to anyone currently participating in the conversation.

[0025] The outfield refers to the position of a person who is not paying attention to the local conversation at that time, such as a remote person having a conversation with another remote person, relative to the local speaker. If the listener for the remote speaker is also remote, that is, if the conversation is between two remote speakers, the remote listener is treated as an "outfield" for the local speaker. An outfield person can also be said to be someone who is not participating in the conversation at that time. However, a remote speaker can become an infield or outfield depending on changes in that person's behavior over time.

[0026] <1.1.4. Multiple Opposite Speakers> The multiple opposed speakers 4 are a speaker system that includes a pair of a speaker array 41 made up of a plurality of speaker units 411 and a speaker array 42 made up of a plurality of speaker units 412. In this embodiment, the speaker array 41 and the speaker array 42 are arranged side by side in the vertical direction, sandwiching the screen of the display 6 therebetween. For example, the speaker array 41 is installed above the speaker array 42 in the image projected on the screen. The following describes an example in which the speaker array 41 is installed above the speaker array 42.

[0027] The speaker array 41 and the speaker array 42 have the same number of speaker units 411 and 412. In the speaker array 41, a plurality of speaker units 411 are arranged horizontally. Similarly, in the speaker array 42, a plurality of speaker units 412 are arranged horizontally. Each speaker unit 411 and each speaker unit 412 are arranged at a corresponding position in the vertical direction. Hereinafter, a pair of speaker units 411 and 412 arranged at corresponding positions in the vertical direction will be referred to as speaker units 411 and 412 in the same row. By reproducing the same sound from the speaker units 411 and 412, a sound source is virtually localized in the center between the speaker array 41 and the speaker array 42.

[0028] The multiple opposed speakers 4 have the following advantages over a system that simply has speakers lined up. By reproducing the voices of separate speakers on individual channels, the multiple opposed speakers 4 make it easier to hear each individual voice than when multiple voices are mixed and reproduced on a single channel. Furthermore, by localizing the sound source in the vertical center of the display 6 screen, the multiple opposed speakers 4 can make the voice sound as if it is coming from approximately the position of a person's face.

[0029] However, in the first embodiment, the output direction of speaker array 41 and the output direction of speaker array 42 are opposed to each other to achieve point localization and enhance the effect of improving the sense of realism, but it is also possible to use a general array speaker to achieve surface localization without opposing them.

[0030] 1.2. Transmission Unit The transmission unit 10 is a function used in the transmitting remote communication device 1. In the following explanation, the case where sound is sent from the remote to the local will be explained, that is, the remote will be the transmitting side and the local will be the receiving side. As shown in FIG. 1 , the transmission unit 10 has a human sensing unit 11, a voice separation unit 12, an acoustic signal processing unit 13, an output integration unit 14, and a transmission unit 15.

[0031] <1.2.1. Person Position Detection Unit> The person sensing unit 11 receives input of video captured by the camera 2 in a specific space on the remote side where the camera 2 is located. The person sensing unit 11 then executes person sensing processing including a person position determination processing and an infield / outfield determination processing, which will be described below. Hereinafter, the space on the remote side captured by the camera 2 will be referred to as the "remote space."

[0032] The human sensing unit 11 detects the position of the horizontal origin in the remote space from the image of the camera 2. FIG. 3 is a diagram for explaining the processing of the human sensing unit. Here, as shown in the image in the upper row of FIG. 3 facing the paper surface, an example will be described in which an image 60 showing three speakers 61 to 63 is displayed on the display 6. For example, the human sensing unit 11 assumes that the origin is located at the horizontal center 600 of the screen when the image of the camera 2 in the remote space is projected onto the screen of the display 6. Next, the human sensing unit 11 sets normalized coordinates that indicate the horizontal position of the screen of the display 6 from the origin. For example, the human sensing unit 11 sets the distance from the horizontal edge of the screen of the display 6 to the center 600 as a distance of 0.5 in normalized coordinates. In FIG. 3, the human sensing unit 11 uses the x-coordinate as a normalized coordinate, and sets the normalized coordinate of the right edge of the screen of the display 6 to 0.5 and the normalized coordinate of the left edge to -0.5. That is, in this embodiment, the human sensing unit 11 sets normalized coordinates with the horizontal width of the screen set to 1.

[0033] Next, the human sensing unit 11 performs the following human position determination process. The human sensing unit 11 performs skeletal recognition of each speaker who appears in the image captured by camera 2. Then, the human sensing unit 11 acquires the normalized coordinates of the neck of each speaker who appears in the image captured by camera 2. As a result, the human sensing unit 11 can acquire normalized coordinates that indicate the position of each speaker, and can determine the lateral positional relationship between the speakers. The human sensing unit 11 can perform skeletal recognition using a general tool that has already been commercially released. Then, the human sensing unit 11 generates a person list consisting of elements equal to the number of speakers whose normalized coordinates have been acquired. In this embodiment, the human sensing unit 11 arranges the elements in the person list in ascending order of normalized coordinate values.

[0034] The images at the bottom of the page in Figure 3 show details of the person position determination process and the infield / outfield determination process. For example, when video 60 is acquired, the person sensing unit 11 performs skeletal recognition to obtain the normalized coordinates of the neck 601 of speaker 61, the normalized coordinates of the neck 602 of speaker 62, and the normalized coordinates of the neck 603 of speaker 63. Then, from the acquired normalized coordinates, the person sensing unit 11 confirms that speakers 61, 62, and 63 are lined up in this order from the left of the video 60. Then, the person sensing unit 11 sets elements in the person list in the order of speaker 61, speaker 62, and speaker 63 in accordance with the video 60.

[0035] Here, the human sensing unit 11 detects skeletons from the video, but the coordinates obtained from the image are usually expressed as screen coordinates using the number of pixels in the video, rather than a normalized coordinate system. Therefore, in practice, the human sensing unit 11 acquires the screen coordinates of the speaker's position from the video in the human position determination process, and then converts the screen coordinates into normalized coordinates. Figure 4 is a diagram showing the coordinate conversion from screen coordinates to normalized coordinates. Here, the axes of the screen coordinates 621 are the u-axis and v-axis, and the axes of the normalized coordinates 622 are the x-axis and y-axis.

[0036] The human sensing unit 11 determines the position of each speaker from the image projected on the screen 620 using a screen coordinate system 621 having a u-axis and a v-axis, and then converts the determined position of each speaker into a normalized coordinate system 622 having an x-axis and a y-axis. Scr The screen height is H Scr Then, the range of each coordinate system is 0≦u≦W Scr and 0≦v≦H Scr , and -0.5≦x≦0.5 and -0.5≦y≦0.5. Here, the vertical normalized coordinate system is expressed in the range of -0.5 to 0.5, with the vertical center of the image set as 0. In this case, the human sensing unit 11 performs coordinate transformation using the following equation (1). By performing coordinate transformation in this way, processing can be performed without being affected by the camera's angle of view or aspect ratio.

[0037]

[0038] Next, the human sensing unit 11 executes the following infield / outfield determination process. The human sensing unit 11 estimates the head direction of each person in the video, i.e., the direction in which the speaker is facing. For example, if normalized coordinates for p people are obtained in the human position determination process, the human sensing unit 11 detects the faces of the speakers located at each of the normalized coordinates for the p people and estimates the head direction by image analysis, etc. More specifically, the human sensing unit 11 estimates the head direction by defining the direction of the vector in the positive direction of the normalized coordinates as 0 degrees and estimating the angle Φ of the head direction vector from the 0-degree vector.

[0039] Here, the human sensing unit 11 has an infield range in advance for determining whether each speaker is in the infield or outfield. For example, the human sensing unit 11 defines a state in which the speaker faces the camera 2 as 90 degrees, and defines a range around a predetermined angle of 90 degrees as the infield range. The human sensing unit 11 then determines a speaker whose head direction is within the infield range to be in the infield, and determines a speaker whose head direction is not within the infield range to be in the outfield. Next, the human sensing unit 11 provides an infield flag item for each element in the person list and sets the initial value to False. Then, the human sensing unit 11 sets the infield flag to True for speakers determined to be in the infield. Furthermore, the human sensing unit 11 sets the infield flag to False for speakers determined to be in the outfield. As a result, the human sensing unit 11 can set the infield flag of each speaker in the person list to Tre if the speaker is in the infield, to False if the speaker is in the outfield, and further to False if the head direction is not detected.

[0040] For example, as shown in the image in the lower part of Figure 3, the human sensing unit 11 estimates head directions 611 to 613 for speakers 61 to 63. Furthermore, here, the human sensing unit 11 performs an infield / outfield determination process, with the infield range ψ set to 45 degrees forward or backward from the front direction 610, which is 90 degrees. In this case, because head direction 611 is included in the infield range, the human sensing unit 11 determines that speaker 61 is in the infield. Furthermore, because head directions 612 and 613 are not included in the infield range, the human sensing unit 11 determines that speakers 62 and 63 are people in the outfield. Then, the human sensing unit 11 sets the infield flag of speaker 61 in the person list to True and the infield flags of speakers 62 and 63 to False.

[0041] Continuing the explanation, returning to Fig. 1, the human sensing unit 11 then outputs the created person list to the acoustic signal processing unit 13 and the output integration unit 14.

[0042] Here, the remote space is an example of a "predetermined space," and determining whether a speaker is in the infield or outfield is an example of "determining whether a speaker is participating in a conversation." That is, the human sensing unit 11 determines whether each speaker is participating in a conversation based on video captured by a camera 2 that captures multiple speakers present in the predetermined space. The human sensing unit 11 also determines the position of each speaker based on the video.

[0043] <1.2.2. Audio Separation Unit> The audio separation unit 12 receives audio input collected by the multiple microphones 3. The input to the microphone array 30 includes not only speech sounds generated by speech but also background sounds generated in the environment, such as background noise, background music (BGM), and other sounds. Therefore, the audio separation unit 12 performs audio separation processing to separate the speech sounds input from the microphone array 30 from the background sounds so that the speech sounds can be processed later to adjust their audibility. The audio separation processing can also be considered speech extraction processing. For example, the audio separation unit 12 performs audio separation using vocal extraction or technology to extract only the singing voice from music with singing voices.

[0044] In this way, the audio separator 12 separates the background sound from the audio in the predetermined space that corresponds to the remote space. By separating the speech sound from the background sound, it becomes easier to adjust the clarity of the speech sound and the background sound.

[0045] In this embodiment, an omnidirectional microphone is used as the microphone 3, but alternatively, each speaker may wear a pin microphone and an omnidirectional microphone may be used to collect background sound as the microphone 3. In this case, the audio separator 12 may use the audio collected by the pin microphone as the speech of each speaker and the audio collected by the omnidirectional microphone as the background sound.

[0046] 1.2.3. Acoustic Signal Processing Unit The acoustic signal processing unit 13 includes an individual sound separation unit 131 and a clarity adjustment unit 132 .

[0047] The individual sound separation unit 131 receives input of speech sounds extracted by the speech separation process by the speech separation unit 12. The individual sound separation unit 131 also acquires the person list created by the person sensing unit 11. Then, the individual sound separation unit 131 executes the following individual sound separation process using the person list to separate the speech sounds of each person appearing in the video.

[0048] FIG. 5 is a diagram illustrating the individual sound separation process. Here, the case where the following conditions are satisfied will be described. The camera 2 and the microphone array 30 are arranged in series in the shooting direction of the camera 2. That is, the camera 2 and the microphone array 30 are arranged in the normal direction of the two-dimensional image captured by the camera 2, and when considering the coordinates formed by the vertical, horizontal, and depth directions in the image, the vertical and horizontal coordinates of the camera 2 and the microphone array 30 match. Furthermore, speakers in the remote space line up horizontally at positions where the vertical direction of the two-dimensional image captured by the camera 2 matches. Here, the planes corresponding to the vertical and horizontal positions of the speakers in the remote space are referred to as the "virtual screen."

[0049] The normalized coordinates shown in the range of -0.5 to 0.5 in FIG. 5 coincide with the horizontal coordinates of the virtual screen. The distance L between the camera 2 and the microphone 3 is known. Furthermore, the x ph is the physical distance from the origin of the coordinates on the virtual screen. ph can be calculated from the angle of view of the camera 2 and the value of the distance L.

[0050] In this case, the individual sound separation unit 131 creates threads for the number of elements registered in the acquired person list, and sets an ID (Identifier) ​​for identifying each speaker as an attribute to each thread. For example, the individual sound separation unit 131 assigns thread IDs 1, 2, 3, ... in the order of arrangement from the left in the person list. Furthermore, the individual sound separation unit 131 stores the normalized coordinate values ​​registered in the person list in each thread as thread-specific values. Then, the individual sound separation unit 131 uses the normalized coordinates obtained in the person position determination process as coordinates on the virtual screen, and calculates x, which is the physical distance from the origin of each speaker. phThen, the individual sound separation unit 131 calculates the calculated physical distance x ph Using the distance L from the virtual screen to the camera 2, the angle θ to a specific person can be calculated using the following equation (2).

[0051]

[0052] For example, for the speaker 101 in FIG. 5, the individual sound separation unit 131 calculates x 1 Then, the individual sound separation unit 131 obtains the normalized coordinate x 1 The physical distance from the origin to the speaker 101 is calculated from the equation (2), and the angle θ from the origin centered on the camera 2 and microphone 3 of the speaker 101 is calculated using the calculated physical distance and the distance L. 1 The individual sound separation unit 131 can obtain the angle of each speaker for each thread by performing calculations to obtain the angle θ for each thread.

[0053] Next, the individual sound separation unit 131 performs individual sound separation processing on the speech sound extracted by the speech separation unit 12 through the speech separation for each speaker using beamforming. The speech sound extracted through the speech separation by the speech separation unit 12 is sound obtained from the microphone array 30, and sensitivity in the direction of each speaker is ensured. Therefore, in the case of speaker 101, the individual sound separation unit 131 calculates an angle θ with respect to speaker 101. 1 is the beam direction, and the angle θ 1 The beam width is set to a range 102 of 15 degrees centered on the target angle. The individual sound separation unit 131 then specifies the beam direction and beam width to suppress sounds from directions outside the specified angular range among the acquired speech sounds, thereby ensuring sensitivity in the target direction and acquiring the individual sounds of the speaker 101.

[0054] In this way, the individual sound separation unit 131 separates individual sounds, which are the speech sounds of each speaker, from the sound in a specified space corresponding to the remote space, based on the position of each speaker determined by the human sensing unit 11.

[0055] In this regard, it has been difficult in the past to acquire individual sounds of speakers without wearing a microphone while associating them with the positions of people on the video. On the other hand, the individual sound separation unit 131 according to this embodiment uses a person list in which the positions of people on the video are registered in the order in which they appear on the video, and performs beamforming processing while maintaining that order, thereby easily creating pairs of the position of each speaker on the video and the separated individual sounds. This makes it easy to match the video and sound output positions when playing back the speech locally after transmission.

[0056] 1 , the explanation will be continued. The clarity adjustment unit 132 receives input of the individual sounds of each person present in the remote space generated by the individual sound separation unit 131. The clarity adjustment unit 132 also receives input of a person list from the individual sound separation unit 131. Furthermore, the clarity adjustment unit 132 receives input of background sounds extracted by the sound separation by the sound separation unit 12.

[0057] The clarity adjustment unit 132 determines, for each individual sound, whether each speaker is in the infield or outfield using the infield flag in the person list. Then, the clarity adjustment unit 132 performs clarity adjustment processing on the individual sounds and background sounds according to the determination result. Specifically, the clarity adjustment unit 132 performs an enhance processing on the individual sounds of people in the infield to make them easier to hear. The clarity adjustment unit 132 also performs a degrade processing on the individual sounds and background sounds of people in the outfield to make them harder to hear. The clarity adjustment unit 132 then outputs the individual sounds and background sounds that have been subjected to the clarity adjustment processing to the output integration unit 14.

[0058] The clarity adjustment unit 132 can use formant weighting, an equalizing filter, or a reverberation filter for the clarity adjustment process. In the case of formant weighting, the clarity adjustment unit 132 reinforces the second formant and above as an enhancement process, and suppresses the second formant and above as a degrading process. Furthermore, the clarity adjustment unit 132 performs a process using an equalizing filter, such as a degrading process of background sound using a low-pass filter or a degrading process by adding reverberation.

[0059] In this way, the clarity adjustment unit 132 adjusts the clarity of the speech of a speaker who is participating in the conversation determined by the human sensing unit 11, which is included in the audio of the predetermined space corresponding to the remote space collected by the microphone 3. The clarity adjustment unit 132 also performs clarity reduction adjustment on the speech of a speaker who is not participating in the conversation, which is included in the audio of the predetermined space corresponding to the remote space, to reduce the clarity below that of the audio collected by the microphone 3. The clarity adjustment unit 132 also performs clarity reduction adjustment on the background sound separated by the audio separation unit 12 to reduce the clarity below that of other audio.

[0060] In this regard, it has been difficult to appropriately handle background sounds and adjust clarity according to whether a speaker is participating in a conversation. In contrast, the clarity adjustment unit 132 can change the ease of listening by determining whether the transmitted sound is background sound or speech, and also, among speech, whether the speaker is in the infield or outfield, thereby promoting smooth dialogue.

[0061] <1.2.4. Output Integration Unit> The output integration unit 14 receives input of the individual sounds and background sounds that have been subjected to clarity adjustment processing from the clarity adjustment unit 132. The output integration unit 14 also receives input of the person list from the person sensing unit 11. Then, the output integration unit 14 executes the clarity adjustment processing described below.

[0062] Audio processing of background sounds and speech sounds is performed on one channel at a time, corresponding to each sound. That is, the output integrating unit 14 receives input data equal to the number of speakers plus one channel of background sounds. In other words, if the background sounds are one channel of data and the number of speakers is p, the output integrating unit 14 acquires (1 + P) channels of data.

[0063] Then, in order to transmit the audio locally, the output integrator 14 integrates the sounds of all channels into a single piece of audio data. For example, the output integrator 14 can generate a single piece of audio data using channel interleaving, which includes elements of each channel in each frame. The output integrator 14 then outputs the generated single piece of audio data to the transmitter 15. The output integrator 14 also outputs the personal list to the transmitter 15, along with information associating each element of the personal list and the background sound with each channel included in the audio data. Here, the output integrator 14 may add the data of the personal list to the audio data to create a single piece of data.

[0064] 1.2.5. Transmitting Unit The transmitting unit 15 receives an input of a single piece of sound data including the individual sounds of each speaker and background sound from the output integration unit 14. The transmitting unit 15 also receives an input of a personal list from the output integration unit 14. The transmitting unit 15 also receives an input of captured video from the camera 2. The transmitting unit 15 then transmits the single piece of sound data including the individual sounds of each speaker and background sound, the personal list, and the video from the camera 2 to the receiving unit 20 of the remote communication device 1 on the local side via the network 7. The network 7 is, for example, the Internet or a local area network (LAN). In this way, the transmitting unit 15 transmits the audio of a predetermined space corresponding to the remote space processed by the clarity adjustment unit 132.

[0065] 1.3 Receiving Unit The receiving unit 20 is a function used in the receiving-side remote communication device 1, which is the local side in this embodiment. As shown in FIG. 1 , the receiving unit 20 has a receiving section 21 and an audio output control section 22.

[0066] <. 3.1. Receiving Unit> The receiving unit 21 receives one piece of sound data including individual sounds of each speaker in the remote space and background sounds, a person list, and video transmitted from the transmitting unit 15 of the remote communication device 1 on the remote side. The receiving unit 21 outputs the received video to the display 6. The display 6 displays the video output from the receiving unit 21 on a screen. The receiving unit 21 also outputs the sound data and person list to the audio output control unit 22.

[0067] In this way, the receiving unit 21 receives the audio of the predetermined space corresponding to the remote space transmitted by the transmitting unit 15. The receiving unit 21 also receives the video together with the audio of the predetermined space corresponding to the remote space, and projects the video on a screen on which the speaker arrays 41 and 42 are arranged, one at each end of the video in the vertical direction when the video is projected.

[0068] <1.3.2. Audio output control unit> The audio output control unit 22 performs a multiple facing speaker balancing process described below to specify the speaker units 411 and 412 that will output the background sound, infield speech sound, and outfield speech sound, respectively, and play them. The audio output control unit 22 receives sound data including the individual sounds of each speaker and the background sound, as well as an input of a person list, from the receiving unit 21. Hereinafter, speech sounds from people in the infield will be referred to as "infield speech sound," and speech sounds from people in the outfield will be referred to as "outfield speech sound."

[0069] Fig. 6 is a diagram showing the multiple facing speaker balancing process. The image at the top of Fig. 6 shows the lengths of each part of the speaker arrays 41 and 42, and a graph 200 at the bottom of the page shows the sounds output from the speaker arrays 41 and 42 of the background sound, infield speech sounds, and outfield speech sounds after processing. In graph 200, the horizontal axis indicates the positions corresponding to the speaker arrays 41 and 42, and the vertical axis indicates the volume. The vertical axis of graph 200 indicates a value normalized with the peak volume of the infield speech sounds set to 1.

[0070] The audio output control unit 22 determines the physical speaker array width W spk and display width W dispThe audio output control unit 22 also has information about the number u of speaker units 411 in the speaker array 41. The number of speaker units 412 in the speaker array 42 is also u. Here, W spk represents the distance between the leftmost speaker unit 411 and the rightmost speaker unit 411. spk The same applies to the speaker unit 412. In addition, the speaker distances W between the speaker units 411 and between the speaker units 412 are u are all equal, and the display 6, the speaker array 41, and the speaker array 42 are aligned at the center of the left and right sides.

[0071] As a pre-processing common to the output adjustment processing for the background sound, the infield speech sound, and the outfield speech sound, the audio output control unit 22 spk and u, W u =W spk / u, and the speaker distance W between the speaker units 411 u Here, W u is the actual physical distance, so the audio output control unit 22 sets the processing distance as W to match the position of the speaker on the video with the audio. u_n =W u / W disp The audio output control unit 22 adjusts the size of the speaker array width W spk Similarly, W spk_n Furthermore, the audio output control unit 22 calculates the normalized coordinates xu of the speaker units 411 and 412.

[0072] Here, if the amplitude ratio is x, the sound pressure level is y, and the loudness ratio is z, then y = 20 log 10 x and z=2 y/10 The following relation holds. When we solve this for x, we get x = z 1/2log 10 2 This means that if you want to increase the volume by z times, you need to increase the amplitude by z. 1.66 This means that you should double it.

[0073] The following describes the multiple opposing speaker balancing process for background sound. In the remote communication device 1 according to this embodiment, in order to facilitate communication with a remote conversation partner while maintaining a sense of spatial connection, background sound is transmitted locally rather than being completely removed. However, if the volume of the background sound is made too high, it becomes difficult to hear infield speech. Therefore, the audio output control unit 22 further adjusts the volume of the background sound, whose clarity has been reduced by the degradation process, at the timing of playback to blur the sense of localization of the background sound.

[0074] Specifically, the audio output control unit 22 acquires the background sound from the sound data. Next, the audio output control unit 22 outputs the background sound from all the speaker units 411 and 412 so that the loudness of the background sound becomes 1 / u. That is, the audio output control unit 22 reduces the amplitude of the background sound to (1 / u) 1.66 The processing is performed so that the number of speakers is doubled, and sounds are output from all speaker units 411 of the speaker array 41 and all speaker units 412 of the speaker array 42.

[0075] By performing the above processing, background sound is output at the same reduced volume from all speaker units 411 and 412, as shown by curve 201 in graph 200. Here, background sound refers to sounds whose source location is unclear, such as natural environmental sounds, background music in a cafe, and office noise. By blurring the sense of localization of the background sound, the audio output control unit 22 can reproduce the remote space with a higher sense of realism than when the background sound is played from a single speaker unit 411 or 412.

[0076] Next, we will explain the multiple facing speaker balancing process for infield speech sounds. Since infield speech sounds are intended for local people, it is preferable to make them as easy to hear as possible. Therefore, the audio output control unit 22 performs the following process on individual infield sounds that have been remotely enhanced.

[0077] The audio output control unit 22 acquires infield speech from the sound data using the person list. Next, the audio output control unit 22 reproduces the acquired infield speech directly from the speaker arrays 41 and 42 without blurring the sense of localization of the acquired infield speech and without adjusting the volume or volume. Here, the audio output control unit 22 identifies the speaker units 411 and 412 that are closest to the normalized coordinates of the speaker of the infield speech. That is, the normalized coordinates of the speaker are set to x p In this case, the audio output control unit 22 calculates the normalized coordinate x u The speaker units 411 and 421 present in the

[0078]

[0079] Then, the audio output control unit 22 outputs the infield speech sound from the specified speaker units 411 and 412. In this case, the audio output control unit 22 does not output the infield speech sound from the other speaker units 411 and 412. By performing the above processing, the infield speech sound is output from one of the specified speaker units 411 and 412 at a volume louder than the other sounds, as shown by curve 202 in graph 200.

[0080] Next, we will explain the multiple facing speaker balancing process for outfield speech. Outfield speech does not need to be clearly audible because the local person is not directly speaking with them. However, as with background sounds, it is undesirable to completely eliminate it from the perspective of spatial connectivity. Furthermore, by lowering the clarity but still being audible if one listens closely, there is an opportunity for the local person to talk to someone in the outfield, develop communication, and transition to the infield. Therefore, to make infield speech relatively easier to hear and outfield speech relatively harder to hear, a degradation process is performed on individual outfield sounds in the remote environment. Then, the audio output control unit 22 performs a process to blur the sense of localization of the acquired outfield speech.

[0081] Specifically, since the sound source position of outfield speech is clear, the audio output control unit 22 blurs the sense of localization of outfield speech by processing using a normal distribution function. Here, the normal distribution function is expressed by the following equation (4), where μ is the mean value and σ is the standard deviation.

[0082]

[0083] In order to match the peak portion of the normal distribution function with the position of the speaker of the outfield speech, the audio output control unit 22 calculates the normalized coordinates x of the speaker units 411 and 412 that are closest to the normalized coordinates of the speaker that satisfy the formula (3). u The position of the sound output control unit 22 is set to the value of μ. u 7 is a diagram showing processing using a normal distribution function for outfield speech. The audio output control unit 22 creates a function f(x) represented by a graph 211 whose peak is the position of the speaker of the outfield speech. Here, f(x) is expressed by the following mathematical formula (5).

[0084]

[0085] Although the x-coordinate range of the normalized distribution function is normally from -∞ to ∞, the audio output control unit 22 imposes a restriction on the range of reproduction by the speaker arrays 41 and 42, for example, by reproducing using the speaker units 411 and 412 corresponding to the range up to a 95% confidence interval. This allows the audio output control unit 22 to output sound within a range of μ-2σ to μ+2σ=2, thereby fining the range of sound output. In this case, the audio output control unit 22 reproduces one person's outfield speech from five speaker units 411 and 412 each. By narrowing the range of the speaker units 411 and 412 to be reproduced in this way, the audio output control unit 22 can reduce the processing load.

[0086] Furthermore, depending on the value of σ, the maximum value may be greater than 1. In that case, using f(x) to adjust the volume may result in outfield speech sounds being louder than infield speech sounds. Therefore, the audio output control unit 22 uses a scaling factor sf to reduce the overall volume as shown in graph 212 of FIG. 7 by using g(x) = sf / f(μ) × f(x), thereby preventing outfield speech sounds from being reproduced louder than infield speech sounds. sf can be set to, for example, 0.8.

[0087] Then, the audio output control unit 22 adjusts the amplitude of the outfield speech sound to {g(x)} in accordance with the normalized coordinates. 1.66 The outfield speech is processed so that the volume is doubled, and the outfield speech is played back by a specified number of speaker units 411 and 412 centered around the position of the speaker of the outfield speech. Here, the value of the scaling factor sf (0≦sf≦1) that defines the peak value and the speaker playback range can be changed by the user using a configuration file, etc. By performing the above processing, the outfield speech is output at a lower volume than the infield speech from speaker units 411 and 412 in a specified range centered around the position of the speaker, as shown by curve 203 of graph 200 in Figure 5.

[0088] In this way, the audio output control unit 22 executes the following process in the multiple opposed speaker 4 having two speakers: a speaker array 41 in which a plurality of speaker units 411 are arranged in a row, and a speaker array 42 in which a plurality of speaker units 412 are arranged in a row. Based on the position of each speaker, the audio output control unit 22 selects speaker units 411 and 412 that will reproduce the speech of each speaker from the audio received by the receiving unit 21 in a predetermined space corresponding to the remote space. Then, the audio output control unit 22 causes the selected speaker units 411 and 412 to reproduce the speech of each speaker. Furthermore, for each speaker displayed on the screen, the audio output control unit 22 selects speaker units 411 and 412 to reproduce the speech based on the position of the speaker so that the speech is reproduced near the position on the screen.

[0089] In this regard, it has been difficult to control the playback method of multi-channel speakers depending on whether the speaker is conversing with the other party on the other side of the screen, i.e., whether they are participating in the conversation. In contrast, the audio output control unit 22 makes the voice of the infield speaker, who is the local conversation partner, relatively easy to hear, while maintaining a sense of connection with the remote space by leaving background sounds and outfield speech. Furthermore, the audio output control unit 22 blurs the sense of sound localization by reproducing background sounds and outfield speech from multiple speaker units 411 and 412, making them more difficult to hear, while making the infield speech relatively easy to hear, allowing the user to better concentrate on the conversation with the remote party.

[0090] 8 is a diagram showing an outline of the data flow between the local remote communication devices according to the first embodiment. Here, too, the case where data is transmitted from the transmitting unit 10 of the remote remote communication device 1 to the receiving unit 20 of the local remote communication device 1 will be described. Here, the transmission of audio will be described.

[0091] The sending side process 221 is a process executed by the remote communication device 1 located on the remote side, and the receiving side process 222 is a process executed by the remote communication device 1 located on the local side.

[0092] The microphone 3 inputs the sound picked up in the remote space to the remote communication device 1 (step S1). Here, for example, data of the s channel exists as data of the microphone array 30. The camera 2 also inputs the video generated by capturing the remote space to the remote communication device 1 (step S2).

[0093] The image input from the camera 2 is sent to the human sensing unit 11. The human sensing unit 11 executes human sensing processing using the image captured by the camera 2 to determine the position of each person in the remote space and to determine whether the field is infield or outfield (step S3).

[0094] Specifically, the human sensing unit 11 sets normalized coordinates in the horizontal direction of the image projected on the screen. Then, the human sensing unit 11 performs skeletal recognition on speakers in the image and obtains normalized coordinates of the position of each speaker. After that, the human sensing unit 11 generates a person list having elements corresponding to the speaker sequence and registers the normalized coordinates of each person's position, thereby executing the human position determination process (step S31). Here, the human sensing unit 11 extracts p speakers present in the image. In other words, p elements are registered in the person list.

[0095] The human sensing unit 11 also estimates the head direction of each speaker in the video. Then, the human sensing unit 11 determines whether each speaker is in the infield or outfield depending on whether the estimated head direction is included in a predetermined infield range, and executes infield / outfield determination processing by registering an infield flag indicating the infield or outfield in the person list based on the determination result (step S32).

[0096] The audio input from the microphone 3 is sent to the audio separator 12. The audio separator 12 performs an audio separation process to separate speech audio and background sound contained in the audio input from the microphone 3 (step S4). The audio separator 12 generates one channel of data as background sound and another channel of data as speech audio from the s channel data input from the microphone array 30.

[0097] The background sound, which is one-channel data extracted by the audio separation unit 12, is sent to the clarity adjustment unit 132 of the audio signal processing unit 13. The clarity adjustment unit 132 performs a degrading process as a clarity adjustment process to make the background sound more difficult to hear (step S5).

[0098] Furthermore, the speech voice, which is s-channel data extracted by the voice separation unit 12, is sent to the individual sound separation unit 131 of the acoustic signal processing unit 13. Furthermore, the person list is sent to the individual sound separation unit 131. The acoustic signal processing unit 13 performs speech voice signal processing on the speech voice using the person list (step S6).

[0099] Specifically, the individual sound separation unit 131 creates p threads, which is the number of elements registered in the acquired personal list. Furthermore, the individual sound separation unit 131 stores the normalized coordinate values ​​registered in the personal list as thread-specific values ​​for each thread. The individual sound separation unit 131 then calculates the physical distance for each speaker using the normalized coordinates of the virtual screen as coordinates. The individual sound separation unit 131 then calculates the angle θ to each speaker using the calculated physical distance and the distance from the virtual screen to camera 2. Next, the individual sound separation unit 131 performs individual sound separation processing on the speech sounds extracted by the speech separation unit 12 using beamforming according to the angle to each speaker (step S61). As a result, the individual sound separation unit 131 generates individual sounds, which are one-channel data for each speaker. That is, one-channel individual sound data is generated for each of the p threads.

[0100] Next, the clarity adjustment unit 132 determines whether the speaker corresponding to the individual sound for each thread is infield or outfield using the infield flag in the person list. Then, as a clarity adjustment process, the clarity adjustment unit 132 performs an enhancement process to make the individual sounds of infield speakers easier to hear, and a degradation process to make the individual sounds of outfield speakers harder to hear (step S62).

[0101] The background sound that has been subjected to the clarity adjustment processing is input to the output integrating unit 14. The individual sounds that have been subjected to the speech signal processing are also input to the output integrating unit 14. As a result, the output integrating unit 14 acquires data of (1+P) channels. The output integrating unit 14 then executes output integrating processing to integrate the sounds of all channels to generate one piece of sound data (step S7).

[0102] The sound data and the person list generated by the output integration unit 14 are sent by the transmission unit 15 to the local-side remote communication device 1 via the network 7 (step S8). By sending the person list, the normalized coordinates of each speaker determined by the person position determination process and information on whether each speaker is in the infield or outfield determined by the infield / outfield determination process are sent to the local-side remote communication device 1.

[0103] The local remote communication device 1 receives the sound data and the person list. The sound data and the person list are sent to the audio output control unit 22 via the receiving unit 21. The audio output control unit 22 performs a multiple facing speaker balancing process on the sound data using the person list, and specifies the speaker units 411 and 412 to output the background sound, infield speech sound, and outfield speech sound (step S9). In detail, the audio output control unit 22 adjusts the loudness of the background sound by dividing it by the number of pairs of speaker units 411 and 412, and plays the adjusted loudness on all speaker units 411 and 412. The audio output control unit 22 also controls the speaker units 411 and 412 closest to the speaker of the infield speech sound to play the infield speech sound as is. In addition, the audio output control unit 22 processes the outfield speech sound using a normal distribution function with a scaling factor, and limits the range of the speaker units 411 and 412 to be used, and causes the speaker units 411 and 412 to reproduce the outfield speech sound.

[0104] The speaker arrays 41 and 42 output background sounds, infield speech sounds, and outfield speech sounds using the speaker units 411 and 412 specified by the audio output control unit 22 (step S10).

[0105] 3. Remote Communication Processing Next, the flow of the remote communication processing will be described. Here, the description will be divided into a transmitting-side processing by the transmitting remote communication device 1 and a multiple opposing speaker balancing processing by the receiving remote communication device 1.

[0106] 9 is a flowchart of the transmission-side process. The overall flow of the transmission-side process will be described with reference to FIG.

[0107] The human sensing unit 11 acquires an image of the remote space captured by the camera 2. The audio separation unit 12 acquires audio of the remote space collected by the microphone 3 (step S11).

[0108] The person sensing unit 11 performs a person determination process using the video captured by the camera 2 to determine the normalized coordinates of the position of each speaker and to create a person-specific list having elements equal to the number of people and registering the normalized coordinates of each speaker (step S12).

[0109] Next, the human sensing unit 11 executes an infield / outfield determination process using the video and the person list to determine whether each speaker is in the infield or the outfield (step S13).

[0110] The voice separation unit 12 performs a voice separation process to separate the speech voice and the background sound contained in the voice input from the microphone 3 (step S14).

[0111] The clarity adjustment unit 132 receives the input of the background sound extracted by the audio separation unit 12. Then, the clarity adjustment unit 132 performs a degrading process as clarity adjustment processing to make the background sound more difficult to hear (step S15).

[0112] The individual sound separation unit 131 receives input of the speech extracted by the audio separation unit 12. The individual sound separation unit 131 also acquires a person-specific list from the person sensing unit 11. Next, the individual sound separation unit 131 generates threads equal to the number of elements registered in the acquired person-specific list (step S16). The individual sound separation unit 131 assigns IDs to each thread, sequentially numbered from 1, equal to the number of elements registered in the person-specific list. Here, the process is performed when the number of elements registered in the person-specific list is p.

[0113] Next, the individual sound separation unit 131 executes an individual sound separation process to separate the speech voice into individual sounds for each speaker (step S17).

[0114] The clarity adjustment unit 132 receives the input of the individual sounds from the individual sound separation unit 131. Next, the clarity adjustment unit 132 initializes i to 0 (step S18).

[0115] Next, the clarity adjustment unit 132 performs clarity adjustment processing on the individual sounds in the thread whose ID is i (step S19).

[0116] Next, the clarity adjustment unit 132 determines whether i is less than p (step S20). If i is less than p (step S20: Yes), the clarity adjustment unit 132 increments i by 1 (step S21). Then, the clarity adjustment unit 132 returns to step S19.

[0117] On the other hand, if i is greater than or equal to p (step S20: No), the clarity adjustment unit 132 outputs the background sound and individual sounds that have been subjected to clarity adjustment processing to the output integrating unit 14. The output integrating unit 14 also receives video input from the camera 2. Then, the output integrating unit 14 executes output integration processing to integrate the background sound and the individual sounds to generate one piece of sound data (step S22).

[0118] Thereafter, the transmitting unit 15 transmits the sound data generated by the output integration unit 14 and the video captured by the camera 2 to the local remote communication device 1 via the network 7 (step 23).

[0119] <3.1.1. Person Position Determination Processing> Fig. 10 is a flowchart of the person position determination processing. The processing shown in Fig. 10 is an example of the processing executed in step S12 in Fig. 9. Next, the flow of the person position determination processing will be described with reference to Fig. 10.

[0120] The human sensing unit 11 sets normalized coordinates in the horizontal direction of the image projected on the screen. Next, the human sensing unit 11 performs skeletal recognition on the speakers in the image to identify the position of the neck of each speaker and obtain the screen coordinates of the neck (step S101).

[0121] Next, the human sensing unit 11 converts the screen coordinates into normalized coordinates (step S102).

[0122] Thereafter, the human sensing unit 11 registers the screen coordinates and normalized coordinates of each speaker in the person list (step S103). Here, the human sensing unit 11 stores the horizontal coordinates when the screen coordinates and normalized coordinates of each image are projected onto the screen of the display 6.

[0123] <3.1.2. Infield / Outfield Determination Process> Figure 11 is a flowchart of the infield / outfield determination process. The process shown in Figure 11 is an example of the process executed in step S13 of Figure 9. Next, the flow of the infield / outfield determination process will be described with reference to Figure 11.

[0124] The human sensing unit 11 provides an infield flag item in the person list, initializes the list, and sets False to all items equal to the number of people present in the remote space (step S111).

[0125] Next, the human sensing unit 11 initializes i to 0 (step S112).

[0126] Next, the human sensing unit 11 detects the head direction of the ith speaker, with the top from the left side of the video as number 0 (step S113).

[0127] The human sensing unit 11 determines whether the head direction of the i-th speaker can be detected (step S114). If the head direction cannot be detected (step S114: No), the human sensing unit 11 proceeds to step S118.

[0128] On the other hand, if the head direction can be detected (step S114: Yes), the human sensing unit 11 determines whether the head direction is within the infield range (step S115).

[0129] If the head direction is within the infield range (step S115: Yes), the human sensing unit 11 stores True in the infield flag of the element corresponding to the i-th speaker in the person list (step S116).

[0130] On the other hand, if the head direction is outside the infield range (step S115: No), the human sensing unit 11 stores False in the infield flag of the element corresponding to the i-th speaker in the person list (step S117).

[0131] Thereafter, the human sensing unit 11 determines whether i is less than p (step S118). If i is less than p (step S118: Yes), the human sensing unit 11 increments i by 1 (step S119). Thereafter, the human sensing unit 11 returns to step S113.

[0132] On the other hand, if i is equal to or greater than p (step S118: No), the human sensing unit 11 ends the infield / outfield determination process.

[0133] <3.1.3. Individual Sound Separation Processing> Fig. 12 is a flowchart of the individual sound separation processing. The processing shown in Fig. 12 is an example of the processing executed in step S17 in Fig. 9. Next, the flow of the individual sound separation processing will be described with reference to Fig. 12.

[0134] The individual sound separation unit 131 creates threads for the number of people registered in the personal list. Next, the individual sound separation unit 131 stores the normalized coordinate values ​​registered in the personal list as thread-specific values ​​for each thread. Then, the individual sound separation unit 131 initializes i to 0 (step S121).

[0135] Next, the individual sound separation unit 131 calculates the beam direction of the i-th speaker, with the top from the left side of the video as 0th (step S122).

[0136] Next, the individual sound separation unit 131 acquires the individual sound of the i-th speaker by performing beamforming processing in the beam direction (step S123).

[0137] Thereafter, the individual sound separation unit 131 determines whether i is less than p (step S124). If i is less than p (step S124: Yes), the individual sound separation unit 131 increments i by 1 (step S125). Thereafter, the individual sound separation unit 131 returns to step S122.

[0138] On the other hand, if i is equal to or greater than p (step S124: No), the individual sound separation unit 131 ends the individual sound separation process.

[0139] <3.1.4. Clarity Adjustment Process> Fig. 13 is a flowchart of the clarity adjustment process for individual sounds. The process shown in Fig. 13 is an example of the process executed in step S19 in Fig. 9. Next, the flow of the clarity adjustment process will be described with reference to Fig. 13.

[0140] The clarity adjustment unit 132 initializes i to 0 (step S131).

[0141] Next, the clarity adjustment unit 132 determines whether the infield flag of the person list for the ith speaker, with the top speaker counted from the left side of the video as 0th, is True (step S132).

[0142] If the infield flag is True (step S132: Yes), the clarity adjustment unit 132 performs enhancement processing on the individual sounds of the i-th speaker (step S133).

[0143] If the infield flag is False (step S132: No), the clarity adjustment unit 132 performs a degrading process on the individual sound of the i-th speaker (step S134).

[0144] Thereafter, the clarity adjustment unit 132 determines whether i is less than p (step S135). If i is less than p (step S135: Yes), the clarity adjustment unit 132 increments i by 1 (step S136). Thereafter, the clarity adjustment unit 132 returns to step S132.

[0145] On the other hand, if i is equal to or greater than p (step S135: No), the clarity adjustment unit 132 ends the individual sound separation process.

[0146] 14 is a flowchart of the multiple opposed speaker balancing process. The overall flow of the multiple opposed speaker balancing process will be described with reference to FIG. 14. Here, the case will be described in which the audio output control unit 22 controls playback using a list for each combination of speaker units 411 and 412.

[0147] The audio output control unit 22 receives input of the sound data and the person list from the receiving unit 21. The audio output control unit 22 calculates the distance between the speaker units 411 (step S321). For example, if the distance between the speaker units 411 at both ends is Wspk and there are u speaker units 411 and 412, the audio output control unit 22 can calculate the inter-speaker distance Wu between the speaker units 411 as Wu = Wspk / u.

[0148] Next, the audio output control unit 22 calculates the normalized coordinates of each pair of speaker units 411 and 412, and stores the normalized coordinates in ascending order in a speaker unit list having elements equal to the number of pairs of speaker units 411 and 412 (step S322).

[0149] Next, the audio output control unit 22 provides an area for storing output sounds for each element of the speaker unit list and initializes each area with the number of channels of the background sound and the individual sounds (step S33). For example, if the background sound is data of one channel and there are p individual sounds, (1 + P) areas for storing sounds of (1 + P) channels are set in the area for storing output sounds.

[0150] Next, the audio output control unit 22 initializes i to 0 (step S34).

[0151] Next, the audio output control unit 22 determines whether the audio of the i-th channel, with the first of the (1+P) channels arranged in order in the audio data being numbered 0, is an individual sound of the speech (step S35).

[0152] If the audio of the i-th channel is background sound (step S35: No), the audio output control unit 22 adjusts the background sound (step S36). After that, the audio output control unit 22 proceeds to step S40.

[0153] On the other hand, if the audio of the i-th channel is an individual sound of speech (step S35: Yes), the audio output control unit 22 determines whether the speaker of the individual sound is Uchino using the personal list (step S37).

[0154] If the speaker is in the outfield (step S37: No), the audio output control unit 22 adjusts the outfield speech sound (step S38). After that, the audio output control unit 22 proceeds to step S40.

[0155] If the speaker is an infield speaker (step S37: Yes), the audio output control unit 22 performs infield speech sound adjustment (step S39).

[0156] Thereafter, the audio output control unit 22 stores the audio of the i-th channel in the area for storing the output audio of the speaker unit list (step S40).

[0157] Next, the audio output control unit 22 determines whether i is less than p (step S41). If i is less than p (step S41: Yes), the audio output control unit 22 increments i by 1 (step S42). Then, the audio output control unit 22 returns to step S35.

[0158] On the other hand, if i is equal to or greater than p (step S41: No), the audio output control unit 22 mixes the sounds stored in the area for storing output sounds in the speaker unit list, and then outputs the mixed sound from the speaker units 411 and 412 corresponding to the speaker unit list (step S43).

[0159] <3.2.1. Background Sound Adjustment> Fig. 15 is a flowchart of the background sound adjustment process. The process shown in Fig. 15 is an example of the process executed in step S36 in Fig. 14. Next, the flow of the background sound adjustment process will be described with reference to Fig. 15.

[0160] The audio output control unit 22 adjusts the amplitude of the background sound by a factor of (1 / u) (step S201), where u is the number of pairs of speaker units 411 and 412.

[0161] Then, the audio output control unit 22 assigns waveforms to the channels of all pairs of speaker units 411 and 412 (step S202).

[0162] 16 is a flowchart of the infield speech sound adjustment process. The process shown in FIG. 16 is an example of the process executed in step S39 in FIG. 14. Next, the flow of the infield speech sound adjustment process will be described with reference to FIG. 16.

[0163] The audio output control unit 22 executes localization center speaker unit derivation to identify a pair of speaker units 411 and 412 that are closest to the speaker (step S211). This localization center speaker unit derivation will be described in detail later.

[0164] Next, the audio output control unit 22 assigns a waveform to the pair of channels of the speaker units 411 and 412 that is closest to the normalized coordinates of the speaker (step S212).

[0165] 17 is a flowchart of the process of adjusting the outfield speech sound. The process shown in FIG. 17 is an example of the process executed in step S38 of FIG. 14. Next, the flow of the process of adjusting the outfield speech sound will be described with reference to FIG. 17.

[0166] The audio output control unit 22 performs localization center speaker unit derivation to identify the pair of speaker units 411 and 412 closest to the speaker (step S221). This localization center speaker unit derivation is the same process as the process in infield speech sound adjustment, and will be described in detail later. Here, when the positions of the pair of speaker units 411 and 412 are represented sequentially from the left, starting with 0, the position of the pair closest to the speaker is referred to as the "closest speaker unit position." Here, the number of pairs of speaker units 411 and 412 is u, and the closest speaker unit position is any one of 0 to u.

[0167] Next, the audio output control unit 22 initializes k to 0 (step S222), where k is a parameter for controlling the repetition of the process of determining the playback range.

[0168] Next, the audio output control unit 22 determines whether k=0 (step S223).

[0169] If k=0 (step S223: Yes), the audio output control unit 22 spk The audio output control unit 22 then proceeds to step S229.

[0170] If k is not 0 (step S223: No), the audio output control unit 22 determines whether the value obtained by adding k to the nearest speaker unit position is equal to or smaller than u (step S225). That is, the audio output control unit 22 determines whether the position obtained by adding k to the nearest speaker unit position extends beyond the right ends of the speaker arrays 41 and 42.

[0171] If the value obtained by adding k to the nearest speaker unit position is greater than u (step S225: No), the audio output control unit 22 proceeds to step S227. On the other hand, if the value obtained by adding k to the nearest speaker unit position is equal to or less than u (step S225: Yes), the audio output control unit 22 spk is set to the position obtained by adding k to the nearest speaker unit (step S226). After that, the audio output control unit 22 proceeds to step S227.

[0172] Next, the audio output control unit 22 determines whether the value obtained by subtracting k from the nearest speaker unit position is equal to or greater than 0 (step S227). That is, the audio output control unit 22 determines whether the position obtained by subtracting k from the nearest speaker unit position extends beyond the left ends of the speaker arrays 41 and 42.

[0173] If the value obtained by subtracting k from the nearest speaker unit position is less than 0 (step S227: No), the audio output control unit 22 proceeds to step S229. On the other hand, if the value obtained by subtracting k from the nearest speaker unit position is 0 or more (step S227: Yes), the audio output control unit 22 spk is set to the position obtained by subtracting k from the nearest speaker unit (step S228). After that, the audio output control unit 22 proceeds to step S229.

[0174] After that, the audio output control unit 22 calculates the amplitude of the audio at the position k as {g(x spk ) 1.66 The adjustment is performed by multiplying the x spk and the other x spk If both k positions exist, the audio output control unit 22 adjusts the amplitude of the audio for both k positions.

[0175] Thereafter, the audio output control unit 22 determines whether k is less than a predetermined range width (step S230).

[0176] If k is equal to or less than the range width (step S230: Yes), the audio output control unit 22 increments k by 1 (step S231), and then returns to step S233.

[0177] On the other hand, if k is greater than the range width (step S230: No), the audio output control unit 22 assigns the waveform to channels of the ±range width centered on the position of the closest speaker unit among the u channels (step S232).

[0178] 18 is a flowchart of the process of deriving the localization center speaker unit. The process shown in Fig. 18 corresponds to an example of the process executed in step S211 of Fig. 16 and step S221 of Fig. 17. Next, the flow of the process of deriving the localization center speaker unit will be described with reference to Fig. 18.

[0179] The audio output control unit 22 sets the initial value of the shortest distance to the speaker unit 411 as the distance between the speaker units 411 (step S241).

[0180] Next, the audio output control unit 22 initializes j to 0 (step S242), where j is a parameter for controlling the repetition of the process of determining whether the distance to the speaker for each pair of speaker units 411 and 412 is to be the nearest speaker unit distance.

[0181] Next, the audio output control unit 22 subtracts the position of the j-th speaker unit 411 when the sets of speaker units 411 are numbered consecutively from the left, starting with number 0, from the position of the speaker (step S243).

[0182] Next, the audio output control unit 22 determines whether the subtraction result is less than the shortest distance (step S244). If the subtraction result is equal to or greater than the shortest distance (step S244: No), the audio output control unit 22 proceeds to step S246.

[0183] If the subtraction result is less than the shortest distance (step S244: Yes), the audio output control unit 22 updates the shortest distance as the subtraction result. Furthermore, the audio output control unit 22 sets j as the nearest speaker unit position (step S245).

[0184] Next, the audio output control unit 22 determines whether j is less than u, which is the number of speaker units 411 (step S246).

[0185] If j is less than u (step S246: Yes), the audio output control unit 22 increments j by 1 (step S247), and then returns to step S243.

[0186] On the other hand, if j is equal to or greater than u (step S246: No), the audio output control unit 22 ends the derivation of the localization center speaker unit.

[0187] <4. Effects> As described above, the remote communication device 1 according to this embodiment identifies the positions of speakers from the video, estimates the head direction of each identified speaker, and determines whether each speaker is in the infield or outfield. The remote communication device 1 also separates background sound from the collected audio and acquires the individual sounds of the speakers based on the identified positions. The remote communication device 1 then reduces the volume of the background sound and outfield speech sounds and outputs them from a pair of speaker units 411 and 412 within a certain range. The remote communication device 1 also outputs the infield speech sounds from the speaker units 411 and 412 closest to the speaker's position.

[0188] In this way, by classifying speech sounds into infield speech sounds and outfield speech sounds, separate acoustic processing can be performed for each. Furthermore, by blurring the sense of sound localization of background sounds and outfield speech sounds to make them harder to hear, the infield speech sounds become relatively easier to hear. This makes it easier to hear the voice of the conversation partner while preserving the sense of presence and atmosphere, allowing for concentration on the conversation. Furthermore, by determining conversation participation through head direction estimation, it is possible to automatically determine speakers participating in communication. Furthermore, by processing the background sounds, infield speech sounds, and outfield speech sounds separately, signal processing based on participation levels becomes possible, and multi-channel speaker control can be performed in accordance with the background sounds, infield speech sounds, and outfield speech sounds. Therefore, smoother dialogue can be promoted.

[0189] 5. Remote communication device according to a second embodiment Next, a remote communication device 1 according to a second embodiment will be described. The remote communication device 1 according to the second embodiment outputs a sound without blurring the sense of position of an unexpected background sound among background sounds.

[0190] Fig. 19 is a block diagram of a remote communication device according to the second embodiment. In Fig. 19, the same components as those in Fig. 1 are denoted by the same reference numerals, and the following description will focus on the differences from Fig. 1, and may omit a description of the same components as those in Fig. 1.

[0191] 5.1. Acoustic Signal Processing Unit The acoustic signal processing unit 13 according to this embodiment includes a background sound separation unit 133 in addition to an individual sound separation unit 131 and a clarity adjustment unit 132 .

[0192] 5.1.1. Background Sound Separation Unit The background sound separation unit 133 receives input of the background sound separated by the audio separation unit 12. The background sound separation unit 133 then executes peak detection processing to detect peaks of the background sound. Next, the background sound separation unit 133 executes sudden sound extraction processing to extract the sudden sound from the background sound by estimating the direction of arrival (DoA) of the sudden sound to estimate the direction from which the sudden sound occurred. Here, the background sound separation unit 133 executes DoA estimation of the sudden sound by using a technique for visualizing audio, a method using a cross-correlation function of audio signals obtained from the microphone array 30, or the like.

[0193] Then, the background sound separation unit 133 outputs information about the estimated position of the sudden sound to the output integration unit 14. In addition, the background sound separation unit 133 outputs the extracted sudden sound to the output integration unit 14.

[0194] Furthermore, the background sound separation unit 133 executes a background sound extraction process to extract sounds other than sudden sounds from the background sound as steady sounds, and outputs the extracted steady sounds to the clarity adjustment unit 132. In this way, the background sound separation unit 133 separates the background sound into sudden sounds and steady sounds.

[0195] <5.1.2. Clarity Adjustment Unit> The clarity adjustment unit 132 receives an input of a stationary sound from the background sound from the background sound separation unit 133. The clarity adjustment unit 132 then performs a degradation process on the stationary sound from the background sound to make it harder to hear. The clarity adjustment unit 132 then outputs the clarity-adjusted stationary sound from the background sound to the output integration unit 14.

[0196] 5.2. Output Integration Unit The output integration unit 14 receives input of information about sudden sounds and the positions of the sudden sounds in the background sound from the background sound separation unit 133. The output integration unit 14 also receives input of steady sounds in the background sound from the clarity adjustment unit 132. In this case, the output integration unit 14 acquires individual sounds of the same number of channels as the number of people present in the remote space, as well as one channel of data for each of the sudden sounds and steady sounds in the background sound. In other words, if the number of people present in the remote space is p, the output integration unit 14 acquires (2+p) channels of data.

[0197] The output integrating unit 14 then integrates each individual sound, the sudden sound from the background sound, and the steady sound from the background sound to generate one piece of audio data. The output integrating unit 14 then adds speaker information in the person list corresponding to each individual sound. The output integrating unit 14 also associates the sudden sound with positional information of the sudden sound. The output integrating unit 14 then outputs the audio data, the person list, and the positional information of the sudden sound to the transmitting unit 15, causing them to be transmitted to the local remote communication device 1.

[0198] 5.3. Audio Output Control Unit The audio output control unit 22 in the local remote communication device 1 receives input of audio data including each individual sound, a sudden sound from the background sound, and a steady sound from the background sound. The audio output control unit 22 also receives input of a person list and position information of the sudden sound.

[0199] The audio output control unit 22 acquires the sudden sound of the background sound and the steady sound of the background sound from the sound data. Next, the audio output control unit 22 calculates the amplitude of the steady sound of the background sound by (1 / u) 1.66 Then, the audio output control unit 22 controls all the speaker units 411 of the speaker array 41 and all the speaker units 412 of the speaker array 42 to output steady sounds with reduced loudness.

[0200] On the other hand, for a sudden sound in the background sound, the audio output control unit 22 identifies the speaker units 411 and 412 that are closest to the position of the sudden sound. Then, the audio output control unit 22 causes the identified speaker units 411 and 412 to reproduce the sudden sound in the background sound as is.

[0201] Here, in this embodiment, the audio output control unit 22 applies the same processing to the sudden sound as to the infield speech sound in the first embodiment to output the sound, but it may also apply the same processing to the sudden sound as to the outfield speech sound in the first embodiment to output the sound.

[0202] 20 is a diagram showing an outline of the data flow between the local remote communication devices according to the second embodiment. Here, too, the case where data is transmitted from the transmitting unit 10 of the remote remote communication device 1 to the receiving unit 20 of the local remote communication device 1 will be described. Here, the transmission of audio will be described.

[0203] The microphone 3 inputs the sound picked up in the remote space to the remote communication device 1 (step S301). Here, for example, data of the s channel exists as data of the microphone array 30. The camera 2 also inputs the video generated by capturing the remote space to the remote communication device 1 (step S302).

[0204] The image input from camera 2 is sent to the human sensing unit 11. Using the image captured by camera 2, the human sensing unit 11 executes sensing processing (step S303) including human position determination processing (step S331) and infield / outfield determination processing (step S332).

[0205] The sound input from the microphone 3 is sent to the sound separation unit 12. The sound separation unit 12 performs a sound separation process to separate the speech sound and background sound contained in the sound input from the microphone 3 (step S304).

[0206] The background sound extracted by the audio separation unit 12 is sent to the background sound separation unit 133 of the audio signal processing unit 13. The background sound separation unit 133 executes peak detection processing to detect peaks of the background sound (step S305).

[0207] Next, the background sound separation unit 133 executes a sudden sound extraction process to estimate the DoA of the sudden sound and extract the sudden sound from the background sound (step S306). The sudden sound in the background sound and position information of the sudden sound are sent to the output integration unit 14.

[0208] The background sound separator 133 then executes a steady sound extraction process to extract steady sounds from the background sound (step S307). The steady sounds from the background sound are sent to the clarity adjuster 132.

[0209] The clarity adjustment unit 132 performs a degrading process to make the background sound less audible as the clarity adjustment process (step S308).

[0210] Furthermore, the individual sounds and person list, which are s-channel data extracted by the voice separation unit 12, are sent to the audio signal processing unit 13. The audio signal processing unit 13 performs speech signal processing (step S309), including individual sound separation processing by the individual sound separation unit 131 (step S391) and clarity adjustment processing by the clarity adjustment unit 132 (step S392).

[0211] The steady sound, which is one channel of data, the sudden sound, and the position information of the sudden sound, which is one channel of data, are input to the output integrating unit 14. In addition, the individual sounds, which are p channel of data that have been subjected to speech signal processing, are also input to the output integrating unit 14. As a result, the output integrating unit 14 obtains (2+P) channel of data. The output integrating unit 14 then executes output integrating processing to integrate the sounds of all channels to generate one sound data (step S310).

[0212] The sound data, the person list, and the position information of the sudden sound generated by the output integration unit 14 are sent by the transmission unit 15 to the local-side remote communication device 1 via the network 7 (step S311).

[0213] The local remote communication device 1 receives the sound data, the person list, and the position information of the sudden sound. The sound data, the person list, and the position information of the sudden sound are sent to the audio output control unit 22 via the receiving unit 21. The audio output control unit 22 performs a multiple facing speaker balancing process on the sound data using the person list and the position information of the sudden sound. The audio output control unit 22 then designates and plays back the speaker units 411 and 412 that will output the background sound, the infield speech sound, and the outfield speech sound (step S312). At this time, the audio output control unit 22 adjusts the loudness of the steady sound among the background sound to a level obtained by dividing the loudness by the number of pairs of speaker units 411 and 412, and plays it back through all speaker units 411 and 412. The audio output control unit 22 also plays back the sudden sound among the steady sound as is through the speaker units 411 and 412 closest to the position of the sudden sound.

[0214] The speaker arrays 41 and 42 use the speaker units 411 and 412 specified by the audio output control unit 22 to output sudden sounds from the background sounds, steady sounds from the background sounds, infield speech sounds, and outfield speech sounds (step S313).

[0215] 7. Effects As described above, the remote communication device 1 according to the present embodiment divides background sounds into two types, sudden sounds and steady sounds, and transmits position information for the background sounds. The remote communication device 1 then outputs the sudden sounds among the background sounds without blurring their positioning, and outputs the steady sounds among the background sounds with blurred positioning.

[0216] This allows background sounds to be played back louder than steady sounds from the vicinity of the location where they were generated, preventing sudden background sounds from losing their sense of position and causing discomfort, making conversations more natural and facilitating smoother conversations.

[0217] 8. Remote Communication Device According to a Third Embodiment Next, a remote communication device 1 according to a third embodiment will be described. The remote communication device 1 according to the third embodiment converts the speech voice picked up by the microphone 3 into a mechanical voice to make it easier to hear.

[0218] Fig. 21 is a block diagram of a remote communication device according to the third embodiment. In Fig. 21, the same components as those in Fig. 1 are denoted by the same reference numerals, and the following description will focus on the differences from Fig. 1, and may omit a description of the same components as those in Fig. 1.

[0219] 8.1. Acoustic Signal Processing Unit The acoustic signal processing unit 13 according to this embodiment includes an individual sound separation unit 131, an articulation adjustment unit 132, and a voice synthesis unit 134.

[0220] 8.1.1. Speech Synthesis Unit The speech synthesis unit 134 receives input of the individual sounds of each speaker separated by the individual sound separation unit 131. The speech synthesis unit 134 then performs speech recognition processing to recognize the speech and transcribe the speech into text. Next, the speech synthesis unit 134 performs speech synthesis processing to synthesize the textual speech and regenerate the individual sounds. The speech synthesis unit 134 then outputs the regenerated individual sounds of each speaker to the clarity adjustment unit 132.

[0221] In this way, the voice synthesis unit 134 performs voice recognition for each individual sound to generate characters indicating the speech content of the individual sound, and performs voice synthesis based on the characters indicating the speech content to reproduce the speech of the speaker. In this case, the clarity adjustment unit 132 makes adjustments for each individual sound reproduced by the voice synthesis unit 134 based on whether the speaker is participating in the conversation.

[0222] 22 is a diagram showing an outline of data flow between local remote communication devices according to the third embodiment. Here, too, the case where data is transmitted from the transmitting unit 10 of the remote remote communication device 1 to the receiving unit 20 of the local remote communication device 1 will be described. Here, the transmission of audio will be described.

[0223] The microphone 3 inputs the sound picked up in the remote space to the remote communication device 1 (step S401). Here, for example, data of the s channel exists as data of the microphone array 30. The camera 2 also inputs the video generated by capturing an image of the remote space to the remote communication device 1 (step S402).

[0224] The video input from camera 2 is sent to the human sensing unit 11. Using the video captured by camera 2, the human sensing unit 11 executes sensing processing (step S403), including human position determination processing (step S431) and infield / outfield determination processing (step S432).

[0225] The audio input from the microphone 3 is sent to the audio separation unit 12. The audio separation unit 12 performs an audio separation process to separate the audio input from the microphone 3 into speech and background sounds (step S404).

[0226] The background sound extracted by the audio separation unit 12 is sent to the clarity adjustment unit 132 of the audio signal processing unit 13. The clarity adjustment unit 132 performs a degrading process to make the background sound more difficult to hear as clarity adjustment processing (step S405).

[0227] Furthermore, the individual sounds and the person list, which are s-channel data extracted by the voice separation unit 12, are sent to the acoustic signal processing unit 13. The acoustic signal processing unit 13 performs speech signal processing on the speech voice using the person list (step S406).

[0228] In detail, the individual sound separation unit 131 creates p threads, which is the number of elements registered in the acquired personal list. Furthermore, the individual sound separation unit 131 stores the normalized coordinate values ​​registered in the personal list as thread-specific values ​​for each thread. Next, the individual sound separation unit 131 calculates the angle θ to each speaker. Then, the individual sound separation unit 131 performs individual sound separation processing on the speech sounds extracted by speech separation by the speech separation unit 12 using beamforming according to the angle to each speaker, thereby generating individual sounds for each speaker for each thread (step S461).

[0229] The individual sounds generated by the individual sound separation unit 131 are sent to the speech synthesis unit 134. The speech synthesis unit 134 executes speech recognition processing to recognize the individual sounds and transcribe the speech content into text (step S462).

[0230] Next, the speech synthesis unit 134 executes a speech synthesis process to synthesize the text of the speech content and regenerate the individual sounds (step S463). The individual sounds of each speaker regenerated by the speech synthesis unit 134 are sent to the clarity adjustment unit 132.

[0231] The clarity adjustment unit 132 determines whether the speaker corresponding to the individual sound for each thread is infield or outfield using the infield flag in the person list. Then, the clarity adjustment unit 132 performs clarity adjustment processing to enhance the individual sounds of infield speakers to make them easier to hear, and to degrade the individual sounds of outfield speakers to make them harder to hear (step S464).

[0232] The background sound, which is one channel of data, is input to the output integrating unit 14. The individual sounds, which are p channel of data that have been subjected to speech signal processing, are also input to the output integrating unit 14. As a result, the output integrating unit 14 acquires (1+P) channel of data. The output integrating unit 14 then executes output integration processing to integrate the sounds of all channels to generate one piece of sound data (step S407).

[0233] The sound data and the person list generated by the output integration unit 14 are sent by the transmission unit 15 to the local remote communication device 1 via the network 7 (step S408).

[0234] The local remote communication device 1 receives the sound data and the person list. The sound data and the person list are sent to the audio output control unit 22 via the receiving unit 21. The audio output control unit 22 performs a multiple facing speaker balancing process on the sound data using the person list, and specifies the speaker units 411 and 412 to output the background sound, the infield speech sound, and the outfield speech sound, respectively, and plays them (step S409).

[0235] The speaker arrays 41 and 42 use the speaker units 411 and 412 specified by the audio output control unit 22 to output sudden sounds from the background sounds, steady sounds from the background sounds, infield speech sounds, and outfield speech sounds (step S410).

[0236] <10. Effects> As described above, the remote communication device 1 according to the present embodiment performs speech recognition to transcribe the spoken content into text, and then synthesizes the text to produce data of the spoken voice to be played back.

[0237] When speech recorded by the microphone 3 is played back as is, it can be difficult to hear due to differences in voice volume and articulation between people. In contrast, the remote communication device 1 according to this embodiment can suppress variations in voice volume and articulation, making speech easier to hear overall, thereby facilitating smooth communication.

[0238] <11. Remote Communication Device According to a Fourth Embodiment> Next, a remote communication device 1 according to a fourth embodiment will be described. The remote communication device 1 according to the fourth embodiment performs voice recognition, such as infield / outfield determination, using a method other than head direction estimation. The remote communication device 1 according to this example is also represented by the block diagram of FIG. 1. The following description will focus on parts that are different from FIG. 1, and descriptions of parts that are the same as those in FIG. 1 may be omitted.

[0239] <11.1. Person Sensing Unit> The person sensing unit 11 performs depth sensing to detect the distance of each speaker from the video. Then, under the assumption that people far from camera 2 will not converse with remote people, the person sensing unit 11 determines that people far from camera 2 are people in the outfield. In this way, the person sensing unit 11 determines the distance of the speaker based on the video and determines whether the speaker is participating in the conversation based on the distance.

[0240] In this way, by determining whether the location is infield or outfield using the distance from camera 2, it is possible to avoid a situation where someone who happens to be passing behind you and is talking while looking at the screen is determined to be a participant in a conversation with a local person, thereby facilitating smooth communication.

[0241] 11.2. Acoustic Signal Processing Unit The acoustic signal processing unit 13 performs speech recognition on each individual sound generated by the individual sound separation unit 131 and transcribes the speech content into text. The acoustic signal processing unit 13 then determines whether the speech is directed to a person on the local side based on the context of the transcribed speech content. The acoustic signal processing unit 13 can make this determination, for example, by performing AI analysis of emotions, empathy level, and listening level.

[0242] If the acoustic signal processing unit 13 determines that the utterance is directed to a person on the local side, it determines that the speaker is an infield person.The acoustic signal processing unit 13 then updates the person list by setting the infield flag of the speaker determined to be an infield person to True.The acoustic signal processing unit 13 then outputs the updated person list to the clarity adjustment unit 132.

[0243] Here, for example, the voice of a person giving a presentation while looking at slides, or the interjections of a person listening while taking notes, are also messages sent to the remote, and are preferably processed as infield. Therefore, by determining infield / outfield based on the context, the speech of a speaker who is actually participating in the conversation but who would be determined to be infield in head direction estimation is transmitted to the remote and played back as the speech of an infield person. This allows for processing so that speech is easier to hear as infield speech, even when the speaker is not facing the remote but is speaking to it. This facilitates communication.

[0244] <11.3. Other Voice Recognition> Additionally, when a local person points to a specific position on the screen of the display 6, the local-side human sensing unit 11 extracts the specific position using motion capture, gaze estimation, or the like. Then, the local-side human sensing unit 11 transmits information about the specific position pointed to via the transmission unit 15 to the remote-side remote communication device 1.

[0245] In this case, for example, the individual sound separation unit 131 extracts the sound at the specific position as an individual sound using beamforming or the like. The clarity adjustment unit 132 applies appropriate processing to the individual sound at the specific position, such as the same processing as that applied to infield speech. The output integration unit 14 then combines the extracted sound and the other individual sounds into one data piece, and transmits the combined data together with information about the specific position to the local remote communication device 1 via the transmission unit 15.

[0246] The audio output control unit 22 of the local remote communication device 1 extracts individual sounds at specific positions from the sound data and performs appropriate multiple opposing speaker balancing processing, such as the same processing as for infield speech, using the information on the specific positions.The audio output control unit 22 then causes the speaker units 411 and 412 to reproduce the processed individual sounds at the specific positions.

[0247] <11.4. Effects> In this way, by emphasizing and reproducing the sound at a specific point compared to background sounds, it is possible to make sounds that local people want to hear easier to hear, in addition to spoken voices, and to facilitate smooth communication.

[0248] <12. Remote Communication Device According to Fifth Embodiment> Next, a remote communication device 1 according to a fifth embodiment will be described. In the first embodiment, the camera 2 and microphone 3 are aligned in the normal direction of the virtual screen, and the human sensing unit 11 determines the angle of the direction of each speaker based on normalized coordinates indicating the position of each speaker in the image captured by the camera 2. In contrast, the remote communication device 1 according to the fifth embodiment identifies the direction of a sound source using the sound captured by the microphone 3. The remote communication device 1 according to this example is also represented by the block diagram of FIG. 1. The following description will focus on parts that are different from FIG. 1, and descriptions of parts that are the same as those in FIG. 1 may be omitted.

[0249] Fig. 23 is a block diagram of a remote communication device according to the fifth embodiment, and Fig. 24 is a diagram for explaining audio signal processing by the remote communication device according to the fifth embodiment.

[0250] 12.1. Sound Source Direction Estimation Unit The remote communication device 1 according to this embodiment has a sound source direction estimation unit 16 as shown in Fig. 23. The sound source direction estimation unit 16 acquires sounds obtained by directing beams in each direction in turn using the microphone array 30. The direction to direct the beams may be preset in the microphone 3, or the sound source direction estimation unit 16 may specify the direction to direct the beams in turn to the microphone 3.

[0251] Here, we assume that beamforming detects only speech and sudden background sounds, and does not pick up steady background sounds. We also assume that no sounds are generated other than the speaker appearing on the video and any sudden background sounds occurring within that range, i.e., no sounds are heard from outside the field of view.

[0252] Next, as shown in Fig. 24, the sound source direction estimation unit 16 selects a sound whose output is greater than a threshold from among the sounds obtained by directing a beam in each direction, and estimates the sound source direction by assuming that the sound source is located in the direction of the beam that obtained the selected sound (step S501). Step S501 in Fig. 24 shows that a large number of sound source directions have been estimated. The sound source direction estimation unit 16 then determines the sound source direction as the position where the speaker is located. In this way, the sound source direction estimation unit 16 determines the position of the speaker based on the sound in a predetermined space that corresponds to the remote space.

[0253] Thereafter, the sound source direction estimation unit 16 notifies the human sensing unit 11 of the normalized coordinates of the positions of the determined speakers, and creates a person list based on the normalized coordinates of the positions of the speakers. In this case, the human sensing unit 11 identifies speakers from among people appearing in the video using the notified normalized coordinates, performs infield / outfield determination for each identified speaker, and registers them in the person list.

[0254] 12.2. Acoustic signal processing unit The acoustic signal processing unit 13 acquires a person list in which normalized coordinates of the position of each speaker estimated from the voice by the sound source direction estimation unit 16 are registered from the person sensing unit 11. Then, the acoustic signal processing unit 13 executes speech signal processing using the acquired person list (step S502).

[0255] More specifically, the individual sound separation unit 131 generates threads for the number of speakers estimated from the voices registered in the person list. Then, the individual sound separation unit 131 performs individual sound separation processing using normalized coordinates of the position of each speaker for each thread to generate individual sounds for each speaker (step S521). In this way, the individual sound separation unit 131 separates individual sounds, which are the speech sounds of each speaker, from the voice in the specified space corresponding to the remote space, based on the positions of the speakers determined by the sound source direction estimation unit 16.

[0256] The individual sound separation unit 131 also performs individual sound separation for the sound source direction detected using the microphone array 30, starting from the angle on the left side of the image displayed on the display 6, thereby associating the individual sounds with normalized coordinates indicating the position of each speaker registered in the person list. This makes it possible to match the image and sound even after transmission to the local side.

[0257] The clarity adjustment unit 132 performs clarity adjustment processing for the individual sounds of the speakers estimated from the speech, for each thread, using the infield flag of each speaker registered in the person list (step S522). In this way, the clarity adjustment unit 132 adjusts the clarity for each individual sound separated by the individual sound separation unit 131, based on whether the speaker is participating in the conversation.

[0258] <12.3. Effects> As a result, even if it is practically difficult to install the camera 2 and the microphone 3 side by side in the normal direction of the virtual screen, individual sounds can be generated without any restrictions on the positional relationship between the camera 2 and the microphone 3.

[0259] 13. Remote Communication Device According to a Sixth Embodiment Next, a remote communication device 1 according to a sixth embodiment will be described. When the width of the speaker arrays 41 and 42 is shorter than the width of the screen of the display 6, the remote communication device 1 according to this embodiment remaps the audio playback position to match the width of the display 6. The remote communication device 1 according to this example is also represented by the block diagram of FIG. 1. The following description will focus on the parts that are different from FIG. 1, and descriptions of parts that are the same as those in FIG. 1 may be omitted.

[0260] 25 is a diagram for explaining audio reproduction by a remote communication device according to the sixth embodiment. In this embodiment, as shown in FIG. 25, the speaker array width W spk is the display width W, which is the width of the screen of the display 6 disp is shorter than.

[0261] In the first embodiment, the audio output control unit 22 adjusts the speaker array width W spk and speaker distance W u the display width W disp However, in this method of balancing multiple opposing speakers, the speaker array width W spk is the display width W disp If the time is shorter than 100 ms, it becomes difficult to reproduce the voice of a speaker who appears at the edge of the screen of the display 6.

[0262] Therefore, the audio output control unit 22 according to the present embodiment uses the speaker array width W spk and speaker distance W u is the speaker array width W spk In this case, the width of the speaker arrays 41 and 42 is 1 in the normalized coordinate system, and the width between the pair of speaker units 411 and 412 is the speaker distance W u is the speaker array width W spk That is, the audio output control unit 22 sets the speaker array width in processing to -0.5 to 0.5 of the normalized coordinates. Then, the audio output control unit 22 executes a multiple facing speaker balancing process using this processing distance to determine the pair of speaker units 411 and 412 that will output each sound.

[0263] In this way, the audio output control unit 22 selects the speaker units 411 and 412 to be played back based on the position of the speaker so that the playback position of the spoken voice in the speaker arrays 41 and 42 maintains the distance relationship between the positions of each speaker on the screen.

[0264] <13.2. Effects> In this way, the remote communication device 1 according to the present embodiment remaps the audio playback position to match the width of the speaker arrays 41 and 42. This makes it possible to play back the voices of all speakers appearing on the screen. Therefore, even if the width of the speaker arrays 41 and 42 is significantly shorter than the screen of the display 6, it is possible to prevent the voices of speakers speaking at the edges of the screen from being cut off, and all speakers appearing on the screen can be played back from the speaker arrays 41 and 42.

[0265] 14. Remote Communication Device According to Seventh Embodiment Next, a remote communication device 1 according to a seventh embodiment will be described. The multiple opposed speakers 4 virtually localize the sound source in the vertical center of the image by reproducing the same monaural sound from both the top and bottom. However, this may result in an unnatural sound when reproducing the speech of a short speaker such as a child or a tall speaker.

[0266] Therefore, the remote communication device 1 according to this embodiment changes the vertical position of the sound source according to the height of the sound source and plays back the sound. The remote communication device 1 according to this embodiment is represented by the block diagram of Fig. 23. The following description will focus on the parts that are different from Fig. 23, and descriptions of the parts that are the same as those in Fig. 1 may be omitted. However, in this embodiment, the sound source direction estimation unit 16 does not need to estimate the position of each speaker in the normalized coordinate direction. Fig. 26 is a diagram for explaining sound playback by a remote communication device according to the seventh embodiment.

[0267] <14.1. Person Sensing Unit> The person sensing unit 11 acquires the vertical position of the mouth of each speaker shown in the video. For example, the person sensing unit 11 can estimate the position of the mouth through video analysis. The person sensing unit 11 then stores the normalized vertical coordinate Y in FIG. 26 for each speaker in the person list. In this case, as shown in FIG. 26, the person sensing unit 11 sets the vertical normalized coordinate Y with the center of the screen as the origin and the vertical coordinate between -0.5 and 0.5.

[0268] <14.2. Individual Sound Separation Unit> Furthermore, the sound source of background sounds other than speech, such as sudden sounds, is not necessarily located at the center of the image. In this regard, when identifying a sudden sound as in the second embodiment, the background sound separation unit 133 identifies the vertical position of the sudden sound. In this case, for example, the background sound separation unit 133 transmits and receives signals using a signal line connected to the microphone array 30 via the individual sound separation unit 131 and the audio interface 5. The background sound separation unit 133 can cause the microphone array 30 to scan a beam in the vertical direction and estimate the vertical position of the sound source of the sudden sound from the vertical sound intensity. The background sound separation unit 133 then represents the position of the sudden sound in the background sound using normalized coordinates Y and transmits this information to the remote communication device 1 via the output integration unit 14 and the transmission unit 15.

[0269] Additionally, although the configuration has been described here in which the background sound separation unit 133 operates the microphone array 30, the remote communication device 1 of the second embodiment may also be equipped with the sound source direction estimation unit 16 shown in Fig. 23. In this case, the vertical position of the sound source estimated by the sound source direction estimation unit 16 using the microphone array 30 is sent to the person sensing unit 11 and registered in the person list.

[0270] 14.3 Audio Output Control Unit The audio output control unit 22 stereophonizes the audio signal, for example, by using the speaker array 41 as the L channel and the speaker array 42 as the R channel. The audio output control unit 22 then pans the stereophonized audio signal using the L channel and the R channel in accordance with the vertical normalized coordinate Y registered in the person list, and changes the sound source position so that the speech sound is reproduced at a specified vertical position.

[0271] For example, the audio output control unit 22 reproduces the speech of the speaker P1 in Fig. 26 from a position below the center in the vertical direction as the sound source. Also, the audio output control unit 22 reproduces the speech of the speaker P1 from a position above the center in the vertical direction as the sound source.

[0272] Furthermore, when a sudden sound in the background sound is reproduced separately from the steady sound as in the second embodiment, the audio output control unit 22 may perform the following process. That is, the audio output control unit 22 also pans the stereo audio signal for the vertical position of the sudden sound, according to the normalized coordinate Y of the vertical position of the sudden sound notified by the individual sound separation unit 131. As a result, the audio output control unit 22 changes the sound source position so that the sudden sound is reproduced at the specified vertical position. In this way, the audio output control unit 22 stereophonically reproduces the sound in a predetermined space that corresponds to the remote space, and adjusts the position of the sound source between the speaker array 41 and the speaker array 42.

[0273] <14.4. Effects> In monaural playback, a virtual sound source is positioned in the vertical center of the screen. In contrast, the remote communication device 1 according to this embodiment can move the sound source position to an appropriate position in the vertical direction when there is a difference in the height of speakers or when the sound being emitted is far from the vertical center. This makes the conversation more natural, and facilitates smooth communication.

[0274] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.

[0275] 15. Hardware Configuration FIG. 27 is a hardware configuration diagram showing an example of a computer that realizes the arithmetic unit of the remote communication device that is the information processing device according to the first to seventh embodiments.

[0276] The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected to each other via a bus 1050.

[0277] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0278] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0279] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an application program according to the present disclosure, which is an example of program data 1450.

[0280] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0281] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display 6, an audio interface 5, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, or semiconductor memory.

[0282] Although the CPU 1100 reads and executes the program data 1450 from the HDD 1400, as another example, the CPU 1100 may obtain these programs from other devices via an external network 1550.

[0283] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.

[0284] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0285] The present technology can also be configured as follows.

[0286] (1) An information processing device comprising: a person sensing unit that determines whether each speaker is participating in a conversation based on video of multiple speakers present in a predetermined space captured by a camera, a clarity adjustment unit that adjusts clarity of speech sounds of the speakers participating in the conversation determined by the person sensing unit and included in audio of the predetermined space collected by a microphone, and a transmission unit that transmits the audio of the predetermined space processed by the clarity adjustment unit. (2) The person sensing unit further comprises: an individual sound separation unit that determines a position of each speaker based on the video, and separates individual sounds that are speech sounds of each speaker from the audio of the predetermined space based on the positions of each speaker determined by the person sensing unit, and the clarity adjustment unit adjusts clarity for each of the individual sounds separated by the individual sound separation unit based on whether each speaker is participating in the conversation. (3) The information processing device according to (2), further comprising: a receiving unit that receives the audio of the predetermined space transmitted by the transmitting unit; and an audio output control unit that selects, in a speaker having two speaker arrays in which a plurality of speaker units are arranged in a row, a speaker unit that reproduces the speech of each speaker from the audio of the predetermined space received by the receiving unit based on the position of each speaker, and causes the selected speaker unit to reproduce the speech of each speaker. (4) The information processing device according to (3), wherein the receiving unit receives the video together with the audio of the predetermined space and projects the video on a screen on which the speaker arrays are arranged, one at each end of the video in a vertical direction when the video is projected, and the audio output control unit selects, for each speaker projected on the screen, the speaker unit to reproduce the speech of the speaker based on the position of the speaker so that the speech is reproduced near the position on the screen.(5) The information processing device according to (3), wherein the receiving unit receives the video together with audio from the predetermined space and projects the video on a screen on which the speaker array is arranged, one at each end of the video in a vertical direction when the video is projected, and the audio output control unit selects the speaker unit to be reproduced based on the position of the speaker so that the reproduction position of the speech audio in the speaker array maintains a distance relationship between the positions of the speakers on the screen. (6) The information processing device according to any one of (1) to (5), wherein the clarity adjustment unit performs clarity reduction adjustment on speech audio from speakers not participating in the conversation, which is included in the audio from the predetermined space, to reduce the clarity compared to a state where the microphones are collecting the speech audio. (7) The information processing device according to any one of (1) to (6), further comprising an audio separation unit that separates background audio from the audio from the predetermined space, and the clarity adjustment unit performs clarity reduction adjustment on the background audio separated by the audio separation unit to reduce the clarity of the background audio relative to other audio. (8) The information processing device according to any one of (7), further comprising a background sound separation unit that separates the background sound into a sudden sound and a steady sound, and the clarity adjustment unit performs the clarity reduction adjustment on the steady sound of the background sound. (9) The information processing device according to any one of (2) to (5), further comprising a voice synthesis unit that performs voice recognition for each of the individual sounds to generate characters indicating the speech content of the individual sounds and performs voice synthesis based on the characters indicating the speech content to reproduce the speech of a speaker, and the clarity adjustment unit adjusts the clarity for each of the individual sounds reproduced by the voice synthesis unit based on whether the speaker is participating in a conversation. (10) The information processing device according to any one of (1) to (9), further comprising a voice synthesis unit that performs voice recognition for each of the individual sounds to generate characters indicating the speech content of the individual sounds and performs voice synthesis based on the characters indicating the speech content to reproduce the speech of a speaker, and the clarity adjustment unit adjusts the clarity for each of the individual sounds reproduced by the voice synthesis unit based on whether the speaker is participating in a conversation.(11) The information processing device according to any one of (1) to (10), further comprising a sound source direction estimation unit that determines a position of a speaker based on a sound in the predetermined space, wherein the individual sound separation unit separates individual sounds that are speech sounds of each speaker from the sound in the predetermined space based on the position of the speaker determined by the sound source direction estimation unit, and the clarity adjustment unit adjusts clarity for each of the individual sounds separated by the individual sound separation unit based on whether the speaker is participating in a conversation. (12) The information processing device according to (3), wherein the audio output control unit stereophonically converts the sound in the predetermined space and adjusts the position of the sound source between the speaker arrays. (13) An information processing method in which an information processing device executes the following steps: a person sensing step of determining whether each speaker is participating in a conversation based on video captured by a camera that captures multiple speakers present in a predetermined space, a clarity adjustment step of adjusting the clarity of speech sounds of the speakers participating in the conversation determined in the person sensing step, which speech sounds are included in audio of the predetermined space collected by a microphone that collects audio of the predetermined space, and a transmission step of transmitting the audio of the predetermined space processed in the clarity adjustment step. (14) An information processing program that causes a computer to execute processes of: determining whether each speaker is participating in a conversation based on video captured by a camera that captures multiple speakers present in a predetermined space, adjusting the clarity of speech sounds of the speakers participating in the conversation, which speech sounds are included in audio of the predetermined space collected by a microphone that collects audio of the predetermined space, and transmitting the audio of the predetermined space with the adjusted speech sounds of the speakers participating in the conversation.(15) An information processing system comprising: a transmitting device having a human sensing unit that determines whether each speaker in a specified space is participating in a conversation based on images of the speakers captured by a camera; a clarity adjustment unit that adjusts clarity of the speech of the speakers who are participating in the conversation and that is included in the audio of the specified space picked up by a microphone; and a transmitting unit that transmits the audio of the specified space processed by the clarity adjustment unit; and a receiving device having a receiving unit that receives the audio of the specified space transmitted by the transmitting unit; and an audio output control unit that selects, in a speaker having two speaker arrays in which a plurality of speaker units are arranged in a row, a speaker unit that will play the speech of each speaker from the audio of the specified space received by the receiving unit based on the positions of each speaker, and causes the selected speaker unit to play the speech of each speaker.

[0287] REFERENCE SIGNS LIST 1 Remote communication device 2 Camera 3 Microphone 4 Multiple opposing speakers 5 Audio interface 6 Display 7 Network 10 Transmission unit 11 Human sensing unit 12 Voice separation unit 13 Acoustic signal processing unit 14 Output integration unit 15 Transmission unit 16 Sound source direction estimation unit 20 Receiving unit 21 Receiving unit 22 Voice output control unit 30 Microphone array 41, 42 Speaker array 131 Individual sound separation unit 132 Clarity adjustment unit 133 Background sound separation unit 134 Voice synthesis unit 411, 412 Speaker unit

Claims

1. An information processing device having a human sensing unit that determines whether each speaker is participating in a conversation based on images of multiple speakers present in a specified space captured by a camera; a clarity adjustment unit that adjusts the clarity of the speech of the speakers who are participating in the conversation and determined by the human sensing unit to be included in the audio of the specified space picked up by a microphone; and a transmission unit that transmits the audio of the specified space processed by the clarity adjustment unit.

2. The information processing device of claim 1, further comprising: an individual sound separation unit that determines the position of each speaker based on the video; and separates individual sounds, which are the speech sounds of each speaker, from the sound in the specified space based on the position of each speaker determined by the individual sound separation unit; and the clarity adjustment unit adjusts the clarity of each individual sound separated by the individual sound separation unit based on whether each speaker is participating in a conversation.

3. An information processing device as described in claim 2, further comprising: a receiving unit that receives the audio of the specified space transmitted by the transmitting unit; and an audio output control unit that selects, in a speaker having two speaker arrays in which a plurality of speaker units are arranged in a row, a speaker unit that will play the speech of each speaker from the audio of the specified space received by the receiving unit based on the position of each speaker, and plays the speech of each speaker on the selected speaker unit.

4. The information processing device described in claim 3, wherein the receiving unit receives the video together with the audio from the specified space and projects the video on a screen on which one speaker array is arranged at each end of the video in the vertical direction when the video is projected, and the audio output control unit selects the speaker unit to be played back based on the position of each speaker projected on the screen so that the spoken audio is played back near the position on the screen.

5. The information processing device described in claim 3, wherein the receiving unit receives the video together with the audio from the specified space and projects the video on a screen on which one speaker array is arranged at each end of the video in the vertical direction when the video is projected, and the audio output control unit selects the speaker unit to be played back based on the position of the speaker so that the playback position of the spoken audio in the speaker array maintains the distance relationship between the positions of each speaker on the screen.

6. An information processing device as described in claim 1, wherein the clarity adjustment unit performs a clarity reduction adjustment to reduce the clarity of the speech of speakers not participating in the conversation contained in the audio of the specified space compared to the state of sound pickup by the microphone.

7. The information processing device according to claim 1, further comprising an audio separation unit that separates background sound from the audio in the specified space, and wherein the clarity adjustment unit performs clarity reduction adjustment to reduce the clarity of the background sound separated by the audio separation unit compared to other audio.

8. An information processing device according to claim 7, further comprising a background sound separation unit that separates the background sound into a sudden sound and a steady sound, and wherein the clarity adjustment unit performs the clarity reduction adjustment on the steady sound of the background sound.

9. An information processing device as described in claim 2, further comprising a voice synthesis unit that performs voice recognition for each of the individual sounds to generate characters indicating the spoken content of the individual sounds, and performs voice synthesis based on the characters indicating the spoken content to reproduce the speaker's spoken voice, and the clarity adjustment unit adjusts the clarity for each of the individual sounds reproduced by the voice synthesis unit based on whether the speaker is participating in the conversation.

10. The information processing device according to claim 1, wherein the human sensing unit determines the distance of the speaker based on the video and determines whether the speaker is participating in the conversation based on the distance.

11. An information processing device as described in claim 1, further comprising a sound source direction estimation unit that determines the position of a speaker based on the sound in the specified space, wherein the individual sound separation unit separates individual sounds, which are the speech sounds of each speaker, from the sound in the specified space based on the position of the speaker determined by the sound source direction estimation unit, and the clarity adjustment unit adjusts the clarity of each individual sound separated by the individual sound separation unit based on whether the speaker is participating in a conversation.

12. The information processing device according to claim 3, wherein the audio output control unit converts the audio in the predetermined space into stereo and adjusts the position of the sound source between the speaker arrays.

13. An information processing method in which an information processing device executes: a person sensing step of determining whether each speaker is participating in a conversation based on video captured by a camera that captures multiple speakers present in a specified space; a clarity adjustment step of adjusting the clarity of the speech of the speakers who are participating in the conversation determined in the person sensing step, which speech is included in the audio of the specified space picked up by a microphone that picks up the audio of the specified space; and a transmission step of transmitting the audio of the specified space processed in the clarity adjustment step.

14. An information processing system comprising: a transmitting device having a human sensing unit that determines whether each speaker is participating in a conversation based on images of multiple speakers present in a specified space captured by a camera; a clarity adjustment unit that adjusts the clarity of the speech of the speakers who are participating in the conversation and determined by the human sensing unit to be included in the audio of the specified space picked up by a microphone; and a transmitting unit that transmits the audio of the specified space processed by the clarity adjustment unit; and a receiving device having a receiving unit that receives the audio of the specified space transmitted by the transmitting unit; and an audio output control unit that, in a speaker having two speaker arrays in which multiple speaker units are arranged in a row, selects a speaker unit that will play the speech of each speaker from the audio of the specified space received by the receiving unit based on the position of each speaker, and causes the selected speaker unit to play the speech of each speaker.

Citation Information

Patent Citations

  • Conference voice data processing method and device and storage medium

    CN113140223A

  • Information processing device, information processing method, and information processing program

    JP2023047956A

  • Information processing device, information processing method, and program

    WO2023100594A1