Information processing apparatus, display device, presentation method, and program
The information processing device addresses the challenge of determining sound direction for hearing aid users by generating text images based on audio input, improving conversation clarity in multi-speaker scenarios.
Patent Information
- Application Number
- JP2025181402
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-11
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-27
AI Technical Summary
Hearing aid users struggle to determine the direction of sound, especially when multiple speakers are talking simultaneously, making it difficult to maintain conversations.
An information processing device that acquires audio using multiple microphones, estimates the arrival direction, generates a text image corresponding to the audio, and presents it in a determined mode to help users recognize the sound direction.
Enables users to easily recognize the direction from which sound is coming by presenting text images corresponding to the sound direction, enhancing conversation clarity in multi-speaker environments.
Smart Images

Figure 2026012872000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, a display device, a presentation method, and a program. [Background technology]
[0002] Hearing aids are widely used as devices to assist hearing. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-236396 Summary of the Invention [Problem to be solved by the invention]
[0004] Hearing aid wearers may have a reduced ability to determine the direction of sound due to a decline in their hearing function. When such wearers try to hold a conversation with multiple people, they are unable to determine the direction from which the sound is coming, making it difficult to maintain a conversation.
[0005] For example, Patent Document 1 proposes a hearing aid that reproduces the direction of arrival of sound and improves the clarity of the sound uttered by a speaker (hereinafter referred to as "speech sound"). However, reproducing the direction of arrival of sound alone is not sufficient for the wearer of the hearing aid to recognize the direction of arrival. In particular, when multiple speakers are speaking simultaneously, reproducing the direction of arrival of sound alone makes it difficult for the wearer to recognize the direction of arrival of each speaker's speech sound.
[0006] An object of the present disclosure is to make it easy to recognize the direction from which sound is coming. [Means for solving the problem]
[0007] According to one aspect of the present disclosure, there is provided an information processing device. The information processing device includes means for acquiring audio collected by a plurality of microphones. The information processing device includes means for estimating an arrival direction of the acquired audio. The information processing device includes means for generating a text image corresponding to the acquired audio. The information processing device includes means for determining a presentation mode of the text image by referring to the estimated arrival direction. The information processing device includes means for presenting the text image in the determined presentation mode. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram illustrating a configuration of a display device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram of a glass-type display device, which is an example of the display device shown in FIG. 1. [Figure 3] FIG. 1 is an explanatory diagram of an overview of the present embodiment. [Figure 4] 10 is a flowchart illustrating an example of a presentation process according to the present embodiment. [Figure 5] FIG. 2 is a diagram for explaining collection of speech sounds emitted by a speaker. [Figure 6] FIG. 2 is a diagram for explaining the direction from which a speech sound arrives. [Figure 7] FIG. 1 is a schematic diagram showing an example of a display device using a glass display. [Figure 8] FIG. 10 is a diagram for explaining the field of view of a wearer. [Figure 9] FIG. 10 is a schematic diagram showing the configuration of a display device according to a first modification. [Figure 10] 10 is a schematic diagram illustrating a display device according to a second modification and a presentation example of the display device. FIG. [Figure 11] FIG. 10 is a schematic diagram showing the configuration of a display device according to a third modification. [Figure 12] 10 is a schematic diagram illustrating a first example of a display device according to Modification 3 and a presentation example of the display device. FIG. [Figure 13]FIG. 12 is a schematic diagram showing the photographing range of the camera shown in FIG. [Figure 14] 13 is a schematic diagram showing a second example of a display device according to Modification 3 and a presentation example of the display device. FIG. [Figure 15] 10 is a schematic diagram illustrating a third example of a display device according to Modification 3 and a presentation example of the display device. FIG. [Figure 16] 10 is a schematic diagram illustrating a fourth example of a display device according to Modification 3 and a presentation example of the display device. FIG. [Figure 17] FIG. 13 is a schematic diagram showing the configuration of a display device according to a fourth modification. [Figure 18] FIG. 13 is a schematic diagram showing the configuration of a display device according to a fifth modification. [Figure 19] FIG. 19 is a schematic diagram of a conference system that is an example of the display device shown in FIG. 18. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. In the drawings for explaining the embodiment, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.
[0010] (1) Configuration of the information processing device The configuration of a display device 1 of this embodiment will be described. Fig. 1 is a schematic diagram showing the configuration of a display device of this embodiment. Fig. 2 is a schematic diagram of a glass-type display device, which is an example of the display device shown in Fig. 1.
[0011] The display device 1 shown in FIG. 1 is configured to collect sound and display (an example of "presentation") a text image corresponding to the collected sound in a presentation mode according to the direction from which the sound comes. The display device 1 may have at least one of the following configurations: Glass-type display device Mobile devices Conference system
[0012] As shown in FIG. 1, the display device 1 includes a plurality of microphones 101, a display 102, and a controller 10. The microphones 101 are arranged at predetermined distances from each other.
[0013] As shown in FIG. 2, when the display device 1 is a glasses-type display device, the display device 1 includes a right temple 21, a right end piece 22, a bridge 23, a left end piece 24, a left temple 25, and a rim 26.
[0014] The microphone 101-1 is disposed on the right temple 21. The microphone 101-2 is disposed on the right armor 22. The microphone 101-3 is disposed on the bridge 23. The microphone 101-4 is disposed on the left end piece 24. The microphone 101-5 is disposed on the left temple 25. The microphone 101 collects at least one of the following sounds, for example: -Speech by a person Sounds of the environment in which the display device 1 is used (hereinafter referred to as "environmental sounds")
[0015] When the display device 1 is a glasses-type display device, the display 102 is a transparent member (for example, at least one of glass, plastic, and a half mirror). In this case, the display 102 is placed in a position where it can be seen by a user wearing the glasses-type display device.
[0016] Displays 102-1 to 102-2 are supported by a rim 26. Display 102-1 is disposed so as to be located in front of the right eye of a user when the user wears display device 1. Display 102-2 is disposed so as to be located in front of the left eye of a user when the user wears display device 1.
[0017] The display 102 presents (for example, displays) an image under the control of the controller 10. The method by which the display 102 presents an image is not limited, and any existing method may be used.
[0018] 2, an image corresponding to image light from a projector (not shown) disposed on the rear side of the right temple 21 is projected onto the display 102-1. An image corresponding to image light from a projector (not shown) disposed on the rear side of the left temple 25 is projected onto the display 102-2. The display 102-1 and the display 102-2 present an image, and the user can simultaneously view the image and the scenery passing through the display 102-1 and the display 102-2.
[0019] The controller 10 is an information processing device that controls the display device 1. The controller 10 is connected to a microphone 101 and a display 102 by wire or wirelessly. When the display device 1 is a glasses-type display device as shown in FIG. 2, the controller 10 is disposed, for example, inside the right temple 21.
[0020] As shown in FIG. 1, the controller 10 includes a storage device 11, a processor 12, an input / output interface 13, and a communication interface 14.
[0021] The storage device 11 is configured to store programs and data, and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and a storage (for example, a flash memory or a hard disk).
[0022] The programs include, for example, the following programs: OS (Operating System) programs Application programs that perform information processing
[0023] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)
[0024] The processor 12 is configured to realize the functions of the controller 10 by activating a program stored in the storage device 11. The processor 12 is an example of a computer. For example, by activating a program stored in the storage device 11, the processor 12 realizes the function of presenting an image (hereinafter referred to as a "text image") representing text corresponding to speech sounds collected by the microphone 101 at a predetermined position on the display 102.
[0025] The input / output interface 13 acquires at least one of the following: Audio signal collected by microphone 101 A user's instruction input from an input device connected to the glass-type display device 1 The input device is, for example, a drive button, a keyboard, a pointing device, a touch panel, a remote controller, a switch, or a combination thereof. The input / output interface 13 is also configured to output information to an output device connected to the display device 1. The output device is, for example, a display 102.
[0026] The communication interface 14 is configured to control communication between the display device 1 and an external device (not shown) (for example, a server or a mobile terminal).
[0027] (2) Overview of the embodiment An outline of this embodiment will be explained below with reference to Fig. 3, which is an explanatory diagram of the outline of this embodiment.
[0028] In FIG. 3, a wearer P1 wearing the display device 1 is having a conversation with speakers P2 to P4. The microphone 101 collects the speech sounds of the speakers P2 to P4. The controller 10 estimates the direction from which the collected speech sound comes. The controller 10 determines text corresponding to the speech sounds by analyzing the audio signals corresponding to the collected speech sounds. The controller 10 generates text images T1 to T3 corresponding to the determined text. The controller 10 determines the presentation mode for each of the text images T1 to T3 according to the direction from which the speech sound comes. Controller 10 presents text images T1 to T3 on displays 102-1 to 102-2 in the determined presentation format.
[0029] (3) Presentation process The presentation process of this embodiment will be described. Fig. 4 is a flowchart showing an example of the presentation process of this embodiment. Fig. 5 is a diagram for explaining collection of speech sounds emitted by a speaker. Fig. 6 is a diagram for explaining the direction from which the speech sounds come. Fig. 7 is a schematic diagram showing a presentation example of the glasses-type display device of Fig. 2. Fig. 8 is a diagram for explaining the field of view of a wearer.
[0030] Each microphone 101 collects the speech sound emitted by the speaker. For example, in the example shown in Fig. 2, microphones 101-1 to 101-5 arranged at the right temple 21, right end piece 22, bridge 23, left end piece 24, and left temple 25 of the display device 1, respectively, collect the speech sound arriving via the path shown in Fig. 5. The microphones 101-1 to 101-5 convert the collected speech sound into an audio signal.
[0031] The controller 10 acquires the audio signal converted by the microphone 101 (S110).
[0032] Specifically, processor 12 acquires audio signals transmitted from microphones 101-1 to 101-5, the audio signals including speech sounds emitted by at least one of speakers P2, P3, and P4. The audio signals transmitted from microphones 101-1 to 101-5 include spatial information based on the path along which the speech sounds have traveled.
[0033] After step S110, the controller 10 performs direction of arrival estimation (S111).
[0034] Specifically, an arrival direction estimation model is stored in the storage device 11. The arrival direction estimation model describes the correlation between spatial information included in the audio signal and the arrival direction of the speech sound.
[0035] The direction of arrival estimation model may use any existing method for estimating the direction of arrival, such as MUSIC (Multiple Signal Classification) using eigenvalue expansion of an input correlation matrix, the minimum norm method, or ESPRIT (Estimation of Signal Parameters via Rotational Invariance Techniques).
[0036] Processor 12 estimates the arrival direction of the speech sound collected by microphones 101-1 to 101-5 by inputting the audio signals received from microphones 101-1 to 101-5 into an arrival direction estimation model stored in storage device 11. At this time, processor 12 estimates the arrival direction of the speech sound as, for example, an angle from an axis with zero degrees in front. In the example shown in FIG. 6, processor 12 estimates the arrival direction of the speech sound emitted from speaker P2 as angle A1 to the right from the axis. Processor 12 estimates the arrival direction of the speech sound emitted from speaker P3 as angle A2 to the left from the axis. Processor 12 estimates the arrival direction of the speech sound emitted from speaker P4 as angle A3 to the left from the axis.
[0037] After step S111, the controller 10 extracts the audio signal (S112).
[0038] Specifically, a beamforming model is stored in the storage device 11. The beamforming model describes the correlation between a predetermined direction and parameters for forming a directivity having a beam in this direction. Here, the parameters for forming the directivity are parameters related to amplifying or attenuating multiple audio signals.
[0039] The processor 12 inputs the estimated arrival direction into the beamforming model stored in the storage device 11, and calculates parameters for forming a directivity having a beam in the arrival direction.
[0040] 6, processor 12 inputs the calculated angle A1 into the beamforming model and calculates parameters for forming a directivity having a beam in the direction of angle A1 in a right direction from the axis. Processor 12 inputs the calculated angle A2 into the beamforming model and calculates parameters for forming a directivity having a beam in the direction of angle A2 in a left direction from the axis. Processor 12 inputs the calculated angle A3 into the beamforming model and calculates parameters for forming a directivity having a beam in the direction of angle A3 in a left direction from the axis.
[0041] Processor 12 amplifies or attenuates the audio signals transmitted from microphones 101-1 to 101-5 using the parameters calculated for angle A1. Processor 12 synthesizes the amplified or attenuated audio signals to extract an audio signal for the speech sound arriving from angle A1 from the received audio signals.
[0042] Processor 12 amplifies or attenuates the audio signals transmitted from microphones 101-1 to 101-5 using the parameters calculated for angle A2. Processor 12 synthesizes the amplified or attenuated audio signals to extract an audio signal for the speech sound arriving from angle A2 from the received audio signals.
[0043] Processor 12 amplifies or attenuates the audio signals transmitted from microphones 101-1 to 101-5 using the parameters calculated for angle A3. Processor 12 synthesizes the amplified or attenuated audio signals to extract an audio signal for the speech sound arriving from angle A3 from the received audio signals.
[0044] After step S112, the controller 10 performs voice recognition (S113).
[0045] Specifically, a speech recognition model is stored in the storage device 11. The speech recognition model describes a correlation between a speech signal and text corresponding to the speech signal. The speech recognition model is, for example, a trained model trained by machine learning.
[0046] The processor 12 inputs the extracted voice signal into a voice recognition model stored in the storage device 11, thereby determining the text corresponding to the input voice signal.
[0047] In the example shown in FIG. 6, the processor 12 inputs the voice signals extracted for the angles A1 to A3 into a voice recognition model, respectively, to determine the text corresponding to the input voice signals.
[0048] After step S113, the controller 10 executes image generation (S114).
[0049] Specifically, processor 12 generates a text image based on the determined text.
[0050] After step S114, the controller 10 determines the presentation mode (S115).
[0051] In particular, processor 12 determines how the text image is to be presented on display 102 .
[0052] In a first example of step S115, the processor 12 determines a position corresponding to the direction of arrival of the audio signal related to the text image as the presentation position of the text image. Processor 12 determines the type of text image to be presented (an example of a "presentation style") according to the arrival direction.
[0053] More specifically, the processor 12 determines the presentation position of the text image T1, which is generated based on the audio signal extracted in the direction of angle A1 to the right from the axis, to be a position in the direction corresponding to angle A1 and in the direction of a predetermined elevation angle. In an example adopted for a glasses-type display device, the processor 12 determines the presentation position of the text image T1 to be a position on the right display 102-1 of the glasses-type display device in the direction corresponding to angle A1 and in the direction of a predetermined elevation angle. The processor 12 also determines to present the text image T1 so that the text image T1 is imaged at a predetermined distance from the wearer P1. The processor 12 determines the presentation position of the text image T2, which is generated based on the audio signal extracted in the direction at angle A2 to the left of the axis, to be a position in the direction corresponding to angle A2 and at a predetermined elevation angle. In an example adopted for a glasses-type display device, the processor 12 determines the presentation position of the text image T2 to be a position on the left display 102-2 of the glasses-type display device in the direction corresponding to angle A2 and at a predetermined elevation angle. The processor 12 also determines to present the text image T2 so that the text image T2 is focused at a predetermined distance from the wearer P1. The processor 12 determines the presentation position of the text image T3, which is generated based on the audio signal extracted in the direction of angle A3 to the left of the axis, to be a position in the direction corresponding to angle A3 and at a predetermined elevation angle. In an example adopted for a glasses-type display device, the processor 12 determines the presentation position of the text image T3 to be a position on the left display 102-2 of the glasses-type display device in the direction corresponding to angle A3 and at a predetermined elevation angle. The processor 12 also determines to present the text image T3 so that the text image T3 is focused at a predetermined distance from the wearer P1.
[0054] In the second example of step S115, the processor 12 determines predetermined positions as presentation positions of the text images T1 to T3. The processor 12 determines to present the text images T1 to T3 in a format (an example of a "presentation format") that includes at least one of a character string and a symbol corresponding to the arrival direction of the audio signal related to the text image.
[0055] After step S115, the controller 10 executes image presentation (S116).
[0056] Specifically, the processor 12 presents the text image on the display 102 in the determined presentation manner.
[0057] According to a first example of step S115 (FIG. 7), processor 12 presents text image T1 on display 102-1 in a direction corresponding to angle A1 and at a position at a predetermined elevation angle. Processor 12 presents text image T2 on display 102-2 in a direction corresponding to angle A2 and at a position at a predetermined elevation angle. Processor 12 presents text image T3 on display 102-2 in a direction corresponding to angle A3 and at a position at a predetermined elevation angle. The human figures shown by dashed lines on displays 102-1 to 102-2 in FIG. 7 supplementarily represent speakers visible to wearer P1 through displays 102-1 to 102-2, and are not actually presented on displays 102-1 to 102-2.
[0058] According to a second example of step S115, processor 12 presents text image T1 at a predetermined position on display 102-1 in a format including at least one of a character string and a symbol corresponding to a direction corresponding to angle A1. Processor 12 presents text image T2 at a predetermined position on display 102-2 in a format including at least one of a character string and a symbol corresponding to a direction corresponding to angle A2. Processor 12 presents text image T3 at a predetermined position on display 102-2 in a format including at least one of a character string and a symbol corresponding to a direction corresponding to angle A3. As an example, a text image based on speech sounds from a speaker on the left includes, for example, the character for "left" or a symbol reminiscent of "left," and a text image based on speech sounds from a speaker on the right includes, for example, the character for "right" or a symbol reminiscent of "right."
[0059] By presenting the text images T1 to T3 on the displays 102-1 to 102-2 in this way, the wearer P1 of the glasses-type display device 1 is presented with a text image T1, which is the content of the conversation spoken by speaker P2, along with speaker P2 being visible through the display 102-1, as shown in Fig. 8. The wearer P1 is presented with a text image T2, which is the content of the conversation spoken by speaker P3, along with speaker P3 being visible through the display 102-2. The wearer P1 is presented with a text image T3, which is the content of the conversation spoken by speaker P4, along with speaker P4 being visible through the display 102-2.
[0060] (4) Summary According to this embodiment, a text image corresponding to a speech sound is presented in a presentation format according to the direction from which the speech sound comes, thereby allowing the wearer of display device 1 to easily recognize the direction from which the speech sound comes.
[0061] Furthermore, according to the present embodiment, an image is presented at a position corresponding to the direction from which the speech sound comes, which allows the user to more easily recognize the direction from which the speech sound comes.
[0062] Furthermore, according to this embodiment, a voice signal corresponding to the estimated arrival direction is extracted from the acquired voice signal, thereby enabling the user to accurately recognize the arrival direction of the speech sound.
[0063] Furthermore, according to the present embodiment, the display device is applied to at least one of a glasses-type display device, a mobile terminal, and a conference system, thereby enabling the user to easily recognize the direction from which speech sounds are coming in various applications.
[0064] (5) Variations A modification of this embodiment will now be described.
[0065] (5.1) Variation 1 A first modification of this embodiment will be described. In the first modification, an example is shown in which a display device 1 is connected to a microphone module including a plurality of microphones 101. Fig. 9 is a schematic diagram showing the configuration of a display device of the first modification.
[0066] As shown in FIG. 9, in the display device 1 of the first modification, the communication interface 14 is connected to a microphone module 101a. In this case, the microphone 101 is not arranged on the frame of the glass type display device 1.
[0067] The microphone module 101a includes a plurality of microphones 101. The microphones 101 are arranged at predetermined distances from each other. The microphone module 101a is worn on any of the following body parts: ·head Collar Chest ·female servant Other parts that pass through the center of the wearer When the microphone module 101a is worn by a wearer, the microphone module 101a communicates with the controller 10 via the communication interface 14.
[0068] Controller 10 executes steps S110 to S116 in the same manner as in FIG. 4, and presents text images T1 to T3 on displays 102-1 and 102-2.
[0069] According to variant example 1, even in a glass-type display device 1 in which a microphone 101 is not installed, it is possible to present a text image corresponding to the sound collected by the microphone 101 in a manner according to the direction of arrival.
[0070] (5.2) Variation 2 Modification 2 of this embodiment will be described. Modification 2 shows an example in which display device 1 includes a mobile terminal. Fig. 10 is a schematic diagram showing a display device of Modification 2 and a presentation example of the display device.
[0071] In the second modification, the mobile terminal of Fig. 10 is an example of the display device 1. The mobile terminal includes, for example, any of the following: Smartphone Tablet devices -Mobile devices with a display Personal computers (e.g., laptop computers)
[0072] In the second modification, the controller 10 executes steps S110 to S116 in the same manner as in FIG.
[0073] As a result, as shown in FIG. 10, text images T1 to T3 are presented on display 102 at positions in the direction corresponding to the direction from which the speech sound comes.
[0074] According to the second modification, by connecting the microphone module 101a to the mobile terminal, it becomes possible to present a text image corresponding to the speech sound collected by the microphone 101 in a presentation style according to the direction of arrival.
[0075] (5.3) Variation 3 A third modification of this embodiment will be described below. In the third modification, an example is shown in which the display device 1 is equipped with a camera. Fig. 11 is a schematic diagram showing the configuration of the display device of the third modification.
[0076] As shown in FIG. 11, the display device 1a includes a microphone 101, a display 102, a camera 103, and a controller 10a. The camera 103 is positioned so that the speaker is included in the capture area. The camera 103 captures an image in a predetermined direction and generates an image signal.
[0077] The controller 10a is an information processing device that controls the display device 1a. The controller 10a is connected to a microphone 101, a display 102, and a camera 103 via wired or wireless connections.
[0078] As shown in FIG. 11, the controller 10a includes a storage device 11, a processor 12a, an input / output interface 13a, and a communication interface .
[0079] The processor 12a is configured to realize the functions of the controller 10a by starting a program stored in the storage device 11. The processor 12a is an example of a computer. For example, by starting a program stored in the storage device 11, the processor 12a realizes a function of presenting a text image of speech sounds collected by the microphone 101 at a predetermined position on the display 102 by superimposing the text image on an image corresponding to a captured signal generated by the camera 103 (hereinafter referred to as a "captured image").
[0080] The input / output interface 13a acquires at least one of the following: Audio signal collected by microphone 101 A photographic signal captured by the camera 103 A user instruction input from an input device connected to the display device 1 The input device is, for example, a drive button, a keyboard, a pointing device, a touch panel, a remote controller, a switch, or a combination thereof. The input / output interface 13a is also configured to output information to an output device connected to the display device 1. The output device is, for example, a display 102.
[0081] (5.3.1) Presentation Processing The presentation process of the third modification will be described with reference to the flowchart shown in FIG.
[0082] The controller 10a executes steps S110 to S113 in the same manner as in FIG.
[0083] After step S113, the controller 10a executes image generation (S114).
[0084] Specifically, the controller 10a converts the image signal generated by the camera 103 into a captured image. The controller 10a generates a text image in the same manner as in FIG.
[0085] After step S114, the controller 10a determines the presentation mode (S115).
[0086] Specifically, the processor 12a determines how the text image and the captured image are to be presented on the display 102. For example, similar to FIG. 4, the processor 12a determines the position corresponding to the arrival direction of the audio signal related to the text image as the presentation position of the text image, and determines the type of text image to be presented depending on the arrival direction. The processor 12a determines the presentation position of the captured image and the type of captured image to be presented according to the direction of arrival.
[0087] After step S115, the controller 10a executes image presentation (S116).
[0088] Specifically, the processor 12a superimposes the text image generated in step S114 on the captured image in the determined presentation format and presents it on the display 102.
[0089] (5.3.2) First Example of Display Device of Modification Example 3 A first example of a display device according to Modification 3 will be described. In the first example of the display device according to Modification 3, an example is shown in which the display device 1a includes a glass-type display device. Fig. 12 is a schematic diagram showing the first example of the display device according to Modification 3 and a presentation example of the display device. Fig. 13 is a schematic diagram showing the shooting range of the camera shown in Fig. 11.
[0090] 12, the camera 103 is disposed on the bridge 23 so as to capture an image of an area including the field of view of the wearer. The camera 103 is set so that the capture range is a range that includes the field of view of the wearer. In Fig. 13, the solid line represents the range of photography by camera 103, and the dashed line represents the field of view of the wearer. According to the example shown in Fig. 13, camera 103 is capable of photographing the scenery within the field of view of the wearer. As a result, if speakers P2 to P4 are included in the field of view of the wearer, camera 103 will photograph speakers P2 to P4.
[0091] The controller 10a executes steps S110 to S114 shown in FIG.
[0092] After step S114, the controller 10a determines the presentation mode (S115).
[0093] Specifically, the processor 12a determines the presentation position of the text image T1, which is generated based on the audio signal extracted for a predetermined arrival direction, to be a position in a direction corresponding to the arrival direction and at a predetermined elevation angle. That is, the processor 12a determines the presentation position of the text image T1 to be a position on the display 102-1 in a direction corresponding to the arrival direction and at a predetermined elevation angle. The processor 12a also determines to present the text image T1 so that the text image T1 is formed at a predetermined distance from the wearer. The processor 12a determines the presentation positions of the text images T2 and T3 generated based on the audio signal extracted for a predetermined arrival direction to be positions in a direction corresponding to the arrival direction and at a predetermined elevation angle. That is, the processor 12a determines the presentation positions of the text images T2 and T3 to be positions on the display 102-2 in a direction corresponding to the arrival direction and at a predetermined elevation angle. The processor 12a also determines to present the text images T2 and T3 so that they are imaged at a predetermined distance from the wearer. The processor 12a determines the presentation position of the captured image based on the shooting direction of the camera 103. The processor 12a also determines to present the captured image so that the captured image is formed at a predetermined distance from the wearer.
[0094] After step S115, the controller 10a executes image presentation (S116). Specifically, the processor 12a superimposes the text image generated in step S114 on the captured image in the determined presentation format and presents it on the display 102.
[0095] 12, the processor 12a presents the captured images on the displays 102-1 and 102-2. As a result, for example, as shown in FIG. 12, the captured image I1 of the speaker P2 is presented on the display 102-1, and the captured images I2 and I3 of the speakers P3 and P4 are presented on the display 102-2. Processor 12a presents text image T1 superimposed on the captured image at a position on display 102-1 in a direction corresponding to the direction from which the speech sound is coming and at a predetermined elevation angle. Processor 12a presents text images T2 to T3 superimposed on the captured image on display 102-2 in a direction corresponding to the direction from which the speech sound is coming and at a predetermined elevation angle.
[0096] In this way, by presenting images I1 to I3 and text images T1 to T3 on displays 102-1 to 102-2, a text image T1 representing the conversation spoken by speaker P2 is presented to a wearer of display device 1a together with an image I1 representing speaker P2. A text image T2 representing the conversation spoken by speaker P3 is presented to wearer P1 together with an image I2 representing speaker P3. A text image T3 representing the conversation spoken by speaker P4 is presented to wearer P1 together with an image I3 representing speaker P4.
[0097] (5.3.3) Second Example of Display Device of Modification Example 3 A second example of the display device of Modification 3 will be described. In the second example of the display device of Modification 3, an example is shown in which the display device 1a is connected to a microphone module including a plurality of microphones 101. Fig. 14 is a schematic diagram showing the second example of the display device of Modification 3 and an example of presentation of the display device.
[0098] As shown in FIG. 14, in the second example of the display device of the third modified example, the microphone 101 is not arranged on the frame of the glasses-type display device 1a.
[0099] The controller 10a executes steps S110 to S116 shown in FIG. 4 as described in the first example of the display device of the third modified example.
[0100] 14, image I1 is presented on display 102-1, and images I2 and I3 are presented on display 102-2. Also, text image T1 is presented superimposed on display 102-1 at a position corresponding to the arrival direction. Also, text images T2 and T3 are presented superimposed on display 102-2 at positions corresponding to the arrival direction.
[0101] (5.3.4) Third Example of Display Device of Modification Example 3 A third example of the display device of Modification Example 3 will be described. In the third example of the display device of Modification Example 3, an example is shown in which the display device 1a includes a mobile terminal. Fig. 15 is a schematic diagram showing the third example of the display device of Modification Example 3 and an example of presentation of the display device.
[0102] In the example shown in FIG. 15, the camera 103 is a camera placed on the backside of the display 102 so as to capture an image of an area including the field of view of the user P1.
[0103] The controller 10a executes steps S110 to S114 shown in FIG. 4 as described in the first example of the display device of the third modification.
[0104] After step S114, the controller 10a determines the presentation mode (step S115).
[0105] Specifically, processor 12a determines the presentation positions of text images T1 to T3 generated based on the audio signal extracted for a predetermined arrival direction as positions in a direction corresponding to the arrival direction. In an example adopted in a mobile terminal, processor 12a sets the presentation positions of text images T1 to T3 to positions on display 102 of the mobile terminal in a direction corresponding to the arrival direction. Processor 12a also determines to present text images T1 to T3. The processor 12a determines the presentation position of the captured image based on the shooting direction of the camera 103. The processor 12a also determines to present the captured image.
[0106] After step S115, the controller 10a executes image presentation (S116).
[0107] Specifically, the processor 12a superimposes the text image generated in step S114 on the captured image in the determined presentation format and presents it on the display 102.
[0108] According to the example shown in Fig. 15, the processor 12a presents the captured images on the display 102. As a result, for example, images I1 to I3 of the captured speakers are presented on the display 102 as shown in Fig. 15. The processor 12a presents the text images T1 to T3 on the display 102 of the mobile terminal at positions in the direction corresponding to the direction from which the speech sound is coming.
[0109] In this way, by presenting images I1-I3 and text images T1-T3 on display 102, user P1 of display device 1a is presented with text image T1, which is the content of the conversation spoken by speaker P2, along with image I1 representing speaker P2. User P1 is presented with text image T2, which is the content of the conversation spoken by speaker P3, along with image I2 representing speaker P3. User P1 is presented with text image T3, which is the content of the conversation spoken by speaker P4, along with image I3 representing speaker P4.
[0110] (5.3.5) Fourth Example of Display Device of Modification Example 3 A fourth example of the display device of Modification 3 will be described. In the fourth example of the display device of Modification 3, an example is shown in which the display device 1a is adopted in a conference system. Fig. 16 is a schematic diagram showing the fourth example of the display device of Modification 3 and an example of presentation of the display device.
[0111] In a fourth example of the display device of the third modification, the conference system is a system that presents speech sounds collected during a conference on a display as a text image at a position according to the direction of arrival.
[0112] The display 102 is positioned so that it can be seen by the conference participants.
[0113] Camera 103 is placed in a position where it can capture images of the conference participants. In the example shown in Fig. 16, camera 103 is placed above display 102. Camera 103 captures images of conference participants P2 to P4 who are participating in the conference.
[0114] The microphone module 101a is placed in one of the following positions: Conference table ·Hollow position suspended from the ceiling When the microphone module 101a is placed in a predetermined position, it performs regulation with the controller 10a.
[0115] The controller 10a executes steps S110 to S116 shown in FIG. 4 as described in the third example of the display device of the third modified example.
[0116] 16, the processor 12a presents captured images on the display 102. As a result, images I1 to I3 of the conference participants P2 to P4 are presented on the display 102. The processor 12a presents text images T1 to T3 at positions on the display 102 in the direction corresponding to the direction from which the speech sound is coming.
[0117] In this way, by presenting images I1-I3 and text images T1-T3 on display 102, text image T1, which is the content of the conversation spoken by conference participant P2, is presented together with image I1 representing conference participant P2. Text image T2, which is the content of the conversation spoken by conference participant P3, is presented together with image I2 representing conference participant P3. Text image T3, which is the content of the conversation spoken by conference participant P4, is presented together with image I3 representing conference participant P4.
[0118] According to the third modification, it is possible to present a captured image and, in addition, to present a text image corresponding to the speech sound collected by the microphone 101 in a presentation format according to the direction of arrival, in accordance with the speaker image included in the captured image. This makes it possible to improve the visibility of the relationship between the source of the sound (for example, the speaker) and the text image.
[0119] (5.4) Variation 4 A fourth modification of this embodiment will be described. In the fourth modification, the functions of the controller are implemented by a server device. Fig. 17 is a schematic diagram showing the configuration of a display device in the fourth modification.
[0120] As shown in FIG. 17, a display device 1b includes a plurality of microphones 101, a display 102, and a server device 10b.
[0121] The server device 10b is an information processing device that controls the display device 1b, and is connected to a network via a wired or wireless connection.
[0122] As shown in FIG. 17, the server device 10b includes a storage device 11, a processor 12b, an input / output interface 13, and a communication interface 14b.
[0123] The processor 12b is configured to realize the functions of the server device 10b by starting a program stored in the storage device 11. The processor 12b is an example of a computer. For example, the processor 12b realizes the function of presenting a text image based on speech sounds collected by the microphone 101 at a predetermined position on the display 102 by starting a program stored in the storage device 11.
[0124] The communication interface 14b is configured to control communication between the display device 1b and the microphone 101 and display 102 via a network.
[0125] In the fourth modification, the server device 10b executes steps S110 to S116 in the same manner as in FIG.
[0126] According to variant example 4, even if the terminal does not have a processor capable of complex calculations, it is possible to present a text image corresponding to the speech sound collected by microphone 101 in a presentation format that corresponds to the direction of arrival.
[0127] (5.5) Variation 5 Modification 5 of this embodiment will be described. Modification 5 shows an example in which the display device of Modification 4 is equipped with a camera. Fig. 18 is a schematic diagram showing the configuration of the display device of Modification 5. Fig. 19 is a schematic diagram of a conference system that is an example of the display device shown in Fig. 18.
[0128] As shown in FIG. 18, a display device 1c includes a plurality of microphones 101, a display 102, a camera 103, and a server device 10c.
[0129] The server device 10c is a device that controls the display device 1c and is connected to a network via a wired or wireless connection.
[0130] As shown in FIG. 18, the server device 10c includes a storage device 11, a processor 12c, an input / output interface 13, and a communication interface 14c.
[0131] The processor 12c is configured to realize the functions of the server device 10c by starting a program stored in the storage device 11. The processor 12c is an example of a computer. For example, the processor 12c realizes the function of presenting a text image based on speech sounds collected by the microphone 101 at a predetermined position on the display 102 by starting a program stored in the storage device 11.
[0132] The communication interface 14c is configured to control communication between the display device 1c and the microphone 101, the display 102, and the camera 103 via a network.
[0133] In the conference system shown in Fig. 19, a conference held remotely is photographed and speech sounds from the conference are collected. The conference system presents the photographed image on a display and also presents a text image based on the speech sounds at a position on the display that corresponds to the direction from which the speech sounds are coming. Hereinafter, a conference held remotely will be referred to as a remote conference.
[0134] The display 102 is disposed in a position that is visible to at least one of the following people: Participants in remote conferences · Person who monitors remote conferences
[0135] Camera 103 is placed in a position where it can capture images of the remote conference. In the example shown in Fig. 19, camera 103 captures images of conference participants P2 to P4 who are participating in the remote conference. Camera 103 captures images and generates a capture signal. Camera 103 transmits the capture signal to server device 10c via the network.
[0136] The microphone module 101a is placed at one of the following positions where it can collect the speech sounds of the remote conference: Conference table ·Hollow position suspended from the ceiling When the microphone module 101a is placed in a predetermined position, it performs regulation with the server device 10c.
[0137] In FIG. 19, the server device 10c executes steps S110 to S116 in the same manner as in FIG.
[0138] 19, the processor 12c presents the captured images on the display 102. As a result, images I1 to I3 of the conference participants P2 to P4 are presented on the display 102. The processor 12c presents the text images T1 to T3 at positions on the display 102 in the direction corresponding to the direction from which the speech sound is coming.
[0139] In this way, by presenting images I1-I3 and text images T1-T3 on display 102, text image T1, which is the content of the conversation spoken by conference participant P2, is presented together with image I1 representing conference participant P2. Text image T2, which is the content of the conversation spoken by conference participant P3, is presented together with image I2 representing conference participant P3. Text image T3, which is the content of the conversation spoken by conference participant P4, is presented together with image I3 representing conference participant P4.
[0140] According to variant example 5, it is possible to present a captured image and, in addition, to match the speaker image contained in the captured image, present a text image corresponding to the speech sound collected by microphone 101 in a presentation format that corresponds to the direction of arrival.
[0141] (6) Other variations In this embodiment, we have described a case where a user's instructions are input from an input device connected to the input / output interface 13, but this embodiment is also applicable to a case where a user's instructions are input from a drive button object presented by an application on a computer (e.g., a smartphone) connected to the communication interface 14.
[0142] The display device 1 may be realized by any method as long as it can present an image to the user. The display device 1 can be realized by, for example, the following methods. HOE (Holographic optical element) or DOE (Diffractive optical element) using optical elements (e.g., light guide plates) LCD display Retina projection display LED (Light Emitting Diode) display Organic EL (Electro Luminescence) display Laser display A display that uses optical elements (such as lenses, mirrors, diffraction gratings, liquid crystals, MEMS mirrors, and HOEs) to guide light emitted from a light emitter. In particular, retinal projection displays allow even people with weak eyesight to easily observe images, making it easier for people with both hearing loss and weak eyesight to recognize the direction from which speech sounds are coming.
[0143] In this embodiment, the display device 1a is described as having the camera 103, but this embodiment is also applicable to a case where the display device 1a is provided with a sensor configured to perform sensing. The sensor is, for example, at least one of the following: -Human sensor TOF (Time Of Flight) sensor Millimeter wave radar ·LiDAR(Light Detection And Ranging) Image sensor If the display device 1 is equipped with the sensor, for example, the input / output interface 13 acquires a sensing signal generated by the sensor. The processor 12 determines the presentation mode of the text image in step S115 based on the acquired sensing signal. This can improve the accuracy with which the text image is presented. The sensing signal is, for example, a photographic signal obtained by photographing an area where sound is collected by a plurality of microphones using a camera equipped with an image sensor.
[0144] In this embodiment, we have described a case where the presentation position of a text image is determined based on the direction from which the speech sound comes, even when a captured image is present. However, this embodiment is also applicable to a case where the processors 12a and 12c determine the presentation position of a text image in association with an image of a speaker located within a predetermined range from the direction from which the speech sound comes. Specifically, for example, processors 12a and 12c determine the presentation position of the captured image based on the shooting direction of camera 103. Processors 12a and 12c associate the arrival direction of the speech sound with the position of the speaker included in the captured image. Processors 12a and 12c determine the presentation positions of text images T1 to T3 generated based on the audio signal extracted for a predetermined arrival direction as positions near the speaker associated with the arrival direction.
[0145] In the present embodiment, an example of extracting an audio signal by beamforming has been described as a method of extracting an audio signal, but the scope of the present embodiment is not limited to this. The extraction of an audio signal in the present embodiment can also be achieved by the following method. Frost Beamformer Adaptive filter beamforming (for example, generalized sidelobe canceller)
[0146] In this embodiment, an example has been described in which the presentation mode of a text image includes the presentation position and the type of text image, but this embodiment can also be applied to cases in which the presentation mode includes, for example, the following modes. ·font Text color Emojis If the presentation manner includes manners such as font, character color, emoticons, etc., processor 12 may present the text image on display 102 in a color or font, etc., that corresponds to the direction of arrival of the speech sound, instead of presenting the text image at a position that corresponds to the direction of arrival of the speech sound. In this embodiment, a case has been described in which text is created based on a voice signal by voice recognition. In this embodiment, the processor 12 may estimate speaker attributes (hereinafter referred to as "speaker attributes"), for example, by voice analysis of speech sounds collected by the microphone 101 or image analysis of images captured by the camera 103. Speaker attributes include, for example, the following: ·mood ·sex ·age The processor 12 determines the presentation mode of the text image, such as the font, the color of the characters, and the pictograms, based on the estimated speaker attributes, thereby enabling the wearer of the display device 1 to easily recognize the speaker attributes.
[0147] In this embodiment, a case has been described in which an image captured by the camera 103 is transmitted to the server device 10c via a network, but this embodiment is also applicable to a case in which an image captured by the camera 103 is not transmitted to the server device 10c. In this case, the image captured by the camera 103 is presented on the display 102.
[0148] In this embodiment, the processor 12 may apply a voice analysis process to the input voice signal, the voice signal being processed, or the voice signal after processing to extract speech sounds from the acquired voice, specify the direction from which the extracted voice comes, and present a text image corresponding to the extracted voice. This omits processing of environmental sounds from voices that include sounds other than speech sounds (e.g., environmental sounds), thereby reducing the processing load on the information processing device.
[0149] In this embodiment, a case has been described in which a speech recognition model stored in the storage device 11 is used, but this embodiment can also be applied to a case in which a speech recognition model stored in a server connectable via the communication interface 14 is used. In this case, steps S111 to S115 in Fig. 5 are executed by a processor of the server.
[0150] Although the embodiments of the present invention have been described in detail above, the scope of the present invention is not limited to the above-described embodiments. Furthermore, the above-described embodiments can be improved or modified in various ways without departing from the spirit of the present invention. Furthermore, the above-described embodiments and modifications can be combined.
[0151] (7) Supplementary Notes The matters explained in the embodiment are additionally noted below.
[0152] (Appendix 1) The apparatus includes a means for acquiring sounds collected by a plurality of microphones (for example, a processor (12) that executes step S110), The method includes: estimating the direction of arrival of the acquired sound (for example, the processor 12 that executes step S111); means for generating a text image corresponding to the captured speech (e.g., the processor 12 executing step S114); The method further comprises: determining a presentation mode of a text image by referring to the estimated direction of arrival (for example, a processor 12 that executes step S115); An information processing device (for example, the controller 10) comprising means for presenting a text image in the determined presentation format (for example, the processor 12 that executes step S116).
[0153] According to (Supplementary Note 1), the direction from which the sound is coming can be easily recognized.
[0154] (Appendix 2) The information processing device according to (Supplementary Note 1), wherein the means for determining the presentation mode determines a presentation mode for presenting the text image at a position according to the estimated arrival direction.
[0155] According to (Supplementary Note 2), the direction from which the sound is coming can be more easily recognized.
[0156] (Appendix 3) The method includes: extracting a sound corresponding to the estimated arrival direction from the acquired sound (for example, the processor 12 that executes step S112); The information processing device according to (Supplementary Note 1) or (Supplementary Note 2), wherein the means for generating a text image generates a text image corresponding to the extracted voice.
[0157] According to (Appendix 3), the direction from which the sound is coming can be accurately recognized.
[0158] (Appendix 4) A means for estimating speaker attributes by analyzing the acquired speech is provided; The information processing device according to any one of (Supplementary Note 1) to (Supplementary Note 3), wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the estimated speaker attribute.
[0159] According to (Appendix 4), speaker attributes can be easily recognized.
[0160] (Appendix 5) The device includes a means (e.g., an input / output interface 13) for acquiring, using a sensor, a sensing signal relating to an area where sound is collected by a plurality of microphones; The information processing device according to any one of (Supplementary Note 1) to (Supplementary Note 4), wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the acquired sensing signal.
[0161] According to (Supplementary Note 5), the accuracy with which text images are presented can be improved.
[0162] (Appendix 6) The information processing device according to (Supplementary Note 5), wherein the sensing signal is an imaging signal obtained by imaging an area using an image sensor.
[0163] According to (Supplementary Note 6), the accuracy with which text images are presented can be improved.
[0164] (Appendix 7) The apparatus includes a means (e.g., an input / output interface 13a) for acquiring an image signal obtained by photographing an area, The apparatus includes a means for converting the acquired photographic signal into a photographic image (for example, a processor 12 that executes step S114), The information processing device according to any one of (Supplementary Note 1) to (Supplementary Note 5), wherein the means for presenting the text image presents the text image by superimposing it on the captured image.
[0165] According to (Supplementary Note 7), the visibility of the relationship between the source of the voice (for example, the speaker) and the text image can be improved.
[0166] (Appendix 8) a means for estimating speaker attributes by analyzing the captured signal; The information processing device according to (Supplementary Note 6) or (Supplementary Note 7), wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the estimated speaker attribute.
[0167] According to (Appendix 8), speaker attributes can be easily recognized.
[0168] (Appendix 9) The method includes: extracting speech sounds uttered by a person from the acquired speech sounds; The means for estimating the direction of arrival estimates the direction of arrival of the extracted voice, the means for generating a text image generates a text image corresponding to the extracted audio; An information processing device according to any one of (Supplementary Note 1) to (Supplementary Note 8).
[0169] According to (Supplementary Note 9), among sounds that include sounds other than speech sounds (for example, environmental sounds), processing for environmental sounds is omitted, so that the processing load on the information processing device can be reduced.
[0170] (Appendix 10) The apparatus includes a means for acquiring sounds collected by a plurality of microphones (for example, a processor (12) that executes step S110), The method includes: estimating the direction of arrival of the acquired sound (for example, the processor 12 that executes step S111); means for generating a text image corresponding to the captured speech (e.g., the processor 12 executing step S114); The method includes: determining a presentation mode of a text image by referring to the estimated direction of arrival (for example, a processor 12 that executes step S111); The method includes: a means for presenting the text image in the determined presentation manner (e.g., the processor 12 that executes step S116); Display device 1.
[0171] According to (Supplementary Note 10), the direction from which the sound is coming can be easily recognized.
[0172] (Appendix 11) The display device according to (Supplementary Note 10), wherein the display device is at least one of a glasses-type display device, a mobile terminal, and a conference system.
[0173] According to (Supplementary Note 11), the direction from which the sound is coming can be easily recognized in various applications.
[0174] (Appendix 12) The display device according to (Appendix 10) or (Appendix 11), wherein the display device is a retinal projection display device.
[0175] According to (Appendix 12), people who suffer from both hearing loss and low vision can easily recognize the direction from which sound is coming.
[0176] (Appendix 13) A program for causing a computer (for example, processor 12) to realize the means described in any one of (Supplementary Note 1) to (Supplementary Note 12).
[0177] According to (Supplementary Note 13), the direction from which the sound is coming can be easily recognized.
[0178] (Appendix 14) A presentation method for presenting an image corresponding to a sound, comprising: The method includes a step of acquiring sounds collected by a plurality of microphones (for example, step S110), The method includes a step of estimating the direction of arrival of the acquired sound (for example, step S111), generating a text image corresponding to the captured speech (e.g., step S114); determining a presentation mode of the text image by referring to the estimated direction of arrival (for example, step S115); and presenting the text image in the determined presentation manner (e.g., step S116). method.
[0179] According to (Supplementary Note 14), the direction from which the sound is coming can be easily recognized. [Explanation of symbols]
[0180] 1: Glass display device 1: Display device 10: Controller 11:Storage device 12: Processor 13: Input / output interface 21: Right temple 22: Right Armor 23: Bridge 24: Left armor 25: Left temple 26: Rim 101: Microphone 102: Display 103: Camera
Claims
1. The device includes a means for acquiring sounds collected by a plurality of microphones, means for estimating a direction of arrival of the acquired sound; a means for extracting a sound corresponding to the estimated arrival direction by beamforming processing based on the estimated arrival direction; means for generating a text image corresponding to the extracted speech; a means for determining a presentation mode so as to present the text image at a position according to the estimated arrival direction; means for presenting the text image in the determined presentation manner; the means for presenting the text image is means for presenting the text image in a format including a character string or a symbol according to the estimated direction of arrival. Information processing device.
2. 2. The information processing device according to claim 1, wherein, when the arrival direction estimating means estimates a plurality of arrival directions, the means for extracting the voice extracts the voice corresponding to each of the plurality of arrival directions by beamforming processing based on the plurality of arrival directions.
3. the means for presenting the text image is means for presenting the text image in a format including a character string or a symbol that evokes the estimated direction of arrival.
2. The information processing device according to claim 1.
4. 4. The information processing device according to claim 1, wherein the means for determining the presentation mode determines the presentation mode so as to present the text image in a format including at least one of a character string and a symbol corresponding to the estimated arrival direction.
5. means for estimating speaker attributes by analyzing the acquired speech; 5. The information processing apparatus according to claim 1, wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the estimated speaker attributes.
6. a means for acquiring, using a sensor, a sensing signal relating to an area where sound is collected by the plurality of microphones; The information processing apparatus according to claim 1 , wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the acquired sensing signal.
7. The information processing apparatus according to claim 6 , wherein the sensing signal is an image signal obtained by capturing an image of the area using an image sensor.
8. a means for acquiring an imaging signal obtained by imaging the area, means for converting the acquired photographic signal into a photographic image; The information processing apparatus according to claim 6 , wherein the means for presenting the text image presents the text image by superimposing it on the photographed image.
9. a means for estimating speaker attributes by analyzing the captured signal, 9. The information processing apparatus according to claim 7, wherein the means for determining the presentation mode determines the presentation mode of the text image by referring to the estimated speaker attribute.
10. a means for extracting speech sounds uttered by a person from the acquired speech sounds, the means for estimating the direction of arrival estimates the direction of arrival of the extracted sound; 10. The information processing apparatus according to claim 1, wherein said text image generating means generates a text image corresponding to said extracted voice.
11. The device includes a means for acquiring sounds collected by a plurality of microphones, means for estimating a direction of arrival of the acquired sound; a means for extracting a sound corresponding to the estimated arrival direction by beamforming processing based on the estimated arrival direction; means for generating a text image corresponding to the extracted speech; a means for determining a presentation mode so as to present the text image at a position according to the estimated arrival direction; a display that presents the text image in the determined presentation manner; the display is a means for presenting the text image in a format including a character string or a symbol according to the estimated direction of arrival. Display device.
12. The display device of claim 11 , wherein the display device is at least one of a glasses-type display device, a mobile terminal, and a conference system.
13. 13. A display device according to claim 11 or claim 12, wherein the display device is a retinal projection display device.
14. means for communicating with a microphone module located separately from the display device; the means for acquiring the sound acquires the sound picked up by the plurality of microphones included in the microphone module via the means for communicating. A display device according to any one of claims 11 to 13.
15. A program for causing a computer to implement the means according to any one of claims 1 to 14.
16. A presentation method in which a display device having a display presents an image corresponding to audio, comprising: The display device comprises: Acquiring sounds collected by a plurality of microphones, estimating a direction of arrival of the acquired sound; extracting a sound corresponding to the estimated arrival direction by beamforming processing based on the estimated arrival direction; generating a text image corresponding to the extracted audio; determining a presentation mode so as to present the text image at a position according to the estimated arrival direction; presenting the text image in the determined presentation manner; the step of presenting the text image is a step of presenting the text image on the display in a format including a character string or a symbol according to the estimated direction of arrival. Presentation method.
Citation Information
Patent Citations
Hearing aid
JP2013236396A