Audio signal processing device, audio signal processing method, and program
The audio signal processing device enhances audio listening by estimating the conversation partner's position and adjusting the beam direction in beamforming, addressing the challenge of directing audio in noisy environments for improved conversation clarity.
Patent Information
- Application Number
- PCT/JP2025/008133
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2025-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing audio signal processing methods fail to determine the listening direction in beamforming based on the user's conversation situation, making it difficult for individuals with hearing impairments or elderly users to easily listen to audio in a direction appropriate for their conversation context, especially in noisy environments.
An audio signal processing device that estimates the position of a conversation partner using a sound collection unit and determines a beam direction in beamforming based on this position, incorporating a position estimation unit and a beam information determination unit to automatically adjust the listening direction.
Enables users to easily listen to sounds from a direction suitable for their conversation situation, enhancing audio listening performance by automatically adjusting the beam direction to focus on the conversation partner, even in changing environments.
Smart Images

Figure JP2025008133_02102025_PF_FP_ABST
Abstract
Description
Audio signal processing device, audio signal processing method, and program
[0001] The present technology relates to an audio signal processing device, an audio signal processing method, and a program, and in particular to an audio signal processing device, an audio signal processing method, and a program that enable a user to easily listen to audio from a direction appropriate for the situation of conversation.
[0002] When users who are concerned about their hearing, such as people with hearing impairments or elderly people whose hearing has deteriorated due to aging, listen to conversations or voice announcements in noisy environments such as train stations, there is a demand for them to be able to selectively listen to only specific sounds.
[0003] Therefore, a method has been devised that uses beamforming technology with multiple microphones to extract sounds from a specific direction and output them with emphasis. However, when a user controls their listening, such as specifying a listening direction, by operating a GUI (Graphical User Interface) on a smartphone or other device, the user often loses opportunities to listen during face-to-face conversations where the desired listening direction changes frequently. Therefore, a hearing assistance device has been devised that selectively reconstructs audio signals from a specific sound source based on the user's facial posture (see, for example, Patent Document 1).
[0004] Japanese Patent Application Laid-Open No. 2022-122533
[0005] However, no method has been devised to determine the listening direction, i.e., the beam direction in beamforming, based on the user's conversation situation. Therefore, it is difficult to satisfy the user's demand for easily listening to audio in a direction appropriate for the user's conversation situation. As a result, the audio listening performance expected by users has not yet been achieved.
[0006] The present technology has been made in view of such circumstances, and enables a user to easily listen to a sound from a direction suitable for the situation of the conversation.
[0007] An audio signal processing device or a program according to one aspect of the present technology is an audio signal processing device or a program for causing a computer to function as an audio signal processing device, including: a position estimation unit that estimates a position of a conversation partner who is in a predetermined conversation situation with a user based on an audio signal acquired by a sound collection unit worn by the user; and a beam information determination unit that determines a beam direction in beamforming of the audio signal based on the position estimated by the position estimation unit.
[0008] An audio signal processing method according to one aspect of the present technology is an audio signal processing method including: an audio signal processing device estimating a position of a conversation partner who is in a predetermined conversation situation with a user, based on an audio signal acquired by a sound collection unit worn by the user; and determining a beam direction in beamforming of the audio signal, based on the position.
[0009] In one aspect of the present technology, the position of a conversation partner in a predetermined conversation situation with the user is estimated based on an audio signal acquired by a sound collection unit worn by the user, and the beam direction in beamforming of the audio signal is determined based on the position.
[0010] The audio signal processing device may be a stand-alone device or a module that is incorporated into another device.
[0011] 1 is a diagram showing an example of the external configuration of an eyeglasses-type audio device which is a first embodiment of an audio signal processing device to which the present technology is applied. FIG. 2 is a block diagram showing an example of the hardware configuration of the eyeglasses-type audio device of FIG. 1. FIG. 3 is a block diagram showing an example of the configuration of an automatic beamforming processing unit. FIG. 4 is a diagram showing a first example of generation of beam information. FIG. 5 is a diagram showing a second example of generation of beam information. FIG. 6 is a diagram showing a third example of generation of beam information. FIG. 7 is a diagram showing an example of determination of a direction mode by the direction mode determination unit of FIG. 3. FIG. 8 is a diagram showing an example of selection of final beam information. FIG. 9 is a first flowchart explaining automatic beamforming processing. FIG. 10 is a flowchart explaining selection processing. FIG. 11 is a block diagram showing an example of a search range. FIG. 12 is a block diagram showing a first example of the configuration of a manual beamforming processing unit. FIG. 13 is a diagram showing an example of determination of a beam direction in front mode. FIG. 14 is a diagram showing an example of determination of a beam direction in manual target fixed mode. FIG. 15 is a diagram showing an example of determination of a beam direction in multiple target mode. FIG. 16 is a diagram showing an example of determination of a beam range. FIG. 17 is a flowchart explaining direction mode determination processing. FIG. 18 is a flowchart explaining front mode processing. FIG. 19 is a flowchart explaining manual target fixed mode processing. FIG. 20 is a flowchart explaining multiple target mode processing. FIG. 21 is a block diagram showing a second example of the configuration of a manual beamforming processing unit. 28 is a diagram showing an example of determining the direction and range of a beam. FIG. 29 is a flowchart illustrating direction range mode determination processing. FIG. 29 is a flowchart illustrating target variable direction range mode processing. FIG. 29 is a flowchart illustrating target fixed direction range mode processing. FIG. 29 is a diagram showing an example of the external configuration of a neckband-type audio device that is a second embodiment of an audio signal processing device to which the present technology is applied. FIG. 30 is a diagram showing an example of the external configuration of an audio system that is a third embodiment of an audio signal processing device to which the present technology is applied. FIG. 31 is a diagram showing an example of determining the direction and range of a beam by the audio system of FIG. 28.
[0012] Hereinafter, modes for carrying out the present technology (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. First embodiment (eyeglasses-type audio device) 2. Second embodiment (neckband-type audio device) 3. Third embodiment (audio system)
[0013] 1. First Embodiment Example of External Configuration of Glasses-Type Audio Device FIG. 1 is a diagram showing an example of the external configuration of a glasses-type (glasses add-on) audio device that is a first embodiment of an audio signal processing device to which the present technology is applied.
[0014] 1 is configured by arranging seven microphones 11-1 to 11-7, a motion sensor 12, a focus button 13, and a touch sensor 14 on an eyeglass frame 10. This eyeglass audio device 1 is worn on the face (head) of a user.
[0015] For example, microphone 11-1 is disposed on the right temple 10a of the eyeglass frame 10. Microphones 11-2 and 11-3 are disposed on the right lens rim 10b, microphone 11-4 is disposed at the center of the left and right lens rims 10b, and microphones 11-5 and 11-6 are disposed on the left lens rim 10b. Microphone 11-7 is disposed on the left temple 10a. Microphones 11-1 to 11-7 are sound collection units that acquire audio signals of ambient sounds. Note that, hereinafter, when there is no need to particularly distinguish between the seven microphones 11-1 to 11-7, they will be collectively referred to as microphones 11. The number of microphones 11 can be any number as long as there is more than one. The microphones 11 can be disposed in any arrangement.
[0016] The motion sensor 12 is disposed, for example, on the left rim 10b and is composed of an acceleration sensor, a gyro sensor, etc. The motion sensor 12 detects the movement of the eyeglass-type audio device 1, i.e., the movement of the user's face. The focus button 13 and touch sensor 14 are disposed, for example, on the left and right temples 10a. The focus button 13 and touch sensor 14 are operated by the user to control the direction and range of beams in beamforming for the audio signal acquired by the microphone 11.
[0017] <Example of Hardware Configuration of Eyeglasses-Type Audio Device> FIG. 2 is a block diagram showing an example of the hardware configuration of the glasses-type audio device 1 of FIG.
[0018] In the eyeglasses-type audio device 1, a CPU (Central Processing Unit) 31, a ROM (Read Only Memory) 32, and a RAM (Random Access Memory) 33 are interconnected by a bus .
[0019] An input / output interface 35 is also connected to the bus 34. An input unit 36, an output unit 37, a storage unit 38, a communication unit 39, and a drive 40 are connected to the input / output interface 35.
[0020] The input unit 36 includes the seven microphones 11, the motion sensor 12, the focus button 13, and the touch sensor 14 shown in FIG. 1 . The output unit 37 includes a display, a speaker, and the like. The storage unit 38 includes a hard disk, a non-volatile memory, and the like. The communication unit 39 includes a network interface, and the like. An audio output device (not shown) is connected to the communication unit 39 via a wired or wireless connection. Examples of audio output devices include earphones, headphones, and hearing aids. This audio output device may be integrated with the eyeglass-type audio device 1 and included in the output unit 37. The drive 40 drives removable media 41, such as a semiconductor memory.
[0021] In the eyeglasses-type audio device 1 configured as described above, the CPU 31 performs various processes by, for example, loading programs stored in the storage unit 38 into the RAM 33 via the input / output interface 35 and the bus 34 and executing them. For example, the CPU 31 performs beamforming processing to perform beamforming on audio signals acquired by the microphone 11 depending on the operation mode. The operation modes include an automatic mode in which the direction and range of the beam in beamforming are set according to the user's conversation situation, and a manual mode in which the direction and range of the beam in beamforming are set based on user input. The operation mode can be selected by the user operating the focus button 13, touch sensor 14, etc.
[0022] The program executed by the CPU 31 can be provided by being recorded on a removable medium 41 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0023] In the eyeglasses-type audio device 1, a program can be installed in the storage unit 38 via the input / output interface 35 by inserting the removable media 41 into the drive 40. The program can also be received by the communication unit 39 via a wired or wireless transmission medium and installed in the storage unit 38. Alternatively, the program can be installed in the ROM 32 or the storage unit 38 in advance.
[0024] The program executed by the eyeglass-type audio device 1 may be a program that is processed in chronological order according to the order described in this specification, or may be a program that is processed in parallel or at the required timing, such as when called.
[0025] <Configuration Example of Automatic Beam Forming Processor> FIG. 3 is a block diagram showing a configuration example of an automatic beam forming processor when the CPU 31 functions as an automatic beam forming processor that performs beam forming processing in automatic mode.
[0026] The automatic beam forming processing unit 60 in FIG. 3 is composed of a position estimation unit 61 , a beam information determination unit 62 , a direction mode determination unit 63 , a storage unit 64 , a beam forming unit 65 , a transmission control unit 66 , and a feature extraction unit 67 .
[0027] The position estimation unit 61 is supplied with an audio signal acquired by the microphone 11 from the microphone 11. The position estimation unit 61 monitors speech based on the audio signal and estimates the positions of one or more conversation partners who are in a predetermined conversation situation with the user.
[0028] For example, the position estimation unit 61 estimates, based on an audio signal, the position of a sound source of a response corresponding to a user's utterance as the position of a conversation partner who is responding to the user's utterance. Based on the audio signal, the position estimation unit 61 estimates, based on the audio signal, the position of a sound source of an utterance of a person other than the user before the user's response corresponding to the user's response as the position of the conversation partner who is responding to the user's utterance. When estimating the position of the conversation partner, the position estimation unit 61 calculates an estimation score that indicates the estimation accuracy. The position estimation unit 61 supplies the estimated positions of one or more conversation partners and the estimation scores for those positions to the beam information determination unit 62 as estimated position information. The position of the sound source or the position of the conversation partner is expressed, for example, by the angle of a line connecting that position and the center of the eyeglass frame 10 relative to the front direction of the center of the eyeglass frame 10. Hereinafter, unless otherwise specified, the "front direction" refers to the front direction of the center of the eyeglass frame 10.
[0029] When the estimated position information supplied from the position estimation unit 61 indicates only one position, the beam information determination unit 62 determines the direction and range of the beam used in beamforming by the beam forming unit 65 based on that position. Then, the beam information determination unit 62 generates beam information indicating the direction and range of that beam. When the estimated position information indicates multiple positions, the beam information determination unit 62 determines the direction, range, and volume of the beam for each position based on the position and the estimation score, and generates beam information indicating the direction, range, and volume of the beam.
[0030] The beam direction is expressed, for example, by the angle of the beam center direction relative to the front direction, and the beam range is expressed by the angle relative to the beam direction. Unless otherwise specified, the beam direction below refers to the angle of the beam center direction relative to the front direction.
[0031] The beam information determination unit 62 supplies the determined beam direction to the direction mode determination unit 63. The beam information determination unit 62 is supplied with a direction mode indicating a method for determining the beam direction from the direction mode determination unit 63. The direction modes include an automatic target variable mode and an automatic target fixed mode. The automatic target variable mode is a mode indicating a method for determining the beam direction based on the position of the conversation partner estimated by the position estimation unit 61. The automatic target fixed mode is a mode indicating a method for determining the beam direction as a direction toward a sound source corresponding to a position estimated by the position estimation unit 61 at a predetermined timing.
[0032] When the direction mode is the automatic target variable mode, the beam information determination unit 62 supplies the beam information generated based on the estimated position information as final beam information to the beam forming unit 65. When the direction mode is the automatic target fixed mode, the beam information determination unit 62 generates a candidate for final beam information based on the movement information representing the movement of the user's face supplied from the movement sensor 12 and the previous final beam information stored in the memory unit 64. The beam information determination unit 62 supplies the candidate for final beam information to the beam forming unit 65.
[0033] The beam information determination unit 62 determines whether the position of the conversation partner has moved based on the movement information, the feature corresponding to the previous final beam information stored in the storage unit 64, and the feature corresponding to the candidate final beam information supplied from the feature extraction unit 67. When the beam information determination unit 62 determines that the position of the conversation partner has moved, it selects, as the candidate final beam information, the candidate final beam information that corresponds to the position of the conversation partner after the movement from among the candidates for final beam information.
[0034] On the other hand, if the beam information determination unit 62 determines that the position of the conversation partner has not moved, it selects, from among the candidates for final beam information, a candidate for final beam information that corresponds to the previous position of the conversation partner, as the final beam information. The beam information determination unit 62 supplies the final beam information to the beam forming unit 65 and also to the storage unit 64 for storage.
[0035] The direction mode determination unit 63 calculates the amount of left-right movement of the user during a predetermined period of time while conversing with the same conversation partner, based on the movement information from the movement sensor 12 and the direction of the beam supplied from the beam information determination unit 62. The direction mode determination unit 63 determines a direction mode based on the amount of movement or the direction of the beam, and supplies the determined direction mode to the beam information determination unit 62.
[0036] The storage unit 64 stores the final beam information supplied from the beam information determination unit 62 and the feature quantity corresponding to the final beam information supplied from the feature quantity extraction unit 67 .
[0037] The beamforming unit 65 performs beamforming on the audio signal supplied from the microphone 11 based on the final beam information or candidate final beam information supplied from the beam information determination unit 62. The beamforming unit 65 supplies the audio signal obtained as a result of beamforming based on the final beam information to the transmission control unit 66 and also supplies the audio signal to the feature extraction unit 67. The beamforming unit 65 supplies the audio signal obtained as a result of beamforming based on the candidate final beam information to the feature extraction unit 67.
[0038] The transmission control unit 66 controls the communication unit 39 in FIG. 2 to transmit the audio signal supplied from the beam forming unit 65 to an audio output device (not shown).
[0039] The feature extraction unit 67 extracts features of the sound source in the beam direction based on the audio signal supplied from the beam forming unit 65. Methods for extracting features of the sound source include a method of performing speaker recognition of the sound source and extracting the recognized speaker as features, and a method of analyzing the sound source and extracting pitch, tone, etc. as features. The feature extraction unit 67 supplies features corresponding to candidates for final beam information to the beam information determination unit 62. The feature extraction unit 67 supplies the features corresponding to the final beam information to the storage unit 64 for storage.
[0040] <First Example of Beam Information Generation> FIG. 4 is a diagram showing a first example of beam information generation by the beam information determination unit 62 of FIG.
[0041] The position estimation unit 61 monitors speech based on the audio signal supplied from the microphone 11. As shown in A of Fig. 4, when the user 81 speaks, the position estimation unit 61 detects the speech. Then, the position estimation unit 61 searches for a sound source other than the user 81 in the range ahead of the user 81, estimates the position of the sound source as the position of a conversation partner who is responding to the user's speech, and calculates an estimation score.
[0042] 4B, when the only position of the conversation partner for which the estimated score is equal to or greater than the threshold is position 91, the beam information determination unit 62 determines, for example, direction 101 toward position 91 as the beam direction and a predetermined range 102 as the beam range. Then, the beam information determination unit 62 generates beam information representing the direction and range. As described above, beam information corresponding to position 91 of the conversation partner is generated simply by the user 81 making an utterance and the conversation partner responding to the utterance.
[0043] In the first embodiment, when the position estimation unit 61 detects the user 81 speaking, it starts searching for sound sources other than the user 81, but it may also start searching for sound sources other than the user 81 by the user 81 operating the focus button 13 or the like.
[0044] <Second Example of Beam Information Generation> FIG. 5 is a diagram showing a second example of beam information generation by the beam information determination unit 62 of FIG.
[0045] The position estimation unit 61 monitors utterances based on the audio signal supplied from the microphone 11. As shown in A of Fig. 5 , when a conversation partner speaks at position 112, the position estimation unit 61 detects the utterance. Then, as shown in B of Fig. 5 , when the user 111 responds to the utterance, the position estimation unit 61 detects the response. The position estimation unit 61 estimates position 112 of the sound source of the utterance before the response corresponding to the response as the position of the conversation partner in a situation where the user is responding to the conversation partner's utterance, and calculates an estimation score.
[0046] If the estimated score of position 112 is equal to or greater than the threshold, the beam information determination unit 62 determines, for example, direction 121 toward position 112 as the beam direction and a predetermined range 122 as the beam range. Then, the beam information determination unit 62 generates beam information representing the direction and range. As described above, when the conversation partner speaks and the user 111 simply responds to the utterance, beam information corresponding to the conversation partner's position 112 is generated.
[0047] <Third Example of Beam Information Generation> FIG. 6 is a diagram showing a third example of beam information generation by the beam information determination unit 62 of FIG.
[0048] 6, the positions of the conversation partners of the user 130 whose estimation scores are equal to or greater than the threshold are three positions 131 to 133. Of the estimation scores of positions 131 to 133, position 131 has the highest estimation score, position 132 has the second highest estimation score, and position 133 has the lowest estimation score.
[0049] In this case, the beam information determination unit 62 determines the direction and range of the beams 141 to 143 for each of the positions 131 to 133. The method for determining the direction and range of the beams 141 to 143 is similar to the method described with reference to FIGS.
[0050] The beam information determination unit 62 also determines the volume of the beams 141 to 143 based on the estimation scores of the positions 131 to 133 so that the volume of the beams 141, 142, and 143 is in descending order. As a result, the volume balance of the sound corresponding to the audio signal obtained as a result of beamforming becomes larger for sound sources located at positions that are more likely to be the positions of the conversation partners of the user 130.
[0051] Therefore, in beamforming, the sound of sound sources other than those at positions that are highly likely to be the positions of the conversation partner is completely deleted, thereby preventing the generation of an audio signal with unnatural sound. When a sound that should be perceived other than the speech of the conversation partner occurs, beamforming can generate an audio signal of sound that includes not only the speech of the conversation partner but also the sound that should be perceived.
[0052] As described above, when the estimated scores for the three positions 131 to 133 are equal to or greater than the threshold, the beam information determination unit 62 determines the direction, range, and volume of the beams 141 to 141 for each of the positions 131 to 133. The beam information determination unit 62 then generates beam information indicating the direction, range, and volume. Therefore, even when the user 130 is having a conversation with three conversation partners who are located at each of the positions 131 to 133, the beam information corresponding to the positions 131 to 133 is generated simply by the user 130 having a conversation with the three conversation partners. In other words, the user 130 does not need to perform any input to adjust the direction or range of the beam in order to generate beam information corresponding to the positions 131 to 133 of the three conversation partners.
[0053] <Example of Direction Mode Determination> FIG. 7 is a diagram showing an example of direction mode determination by the direction mode determination unit 63 of FIG.
[0054] 7A, if the beam direction 161 from the beam information determination unit 62 remains the same for a predetermined time or longer, i.e., if the positional relationship between the user 160 and the conversation partner does not change for a predetermined time or longer, the direction mode determination unit 63 determines the direction mode to be the automatic target variable mode. If the direction mode is the automatic target variable mode and the position of the conversation partner does not change, the beam direction 161 relative to the position of the user 160 changes according to the movement of the user's 160's face.
[0055] 7B, when the amount of horizontal movement of user 160 during a conversation with the same conversation partner is equal to or greater than a threshold, the direction mode determination unit 63 determines the direction mode to be the automatic target fixing mode. When the direction mode is the automatic target fixing mode, the beam direction 162 relative to the position of user 160 remains the same regardless of the movement of the user's 160's face.
[0056] As described above, the direction mode determination unit 63 determines the direction mode based on the amount of horizontal movement of the user 160 who is conversing with a conversation partner who has the same direction as the beam supplied from the beam information determination unit 62. Therefore, the user 160 does not need to perform any input to set the direction mode.
[0057] <Example of Final Beam Information Selection> FIG. 8 is a diagram showing an example of final beam information selection by the beam information determination unit 62 when the direction mode is the automatic target fixing mode.
[0058] 8A, when the position of the conversation partner of user 180 estimated initially after the direction mode is set to the automatic target fixing mode is position 181, the beam information determination unit 62 generates beam information of beam 190 including position 181. Then, the beam information determination unit 62 stores the beam information in the storage unit 64 as final beam information and supplies it to the beam forming unit 65. As a result, the feature extraction unit 67 extracts features of the beamformed audio signal based on the beam information of beam 190 and stores the extracted features in the storage unit 64.
[0059] Next, as shown in B of FIG. 8 , when the face of the user 180 rotates to the right, the beam information determination unit 62 generates a candidate for final beam information based on the movement information and the previous final beam information stored in the storage unit 64. Specifically, the beam information determination unit 62 generates beam information for a beam 191 in a direction 191a obtained by rotating the direction 190a of the beam 190 by the rotation angle of the face of the user 180 as a candidate for final beam information. The beam information determination unit 62 also generates beam information for beams 192 and 193 adjacent to the left and right of the central beam, with the beam 191 as the central beam, as candidates for final beam information. The beam information determination unit 62 then supplies the candidate for final beam information to the beam forming unit 65. As a result, the feature extraction unit 67 extracts features of the beamformed audio signal based on the beam information for each of the beams 191 to 193, which are candidates for final beam information.
[0060] At this time, the feature corresponding to beam 191 including position 181 is similar to the feature corresponding to beam 190 including position 181, which is stored in the storage unit 64. Therefore, the beam information determination unit 62 determines that position 181 of the conversation partner has not moved, and selects the beam information of beam 191 including position 181 from among the candidates for final beam information as final beam information. The beam information determination unit 62 then supplies this final beam information to the beam forming unit 65. As a result, the feature extraction unit 67 extracts the feature of the beamformed audio signal based on the final beam information, and stores the extracted feature in the storage unit 64.
[0061] Next, as shown in Fig. 8C, when position 181 of the conversation partner moves to position 201 without moving the face of user 180, the beam information determination unit 62 generates beam information for each of beams 191 to 193 as candidates for final beam information, similar to the case of Fig. 8B. As a result, the feature extraction unit 67 extracts feature amounts corresponding to each of beams 191 to 193, which are candidates for final beam information, similar to the case of Fig. 8B.
[0062] However, because there is no sound source at position 181, the feature corresponding to beam 191 including position 181 is a feature of a silent audio signal. Therefore, the beam information determination unit 62 determines that position 181 of the conversation partner has moved. Then, from among the candidates for final beam information, the beam information determination unit 62 selects, as final beam information, beam information of beam 193 including position 201, which corresponds to a feature similar to the feature corresponding to beam 191 stored in the storage unit 64. Then, the beam information determination unit 62 supplies this final beam information to the beam forming unit 65. As a result, the feature extraction unit 67 extracts a feature of the beamformed audio signal based on the final beam information and stores it in the storage unit 64.
[0063] As described above, when the direction mode is the automatic target fixing mode, the beam information determination unit 62 selects, as final beam information, beam information of a beam that includes the position of the conversation partner after the movement, in accordance with the movement of the position of the conversation partner. That is, the beam information determination unit 62 selects the final beam information so that the beam follows the position of the moving conversation partner. Therefore, the user can make the beam follow the moving conversation partner without performing any operation.
[0064] 9 and 10 are flowcharts illustrating automatic beamforming processing, which is automatic mode beamforming processing performed by the automatic beamforming processing unit 60 in Fig. 3. This automatic beamforming processing is started, for example, when acquisition of an audio signal by the microphone 11 and detection of motion by the motion sensor 12 are started.
[0065] 9, the position estimation unit 61 starts monitoring the audio signal supplied from the microphone 11 for speech by volume detection, sound type recognition such as "engine sound" or "speech," VAD (Voice Activity Detection), voice recognition, etc. In step S12, the position estimation unit 61 determines whether or not the user's speech has been detected by monitoring the speech. If it is determined in step S12 that the user has spoken, the process proceeds to step S13.
[0066] In step S13, the position estimation unit 61 searches for a sound source of the speech in a range in front of the user by monitoring the speech, and determines whether or not the sound source of the speech in front of the user has been detected. If it is determined in step S13 that a sound source of the speech in front of the user has been detected, the position estimation unit 61 estimates the position of the sound source as the position of a conversation partner who is responding to the user's speech, and proceeds to step S14.
[0067] In step S14, the position estimation unit 61 calculates a response estimation score, which is an estimation score for each position estimated as the position of a conversation partner who is responding to the user's utterance. For example, the position estimation unit 61 calculates the response estimation score Sr using the following equation (1):
[0068] Sr=(α1 / tr+Rr)×β1...(1)
[0069] In equation (1), α and β are constants. tr is the time from the end of the user's speech to the detection of the sound source of the speech in front of the user corresponding to the position of the estimated conversation partner. R is a cosine similarity or the like that indicates the relevance between the user's speech and the speech in front of the user corresponding to the position of the estimated conversation partner.
[0070] The response estimation score may be accumulated each time a corresponding position is estimated. The position estimation unit 61 supplies the estimated position of the conversation partner and the response estimation score Sr to the beam information determination unit 62 as estimated position information, and the process proceeds to step S19.
[0071] On the other hand, if it is determined in step S12 that a user's utterance has not been detected, the process proceeds to step S15. In step S15, the position estimation unit 61 determines whether or not a speech to the user has been detected by monitoring the speech. For example, if the position estimation unit 61 detects a specific phrase such as "excuse me" or a phrase related to a question through voice recognition, it determines that a speech to the user has been detected. If it is determined in step S15 that a speech to the user has been detected, the process proceeds to step S16.
[0072] In step S16, the position estimation unit 61 identifies the position of the sound source of the speech to the user based on the audio signal.
[0073] In step S17, the position estimation unit 61 determines whether or not the user's speech has been detected by monitoring the speech. If it is determined in step S17 that the user's speech has been detected, the position estimation unit 61 estimates the position identified by the processing in step S16 as the position of the conversation partner in a situation where the user is responding to the user's speech, and proceeds to step S18.
[0074] In step S18, the position estimation unit 61 calculates an utterance estimation score Sq, which is an estimation score for each position estimated as the position of a conversation partner in a situation where the user is responding to the user's utterance. For example, the position estimation unit 61 calculates the utterance estimation score Sq using the following formula (2).
[0075] Sq=(α2 / tq+Rq)×β2...(2)
[0076] In equation (2), α and β are constants. t is the time from the end of an utterance to a user corresponding to the estimated position of a conversation partner to the detection of the user's utterance. R is a cosine similarity or the like that indicates the relevance between the utterance to a user corresponding to the estimated position of a conversation partner and the user's utterance.
[0077] The utterance estimation score may be accumulated each time a corresponding position is estimated. The position estimation unit 61 supplies the estimated position of the conversation partner and the utterance estimation score Sq to the beam information determination unit 62 as estimated position information, and the process proceeds to step S19.
[0078] In step S19, the beam information determination unit 62 determines whether or not there is a position of a conversation partner whose response estimation score Sr or utterance estimation score Sq is equal to or greater than a threshold value in the estimated position information supplied from the position estimation unit 61. If it is determined in step S19 that there is a position of a conversation partner whose response estimation score Sr or utterance estimation score Sq is equal to or greater than a threshold value, the process proceeds to step S20.
[0079] In step S20, the beam information determination unit 62 determines whether there are multiple positions of conversational partners whose response estimation score Sr or utterance estimation score Sq is equal to or greater than a threshold value. If there are multiple positions of conversational partners, it is not possible to narrow down the position of the conversational partner to one, or multiple conversational partners exist.
[0080] If it is determined in step S20 that the number of positions of the conversation partner is not multiple, the process proceeds to step S21. In step S21, the beam information determination unit 62 determines the direction of the beam based on the positions of the conversation partner and determines a first range as the beam range. The beam information determination unit 62 then generates beam information representing the direction and range of the beam and supplies the beam direction to the direction mode determination unit 63. The process then proceeds to step S24 in FIG. 10.
[0081] On the other hand, if it is determined in step S20 that there are multiple positions of conversational partners, the process proceeds to step S22. In step S22, the beam information determination unit 62 determines the direction, range, and volume of a beam for each position of a conversational partner based on the position and the response estimation score Sr or the utterance estimation score Sq. The beam information determination unit 62 then generates beam information representing the direction, range, and volume of the beam, and supplies the beam direction to the direction mode determination unit 63. The process then proceeds to step S24.
[0082] If it is determined in step S19 that there is no position of a conversation partner whose response estimation score Sr or utterance estimation score Sq is equal to or greater than the threshold, the process proceeds to step S23. In step S23, the beam information determination unit 62 determines the direction of one beam based on the position of the conversation partner and determines a second range wider than the first range as the beam range. The beam information determination unit 62 then generates beam information representing the direction and range of the beam and supplies the beam direction to the direction mode determination unit 63. The process then proceeds to step S24.
[0083] In step S24, the beam information determination unit 62 determines whether or not the direction mode supplied from the direction mode determination unit 63 is the automatic target fixing mode. If it is determined in step S24 that the direction mode is the automatic target fixing mode, the process proceeds to step S25.
[0084] In step S25, the beam information determination unit 62 determines whether or not the final beam information has not yet been stored in the storage unit 64, i.e., whether or not the beam information has been generated for the first time since the direction mode became the automatic target fixing mode. If it is determined in step S25 that the final beam information has not yet been stored, the process proceeds to step S26.
[0085] In step S26, the beam information determination unit 62 outputs the beam information obtained as a result of any one of the processes in steps S21 to S23 to the beam forming unit 65 as final beam information, and also supplies it to be stored in the storage unit 64. Then, the process proceeds to step S31.
[0086] On the other hand, if it is determined in step S25 that the final beam information has already been stored, the process proceeds to step S27. In step S27, the beam information determination unit 62 calculates the rotation angle of the user's face after the previous final beam information was determined, based on the movement information supplied from the movement sensor 12.
[0087] In step S28, the beam information determination unit 62 rotates the beam direction of the previous final beam information stored in the storage unit 64 based on the rotation angle calculated in step S28, and sets it as the direction of the central beam. The beam information determination unit 62 sets the beam information of the central beam and the beam information of the beams adjacent to the left and right of the central beam as candidates for final beam information. Specifically, when the direction of the central beam is angle θ, the beam information determination unit 62 sets beam information with angle θ and angle θ±a (a>0) as the beam directions as candidates for final beam information. The beam information determination unit 62 supplies the candidates for final beam information to the beam forming unit 65.
[0088] In step S29, the automatic beam forming processing unit 60 performs a selection process to select final beam information from the candidates for final beam information. This selection process will be described later with reference to Fig. 11. After the process of step S29, the process proceeds to step S31.
[0089] On the other hand, if it is determined in step S24 that the direction mode is not the automatic target fixing mode, that is, if the direction mode is the automatic target variable mode, the process proceeds to step S30. In step S30, the beam information obtained as a result of the processing of any one of steps S21 to S23 is output as final beam information to the beam forming unit 65. Then, the process proceeds to step S31.
[0090] In step S31, the beamforming unit 65 performs beamforming on the audio signal supplied from the microphone 11 based on the final beam information supplied from the beam information determination unit 62. The beamforming unit 65 supplies the audio signal obtained as a result of the beamforming to the transmission control unit 66 and the feature extraction unit 67.
[0091] In step S32, the transmission control unit 66 controls the communication unit 39 to transmit the audio signal obtained as a result of the processing in step S31 to an audio output device (not shown).
[0092] In step S33, the feature extraction unit 67 extracts features of the sound source in the beam direction based on the audio signal obtained as a result of the processing in step S31, and supplies and stores them in the storage unit 64. Then, the processing proceeds to step S34.
[0093] On the other hand, if it is determined in step S13, step S15, or step S17 that no detection has been made, the process proceeds to step S34.
[0094] In step S34, the automatic beamforming processing unit 60 determines whether to terminate the automatic beamforming process. For example, if the position estimation unit 61 does not detect any speech for a predetermined period of time, or if a specific speech pattern indicating the end of a conversation is detected, or if the communication unit 39 receives emergency information such as an earthquake early warning, the automatic beamforming processing unit 60 determines to terminate the automatic beamforming process.
[0095] If it is determined in step S34 that the process should not be terminated, the process returns to step S12 in FIG. 9, and the subsequent processes are repeated.
[0096] On the other hand, if it is determined in step S34 that the process is to be ended, the position estimation unit 61 ends monitoring the user's speech, and the automatic beamforming process ends.
[0097] <Description of Selection Process> FIG. 11 is a flowchart illustrating the selection process in step S29 of FIG.
[0098] 11 , the beamforming unit 65 performs beamforming on the audio signal based on the final beam information candidates supplied from the beam information determination unit 62. The beamforming unit 65 supplies the audio signal obtained as a result of the beamforming to the feature extraction unit 67.
[0099] In step S52 , the feature extracting unit 67 extracts the feature of the audio signal obtained as a result of the processing in step S51 and supplies it to the beam information determining unit 62 .
[0100] In step S53, the beam information determination unit 62 determines whether the position of the conversation partner has moved based on the feature extracted by the processing of step S52 and the feature corresponding to the previous final beam information stored by the processing of step S33 in Fig. 10. For example, if the feature corresponding to the extracted center beam is a silent feature and the feature corresponding to the beams other than the center beam is similar to the feature corresponding to the previous final beam information, the beam information determination unit 62 determines that the position of the conversation partner has moved.
[0101] If it is determined in step S53 that the position of the conversation partner has moved, the process proceeds to step S54. In step S54, the beam information determination unit 62 selects, from among the candidates for final beam information, beam information that corresponds to the position of the conversation partner after the movement as the final beam information. Specifically, from among the candidates for final beam information, the beam information determination unit 62 selects, as the final beam information, beam information that corresponds to a feature that is similar to that corresponding to the previous final beam information. The process then proceeds to step S56.
[0102] On the other hand, if it is determined in step S53 that the position of the conversation partner has not moved, the process proceeds to step S55. In step S55, the beam information determination unit 62 selects the beam information of the center beam from among the candidates for final beam information as the final beam information. The process then proceeds to step S56.
[0103] In step S56, the beam information determination unit 62 outputs the final beam information selected by the processing in step S54 or S55 to the beam forming unit 65, and also supplies it to and stores it in the storage unit 64. Then, the processing returns to step S29 in Fig. 10, and proceeds to step S31.
[0104] In the above description, the search range for the position of a sound source other than the user by the position estimation unit 61 is set to the range in front of the user, but it may be set based on the movement of the user.
[0105] <Example of Search Range> FIG. 12 is a diagram showing an example of the search range in this case.
[0106] 12 , when the user 220 speaks while rotating his / her face to the right, the position estimation unit 61 sets, based on the movement of the face, a search range of, for example, a range 222 to the right of the front direction before the rotation, within a range 221 in front of the user 220. That is, when the user 220 speaks to a conversation partner, the user 220 turns his / her face toward the conversation partner, and therefore, there is a high probability that the conversation partner is present in the range 222 corresponding to the movement of the user 220's face. Therefore, the position estimation unit 61 sets the range 222 as the search range. By searching for sound sources other than the user 220 in the range 222, the position estimation unit 61 can improve the accuracy of searching for the position of the conversation partner and reduce processing costs compared to when searching in the range 221.
[0107] As described above, in the automatic beamforming processing unit 60, the position estimation unit 61 estimates the position of a conversation partner who is in a predetermined conversation situation with the user based on the audio signal acquired by the microphone 11. The beam information determination unit 62 determines the beam direction in beamforming of the audio signal based on the position. Therefore, an audio output device (not shown) outputs audio based on the audio signal that has been beamformed based on the beam direction, allowing the user to easily hear audio from a direction appropriate for the conversation situation without performing any operation. This prevents the user from losing opportunities due to operations for controlling listening.
[0108] The position estimation unit 61 may detect a user's response, such as a nod, based on the movement detected by the movement sensor 12, rather than detecting the user's response based on an audio signal. The glasses-type audio device 1 may have a camera, and may detect a user's response, such as a nod, based on an image of the user or conversation partner captured by the camera.
[0109] The direction mode determination unit 63 may determine the direction mode based on the user's operation of the focus button 13 or the like.
[0110] <First Configuration Example of Manual Beamforming Processor> FIG. 13 is a block diagram showing a first configuration example of the manual beamforming processor when the CPU 31 functions as a manual beamforming processor that performs beamforming processing in manual mode.
[0111] The manual beamforming processing unit 260 in FIG. 13 is composed of an input determination unit 261, a direction mode determination unit 262, a direction determination unit 263, a range determination unit 264, a storage unit 265, a beamforming unit 266, a transmission control unit 267, and a feature extraction unit 268.
[0112] The input determination unit 261 receives operation information representing a user operation from the focus button 13. The input determination unit 261 also receives movement information from the movement sensor 12. Based on the operation information, the input determination unit 261 supplies the operation information to the direction mode determination unit 262. The input determination unit 261 calculates a rotation angle of the user's face based on the operation information and the movement information, and supplies the calculated angle to the direction determination unit 263. The input determination unit 261 calculates a rotation angle of the user's face based on the operation information and the movement information, and supplies the calculated angle to the range determination unit 264.
[0113] The direction mode determination unit 262 determines the direction mode based on the operation information supplied from the input determination unit 261. These direction modes include a front mode, a multiple target mode, and a manual target fixation mode. The front mode is a mode in which the front direction when the focus button 13 is operated is determined as the beam direction. The multiple target mode is a mode in which the front direction when the focus button 13 is operated multiple times is determined as the beam direction. The manual target fixation mode is a mode in which the direction toward the sound source in the front direction when the focus button 13 is operated is determined as the beam direction. The direction mode determination unit 262 supplies the determined direction mode to the direction determination unit 263 and the beam forming unit 266.
[0114] When the direction mode supplied from the direction mode determination unit 262 is the front mode, the direction determination unit 263 determines the front direction as the direction of the beam. When the direction mode is the multiple target mode, the direction determination unit 263 determines the direction of the beam based on the rotation angle of the user's face after the previous beam direction supplied from the input determination unit 261 was determined and the previous beam direction stored in the storage unit 265.
[0115] When the direction mode is the manual target fixing mode, the direction determination unit 263 determines a candidate beam direction based on the previous beam direction and the rotation angle of the user's face since the previous beam direction was determined. Then, the direction determination unit 263 selects a beam direction from the candidate beam direction based on the feature corresponding to the candidate beam direction supplied from the feature extraction unit 268 and the feature corresponding to the previous beam direction stored in the storage unit 265. The method of determining the candidate beam direction and the method of selecting the beam direction are similar to the method of generating candidates for final beam information and the method of selecting final beam information in the beam information determination unit 62 of FIG. 3 , and therefore detailed description thereof will be omitted.
[0116] The direction determination unit 263 supplies the beam direction and beam direction candidates to the beam forming unit 266. The direction determination unit 263 supplies the beam direction to the storage unit 265 to be stored.
[0117] The range determination unit 264 determines the range of the beam based on the rotation angle supplied from the input determination unit 261 and supplies it to the beam forming unit 266 .
[0118] The storage unit 265 stores the beam direction supplied from the direction determination unit 263 and the feature amount supplied from the feature amount extraction unit 268 .
[0119] The beam forming unit 266 performs beam forming on the audio signal supplied from the microphone 11 based on the beam direction or candidate beam direction supplied from the direction determining unit 263 and the range supplied from the range determining unit 264 .
[0120] When the direction mode supplied from the direction mode determination unit 262 is the front mode or the multiple target mode, the beamforming unit 266 supplies the audio signal obtained as a result of beamforming to the transmission control unit 267. When the direction mode is the manual target fixed mode, the beamforming unit 266 supplies the audio signal obtained as a result of beamforming to the feature extraction unit 268. When the direction mode is the manual target fixed mode, the beamforming unit 266 supplies the audio signal obtained as a result of beamforming based on the direction and range of the beam to the transmission control unit 267.
[0121] The transmission control unit 267 controls the communication unit 39 to transmit the audio signal supplied from the beam forming unit 266 to an audio output device (not shown).
[0122] 3 , the feature extraction unit 268 extracts features of sound sources in the beam directions based on the audio signals supplied from the beam forming unit 266. The feature extraction unit 268 supplies the features corresponding to candidate beam directions to the direction determination unit 263. The feature extraction unit 268 supplies the features corresponding to the beam directions to the storage unit 64 for storage.
[0123] <Example of Beam Direction in Front Mode> FIG. 14 is a diagram showing an example of the beam direction determined by the direction determination unit 263 in FIG. 13 in the front mode.
[0124] When the user 280 determines the direction mode to the front mode by briefly pressing the focus button 13 once, the direction determination unit 263 determines the beam direction 281 to be the front direction, as shown in A of Fig. 14 . Thereafter, even if the user 280 turns his face to the right, as shown in B of Fig. 14 , the beam direction 281 remains the front direction. Therefore, the user 280 can hear audio in the front direction simply by briefly pressing the focus button 13 once. Therefore, the front mode is an optimal mode for short conversations facing various parties, such as one-on-one customer service.
[0125] <Example of Beam Direction in Manual Target Fixation Mode> FIG. 15 is a diagram showing an example of the beam direction determined by the direction determination unit 263 in FIG. 13 in the manual target fixation mode.
[0126] When the user 290 briefly presses the focus button 13 twice to set the direction mode to the manual target fixed mode, the direction determination unit 263 determines the beam direction 291 to be the forward direction, as shown in A of Fig. 15 . If the user 290 then rotates his / her face to the right, as shown in B of Fig. 15 , the direction determination unit 263 determines the beam direction to be a direction 292 obtained by rotating the beam direction 291 to the left by the rotation angle of the rotation. As a result, the beam direction relative to the position of the user 290 does not change before and after the rotation of the user 290's face. Therefore, the manual target fixed mode is an optimal mode for situations such as when the user 290 is walking next to a conversation partner and the relative positions of the user 290 and the conversation partner do not change significantly, but the user 290's face moves during the conversation.
[0127] <Example of Beam Direction in Multiple Target Mode> FIG. 16 is a diagram showing an example of the determination of the beam direction by the direction determination unit 263 in FIG. 13 in multiple target mode.
[0128] When the user 300 presses the focus button 13 twice for short presses to set the direction mode to the manual target fixation mode, the direction determination unit 263 determines the beam direction 301 to be the front direction, as shown in A of Fig. 16. Thereafter, when the user 300 turns his face to the right and presses the focus button 13 twice for short presses, as shown in B of Fig. 16, the direction mode is set to the multiple target mode.
[0129] In this case, the direction determination unit 263 determines the beam direction to be a direction 302 obtained by rotating the beam direction 301 counterclockwise by the rotation angle of the face of the user 300 and a front direction 303 of the user 300 after the rotation. Therefore, even if the user 300 wants to listen to audio from multiple directions, such as when having a conversation with multiple people, the user 300 can listen to audio from all directions simply by facing each direction and briefly pressing the focus button 13 twice.
[0130] <Example of Determining Beam Range> FIG. 17 is a diagram showing an example of determining the beam range by the range determining unit 264 in FIG.
[0131] As shown in A of Fig. 17 , a user 320 presses focus button 13 with their face facing forward, rotates it to the right by rotation angle 321 as shown in B of Fig. 17 , and then releases focus button 13. In this case, range determination unit 264 determines range 324, which is the range of the beam, from predetermined range 322 when focus button 13 was pressed to predetermined range 323 obtained by rotating predetermined range 322 by rotation angle 321.
[0132] As described above, the user 320 can set the desired listening range within the range of the beam by rotating his / her face by the rotation angle corresponding to the desired listening range while pressing the focus button 13 .
[0133] <Description of Direction Mode Determination Process> Fig. 18 is a flowchart illustrating the direction mode determination process by the manual beamforming processing unit 260 of Fig. 13. This direction mode determination process is started when operation information is supplied from the focus button 13, for example.
[0134] 18, the input determination unit 261 determines whether or not the user has short-pressed the focus button 13 once, based on operation information supplied from the focus button 13. If it is determined in step S101 that the user has short-pressed the focus button 13 once, the input determination unit 261 supplies the operation information to the direction mode determination unit 262, and the process proceeds to step S102.
[0135] In step S102, the direction mode determination unit 262 determines the direction mode to be the front mode based on the operation information supplied from the input determination unit 261, and supplies this to the direction determination unit 263 and the beam forming unit 266. Then, the direction mode determination process ends.
[0136] On the other hand, if it is determined in step S101 that the focus button 13 has not been pressed once, the process proceeds to step S103. In step S103, the input determination unit 261 determines, based on the operation information, whether or not the focus button 13 has been pressed twice. If it is determined in step S103 that the focus button 13 has been pressed twice, the input determination unit 261 supplies the operation information to the direction mode determination unit 262, and the process proceeds to step S104.
[0137] In step S104, the direction mode determination unit 262 determines whether the current direction mode is the manual target fixation mode. If it is determined in step S104 that the current direction mode is not the manual target fixation mode, the process proceeds to step S105. In step S105, the direction mode determination unit 262 sets the direction mode to the manual target fixation mode and supplies this to the direction determination unit 263 and the beamforming unit 266. Then, the direction mode determination process ends.
[0138] On the other hand, if it is determined in step S104 that the mode is the manual target fixation mode, the process proceeds to step S106. In step S106, the direction mode determination unit 262 sets the direction mode to the multiple target mode and supplies this to the direction determination unit 263 and the beamforming unit 266. Then, the direction mode determination process ends.
[0139] If it is determined in step S103 that the user has not pressed the focus button 13 twice, the process proceeds to step S107. In step S107, the input determination unit 261 determines, based on the operation information, whether or not the user has pressed and held the focus button 13. If it is determined in step S107 that the user has pressed and held the focus button 13, the input determination unit 261 supplies the operation information to the direction mode determination unit 262, and the process proceeds to step S108.
[0140] In step S108, the direction mode determination unit 262 cancels the current direction mode and deletes the beam direction and feature amount stored in the storage unit 265. The direction mode determination process then ends. As a result, beamforming is no longer performed on the audio signal acquired by the microphone 11, and output of a wide range of sound based on the acquired audio signal from the audio output device begins.
[0141] On the other hand, if it is determined in step S107 that the focus button 13 has not been pressed for a long time, the direction mode determination process ends.
[0142] <Description of Front Mode Processing> FIG. 19 is a flowchart illustrating front mode processing, which is beamforming processing by the manual beamforming processing unit 260 in FIG. 13 when the direction mode is the front mode.
[0143] 19, the direction determination unit 263 determines the front direction as the beam direction and supplies it to the beam forming unit 266. In step S122, the input determination unit 261 determines whether the user has rotated their face while pressing the focus button 13, based on the operation information and movement information.
[0144] If it is determined in step S122 that the user rotated the face while pressing the focus button 13, the direction determining unit 263 supplies the rotation angle of the user's face to the range determining unit 264, and the process proceeds to step S123.
[0145] In step S123, the range determination unit 264 determines the range of the beam based on the rotation angle supplied from the input determination unit 261, and supplies the determined range to the beam forming unit 266. Then, the process proceeds to step S125.
[0146] On the other hand, if it is determined in step S122 that the user is not rotating his or her face while pressing the focus button 13, the process proceeds to step S124. In step S124, the range determination unit 264 determines a predetermined range as the beam range and supplies it to the beam forming unit 266. Then, the process proceeds to step S125.
[0147] In step S125, the beamforming unit 266 performs beamforming on the audio signal based on the direction determined by the processing of step S121 and the range determined by the processing of step S123 or S124. The beamforming unit 266 supplies the audio signal obtained as a result of the beamforming to the transmission control unit 267.
[0148] In step S126, the transmission control unit 267 controls the communication unit 39 to transmit the audio signal supplied from the beam forming unit 266 to an audio output device (not shown), and the front mode processing then ends.
[0149] <Explanation of Manual Target Fixation Mode Processing> FIG. 20 is a flowchart illustrating manual target fixation mode processing, which is beamforming processing by the manual beamforming processing unit 260 in FIG. 13 when the direction mode is the manual target fixation mode.
[0150] 20, the direction determination unit 263 determines whether or not the previous beam direction has already been stored in the storage unit 265. If it is determined in step S141 that the previous beam direction has not yet been stored, that is, if the first beam direction is to be determined after entering the manual target fixation mode, the process proceeds to step S142.
[0151] In step S142, the direction determining unit 263 determines the front direction as the beam direction, and supplies this to the beam forming unit 266 and also to the storage unit 265 for storage. Then, the process proceeds to step S146.
[0152] On the other hand, if it is determined in step S141 that the previous beam direction has already been stored, the process proceeds to step S143. In step S143, the input determination unit 261 calculates the rotation angle of the user's face since the previous beam direction was determined, based on the movement information supplied from the movement sensor 12. The input determination unit 261 supplies the rotation angle to the direction determination unit 263.
[0153] In step S144, the direction determination unit 263 rotates the previous beam direction stored in the storage unit 265 based on the rotation angle calculated in step S143, and determines a candidate beam direction based on the beam direction after rotation. In step S145, the manual beamforming processing unit 260 performs a selection process similar to the selection process in Fig. 11 on the candidate beam directions, and proceeds to step S146.
[0154] The processing in steps S146 to S150 is similar to the processing in steps S122 to S126 in FIG. 19, and therefore a description thereof will be omitted.
[0155] In step S151, the feature extraction unit 268 extracts features of the sound source in the beam direction based on the audio signal obtained as a result of the processing in step S149, and supplies and stores the extracted features in the storage unit 265. Then, the manual target fixing mode processing ends.
[0156] <Explanation of Multiple Target Mode Processing> FIG. 21 is a flowchart illustrating multiple target mode processing, which is beamforming processing by the manual beamforming processing unit 260 in FIG. 13 when the direction mode is the multiple target mode.
[0157] 21 , the input determination unit 261 calculates the rotation angle of the user's face after the previous beam direction was determined, based on the movement information supplied from the movement sensor 12. The input determination unit 261 supplies the rotation angle to the direction determination unit 263.
[0158] In step S162, the direction determination unit 263 rotates the previous beam direction stored in the storage unit 265 based on the rotation angle calculated in step S161, and determines it as the current beam direction. The direction determination unit 263 supplies this beam direction to the beam forming unit 266 and also to the storage unit 265 for storage.
[0159] In step S163, the direction determining unit 263 determines the front direction as the beam direction, and supplies this to the beam forming unit 266 and also to the storage unit 265 for storage.
[0160] The processing of steps S164 to S168 is the same as the processing of steps S146 to S150 in Fig. 20, and therefore a description thereof will be omitted. After the processing of step S168, the multiple target mode processing ends.
[0161] In addition, even in the multiple target mode, when determining the current beam direction based on the previous beam direction, candidate beam directions may be determined and a selection process may be performed, similar to the manual target fixation mode.
[0162] As described above, the manual beamforming processor 260 allows the user to instantly and freely set the user's desired listening direction as the beam direction by rotating his or her face in the desired listening direction and operating the focus button 13. Furthermore, the user can instantly and freely set the user's desired listening range as the beam range by rotating his or her face by an angle corresponding to the desired listening range while pressing the focus button 13. This prevents the user from losing opportunities to communicate in face-to-face conversations where the desired listening direction or listening range changes frequently.
[0163] 3 may determine the beam range based on the user's operation of the touch sensor 14, similar to the range determination unit 264. The direction mode determination unit 63 may determine the direction mode based on the user's operation of the focus button 13, similar to the direction mode determination unit 262.
[0164] <Second Configuration Example of Manual Beamforming Processing Unit> FIG. 22 is a block diagram showing a second configuration example of the manual beamforming processing unit.
[0165] In the manual beamforming processing unit 360 in Fig. 22 , parts corresponding to those in the manual beamforming processing unit 260 in Fig. 13 are assigned the same reference numerals. Therefore, description of those parts will be omitted as appropriate, and the description will focus on parts that differ from the manual beamforming processing unit 260. The manual beamforming processing unit 360 is composed of a transmission control unit 267, a feature extraction unit 268, an input determination unit 361, a direction range mode determination unit 362, a direction range determination unit 363, a beamforming unit 364, and a storage unit 365. The manual beamforming processing unit 360 determines the direction and range of the beam based on the user's operation of the touch sensor 14.
[0166] Specifically, operation information representing a user operation is input to the input determination unit 361 from the focus button 13 and the touch sensor 14. Movement information is also input to the input determination unit 361 from the movement sensor 12. Based on the operation information from the focus button 13, the input determination unit 361 supplies the operation information to the direction range mode determination unit 362. The input determination unit 361 supplies operation information from the touch sensor 14 to the direction range determination unit 363. The input determination unit 361 calculates a rotation angle of the user's face based on the operation information and movement information, and supplies the calculated angle to the direction range determination unit 363.
[0167] The direction range mode determination unit 362 determines a direction mode based on the operation information supplied from the input determination unit 361. These direction range modes include a target variable direction range mode and a target fixed direction range mode. The target variable direction range mode is a mode in which the direction and range of the beam determined based on the operation of the touch sensor 14 are maintained. The target fixed direction range mode is a mode in which the direction toward the sound source of the beam direction determined based on the operation of the touch sensor 14 is determined as the beam direction, and the range of the beam determined based on the operation of the touch sensor 14 is maintained. The direction range mode determination unit 362 supplies the determined direction range mode to the direction range determination unit 363 and the beam forming unit 364.
[0168] When the direction range mode supplied from the direction range mode determination unit 362 is the target variable direction range mode, the direction range determination unit 363 determines the direction and range of the beam based on the operation information of the touch sensor 14 supplied from the input determination unit 361.
[0169] When the direction mode is the target-fixed direction range mode, the direction range determination unit 363 determines beam direction candidates based on the rotation angle supplied from the input determination unit 361 and the previous beam direction stored in the storage unit 365. The direction range determination unit 363 also determines the range of the previous beam stored in the storage unit 365 as the range of the current beam. Then, the direction range determination unit 363 selects a beam direction from the beam direction candidates based on the feature amounts corresponding to the beam direction candidates supplied from the feature amount extraction unit 268 and the feature amounts corresponding to the previous beam direction stored in the storage unit 365. The method of determining beam direction candidates and selecting beam directions in the direction range determination unit 363 is similar to the method of determining beam direction candidates and selecting beam candidates in the direction determination unit 263, and therefore detailed description thereof will be omitted.
[0170] The direction range determination unit 363 supplies the determined beam direction or beam direction candidates and range to the beam forming unit 364. The direction range determination unit 363 supplies the beam direction and range to the storage unit 365 to store them.
[0171] The beam forming unit 364 performs beam forming on the audio signal supplied from the microphone 11 based on the beam direction or beam direction candidates and range supplied from the direction range determination unit 363 .
[0172] When the direction range mode supplied from the direction range mode determination unit 362 is the target variable direction range mode, the beamforming unit 364 supplies the audio signal obtained as a result of beamforming to the transmission control unit 267. When the direction range mode is the target fixed direction range mode, the beamforming unit 364 supplies the audio signal obtained as a result of beamforming to the feature extraction unit 268. When the direction range mode is the target fixed direction range mode, the beamforming unit 364 supplies the audio signal obtained as a result of beamforming based on the direction and range of the beam to the transmission control unit 267.
[0173] The storage unit 365 stores the beam direction and range supplied from the direction range determination unit 363 and the feature amount supplied from the feature amount extraction unit 268 .
[0174] <Example of Determination of Beam Direction and Range> FIG. 23 is a diagram showing an example of determination of the beam direction and range by the direction range determination unit 363 in FIG. 22 in the target variable direction range mode.
[0175] As shown in A of Fig. 23 , for example, the user 380 touches the front end of the touch sensor 14 on the left rim 10b of the eyeglasses-type audio device 1 that the user 380 is wearing with a finger or the like and moves it backward by a certain amount. In this case, as shown in B of Fig. 23 , the direction range determination unit 363 determines the front direction 381 of the user 380 as the direction of the beam, and determines the angle 382 corresponding to the amount and direction of finger movement as the range of the beam. Therefore, the user 380 can set the desired listening direction and listening range as the direction and range of the beam by simply touching the position of the touch sensor 14 that corresponds to the desired listening direction with a finger or the like and moving it by an amount corresponding to the desired listening range.
[0176] Note that the direction and range of the initial beam in the target fixed direction range mode are also determined in the same way as the direction and range of the beam in the target variable direction range mode described with reference to FIG.
[0177] <Description of Direction Range Mode Determination Process> Fig. 24 is a flowchart illustrating the direction range mode determination process performed by the manual beamforming processing unit 360 in Fig. 22. This direction range mode determination process is started when operation information is supplied from the focus button 13, for example.
[0178] 24, the input determination unit 361 determines whether or not the user has short-pressed the focus button 13 once, based on the operation information supplied from the focus button 13. If it is determined in step S101 that the user has short-pressed the focus button 13 once, the input determination unit 361 supplies the operation information to the direction range mode determination unit 362, and the process proceeds to step S182.
[0179] In step S182, the direction range mode determination unit 362 determines the direction range mode to be the target variable direction range mode based on the operation information supplied from the input determination unit 361, and supplies this to the direction range determination unit 363 and the beam forming unit 364. Then, the direction range mode determination process ends.
[0180] On the other hand, if it is determined in step S181 that the focus button 13 has not been pressed once, the process proceeds to step S183. In step S183, the input determination unit 361 determines, based on the operation information, whether or not the user has pressed the focus button 13 twice. If it is determined in step S183 that the user has pressed the focus button 13 twice, the input determination unit 361 supplies the operation information to the direction range mode determination unit 362, and the process proceeds to step S184.
[0181] In step S184, the direction range mode determination unit 362 determines the direction range mode to be the target-fixed direction range mode, and supplies this to the direction range determination unit 363 and the beamforming unit 364. Then, the direction range mode determination process ends.
[0182] On the other hand, if it is determined in step S183 that the user has not pressed the focus button 13 twice briefly, the process proceeds to step S185. In step S185, the input determination unit 361 determines, based on the operation information, whether or not the user has pressed and held down the focus button 13. If it is determined in step S185 that the user has pressed and held down the focus button 13, the input determination unit 361 supplies the operation information to the direction range mode determination unit 362, and the process proceeds to step S186.
[0183] In step S186, the direction range mode determination unit 362 cancels the current direction range mode and deletes the beam direction, range, and feature amount stored in the storage unit 365. The direction range mode determination process then ends. As a result, beamforming is no longer performed on the audio signal acquired by the microphone 11, and output of a wide range of sound based on the acquired audio signal from the audio output device begins.
[0184] On the other hand, if it is determined in step S185 that the focus button 13 has not been pressed and held, the direction range mode determination process ends.
[0185] <Description of Target Variable Direction Range Mode Processing> FIG. 25 is a flowchart illustrating target variable direction range mode processing, which is beamforming processing when the direction range mode by the manual beamforming processing unit 360 in FIG. 22 is the target variable direction range mode.
[0186] 25, the input determination unit 361 determines whether or not operation information has been supplied from the touch sensor 14. If it is determined in step S201 that operation information has been supplied, the input determination unit 361 supplies the operation information to the direction range determination unit 363, and the process proceeds to step S202.
[0187] In step S202, the direction range determination unit 363 determines the direction and range of the beam based on the operation information supplied from the input determination unit 361, and supplies the determined direction and range to the beam forming unit 364. The processing of steps S203 and S204 is similar to the processing of steps S125 and S126 in Fig. 19, and therefore description thereof will be omitted. After the processing of step S204, the target variable direction range mode processing ends.
[0188] On the other hand, if it is determined in step S201 that operation information has not been supplied, the target variable direction range mode process ends.
[0189] <Description of Target Fixed Direction Range Mode Processing> FIG. 26 is a flowchart illustrating target fixed direction range mode processing, which is beamforming processing when the direction range mode by the manual beamforming processing unit 360 in FIG. 22 is the target fixed direction range mode.
[0190] 26, the direction range determination unit 363 determines whether or not the previous beam direction and range have already been stored in the storage unit 365. If it is determined in step S221 that the beam direction and range have not yet been stored, that is, if the first beam direction and range are to be determined after entering the target fixed direction range mode, the process proceeds to step S222.
[0191] The processing in steps S222 and S223 is the same as the processing in steps S201 and S202 in Fig. 25, and therefore a description thereof will be omitted. After the processing in step S223, the process proceeds to step S227.
[0192] On the other hand, if it is determined in step S221 that the direction and range of the previous beam have already been stored, the process proceeds to step S224. The processes of steps S224 to S226 are the same as the processes of steps S143 to S145 in Fig. 20, and therefore a description thereof will be omitted. After the process of step S226, the process proceeds to step S227.
[0193] The processing in steps S227 to S229 is the same as the processing in steps S149 to S151 in Fig. 20, and therefore a description thereof will be omitted. After the processing in step S151, the target fixed direction range mode processing ends.
[0194] On the other hand, if it is determined in step S221 that operation information has not been supplied, the target fixed direction range mode process ends.
[0195] As described above, in the manual beamforming processor 360, the user can instantly and freely set the beam direction and range to the user's desired listening direction and listening range by operating the touch sensor 14. This prevents the user from losing opportunities to change during face-to-face conversations or other situations where the desired listening direction or listening range changes frequently.
[0196] In addition, the manual beam forming processing unit 360 may also be configured to set the directions and ranges of multiple beams in response to the operation of the touch sensor 14.
[0197] The eyeglasses-type audio device 1 may have a generation mode indicating an audio signal generation method, which includes a narrow-range mode in which automatic beamforming processing or manual beamforming processing is performed, and a wide-range mode in which beamforming is not performed. In this case, the eyeglasses-type audio device 1 sets the generation mode to the narrow-range mode or the wide-range mode, for example, by the user operating the focus button 13 or the touch sensor 14. This allows the user to instantly switch the generation mode. As a result, for example, a user serving customers at a restaurant can set the generation mode to the wide-range mode when they need to hear the sounds of the entire restaurant, and can set the generation mode to the narrow-range mode when they need to have a conversation with a customer in front of them.
[0198] If the user is a restaurant employee or the like and the optimal generation mode is determined depending on the location, the eyeglass-type audio device 1 may set the generation mode based on the user's location. For example, if the user is located at a cash register or in a hall, where the user needs to hear the sounds from the entire surrounding area, the wide-range mode is set as the generation mode. On the other hand, if the user is located at a table or other location where face-to-face service is provided, the narrow-range mode is set as the generation mode.
[0199] The eyeglass-type audio device 1 may be equipped with a camera or the like, and may use images captured by the camera to detect the approach of people from all directions (360 degrees) and the movement of the mouths of people in the vicinity. In this case, for example, when a person moves within a predetermined distance from the user or when a person in the vicinity moves their mouth, it is detected that the person is speaking to the user, and the user is notified. Note that if it is detected that a person in an invisible position, such as behind the user, is speaking to the user, the direction of the person may also be notified to the user by vibration, a UI (User Interface), or the like. When a person moves within a predetermined distance from the user, the position estimation unit 61 may estimate the position of the person as the position of the conversation partner.
[0200] The eyeglass-type audio device 1 may use images captured by the camera to determine whether or not there is a person around the user, and if it determines that there is no person present, may decide to terminate the automatic beamforming process in step S34 of Figure 10.
[0201] The audio glasses 1 may be configured to perform only one of the automatic beam forming process and the manual beam forming process. If the audio glasses 1 performs only the automatic beam forming process, the focus button 13 and the touch sensor 14 may not be provided.
[0202] 2. Second Embodiment <External Configuration Example of Neckband-Type Audio Device> FIG. 27 is a diagram showing an external configuration example of a neckband-type audio device that is a second embodiment of an audio signal processing device to which the present technology is applied.
[0203] The neckband type audio device 400 of FIG. 27 is configured by connecting earphones 402-1 and 402-2 for the left and right ears to a neckband 401.
[0204] Neckband 401 includes six microphones 411-1 to 411-6, a motion sensor 412, a focus button 413, and a touch sensor 414. Neckband audio device 400 is worn around the user's neck.
[0205] Microphones 411-1 to 411-3 are arranged, for example, on the inside of the left side of neckband 401, and microphones 411-4 to 411-6 are arranged, for example, on the inside of the right side of neckband 401. Microphones 411-1 to 411-6 are sound pickup units that acquire audio signals of surrounding sounds. Note that, hereinafter, unless there is a need to particularly distinguish between the six microphones 411-1 to 411-6, they will be collectively referred to as microphones 411. The number of microphones 411 can be any number as long as there is more than one. The microphones 411 can be arranged in any manner.
[0206] The motion sensor 412 is disposed, for example, at the left end of the neckband 401, and is constituted by an acceleration sensor, a gyro sensor, etc. The motion sensor 412 detects the movement of the neckband-type audio device 400, i.e., the movement of the user's body. The focus button 413 is disposed, for example, at the right end of the neckband 401. The touch sensor 414 is disposed, for example, on the outside of the left and right sides of the neckband 401. The focus button 413 and the touch sensor 414 are operated by the user when controlling the direction and range of the beam in beamforming based on the audio signal acquired by the microphone 411.
[0207] Earphone 402-1 is an earphone for the right ear and includes a microphone 421-1 and a motion sensor 422-1. Earphone 402-2 is an earphone for the left ear and includes a microphone 421-2 and a motion sensor 422-2. Motion sensor 422-1 (422-2) is configured with an acceleration sensor, gyro sensor, etc., and detects the movement of earphone 402-1 (402-2), i.e., the movement of the user's face. In the example of FIG. 1, earphones 402-1 and 402-2 are connected to neckband 401 by wire, but they may also be connected wirelessly.
[0208] The processing of the neckband-type audio device 400 is similar to that of the eyeglasses-type audio device 1, and therefore a description thereof will be omitted. The microphones 411, 421-1, and 421-2 correspond to the microphone 11. The motion sensor 412 and the motion sensors 422-1 and 422-2 correspond to the motion sensor 12. Specifically, the facial movement of the user in the eyeglasses-type audio device 1 is calculated based on the body movement of the user detected by the motion sensor 412 and the facial movement of the user detected by the motion sensors 422-1 and 422-2. The body movement of the user detected by the motion sensor 412 may be used instead of the facial movement of the user. The focus button 13 corresponds to the focus button 413, and the touch sensor 14 corresponds to the touch sensor 414. Audio output devices (not shown) correspond to the earphones 402-1 and 402-2.
[0209] The audio signals acquired by microphones 421-1 and 421-2 can also be used to remove noise from the audio signals after beamforming. The neck rotation angle may be calculated based on the user's body movements detected by motion sensor 412 and the user's facial movements detected by motion sensors 422-1 and 422-2. In this case, for example, the beam range is determined so that it becomes wider as the neck rotation angle increases.
[0210] 3. Third Embodiment <External Configuration Example of Audio System> FIG. 28 is a diagram showing an external configuration example of an audio system that is a third embodiment of an audio signal processing device to which the present technology is applied.
[0211] 28 is configured with a glasses-type audio device 501 and a ring-type device 502. In the audio system 500, the ring-type device 502 corresponds to the focus button 13 and the touch sensor 14.
[0212] Specifically, the audio glasses 501 differs from the audio glasses 1 in that it does not include the focus button 13 and the touch sensor 14, but is otherwise configured in the same manner as the audio glasses 1.
[0213] The ring-shaped device 502 is worn on the user's finger and includes a focus button 511 and a motion sensor 512. The motion sensor 512 is configured with an acceleration sensor, a gyro sensor, or the like.
[0214] The processing of the audio system 500 is similar to the processing of the eyeglasses-type audio device 1, except for the method of determining the beam direction and range in the target variable direction range mode and the target fixed direction range mode. Therefore, the following description will focus on the method of determining the beam direction and range in the target variable direction range mode and the target fixed direction range mode.
[0215] <Example of Determining Beam Direction and Range> FIG. 29 is a diagram showing an example of determining the beam direction and range in the target variable direction range mode by the audio system 500 of FIG.
[0216] As shown in A of Fig. 29 , for example, a user 530 points the finger wearing the ring-shaped device 502 in a desired listening direction and rotates the hand to the right while pressing the focus button 511. In this case, as shown in B of Fig. 29 , the direction 531 of the finger is determined as the direction of the beam, and an angle 532 corresponding to the rotation angle of the hand is determined as the range of the beam. Therefore, the user 530 can set a desired listening direction and listening range simply by pointing the finger in the desired listening direction and rotating the hand by an amount corresponding to the desired listening range.
[0217] Note that the direction and range of the initial beam in the target fixed direction range mode are also determined in the same way as the direction and range of the beam in the target variable direction range mode described with reference to FIG.
[0218] The eyeglasses-type audio device 501 may include a focus button 13 and a touch sensor 14. The audio system 500 may include a neckband-type audio device 400 instead of the eyeglasses-type audio device 501.
[0219] The above-described series of processes can be executed by software or by hardware.
[0220] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0221] For example, it is possible to adopt a form in which all or part of the above-described embodiments are combined.
[0222] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed jointly by multiple devices via a network. For example, automatic beamforming processing and manual beamforming processing may be performed on the cloud.
[0223] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0224] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0225] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0226] The present technology may have the following configurations. (1) An audio signal processing device comprising: a position estimation unit that estimates a position of a conversation partner who is in a predetermined conversation situation with a user, based on an audio signal acquired by a sound collection unit worn by the user; and a beam information determination unit that determines a beam direction in beamforming of the audio signal, based on the position estimated by the position estimation unit. (2) The audio signal processing device according to (1), wherein the predetermined conversation situation is a situation in which the conversation partner is responding to an utterance of the user, and the position estimation unit is configured to estimate a position of a sound source of the response as the position of the conversation partner, based on the audio signal. (3) The audio signal processing device according to (1), wherein the predetermined conversation situation is a situation in which the user is responding to an utterance of the conversation partner, and the position estimation unit is configured to estimate a position of a sound source of the utterance as the position of the conversation partner, based on the audio signal. (4) The audio signal processing device according to any one of (1) to (3), wherein the beam information determination unit is configured to determine a direction of the beam for each of the positions when a plurality of positions are estimated by the position estimation unit. (5) The audio signal processing device according to any one of (1) to (4), wherein the position estimation unit is also configured to calculate a score representing estimation accuracy of the position, and wherein the beam information determination unit is configured to determine a volume of the beam for each of the positions based on the score when a plurality of positions are estimated by the position estimation unit. (6) The audio signal processing device according to any one of (1) to (4), wherein the beam information determination unit is also configured to determine a range of the beam. (7) The audio signal processing device according to (6), wherein the position estimation unit is also configured to calculate a score representing estimation accuracy of the position, and wherein the beam information determination unit is configured to determine a range of the beam when there is no position where the score is equal to or greater than a threshold value, so that the range of the beam is wider than when there is a position where the score is equal to or greater than the threshold value.(8) The audio signal processing device according to (6), wherein the beam information determination unit is configured to determine a range of the beam based on an operation of the user. (9) The audio signal processing device according to any of (1) to (8), further comprising a direction mode determination unit that determines a direction mode indicating a method for determining a direction of the beam. (10) The audio signal processing device according to (9), wherein the direction mode determination unit is configured to determine the direction mode to be a target variable mode that indicates a determination method for determining the direction of the beam based on the position estimated by the position estimation unit, or a target fixed mode that indicates a determination method for determining the direction of the beam as a direction toward a sound source corresponding to the position estimated by the position estimation unit at a predetermined timing. (11) The audio signal processing device according to (10), wherein the beam information determination unit is configured to determine the direction of the beam such that the beam follows the moving position when the direction mode is the target fixed mode. (12) The audio signal processing device according to any of (9) to (11), wherein the direction mode determination unit is configured to determine the direction mode based on an amount of movement of the user. (13) The audio signal processing device according to any of (9) to (11), further comprising: the direction mode determination unit determines the direction mode based on an operation of the user. (14) The audio signal processing device according to any of (1) to (13), wherein the position estimation unit searches for a sound source other than the user based on the audio signal and estimates the position of the searched sound source as the position of the conversation partner. (15) The audio signal processing device according to (14), wherein the position estimation unit determines a search range for the sound source based on a movement of the user.(16) The audio signal processing device according to any one of (1) to (15), further comprising: a beamforming unit that performs the beamforming on the audio signal based on the beam direction determined by the beam information determination unit; and a transmission control unit that controls transmission of the audio signal obtained as a result of the beamforming by the beamforming unit to an audio output device. (17) The audio signal processing device according to any one of (1) to (16), further comprising: the sound collection unit. (18) An audio signal processing method, including: an audio signal processing device estimating a position of a conversation partner who is in a predetermined conversation situation with the user, based on an audio signal acquired by a sound collection unit worn by the user, and determining a beam direction in beamforming of the audio signal based on the position. (19) A program for causing a computer to function as an audio signal processing device, comprising: a position estimation unit that estimates the position of a conversation partner in a predetermined conversation situation with a user based on an audio signal acquired by a sound collection unit worn by the user; and a beam information determination unit that determines the direction of a beam in beamforming of the audio signal based on the position estimated by the position estimation unit.
[0227] REFERENCE SIGNS LIST 1 eyeglass-type audio device, 11-1 to 11-7 microphones, 61 position estimation unit, 62 beam information determination unit, 63 direction mode determination unit, 65 beam forming unit, 66 transmission control unit, 400 neckband-type audio device, 411-1 to 411-6 microphones, 501 eyeglass-type audio device
Claims
1. An audio signal processing device comprising: a position estimation unit that estimates the position of a conversation partner in a predetermined conversation situation with a user based on an audio signal acquired by a sound collection unit worn by the user; and a beam information determination unit that determines the beam direction in beamforming of the audio signal based on the position estimated by the position estimation unit.
2. The audio signal processing device according to claim 1, wherein the predetermined conversation situation is a situation in which the conversation partner is responding to an utterance of the user, and the position estimation unit is configured to estimate the position of the sound source of the response as the position of the conversation partner based on the audio signal.
3. The audio signal processing device according to claim 1, wherein the predetermined conversation situation is a situation in which the user is responding to an utterance of the conversation partner, and the position estimation unit is configured to estimate the position of the sound source of the utterance as the position of the conversation partner based on the audio signal.
4. An audio signal processing device according to claim 1, wherein the beam information determination unit is configured to determine the beam direction for each of the positions when the number of positions estimated by the position estimation unit is multiple.
5. The audio signal processing device of claim 1, wherein the position estimation unit also calculates a score representing the estimation accuracy of the position, and the beam information determination unit is configured to determine the volume of the beam for each position based on the score when the number of positions estimated by the position estimation unit is multiple.
6. An audio signal processing device according to claim 1, wherein the beam information determination unit is also configured to determine the range of the beam.
7. The audio signal processing device of claim 6, wherein the position estimation unit also calculates a score representing the estimation accuracy of the position, and the beam information determination unit is configured to determine the range of the beam so that it is wider when there is no position where the score is equal to or greater than a threshold value than when there is a position where the score is equal to or greater than the threshold value.
8. An audio signal processing device according to claim 6, wherein the beam information determination unit is configured to determine the range of the beam based on an operation by the user.
9. The audio signal processing device according to claim 1, further comprising: a direction mode determination unit that determines a direction mode indicating a method for determining the direction of the beam.
10. An audio signal processing device as described in claim 9, wherein the direction mode determination unit is configured to determine the direction mode to be a target variable mode indicating a determination method for determining the direction of the beam based on the position estimated by the position estimation unit, or a target fixed mode indicating a determination method for determining the direction of the beam as the direction toward a sound source corresponding to the position estimated by the position estimation unit at a predetermined timing.
11. An audio signal processing device according to claim 10, wherein the beam information determination unit is configured to determine the direction of the beam so that the beam follows the moving position when the direction mode is the target fixed mode.
12. The audio signal processing device according to claim 9, wherein the direction mode determination unit is configured to determine the direction mode based on the amount of movement of the user.
13. The audio signal processing device according to claim 9, further comprising: the direction mode determination unit determines the direction mode based on an operation by the user.
14. The audio signal processing device according to claim 1, wherein the position estimation unit is configured to search for a sound source other than the user based on the audio signal and estimate the position of the searched sound source as the position of the conversation partner.
15. The audio signal processing device according to claim 14, wherein the position estimation unit is configured to determine a search range for the sound source based on the movement of the user.
16. The audio signal processing device according to claim 1, further comprising: a beamforming unit that performs beamforming on the audio signal based on the beam direction determined by the beam information determination unit; and a transmission control unit that controls transmission of the audio signal obtained as a result of the beamforming by the beamforming unit to an audio output device.
17. The audio signal processing device according to claim 1, further comprising the sound pickup unit.
18. An audio signal processing method, comprising: an audio signal processing device estimating the position of a conversation partner in a predetermined conversation situation with a user based on an audio signal acquired by a sound pickup unit worn by the user; and determining a beam direction in beamforming of the audio signal based on the position.
19. A program for causing a computer to function as an audio signal processing device comprising: a position estimation unit that estimates the position of a conversation partner in a given conversation situation with a user based on an audio signal acquired by a sound pickup unit worn by the user; and a beam information determination unit that determines the beam direction in beamforming of the audio signal based on the position estimated by the position estimation unit.
Citation Information
Patent Citations
Audio spatialization and enhancement across multiple headsets
JP2022531067A
Method and apparatus for hearing assistance in multiple-talker settings
US20150016644A1
Signal processing apparatus and signal processing method
WO2011105003A1
Speech processing device and speech processing method
WO2012042768A1
Microphone array control device
WO2014199453A1