Speaker device, sound signal processing method, and program
The speaker device outputs sound from outside the ear canal to simulate spatial sound, addressing the limitations of in-ear headphones by generating a three-dimensional reverberation effect, enabling a more immersive auditory experience.
Patent Information
- Application Number
- JP2024045854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-10-03
AI Technical Summary
In-ear headphones fail to express the spatial spread of sound, causing users to perceive sound images as localized near their heads, rather than in a larger environment.
A speaker device that outputs sound from outside the ear canal, utilizing a microphone to capture sound, a processor to generate a three-dimensional reverberation effect, and speakers positioned near the ears to simulate spatial sound based on head-related transfer functions.
The device allows users to experience a spatial spread of sound, replicating the reverberation of environments like a concert hall, enhancing the auditory experience beyond what in-ear headphones can provide.
Smart Images

Figure 2025145585000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to a speaker device, a sound signal processing method, and a program. [Background technology]
[0002] Patent Document 1 describes an audio processing system using headphones with a microphone. The audio processing system of Patent Document 1 includes left and right speakers provided on both sides of a headband, a microphone, and a smartphone. The sound of a flute played by a user is input to the microphone. The flute sound input to the microphone is subjected to reverberation processing by a reverberation application program on the smartphone. The smartphone transmits a reverberation signal to the left and right speakers. The left and right speakers emit sounds based on the reverberation signal.
[0003] Patent Document 2 describes first to fourth speakers that convert audio signals output from an audio reproduction control device into audio. The first to fourth speakers are arranged around the driver in the vehicle. As an example of localization control, if a localization control signal among the control signals is a signal that localizes a sound image at a short distance to the right of the driver, filter coefficients corresponding to head-related transfer functions from the localization position of each speaker to the driver are set so that the localization position of the sound image output from the first and second speakers, or the third and fourth speakers, is at a short distance to the right of the driver. Each FIR filter performs a convolution operation on the digital signal according to the filter coefficient. This allows localization control to be performed. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2018-74499 [Patent Document 2] Japanese Patent Application Laid-Open No. 2009-4937 Summary of the Invention [Problem to be solved by the invention]
[0005] With audio equipment that outputs sound from within the ear canal, such as in-ear headphones, the user feels as if the sound image is localized close to the user's head. In other words, in-ear headphones cannot express the spatial spread of sound that is produced when listening to sound in a large room.
[0006] An object of one embodiment of the present invention is to provide a speaker device that can express a spatial spread of sound that cannot be achieved with in-ear headphones. [Means for solving the problem]
[0007] A speaker device according to an embodiment of the present invention includes: an acquisition unit that receives a sound based on a user's action as a sound signal; a processing unit that generates a reverberation signal having a three-dimensional reverberation effect based on the sound signal received by the acquisition unit and a head-related transfer function; and an output unit that is placed in contact with or close to the user's body and outputs sound based on the reverberation signal with the three-dimensional reverberation effect from outside the user's ear canal. [Effects of the Invention]
[0008] The speaker device according to one embodiment of the present invention can express a spatial spread of sound that cannot be achieved with in-ear headphones. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a top view of the shoulder speaker 1. FIG. [Figure 2] FIG. 2 is a diagram showing a user wearing the shoulder speaker 1. As shown in FIG. [Figure 3] FIG. 3 is a diagram showing a user wearing the shoulder speaker 1 and playing a musical instrument. [Figure 4] FIG. 4 is a block diagram showing an example of the configuration of the shoulder speaker 1. As shown in FIG. [Figure 5] FIG. 5 is a flowchart showing an example of the process P of the processor 13. [Figure 6] FIG. 6 is a diagram showing discrete waveforms of a sound including a direct sound, an early reflection sound, and a reverberant sound. [Figure 7] FIG. 7 is a diagram showing an example of a GUI in an application program for performing a simulation. [Figure 8] FIG. 8 is a flowchart showing an example of processing by the processor 13 of the shoulder speaker 1a according to the first modification. [Figure 9] FIG. 9 is a diagram showing an example of connections between a shoulder speaker 1b according to the second modification, a near-end PC2, and a far-end PC3. [Figure 10] FIG. 10 is a block diagram showing an example of the configuration of a shoulder speaker 1c according to the third modification. [Figure 11] FIG. 11 is a block diagram showing an example of the configuration of a shoulder speaker 1d according to the fourth modification. [Figure 12] FIG. 12 is a top view of a shoulder speaker 1e according to the fifth modification. [Figure 13] FIG. 13 is a top view of a shoulder speaker 1f according to the sixth modification. [Figure 14] FIG. 14 is a flowchart showing an example of processing by the processor 13 of the shoulder speaker 1g according to the seventh modification. [Figure 15] FIG. 15 is a block diagram showing an example of the configuration of a shoulder speaker 1h according to the eighth modification. [Figure 16] FIG. 16 is a diagram showing a headphone 1i according to the ninth modification. DETAILED DESCRIPTION OF THE INVENTION
[0010] [First embodiment] The shoulder speaker 1 according to the first embodiment will be described below with reference to the drawings. Fig. 1 is a top view of the shoulder speaker 1. Fig. 2 is a diagram showing a user wearing the shoulder speaker 1. Fig. 3 is a diagram showing a user wearing the shoulder speaker 1 and playing a musical instrument. Unless otherwise specified, all signals will be described below as digital signals.
[0011] The shoulder speaker 1 is an example of a speaker device in the present application. As shown in Fig. 1, the shoulder speaker 1 has a symmetrical U-shape when viewed from above. As shown in Figs. 2 and 3, the shoulder speaker 1 is worn on the shoulder of a user. The shoulder speaker 1 does not cover the user's ear, and outputs sound from outside the user's ear canal.
[0012] Fig. 4 is a block diagram showing an example of the configuration of shoulder speaker 1. As shown in Fig. 4, shoulder speaker 1 includes a Bluetooth (registered trademark) interface 10, a user interface 11, a microphone 12, a processor 13, a flash memory 14, a RAM (Random Access Memory) 15, a left speaker 16L, and a right speaker 16R. Processor 13 is, for example, a CPU (Central Processing Unit). Note that shoulder speaker 1 includes two speakers (left speaker 16L and right speaker 16R), but may include three or more speakers.
[0013] The Bluetooth (registered trademark) interface 10 performs short-range communication with an information processing device such as a smartphone or a PC. For example, the Bluetooth (registered trademark) interface 10 receives an audio signal related to the sound of content from the information processing device such as a smartphone or a PC.
[0014] It should be noted that the shoulder speaker 1 and the information processing device do not necessarily have to be connected by the Bluetooth (registered trademark) interface 10. The shoulder speaker 1 and the information processing device may be connected wirelessly, for example, by Wi-Fi or the like, or may be connected by an analog audio cable.
[0015] The user interface 11 includes a volume button for adjusting the volume, a switch button for switching sound reproduction on / off, and the like.
[0016] The microphone 12 is an example of an acquisition unit in the present application. The microphone 12 is, for example, built into the housing of the shoulder speaker 1. Of course, the microphone 12 does not necessarily have to be built into the housing of the shoulder speaker 1, and may be arranged outside the housing. When the microphone 12 is arranged outside the housing of the shoulder speaker 1, it may be connected to the shoulder speaker 1 wirelessly or by wire.
[0017] The microphone 12 picks up sounds around the shoulder speaker 1. The microphone 12 acquires sound signals related to, for example, the sounds of a musical instrument being played or the sounds of a singer being sung. The sounds of a musical instrument being played or the sounds of a singer are examples of sounds based on the user's movements in this application. The microphone 12 outputs the acquired sound signals to the processor 13.
[0018] The flash memory 14 stores various programs. The various programs are, for example, programs that operate the shoulder speaker 1. It should be noted that the flash memory 14 does not necessarily have to store the various programs. The various programs may be stored in a separate device, such as a server. In this case, the shoulder speaker 1 receives the various programs from the separate device, such as the server.
[0019] The processor 13 executes various operations by reading out programs stored in the flash memory 14 into the RAM 15. The processor 13 is an example of a processing unit in the present application. The processor 13 performs signal processing on the sound signal input from the microphone 12. The processor 13 outputs the processed sound signal to the left speaker 16L and the right speaker 16R.
[0020] As shown in FIG. 2, the left speaker 16L is disposed outside the housing of the shoulder speaker 1. The left speaker 16L is disposed on the left part of the housing. The left part of the housing and the left speaker 16L are an example of the left output unit in this application. The left speaker 16L is disposed near the left ear of the user when the user wears the shoulder speaker 1. The left speaker 16L emits sound based on a sound signal input from the processor 13.
[0021] The right speaker 16R is disposed outside the housing of the shoulder speaker 1. The right speaker 16R is disposed on the right part of the housing. The right part of the housing and the right speaker 16R are an example of the right output unit in this application. When the user wears the shoulder speaker 1, the right speaker 16R is disposed near the user's right ear. The right speaker 16R emits sound based on a sound signal input from the processor 13.
[0022] The processor 13 according to this embodiment executes a process (hereinafter referred to as process P) for generating a reverberation signal with a three-dimensional reverberation effect based on a sound signal acquired by the microphone 12 and a head-related transfer function (HRTF). The head-related transfer function is a function that indicates the transfer characteristics of sound from the position of a sound source to the user's ear canal. The processor 13 convolves an impulse response corresponding to the head-related transfer function with the sound signal to generate a reverberation signal with a three-dimensional reverberation effect.
[0023] The process P in the processor 13 will be described in detail below with reference to the drawings. Fig. 5 is a flowchart showing an example of the process P in the processor 13.
[0024] The processor 13 starts the process P when, for example, the shoulder speaker 1 is powered on (FIG. 5: START).
[0025] After starting, the processor 13 acquires an impulse response IR corresponding to the reverberation of a predetermined space such as a concert hall (FIG. 5: step S11). The impulse response IR corresponding to the reverberation of the predetermined space is acquired, for example, using a measurement microphone in a concert hall or the like. The impulse response IR is measured, for example, by placing a dummy head in a predetermined space such as a concert hall and using multiple microphones placed at the left and right ear positions of the dummy head. The impulse response IR includes an impulse response for the left ear (L-side impulse response) acquired by a microphone placed at the left ear of the dummy head and an impulse response for the right ear (R-side impulse response) acquired by a microphone placed at the right ear of the dummy head. The flash memory 14 pre-stores the impulse response IR as data. The processor 13 reads the impulse response IR from the flash memory 14.
[0026] The impulse response IR may be obtained by a method other than the method using a dummy head. For example, the impulse response IR may be obtained and calculated by the close-quarters four-point method using four microphones that are not on the same plane, or may be obtained by simulation using the ray tracing method described below.
[0027] After step S11, the processor 13 removes components corresponding to direct sound and early reflection sound from the impulse response IR (L-side impulse response and R-side impulse response) (FIG. 5: step S12). FIG. 6 is a diagram showing discrete waveforms of sound including direct sound, early reflection sound, and reverberation sound. The processor 13 extracts only the impulse responses corresponding to reverberation sound (L-side impulse response and R-side impulse response), for example, by performing a process of extracting components from a predetermined point in time onward from the impulse response IR (components corresponding only to reverberation sound in FIG. 6).
[0028] The timing at which the reverberation starts (the predetermined time) may be determined based on the impulse response of the listening environment by outputting a test sound from the left speaker 16L or the right speaker 16R, acquiring an impulse response from the test sound using the microphone 12. The impulse response includes the processing time of the shoulder speaker 1, the transmission time, and the delay time of the entire listening environment. This allows the shoulder speaker 1 to accurately determine the extraction time of the reverberation sound from the impulse response IR (the predetermined time).
[0029] The predetermined time point may be adjusted by the user via the user interface 11.
[0030] The processor 13 may receive the L-side impulse response and the R-side impulse response from another information processing device such as a PC.
[0031] After step S12, the processor 13 generates a reverberation signal for the left speaker and a reverberation signal for the right speaker by convolving the L-side impulse response and the R-side impulse response with the sound signal acquired by the microphone 12 (FIG. 5: step S13).
[0032] For example, if the sound signal acquired by the microphone 12 relates to a performance sound, the reverberation signal generated by the processor 13 reproduces the reverberation sound of the performance sound in a concert hall.
[0033] When a user plays an instrument, the user hears the direct sound from the instrument, as well as the early reflections and reverberation of the room in the performance environment. The shoulder speaker 1 in this embodiment reproduces only the reverberation of a concert hall or other environment that differs from the room in the performance environment. This makes it less likely that the user will feel uncomfortable, as if the direct sound and early reflections are excessive or double-heard due to delays.
[0034] After step S13, the processor 13 outputs the reverberation signal for the left speaker to the left speaker 16L and outputs the reverberation signal for the right speaker to the right speaker 16R (FIG. 5: step S14). The processor 13 ends the process P when, for example, the power of the shoulder speaker 1 is turned off (FIG. 5: END).
[0035] The right speaker 16R and the left speaker 16L output binaural sounds (sounds that have been subjected to localization processing based on head-related transfer functions) based on the reverberation signals for the left and right speakers input from the processor 13. The binaural sounds are sounds that allow the user to perceive reverberation sounds in a predetermined space such as a concert hall. Therefore, even if the user is not in the predetermined space such as a concert hall, the user can sense the spatial spread of sounds as if they were playing a musical instrument in the predetermined space.
[0036] With audio devices such as in-ear headphones, sound is output from within the ear canal. Therefore, when a user listens to binaural sound with audio devices such as in-ear headphones, the user feels as if the reverberant sound is localized near the user's head. Therefore, users using audio devices such as in-ear headphones cannot perceive the spatial spread of the sound.
[0037] On the other hand, in the shoulder speaker 1, binaural sound is output from outside the user's ear canal by the right speaker 16R and the left speaker 16L. In other words, binaural sound is output near the user's pinna, away from the user's ear canal. This allows the user to hear sound influenced by the pinna. Therefore, users of the shoulder speaker 1 can enjoy the customer experience of feeling the spatial spread of sound, which is not possible with audio equipment such as in-ear headphones.
[0038] The shoulder speaker 1 may receive a reverberation signal generated by another information processing device such as a PC, and output binaural sounds based on the received reverberation signal.
[0039] Note that the output unit in this application does not necessarily have to be in contact with the user's body, but may be disposed close to the user's body. In other words, the speaker in this application does not necessarily have to be a shoulder speaker. For example, a speaker (headrest speaker) disposed in a headrest (a portion close to the user's head) in a car or theater seat also corresponds to the output unit in this application. In this case, the headrest speaker corresponds to the speaker device in this application.
[0040] The shoulder speaker 1 may perform sound quality adjustment (such as adding bass components) on the generated reverberation signal. Alternatively, the shoulder speaker 1 may perform correction based on the acoustic characteristics, such as the frequency characteristics, of the right speaker 16R and the left speaker 16L, to flatten the frequency characteristics, for example.
[0041] The right speaker 16R and the left speaker 16L do not necessarily have to be disposed on the housing of the shoulder speaker 1. For example, the right speaker 16R and the left speaker 16L may be disposed on the shoulder strap of a saxophone. In this case, the right speaker 16R and the left speaker 16L may be fixed directly to the shoulder strap of the saxophone, or may be fixed to the shoulder strap via a fixing device such as a clip. Furthermore, for example, the right speaker 16R or the left speaker 16L may be disposed on a shoulder pad used in playing the violin.
[0042] The shoulder speaker 1 may prevent the impulse response IR from being copied, for example, by converting the impulse response IR into a hash value.
[0043] Note that multiple users may each use the shoulder speaker 1 to perform a concert or the like. In this case, the impulse response IR used in the shoulder speaker 1 of each of the multiple users may be the same or different. For example, a server (not shown) may store information indicating an impulse response IR common to all users or an individual impulse response IR that differs for each user. The shoulder speaker 1 of each of the multiple users receives information indicating the common impulse response IR or the individual impulse response IR from the server.
[0044] In addition, when multiple users each use shoulder speakers 1 to perform a concert or the like, the performance sounds of each user may be acquired by the microphone 12 on the shoulder speakers 1 of each of the multiple users, or the performance sounds of all users may be acquired by one microphone 12.
[0045] The shoulder speaker 1 may have a structure specifically designed for playing the violin (a structure that prevents the shoulder speaker 1 from getting in the way when playing the violin). For example, the left speaker 16L may be detachable. A violinist may remove the left speaker 16L and wear the left speaker 16L on a shoulder pad when playing the violin. Alternatively, the left speaker 16L may be attached to the shoulder pad in advance. The left part of the housing may be shorter than the right part of the housing.
[0046] The shoulder speaker 1 may have a structure that allows the position of the microphone 12 to be adjusted. For example, the shoulder speaker 1 may be provided with a movable rod, and the microphone 12 may be located at the tip of the movable rod. This makes it possible to adjust the distance between the left speaker 16L or the right speaker 16R and the microphone 12, thereby suppressing the occurrence of howling. It is also possible to position the microphone 12 near the user's mouth.
[0047] Furthermore, for example, the position from which the sound of a trumpet or clarinet is produced is far from the performer's body, but the performer can adjust the position of the microphone 12 to be closer to the position from which the sound is produced by using a structure that allows the position of the microphone 12 to be adjusted.
[0048] [Method of obtaining impulse response using simulation] A method for obtaining an impulse response using a simulation will be described below with reference to the drawings. Figure 7 shows an example of a GUI in an application program for performing a simulation.
[0049] Possible methods for obtaining impulse responses through simulation include, for example, methods based on the virtual image method or the ray method. The ray method is a technique that tracks the trajectories (sound rays) of multiple sound particles radiated from a sound source in space and calculates the energy of the sound rays passing through a sound receiving area. A simulation using the ray method calculates the direction, arrival time, and arrival level from the listening point for each virtual sound source, assuming that each sound ray is a virtual sound image of reverberation sound based on the energy of the sound ray in the sound receiving area. This allows the user to freely set the shape of the space, the position of the sound source, the listening point, etc., and set the corresponding impulse response in the simulation.
[0050] For example, a user inputs information such as the shape of a space, the position of a sound source, and a listening point using an information processing device such as a PC installed with an application program for simulation using the ray tracing method. The information processing device uses the application program to obtain an impulse response IR corresponding to the input shape of the space, the position of the sound source, and the listening point.
[0051] As shown in FIG. 7, the application program accepts settings such as the shape of the virtual space VS (virtual closed space), the position of the sound source SS, and the position of the listening point PP.
[0052] Next, the application program calculates the energy of multiple sound rays that reflect from the sound source SS through the virtual space VS and reach the listening point PP based on the sound ray method, and determines the direction, arrival time, and arrival level from the listening point PP for each virtual sound source. Although two representative virtual sound sources FS1 and FS2 are shown in Figure 7, many virtual sound sources are actually calculated.
[0053] Next, the application program generates impulse responses for each of the virtual sound sources FS1 and FS2 based on head-related transfer functions that represent the direction, arrival time, and arrival level determined by the ray tracing method, and determines the impulse response IR by adding together the impulse responses for FS1 and FS2. The information processing device on which the application program is installed transmits the determined impulse response IR to the shoulder speaker 1. The processing performed after the shoulder speaker 1 receives the impulse response IR is the same as the processing performed by the shoulder speaker 1 in the first embodiment, and therefore will not be described here.
[0054] The processor 13 may acquire a pre-determined impulse response, or may acquire the impulse response IR by performing a simulation in real time using a ray acoustic method or the like.
[0055] The processor 13 may generate a reverberation signal by a level delay filter having a delay amount and an attenuation amount corresponding to each of the virtual sound sources FS1 and FS2 obtained by the ray tracing method, instead of performing a process of convolving the impulse response.
[0056] [Variation 1] The shoulder speaker 1a according to Modification 1 will be described below with reference to the drawings. Fig. 8 is a flowchart showing an example of processing by the processor 13 of the shoulder speaker 1a according to Modification 1.
[0057] The shoulder speaker 1a differs from the shoulder speaker 1 in that it also has a crosstalk cancellation function. Crosstalk is the component of the sound output from the left speaker 16L that reaches the right ear, and the component of the sound output from the right speaker 16R that reaches the left ear. Binaural sound produced by a head-related transfer function is the component of the sound output from the left speaker 16L that reaches the left ear, and the component of the sound output from the right speaker 16R that reaches the right ear. Crosstalk interferes with the binaural sound produced by the head-related transfer function.
[0058] Therefore, after the process of step S13, the processor 13 of the shoulder speaker 1a generates an L-side cancellation signal having a component (for example, an opposite-phase component) that cancels the sound output from the left speaker 16L, and an R-side cancellation signal having a component (for example, an opposite-phase component) that cancels the sound output from the right speaker 16R (FIG. 8: step S20). After step S20, the processor 13 performs the process of step S14.
[0059] After step S14, the processor 13 outputs the L-side cancellation signal to the right speaker 16R and outputs the R-side cancellation signal to the left speaker 16L (FIG. 8: step S21). The right speaker 16R outputs a component that cancels the sound output from the left speaker 16L. The left speaker 16L outputs a component that cancels the sound output from the right speaker 16R.
[0060] Therefore, the sound reaching the user's right ear from the left speaker 16L is canceled by the cancellation component output from the right speaker 16R. Similarly, the sound reaching the user's left ear from the right speaker 16R is canceled by the cancellation component output from the left speaker 16L. This allows the shoulder speaker 1a to properly present binaural sound produced by the head-related transfer function to the user. In other words, the shoulder speaker 1a can prevent the user from having difficulty perceiving the spatial spread of sound.
[0061] The rest of the configuration of the shoulder speaker 1a is the same as that of the shoulder speaker 1, so a description thereof will be omitted.
[0062] Note that the binaural sound based on the head-related transfer function may be generated by another device such as a PC, and the crosstalk cancellation component may be generated by the shoulder speaker 1a.
[0063] [Variation 2] The following describes the shoulder speaker 1b according to Modification 2 with reference to the drawings. Fig. 9 is a diagram showing an example of the connection between the shoulder speaker 1b according to Modification 2, the near-end PC2, and the far-end PC3.
[0064] The shoulder speaker 1b is used, for example, for remote conversation. As shown in FIG. 9 , the shoulder speaker 1b is communicably connected to a PC 2 on the near-end side. The PC 2 on the near-end side communicates with a PC 3 on the far-end side via a communication line 4 such as a LAN (Local Area Network) or the Internet. The PC 3 on the far-end side acquires an audio signal related to the voice of a user on the far-end side (hereinafter referred to as a far-end audio signal) and transmits the far-end audio signal to the PC 2. The voice of the user on the far-end side is an example of a sound based on a user's movement in this application. The PC 2 receives the far-end audio signal from the PC 3. The PC 2 transmits the received far-end audio signal to the shoulder speaker 1b. The shoulder speaker 1b receives the far-end audio signal from the PC 2 on the near-end side. In the second modification, the Bluetooth (registered trademark) interface 10 is an example of a communication interface that receives an audio signal from the far-end side in this application.
[0065] The shoulder speaker 1b performs localization processing on the far-end sound signal based on a head-related transfer function. In particular, the shoulder speaker 1b of the second modification performs processing to localize not only reverberant sound but also the direct sound of the sound source at a predetermined position. The shoulder speaker 1b outputs sound based on the far-end sound signal after localization processing.
[0066] In the second modification, the processor 13 of the shoulder speaker 1b may acquire the impulse response IR by actual measurement (measurement using a dummy head as described in the first embodiment) or by simulation as described in [Method for acquiring an impulse response using a simulation]. When the impulse response IR is acquired by actual measurement, the processor 13 does not remove the component corresponding to the direct sound from the impulse response IR. When the impulse response IR is acquired by simulation, an impulse response (impulse response corresponding to the direct sound) expressing the direction, arrival time, and arrival level of the direct sound determined by the ray tracing method for the sound source SS shown in FIG. 7 is added.
[0067] The processor 13 generates a sound signal for the left speaker and a sound signal for the right speaker by performing a process of convolving the impulse response IR (L-side impulse response and R-side impulse response) with the far-end sound signal. The processor 13 outputs the sound signal for the left speaker and the sound signal for the right speaker to the right speaker 16R and the left speaker 16L. In this case, the direct sound and the reverberant sound output from the right speaker 16R and the left speaker 16L are localized at predetermined positions, respectively.
[0068] The shoulder speaker 1b may perform processing to localize direct sound according to the position of the far-end user. For example, the far-end PC3 acquires an image by photographing the far-end user with a camera (not shown). PC3 transmits the image to PC2 via communication line 4. PC2 receives the image from PC3. PC2 performs image processing (face recognition) on the image to determine the position of the face in the image. Based on the position of the face in the image, PC2 estimates the distance between the far-end user and the camera and the direction of the user. PC2 acquires an impulse response IR that indicates that the direct sound is arriving from the estimated distance and direction, and transmits it to the shoulder speaker 1b. For example, PC2 may acquire (read) an impulse response IR corresponding to the estimated distance and direction from among multiple impulse responses stored in a flash memory or the like, or may calculate and obtain the impulse response IR each time through simulation.
[0069] Note that shoulder speaker 1b may change the localization position of the direct sound when the position of the far-end user changes during the process of localizing the direct sound. PC3 acquires images by repeatedly photographing the far-end user with a camera (not shown) and transmits the images to PC2 each time it acquires them. PC2 estimates the distance between the far-end user and the camera and the user's direction by performing image processing (face recognition) on the images acquired from PC3 each time. PC2 transmits an impulse response IR corresponding to the changed distance and direction to shoulder speaker 1b. For example, PC2 may acquire (read) the impulse response IR corresponding to the changed distance and direction from a flash memory or the like, or may calculate and obtain the impulse response IR corresponding to the changed distance and direction each time through simulation.
[0070] It should be noted that shoulder speaker 1b may perform direct sound localization processing on the far-end sound signal, and at the same time, convolve an impulse response (an impulse response from which components corresponding to direct sound and early reflected sound have been removed) corresponding to the reverberation of the space in which the far-end user is located with the sound signal acquired by microphone 12 (a sound signal related to the voice of the user of shoulder speaker 1b), and output the sound related to the convolved sound signal from right speaker 16R and left speaker 16L. This allows the user of shoulder speaker 1b to feel as if they are in the same space as the far-end user.
[0071] It should be noted that shoulder speaker 1b does not necessarily have to be used for remote conversations, and may be used, for example, to play content such as music or games, play the direct sound of an electronic musical instrument, or play sound distributed from a concert hall or the like. For example, PC3 on the far-end side may be placed in the concert hall where distribution is performed. PC3 acquires a sound signal related to the singer's singing sound and distributes the sound signal to multiple devices such as PC2. Shoulder speaker 1b performs processing to convolve the sound signal with an impulse response corresponding to the reverberation of the space (concert hall, etc.) on the far-end side.
[0072] Note that various processes such as calculation of impulse responses may be performed by the far-end PC3, rather than by the PC2 and the shoulder speaker 1b. The PC3 may transmit impulse response data calculated based on, for example, the ray tracing method. The PC2 may receive impulse response data from the PC3 and perform processing P based on the received impulse response data. The PC3 may also transmit data related to position information of a virtual sound source based on, for example, the ray tracing method. The PC2 may receive data related to position information of a virtual sound source from the PC3 and calculate a corresponding impulse response.
[0073] The shoulder speaker 1b may have an AUX terminal and may be connected to a smartphone or the like via the AUX terminal. In this case, the shoulder speaker 1b may input a sound signal related to content from the smartphone via the AUX terminal, localize a sound image of the direct sound of the sound signal, and generate a reverberation signal by convolving an impulse response IR with the sound signal.
[0074] The shoulder speaker 1b may have a function for switching on / off the sound image localization process of direct sound. When the sound image localization process of direct sound is switched from off to on, the shoulder speaker 1b may execute a process to gradually change the localization position of the sound image of the direct sound from inside the head to outside the head, or may execute a process to gradually change the localization position of the sound image of the direct sound from both ears of the user to in front of the user.
[0075] The shoulder speaker 1b may perform localization processing for sounds played on an electronic musical instrument such as an electronic piano or electronic drums.
[0076] Note that the hardware configuration related to signal processing in shoulder speaker 1b may be built into each unit, such as the bass drum or snare drum. In this case, shoulder speaker 1b receives sound signals processed by the hardware configuration related to signal processing built into each unit, such as the bass drum or snare drum, and mixes the processed sound signals with the sound signals to be output to left speaker 16L and right speaker 16R. Alternatively, the hardware configuration related to signal processing built into each unit, such as the bass drum or snare drum, may output the sound signals to be output to left speaker 16L and right speaker 16R.
[0077] When playing drums, the user may adjust the position (height) of the snare drum or toms, or adjust the number of toms to one or two. For this reason, the shoulder speaker 1b may store a head-related transfer function corresponding to the number and positions of the snare drum and toms for each user as a preset. The shoulder speaker 1b may read out the preset for each user and perform localization processing based on the preset.
[0078] The shoulder speaker 1b may store multiple presets corresponding to one user, so that even if the number and positions of the snare drum and tom-toms are changed depending on the genre of music played by the same user, the shoulder speaker 1b can perform appropriate localization processing according to the changed number and positions of the snare drum and tom-toms.
[0079] The shoulder speaker 1b can also be used simultaneously with VR (Virtual Reality) goggles, and can perform localization processing or reverberation signal generation processing corresponding to VR content. For example, the shoulder speaker 1b may perform processing to generate a reverberation signal corresponding to a virtual space displayed on the VR goggles. The shoulder speaker 1b may also perform processing to localize a sound image at the position of a person or an object emitting sound, such as a musical instrument, in the virtual space.
[0080] [Variation 3] Hereinafter, a shoulder speaker 1c according to Modification 3 will be described with reference to the drawings. Fig. 10 is a block diagram showing an example of the configuration of the shoulder speaker 1c according to Modification 3.
[0081] As shown in FIG. 10 , shoulder speaker 1c according to Modification 3 differs from shoulder speaker 1b in that it further includes a switching button 17. Switching button 17 accepts an operation from the user to switch on / off direct sound localization processing. When "direct sound localization processing: off," processor 13 performs only the process of generating a reverberation signal. On the other hand, when "direct sound localization processing: on," processor 13 performs both the process of generating a reverberation signal and the process of localizing the direct sound at a predetermined position. The process of generating a reverberation signal in Modification 3 is the same as the process of generating a reverberation signal in the first embodiment, and therefore a description thereof will be omitted. The process of localizing a direct sound at a predetermined position in Modification 3 is the same as the process of localizing a direct sound at a predetermined position in Modification 2, and therefore a description thereof will be omitted.
[0082] Depending on the usage situation of the shoulder speaker 1c, the user can decide whether or not to localize the direct sound using the switch button 17. For example, the user can turn off the direct sound localization processing when playing an instrument or singing, and turn on the direct sound localization processing when having a remote conversation.
[0083] In addition, the shoulder speaker 1c may have a function of not performing both the process of generating a reverberation signal and the process of localizing the direct sound at a predetermined position (outputting the direct sound as is), and a function of not outputting any sound.
[0084] The shoulder speaker 1d may be connected to a PC via a USB connection or the like and determine that a remote conference is taking place. The shoulder speaker 1d may automatically turn on direct sound localization processing when it determines that a remote conference is taking place, and may turn off direct sound localization processing when it determines that the remote conference has ended. Alternatively, the user may control the on / off of direct sound localization processing via an application program on a PC connected to the shoulder speaker 1d.
[0085] The rest of the configuration of the shoulder speaker 1c is the same as that of the shoulder speaker 1b, so a description thereof will be omitted.
[0086] [Variation 4] The following describes a shoulder speaker 1d according to Modification 4 with reference to the drawings. Fig. 11 is a block diagram showing an example of the configuration of the shoulder speaker 1d according to Modification 4.
[0087] As shown in FIG. 11 , shoulder speaker 1d differs from shoulder speaker 1 in that it further includes a gyro sensor 18. Gyro sensor 18 is an example of a detection unit that detects the orientation of the user in this application. Gyro sensor 18 is disposed on the housing of shoulder speaker 1d and detects changes in the orientation of the user as angular velocity. Processor 13 of shoulder speaker 1d corrects impulse responses IR (L-side impulse response and R-side impulse response) based on the detection result (user orientation) of gyro sensor 18.
[0088] For example, the processor 13 performs processing to localize a direct sound in front of (in front of) the user. In this state, the user faces diagonally forward to the right. In this case, the processor 13 estimates that "the user is facing diagonally forward to the right" based on the detection result of the gyro sensor 18. At this time, the processor 13 corrects the impulse response IR so that the direct sound is localized to the left front of the user (the position where the direct sound was originally localized). For example, the flash memory 14 stores multiple impulse responses. The processor 13 reads out the impulse response IR corresponding to the estimated result from the multiple impulse responses stored in the flash memory 14. The processor 13 generates a sound signal for the right speaker and a sound signal for the left speaker based on the read impulse response IR. Note that the processor 13 may calculate the impulse response IR corresponding to the estimated result each time by simulation.
[0089] By correcting the impulse response IR, the direct sound is always localized in the same position as seen by the user, even if the user's orientation changes. In other words, the user can perceive the direct sound source as always being located in the same position. As a result, the user no longer feels uncomfortable as if the position of the sound source is changing.
[0090] It should be noted that the shoulder speaker 1d does not necessarily have to detect the orientation of the user using the gyro sensor 18. The shoulder speaker 1d may detect the orientation of the user's head using a sensor other than the gyro sensor 18. For example, a head tracker that detects the orientation of the user's head may be attached to the user's head. The shoulder speaker 1d may receive data related to the orientation of the user's head from the head tracker and correct the impulse response IR based on the data.
[0091] When the shoulder speaker 1d has multiple speakers on each side of the shoulder speaker 1d, the shoulder speaker 1d may output sound from the speaker closest to the user's ear based on the detected orientation of the user's head. Alternatively, the shoulder speaker 1d may output sound from the speaker closest to the user's ear. Alternatively, the shoulder speaker 1d may output the loudest sound from the speaker closest to the user's ear.
[0092] The shoulder speaker 1d may detect the direction or angle of the user's head using, for example, an ultrasonic sensor instead of the gyro sensor 18.
[0093] The shoulder speaker 1d may be connected to the near-end PC2 shown in FIG. 9. The PC2 acquires an image by capturing an image of the near-end user with a camera (not shown). The PC2 estimates the facial orientation of the near-end user by performing image processing on the image. The PC2 transmits the estimation result as data to the shoulder speaker 1d. In this case, the shoulder speaker 1d may correct the impulse response IR based on the data indicating the facial orientation received from the PC2. For example, the shoulder speaker 1d corrects the impulse response IR so that the sound is always localized directly at the position of the PC2 (the device displaying the image of the far-end user in a remote conversation) as seen by the user, even if the orientation of the head of the near-end user changes.
[0094] When the user tilts their head to the right or left, the positional relationship between the user's ears and the right speaker 16R and the left speaker 16L changes. This causes a difference between the arrival time and volume of sound from the right speaker 16R to the right ear and the arrival time and volume of sound from the left speaker 16L to the left ear. For example, when the user tilts their head to the right, the distance between the user's right ear and the right speaker 16R becomes closer, while the distance between the user's left ear and the left speaker 16L becomes farther. This causes the volume of sound from the left speaker 16L to the left ear to decrease and the arrival time to increase. On the other hand, the volume of sound from the right speaker 16R to the right ear to increase and the arrival time to decrease. This causes a difference in the effect of the head-related transfer function between the left and right ears. Therefore, the shoulder speaker 1d may correct the delay amount (delay amount) and volume (level) of the impulse response IR according to the tilt of the user's head to prevent a difference in the effect of the head-related transfer function between the left and right ears.
[0095] For example, shoulder speaker 1d is connected to a PC. The PC acquires an image by photographing the user's head with a camera (not shown), and estimates the position of the user's ears relative to shoulder speaker 1d by analyzing the image. The PC transmits the estimation result as data to shoulder speaker 1d. Processor 13 of shoulder speaker 1d corrects the delay amount and volume of the impulse response IR based on the data indicating the ear position received from the PC. For example, PC2 estimates that "the user is tilting their head to the right." In this case, processor 13 reduces the delay amount of the L-side impulse response and increases the volume. Processor 13 also increases the delay amount of the R-side impulse response and decreases the volume. This reduces the likelihood of a difference in the effect of the head-related transfer function between the left and right.
[0096] The rest of the configuration of the shoulder speaker 1d is the same as that of the shoulder speaker 1, so a description thereof will be omitted.
[0097] [Variation 5] The shoulder speaker 1e according to Modification 5 will be described below with reference to the drawings. Fig. 12 is a top view of the shoulder speaker 1e according to Modification 5.
[0098] The left speaker 16L and the right speaker 16R of the shoulder speaker 1e are disposed at positions asymmetrical with respect to the center of the left and right sides of the housing of the shoulder speaker 1e. For example, as shown in Fig. 12, the right speaker 16R is disposed at the right front part of the housing, and the left speaker 16L is disposed at the left rear part of the housing.
[0099] When a user plays an instrument such as a violin or a flute, the user's head faces diagonally forward and to the left, as shown in FIG. 12. Therefore, the user's left ear is positioned backward and the right ear is positioned forward. At this time, the user's left ear is positioned near the left speaker 16L of the shoulder speaker 1d, and the user's right ear is positioned near the right speaker 16R of the shoulder speaker 1d (see FIG. 12). Therefore, when a user plays an instrument such as a violin or a flute, there is little difference between the arrival time and volume of sound from the right speaker 16R to the right ear and the arrival time and volume of sound from the left speaker 16L to the left ear. Therefore, there is little difference between the left and right head-related transfer functions.
[0100] Furthermore, as long as the right speaker 16R and the left speaker 16L are positioned asymmetrically with respect to the center of the left and right sides of the housing of the shoulder speaker 1e, it is not necessary for the right speaker 16R to be positioned in the right front part of the housing of the shoulder speaker 1e and the left speaker 16L to be positioned in the left rear part of the housing.
[0101] The height of the right front portion and the height of the left front portion of the shoulder speaker 1e may be different.
[0102] The shoulder speaker 1e may also have a structure that allows the height of the right front portion and the height of the left front portion to be changed. The shoulder speaker 1e may automatically determine whether a violin is being played. When the shoulder speaker 1e determines that a violin is being played, the shoulder speaker 1e may increase the height of the right front portion and decrease the height of the left front portion. The shoulder speaker 1e may also have a sensor that measures the height of the user's ears. The shoulder speaker 1e may determine whether a violin is being played based on the measurement results of the sensor.
[0103] The rest of the configuration of the shoulder speaker 1e is the same as that of the shoulder speaker 1, so a description thereof will be omitted.
[0104] [Variation 6] The following describes the shoulder speaker 1f according to Modification 6 with reference to the drawings. Fig. 13 is a top view of the shoulder speaker 1f according to Modification 6.
[0105] As shown in Fig. 13, shoulder speaker 1f according to Modification 6 differs from shoulder speaker 1e in that it further includes a left speaker 16L2 and a right speaker 16R2. Right speaker 16R2 is located in the right rear portion of the housing of shoulder speaker 1f. Left speaker 16L2 is located in the left front portion of the housing.
[0106] The processor 13 of the shoulder speaker 1f changes the output destination of the sound signal for the left speaker and the output destination of the sound signal for the right speaker depending on the orientation of the user's head. For example, a head tracker that detects the orientation of the user's head is attached to the user's head. The shoulder speaker 1f receives data related to the orientation of the user's head from the head tracker and determines the output destination of the sound signal for the left speaker and the sound signal for the right speaker based on the data.
[0107] For example, as shown in Fig. 13, when the user faces diagonally forward to the right, the processor 13 outputs a sound signal for the left speaker to the left speaker 16L2 and a sound signal for the right speaker to the right speaker 16R2. On the other hand, when the user faces diagonally forward to the left, the processor 13 outputs a sound signal for the left speaker to the left speaker 16L and a sound signal for the right speaker to the right speaker 16R. This makes it less likely that a difference will occur between the arrival time and volume of sound from the left speaker to the left ear and the arrival time and volume of sound from the right speaker to the right ear, even if the orientation of the user's head changes. Therefore, it is possible to prevent a difference between the left and right in the effect of the head-related transfer function.
[0108] The processor 13 may correct the localization position by adjusting the output balance between the right speaker 16R and the right speaker 16R2 so that the sound image is localized midway between the right speaker 16R and the right speaker 16R2.
[0109] The rest of the configuration of the shoulder speaker 1f is the same as that of the shoulder speaker 1e, so a description thereof will be omitted.
[0110] [Variation 7] The following describes the shoulder speaker 1g according to the seventh modification with reference to the drawings. Fig. 14 is a flowchart showing an example of processing by the processor 13 of the shoulder speaker 1g according to the seventh modification.
[0111] As shown in FIG. 14, after the process of step S12, the processor 13 of the shoulder speaker 1g performs a process (feedback cancellation process) on the sound signal acquired by the microphone 12 to cancel feedback components from the right speaker 16R and the left speaker 16L to the microphone 12 (FIG. 14: step S30). The feedback cancellation process can be, for example, an echo cancellation method that generates a simulated signal of the feedback component and adds it in antiphase, or an echo suppressor method that reduces gain when the level of the sound signal acquired by the microphone 12 exceeds a threshold. This suppresses feedback that occurs in the shoulder speaker 1g (feedback that occurs when the microphone 12 picks up sounds output from the right speaker 16R and the left speaker 16L). As a result, the user can enjoy a customer experience in which they do not experience discomfort due to feedback while using the shoulder speaker 1g.
[0112] The rest of the configuration of the shoulder speaker 1g is the same as that of the shoulder speaker 1, so a description thereof will be omitted.
[0113] [Variation 8] The following describes the shoulder speaker 1h according to Modification 8 with reference to the drawings. Fig. 15 is a block diagram showing an example of the configuration of the shoulder speaker 1h according to Modification 8.
[0114] As shown in Fig. 15, the shoulder speaker 1h differs from the shoulder speaker 1 in that it further includes a microphone 20. The microphone 20 acquires sound signals related to sounds around the shoulder speaker 1h.
[0115] Feedback occurs when the loop gain of the transmission system in which sound output from the output section (right speaker 16R and left speaker 16L) passes through an acoustic space, reaches a microphone, and then reaches the output section again after signal processing exceeds 1.
[0116] Therefore, the processor 13 of the shoulder speaker 1h switches between the sound signals acquired by the microphone 12 and the microphone 20 at predetermined time intervals (for example, every 5 seconds) to generate a reverberation signal for the left speaker and a reverberation signal for the right speaker. By switching the microphone used to generate the reverberation signal at predetermined time intervals, the transfer system of the acoustic space changes at predetermined time intervals, making it difficult for the loop gain to exceed 1. As a result, the occurrence of howling is suppressed.
[0117] The shoulder speaker 1h may include three or more microphones. In this case, the processor 13 of the shoulder speaker 1h switches between the sound signals acquired by the three or more microphones at predetermined intervals to generate a reverberation signal for the left speaker and a reverberation signal for the right speaker.
[0118] The other configurations of the shoulder speaker 1h are the same as those of the shoulder speaker 1, so a description thereof will be omitted.
[0119] [Variation 9] Hereinafter, headphones 1i according to Modification 9 will be described with reference to the drawings. Fig. 16 is a diagram showing headphones 1i according to Modification 9. The configuration according to the block diagram of headphones 1i is the same as the configuration of shoulder speaker 1, so the block diagram of Fig. 4 will be applied mutatis mutandis in the description.
[0120] The headphone 1i has a reversible housing. As shown in Fig. 9, the headphone 1i can be reversible so that the output section faces upward by changing the housing direction. By hanging the headphone 1i with the output section facing upward on the shoulder, the user can use it in the same way as the above-described shoulder speakers 1, 1a to 1h.
[0121] The headphones 1i may have a function for determining the orientation of the housings. The headphones 1i may turn on / off process P depending on the orientation of the housings. When the headphones 1i determine that the housings are oriented in different directions as shown in FIG. 16, the headphones 1i may turn on process P shown in the first embodiment, and when the headphones 1i determine that the housings are facing each other, the headphones 1i may turn off process P.
[0122] The headphones 1i may have a mechanism for manually or automatically changing the distance between the two housings. In general headphones, the biasing force of the headband shortens the distance between the two housings to bring them into close contact with the ears. Therefore, for example, the user may insert a rod or the like between the two housings of the headphones 1i to widen the headband, thereby increasing the distance between the two housings. Alternatively, the headphones 1i may widen the headband using a mechanical structure (e.g., a motor), thereby increasing the distance between the two housings. This allows the two housings to be positioned below the left and right ears, respectively. It also prevents the headphones 1i from tightening the neck of the user.
[0123] The headphones 1i may have a mechanism that can manually or automatically change the thickness of the top part of the headband. For example, the headphones 1i can position the housing of the headphones 1i near the user's ears by expanding the top part of the headband to increase its thickness.
[0124] The above-described embodiments and modifications should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments or modifications, but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to and within the scope of the claims.
[0125] For example, the first to eighth modifications may be combined as appropriate.
[0126] It should be noted that the shoulder speaker in this application does not necessarily have to have a crosstalk cancellation function. [Explanation of symbols]
[0127] 1, 1a to 1h: shoulder speaker, 1i: headphones, 2, 3: PC, 4: communication line, 10: Bluetooth (registered trademark) interface, 11: user interface, 12, 20: microphone, 13: processor, 14: flash memory, 15: RAM, 16L, 16L2: left speaker, 16R, 16R2: right speaker, 17: switching button, 18: gyro sensor
Claims
1. an acquisition unit that receives a sound based on a user's action as a sound signal; a processing unit that generates a reverberation signal having a three-dimensional reverberation effect based on the sound signal received by the acquisition unit and a head-related transfer function; an output unit that is disposed in contact with or close to the user's body and outputs a sound based on the reverberation signal having the three-dimensional reverberation effect from outside the user's ear canal, Speaker device.
2. the acquisition unit includes a communication interface that receives the sound signal from a far-end side; The processing unit further performs processing to localize a direct sound from a sound source at a predetermined position based on the sound signal and the head-related transfer function.
2. The speaker device according to claim 1.
3. the processing unit receives an operation from a user to switch on / off a process of localizing the direct sound at a predetermined position; 3. The speaker device according to claim 2.
4. The processing unit Obtain an impulse response that corresponds to the reverberation of a given space, generating the reverberation signal by convolving an impulse response, excluding a component corresponding to a direct sound, from the impulse response as the head-related transfer function with the sound signal of the sound based on the user's movement; 4. The speaker device according to claim 1.
5. The head-related transfer function is adapted to ear characteristics or personal characteristics of each user.
4. The speaker device according to claim 1.
6. the acquiring unit acquires the impulse response from an external device communicably connected to the speaker device; 5. The speaker device according to claim 4.
7. a detection unit that detects the position and orientation of the user; the processing unit performs a process of correcting the head-related transfer function based on the position and orientation of the user detected by the detection unit.
4. The speaker device according to claim 1.
8. a detection unit for detecting a tilt of the user's neck, correcting the delay amount and volume of the impulse response in accordance with the tilt of the user's head; 5. The speaker device according to claim 4.
9. The speaker device is a shoulder speaker.
4. The speaker device according to claim 1.
10. the output unit has a right output unit arranged near the right ear of the user and a left output unit arranged near the left ear of the user, the right output unit and the left output unit are disposed at positions asymmetrical with respect to the center of the left and right sides of the housing of the shoulder speaker, 10. The speaker device according to claim 9.
11. the acquisition unit includes a microphone that acquires the sound signal, the processing unit performs a howling cancellation process on the sound signal acquired by the microphone to cancel a feedback component that reaches the microphone from the output unit.
4. The speaker device according to claim 1.
12. the acquisition unit includes a plurality of microphones that respectively acquire the sound signals; the processing unit generates the reverberation signal by switching the sound signals acquired by the plurality of microphones at predetermined time intervals.
4. The speaker device according to claim 1.
13. the output unit includes a right output unit disposed near the right ear of the user and a left output unit disposed near the left ear of the user, the right output unit outputs a component of the sound that is inverse phase to the sound output from the left output unit, The left output unit outputs a component of the sound output from the right output unit that is in the opposite phase to the sound output from the right output unit.
4. The speaker device according to claim 1.
14. A sound based on the user's movement is received as a sound signal. generating a reverberation signal having a three-dimensional reverberation effect based on the received sound signal and a head-related transfer function; a speaker that is in contact with or placed close to the user's body outputs a sound based on the reverberation signal with the three-dimensional reverberation effect from outside the ear canal of the user, or the speaker outputs a sound based on the reverberation signal with the three-dimensional reverberation effect from outside the ear canal of the user; Sound signal processing method.
15. receiving the sound signal from the far-end side; further performing a process of localizing a direct sound from a sound source at a predetermined position based on the sound signal and the head-related transfer function; The sound signal processing method according to claim 14.
16. receiving an operation from a user to switch on / off a process for localizing the direct sound at a predetermined position; 16. The sound signal processing method according to claim 15.
17. Obtain an impulse response that corresponds to the reverberation of a given space, generating the reverberation signal by convolving an impulse response, excluding a component corresponding to a direct sound, from the impulse response as the head-related transfer function with the sound signal of the sound based on the user's movement; 17. A sound signal processing method according to any one of claims 14 to 16.
18. The head-related transfer function is adapted to ear characteristics or personal characteristics of each user.
17. A sound signal processing method according to any one of claims 14 to 16.
19. The impulse response is acquired from an external device communicatively connected to the device equipped with the speaker.
18. A sound signal processing method according to claim 17.
20. Detecting the user's position and the user's orientation; performing a process of correcting the head-related transfer function based on the detected position and orientation of the user; 17. A sound signal processing method according to any one of claims 14 to 16.
21. further detecting a tilt of the user's head; correcting the delay amount and volume of the impulse response in accordance with the tilt of the user's head; 18. A sound signal processing method according to claim 17.
22. The speaker is a shoulder speaker.
17. A sound signal processing method according to any one of claims 14 to 16.
23. the speaker has a right output unit disposed near the right ear of the user and a left output unit disposed near the left ear of the user, the right output unit and the left output unit are disposed at positions asymmetrical with respect to the center of the left and right sides of the housing of the shoulder speaker, 23. A method for processing a sound signal according to claim 22.
24. the speaker includes an output unit that outputs a sound based on the reverberation signal; Acquiring a sound signal based on the user's movement with a microphone; performing a howling cancellation process on the sound signal related to the sound acquired by the microphone to cancel a feedback component that reaches the microphone from the output unit; 17. A sound signal processing method according to any one of claims 14 to 16.
25. Acquiring a sound signal based on the user's movement with each of a plurality of microphones; generating the reverberation signal by switching the sound signals acquired by each of the plurality of microphones at predetermined time intervals; 17. A sound signal processing method according to any one of claims 14 to 16.
26. A process of generating a reverberation signal having a three-dimensional reverberation effect based on a sound signal related to a user's movement and a head-related transfer function; a process of causing a speaker that is in contact with or placed close to the user's body to output a sound based on the reverberation signal with the three-dimensional reverberation effect from outside the user's ear canal, or causing the speaker to output a sound based on the reverberation signal with the three-dimensional reverberation effect from outside the user's ear canal; A program that causes a processor to perform the following.
Citation Information
Patent Citations
Sound reproduction controller
JP2009004937A
Acoustic means built-in wearing object, acoustic processing system, acoustic processing device, acoustic processing method, and program
JP2018074499A