Sound processing system and sound processing method
The sound processing system addresses user intent by switching sound image localization based on head movement, offering high-quality sound when stationary and immersive sound when moving, enhancing user experience.
Patent Information
- Application Number
- JP2024045913
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-10-03
AI Technical Summary
Existing sound processing devices fail to consider user intentions, such as concentrating on content sound or enjoying a realistic sound field, by only reproducing acoustic characteristics without adjusting sound image localization based on user movement.
A sound processing system that includes a sensor to detect user head movement, determining whether the user is stationary or not, and switches between enabling and disabling sound image localization processing accordingly, outputting high-quality sound when stationary and localized sound when moving.
Enables appropriate output of high-quality and realistic sound signals based on user wishes, providing a new customer experience by allowing concentration on content sound or immersive venue atmosphere.
Smart Images

Figure 2025145627000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound processing system and a sound processing method that processes and outputs a sound signal based on information acquired by a sensor. [Background technology]
[0002] Conventionally, there are sound processing devices that enable users to enjoy recorded sounds such as music recorded on a recording medium with a sufficient sense of realism (for example, Patent Document 1). The sound processing device generates sound effects by convolving impulse response data corresponding to the reverberation of an actual acoustic space with the original sound. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-330477 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the sound processing device disclosed in Patent Document 1 is only concerned with reproducing the acoustic characteristics of the acoustic space, and does not take into consideration the user's intentions. For example, there are times when a user wants to concentrate on the sound of the content, and there are times when a user wants to enjoy the realistic sound field of the venue rather than the sound of the content itself.
[0005] An object of the present invention is to appropriately output high-quality sound signals and realistic sound signals in accordance with the user's wishes. [Means for solving the problem]
[0006] A sound processing system according to one embodiment of the present invention is a system including a sound processing device including an output bus, and a sensor connected to the sound processing device and detecting the state of a user of the sound processing device. By executing a program, the system receives a first sound signal, determines whether the user's head is in a stationary state based on information detected by the sensor, and outputs the first sound signal to the output bus if it is determined that the user's head is in a stationary state, and outputs a second sound signal to the output bus, in which a sound image of the first sound signal is localized at a predetermined position based on a head-related transfer function, if it is determined that the user's head is not in a stationary state. [Effects of the Invention]
[0007] According to the sound processing system of the present invention, by switching between enabling and disabling sound image localization processing in accordance with the movement of the user's head, high-quality sound and realistic sound can be appropriately output in accordance with the user's wishes. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing the basic configuration of a sound processing system 100 according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the basic configuration of the sound processing device 1 according to the first embodiment. [Figure 3] FIG. 3 is a block diagram showing the basic configuration of the headphones 3 according to the first embodiment. [Figure 4] 4A and 4B are diagrams showing the relationship between the user's head and the sound image localization position of a sound signal according to the first embodiment. [Figure 5] FIG. 5 is a flowchart showing the operation of the sound processing device 1 according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing the relationship between the user's head and the sound image localization position of a sound signal according to the first modification. [Figure 7] FIG. 7 is a graph showing the relationship between time and gain according to the first modification. [Figure 8]FIG. 8 is a flowchart showing the operation of the sound processing device 1 according to the first modification. [Figure 9] FIG. 9 is a flowchart showing the operation of the sound processing device 1 according to the second modification. [Figure 10] FIG. 10 is a diagram showing the relationship between the user's head and the sound image localization position of a sound signal according to the third modification. [Figure 11] FIG. 11 is a flowchart showing the operation of the sound processing device 1 according to the third modification. [Figure 12] 12(A) and 12(B) are diagrams showing the relationship between the user's head and the sound image localization position of a sound signal according to the fourth modification. [Figure 13] FIG. 13 is a block diagram showing the basic configuration of a sound processing system 100A according to the seventh modification. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, a signal processing device according to an embodiment of the present invention will be described with reference to the drawings. In each drawing, the same parts are assigned the same reference numerals. From Modification 1 onwards, a description of matters common to the first embodiment will be omitted, and only the differences will be described. In particular, similar actions and effects resulting from similar configurations will not be mentioned in each embodiment.
[0010] First Embodiment 1 is a block diagram showing the basic configuration of a sound processing system 100 according to the first embodiment. The sound processing system 100 includes a sound processing device 1, a head tracking sensor 2, and headphones 3.
[0011] The sound processing device 1 is an example of a user terminal in the present invention. The sound processing device 1 plays audio data selected by a user. The sound processing device 1 acquires the audio data, for example, from an external playback device, a network, or a storage medium. The sound processing device 1 may also acquire the audio data from a microphone (not shown). The sound processing device 1 is, for example, a personal computer, a smartphone, or a tablet computer. In the present invention, audio data means a digital signal unless otherwise specified.
[0012] The head tracking sensor 2 is an example of a sensor in the present invention. The head tracking sensor 2 detects the state of the user. Specifically, the head tracking sensor 2 detects information related to the movement of the user's head (hereinafter referred to as movement information). The head tracking sensor 2 is connected to the sound processing device 1 and transmits the detected movement information to the sound processing device 1. The head tracking sensor 2 has a wireless communication function conforming to standards such as Wi-Fi (registered trademark) and Bluetooth (registered trademark). The head tracking sensor 2 also has a wired communication function conforming to standards such as USB. The head tracking sensor 2 is, for example, a three-axis acceleration sensor or a gyro sensor. The head tracking sensor 2 detects changes in the movement of the user's head as acceleration or angular velocity and acquires it as movement information. The head tracking sensor 2 may also be, for example, a camera. In this case, the head tracking sensor 2 acquires an image of the user's head as movement information.
[0013] The headphones 3 are connected to the sound processing device 1 by wire or wirelessly, and output the sound signal received from the sound processing device 1. In this embodiment, the headphones 3 are not limited to overhead headphones. The headphones 3 may be, for example, in-ear earphones or dynamic speakers placed at a distance from the ears.
[0014] 2 is a block diagram showing the basic configuration of a sound processing device 1 according to the first embodiment. The sound processing device 1 includes a communication unit 11, a memory 12, a RAM 13, a CPU 14, an audio interface (I / F) 15, and a user interface (I / F) 16.
[0015] The communication unit 11 receives motion information from the head tracking sensor 2. The communication unit 11 also communicates with other devices such as a server. The communication unit 11 has a wireless communication function conforming to standards such as Wi-Fi (registered trademark) and Bluetooth (registered trademark). The communication unit 11 also has a wired communication function conforming to standards such as USB.
[0016] The CPU 14 functions as a control unit that comprehensively controls the operation of the sound processing device 1 by reading out an operation program from the memory 12 to the RAM 13. The CPU 14 may download an operation program from a server or the like and read it into the RAM 13 each time.
[0017] The functions realized by the CPU 14 are, for example, a sound signal acquisition process, a sensor information acquisition process, a determination process, a localization process, and a sound signal control process. More specifically, the CPU 14 loads programs related to the sound signal acquisition process, the sensor information acquisition process, the determination process, the localization process, and the sound signal control process into the RAM 13. As a result, the CPU 14 configures a sound signal acquisition unit 141, a sensor information acquisition unit 142, a determination unit 143, a localization processing unit 144, and a sound signal control unit 145.
[0018] The sound signal acquisition unit 141 receives audio data. The audio data includes sound signals corresponding to the L channel and the R channel (hereinafter referred to as first sound signals).
[0019] The sensor information acquisition unit 142 acquires the motion information from the head tracking sensor 2 via the communication unit 11 .
[0020] The determination unit 143 determines whether the user's head is in a stationary state based on the motion information acquired by the sensor information acquisition unit 142.
[0021] The localization processing unit 144 performs sound image localization processing on the first sound signal included in the audio data acquired by the sound signal acquisition unit 141. The sound image localization processing is processing for localizing a sound image so that the sound output from the headphones 3 appears as if it were generated at a predetermined position. The localization processing unit 144 achieves the sound image localization processing by convolving a head-related transfer function with the first sound signal. The localization processing unit 144 acquires the head-related transfer function from, for example, the memory 12, a network, an external storage medium, or the like.
[0022] Here, the head-related transfer function will be described in more detail. The head-related transfer function is a function that expresses the transfer characteristics from the position of a sound source to the user's head (specifically, the user's left ear and right ear). There are two head-related transfer functions: one from the sound source to the right ear and one to the left ear. The localization processing unit 144 convolves the sound signals (first sound signals) corresponding to the L channel and R channel with the head-related transfer function to the left ear and the head-related transfer function to the right ear, respectively.
[0023] The sound signal control unit 145 outputs a sound signal (hereinafter referred to as a second sound signal) obtained by convolving the head-related transfer function to the right ear and the head-related transfer function to the left ear to the audio I / F 15, which is an output bus. The audio I / F 15 outputs sound signals corresponding to the R channel and the L channel, respectively, to the headphones 3. The audio I / F 15 is a communication interface such as an analog audio terminal, a digital audio terminal, or USB, or a wireless communication interface such as Bluetooth (registered trademark).
[0024] The user I / F 16 accepts operations from the user, such as operations to change audio data, adjust the volume of the headphones 3, and the like.
[0025] 3 is a block diagram showing the basic configuration of the headphones 3 according to the first embodiment. The headphones 3 include a CPU 31, a memory 32, a RAM 33, an audio I / F 34, an output unit 35, and speaker units 353L and 353R.
[0026] The CPU 31 functions as a control unit that comprehensively controls the operation of the headphones 3 by reading out an operation program from the memory 32 to the RAM 33. The CPU 31 may download the operation program from a server or the like and read it into the RAM 33 each time.
[0027] The audio I / F 34 receives an audio signal from the sound processing device 1 .
[0028] The output unit 35 is connected to the speaker unit 353L and the speaker unit 353R. The output unit 35 has a DA converter 351 (hereinafter referred to as DAC351) and an amplifier 352 (hereinafter referred to as AMP352). The DAC351 converts a digital signal (sound signal) acquired from the sound processing device 1 into an analog signal. The AMP352 amplifies the analog signal to drive the speaker unit 353L and the speaker unit 353R. The output unit 35 outputs the amplified analog signal (sound signal) to the speaker unit 353L and the speaker unit 353R.
[0029] Generally, there is a problem that a sound signal (second sound signal) convolved with a head-related transfer function has a lower sound quality than a sound signal (first sound signal) that is not convolved with a head-related transfer function. This is because the characteristics of the original sound signal (first sound signal) are changed when the head-related transfer function is convolved. In addition, there are cases where a user wants to concentrate on the sound of the content, and cases where a user wants to enjoy the realistic sound field of the venue rather than the sound of the content itself. In other words, there are cases where a user wants to listen to the sound of a high-quality sound source, and cases where a user wants to listen to realistic sound that has been subjected to sound image localization processing.
[0030] Therefore, the sound processing device 1 of this embodiment switches between enabling and disabling the sound image localization process based on the motion information acquired by the head tracking sensor 2. Specifically, if the sound processing device 1 determines that the user's head is stationary, it does not execute the sound image localization process. On the other hand, if the sound processing device 1 determines that the user's head is moving, that is, that the user's head is not stationary, it executes the sound image localization process.
[0031] 4A and 4B, how the sound image localization position of a sound signal changes when the sound processing device 1 switches between enabling and disabling sound image localization processing will be described. Figures 4A and 4B are diagrams showing the relationship between the user's head and the sound image localization position of a sound signal according to the first embodiment. However, in the present invention, the relationship between the user's head and the sound image localization position of a sound signal is not limited to this example.
[0032] The virtual sound source positions VS1-1 and VS1-2 shown in FIGS. 4(A) and 4(B) represent positions where the sound images of the sound signals are localized.
[0033] FIG. 4(A) shows a state in which the user's head is stationary. As shown in FIG. 4(A), when the user's head is stationary, the sound processing device 1 outputs a first sound signal. As described above, the first sound signal is a sound signal to which no head-related transfer function has been convoluted. Therefore, the user feels as if the sound image of the sound signal is localized at the virtual sound source position VS1-1, which is the position of the head. In other words, when the user's head is stationary, the sound image of the sound signal is localized inside the user's head.
[0034] 4(B) shows a state in which the user's head is rotating to the left. As shown in FIG. 4(B), when the user's head is not stationary, the sound processing device 1 generates and outputs a sound signal (second sound signal) by convolving the first sound signal with the first head-related transfer function.
[0035] The first head-related transfer function is a head-related transfer function that realizes the localization of a long-distance sound image. A sound signal (second sound signal) obtained by convolving the first head-related transfer function with the first sound signal is localized at a long distance and in front of the user.
[0036] Here, the "front direction" in the present invention refers to, for example, the direction the user was facing before the user's head started to move, that is, when the user was stationary. The "front direction" may also be the direction relative to a monitor (not shown) provided in the sound processing device 1. The "front direction" may also be arbitrarily set by the user via the user I / F 16.
[0037] The user feels as if the sound image of the sound signal is localized at the virtual sound source position VS1-2 in the front direction. In other words, when the user's head is not stationary, the sound image of the sound signal is localized at a long distance and in the front direction of the user.
[0038] FIG. 5 is a flowchart showing the operation of the sound processing device 1 according to the first embodiment.
[0039] First, the sound processing device 1 receives a first sound signal (S001).
[0040] Next, the sound processing device 1 determines whether the user's head is in a stationary state based on the movement information (S002).
[0041] It is rare for the user's head to be completely still. Therefore, the sound processing device 1 may set a threshold value for the range of movement of the user's head (acceleration, angular velocity), and determine that the user's head is in a stationary state when the user's head is moving within a range that does not exceed the threshold value. Furthermore, the sound processing device 1 may determine that the user's head is in a stationary state unless the threshold value is exceeded for a predetermined continuous period of time.
[0042] If it is determined that the user's head is stationary (S002: YES), the sound processing device 1 outputs the first sound signal (S003) and ends the processing (END). At this time, the user hears only the first sound signal, which is a high-quality sound signal that has not been subjected to acoustic processing.
[0043] If it is determined that the user's head is not stationary (S002: NO), the sound processing device 1 generates a second sound signal (S004).
[0044] Next, the sound processing device 1 outputs the second sound signal (S005) and ends the processing (END). At this time, the user hears only the second sound signal, which is a sound signal with a sense of localization.
[0045] In this way, while the user keeps his or her head still, the user can hear the high-quality first sound signal and concentrate on the sound of the content. Meanwhile, while the user moves his or her head, the user can hear the second sound signal that has been processed for sound image localization, and enjoy the sound of the content with a sense of realism as if the user were actually listening to it at the venue. This allows the user to have a new customer experience in which he or she can hear high-quality sound when he or she wants to concentrate on the sound of the content, and hear immersive sound when he or she wants to enjoy the atmosphere of the venue.
[0046] It should be noted that the determination process executed by the determination unit 143 does not necessarily have to be executed by the sound processing device 1. For example, the head tracking sensor 2 may itself execute the determination process based on the acquired motion information. In this case, the head tracking sensor 2 may transmit the determination result to the sound processing device 1, instead of the motion information.
[0047] The sound processing device 1 may also be equipped with a head tracking sensor 2. In this case, the sound processing device 1 may be equipped with an acceleration sensor, an angular velocity sensor, or a camera to acquire motion information.
[0048] Furthermore, the movement of the user's head is not limited to the horizontal direction (yaw direction). The movement of the user's head may include a rotational direction around the front-to-back axis (roll direction) and a rotational direction around the vertical axis (pitch direction).
[0049] Variation 1 In the first embodiment, the sound processing device 1 outputs only the first sound signal when the user's head is stationary, and outputs only the second sound signal when the user's head is not stationary. In other words, in the first embodiment, the sound image of the sound signal is instantly switched between a state in which it is localized inside the head and a state in which it is localized outside the head in front of the head, based on the motion information. However, if it is desired to smoothly move the sound image localization position of the sound signal, the following processing may be executed.
[0050] Fig. 6 is a diagram showing the relationship between the user's head and the sound image localization position of the sound signal according to Modification 1. Fig. 7 is a graph showing the relationship between time and gain according to Modification 1. However, in the present invention, the relationship between the user's head and the sound image localization position of the sound signal, and the relationship between time and gain are not limited to these examples.
[0051] The virtual sound source positions VS2-1, VS2-2, VS2-3, and VS2-4 shown in Fig. 6 represent the positions where the sound images of the sound signals are localized. Fig. 6 also shows the state in which the user's head, which was in a stationary state, is rotated 90° to the left and then returns to the stationary state.
[0052] In Modification 1, the sound processing device 1 adjusts the gains of the first sound signal and the second sound signal, respectively, and changes the output levels of each sound signal to smoothly move the sound image localization position of the sound signal. Specifically, the sound processing device 1 gradually changes the gains of the first sound signal and the second sound signal, respectively, as time passes after the user's head starts to move. For example, as shown in FIG. 6, the sound processing device 1 decreases the gain of the first sound signal and increases the gain of the second sound signal as the angle increases after the user's head starts to move. Alternatively, as shown in FIG. 7, the sound processing device 1 decreases the gain of the first sound signal linearly and increases the gain of the second sound signal linearly as time passes after the user's head starts to move.
[0053] How the sound image localization position of a sound signal changes will be described with reference to Fig. 6. Note that the front direction in this embodiment is the direction in which the user's head faces when the user's head is in a stationary state.
[0054] First, when the user's head is in a stationary state (angle θ=0°), the sound processing device 1 sets the gain of the first sound signal to 1 and the gain of the second sound signal to 0. At this time, the user hears only the first sound signal. This makes the user feel as if the sound image of the sound signal is localized at the virtual sound source position VS2-1, which is the position of the user's head.
[0055] After the user starts to move their head, the sound processing device 1 starts adjusting the gain of the first sound signal to decrease and the gain of the second sound signal to increase. For example, as shown in Fig. 6, when the user's head is rotated 30° to the left (angle θ = 30°), the sound processing device 1 adjusts the gain of the first sound signal to decrease and the gain of the second sound signal to increase. At this time, the gain of the first sound signal is greater than the gain of the second sound signal, so the user is strongly influenced by the first sound signal. As a result, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS2-2, which is close to their head.
[0056] When the user further moves his / her head (angle θ=60°), the sound processing device 1 further adjusts the gain of the first sound signal to be lower and the gain of the second sound signal to be higher. At this time, the gain of the first sound signal is smaller than the gain of the second sound signal, so the user is strongly influenced by the second sound signal. As a result, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS2-3, which is far from the head.
[0057] When the user moves his / her head further (angle θ=90°), the sound processing device 1 sets the gain of the first sound signal to 0 and the gain of the second sound signal to 1. At this time, the user hears only the second sound signal. This makes the user feel as if the sound image of the sound signal is localized at the farthest virtual sound source position VS2-4.
[0058] In this way, when the user keeps his / her head still, he / she feels that the sound image of the sound signal is localized inside his / her head, and when he / she starts to move his / her head, he / she feels that the sound image gradually moves away in front of him / her. In other words, the sound image that was localized inside his / her head gradually moves away in front of him / her, and finally feels that it is localized at the virtual sound source position VS2-4.
[0059] Alternatively, as shown in FIG. 7, the sound processing device 1 linearly decreases the gain of the first sound signal and linearly increases the gain of the second sound signal as time passes after the user's head starts to move.
[0060] Initially, when the user's head is stationary (angle θ=0°), the sound processing device 1 sets the gain of the first sound signal to 1 and the gain of the second sound signal to 0. After the user starts to move their head, the sound processing device 1 starts adjusting the gain of the first sound signal to lower the gain and the gain of the second sound signal to increase. Then, with the passage of time after the user's head starts to move, the sound processing device 1 linearly decreases the gain of the first sound signal and linearly increases the gain of the second sound signal, finally setting the gain of the first sound signal to 0 and the gain of the second sound signal to 1.
[0061] In this case too, when the user keeps his / her head still, he / she feels as if the sound image of the sound signal is localized inside his / her head, and when he / she starts to move his / her head, he / she feels as if the sound image is gradually moving away in front of him / her.
[0062] Fig. 8 is a flowchart showing the operation of the sound processing device 1 according to Modification 1. The processing of steps S001 to S005 in Fig. 8 is the same as the processing of steps S001 to S005 in Fig. 5, and therefore a description thereof will be omitted.
[0063] After outputting the first sound signal in step S003, the sound processing device 1 again determines whether the user's head is in a stationary state based on the motion information (S011).
[0064] If it is determined that the user's head is in a stationary state (S011: YES), the sound processing device 1 ends the process (END). In other words, the sound signal control unit 145 continues to output only the first sound signal.
[0065] If it is determined that the user's head is not stationary (S011: NO), the sound processing device 1 generates a second sound signal (S012).
[0066] Next, the sound processing device 1 outputs the first sound signal and the second sound signal (S013), and ends the processing (END).
[0067] In this way, when the user's head switches from a stationary state to a non-stationary state, the sound processing device 1 of Modification 1 adjusts and outputs the gains of the first sound signal and the second sound signal according to the change in angle from when the head switched to the non-stationary state or the passage of time. This allows the user to have a new customer experience in which they can properly hear the high-quality first sound signal and the realistic second sound signal without feeling uncomfortable due to the change in the sound image localization position of the sound signal.
[0068] Variation 2 The head tracking sensor 2 according to the second modification detects not only the motion information but also information (hereinafter referred to as direction information) relating to the direction (relative orientation) in which the user's head is facing relative to a certain reference direction (forward direction). Specifically, the head tracking sensor 2 calculates the direction in which the user's head is facing based on the detected acceleration or angular velocity, and acquires the calculation result as the direction information. The head tracking sensor 2 may also acquire, for example, an image of the user's head as the direction information. The direction information may also be obtained from the absolute orientation determined by a geomagnetic sensor.
[0069] Moreover, the sound processing device 1 according to the second modification further has a function of determining whether or not the user's head is facing forward based on the direction information.
[0070] Fig. 9 is a flowchart showing the operation of the sound processing device 1 according to Modification 2. The processes of steps S001, S002, S004, and S005 in Fig. 9 are the same as the processes of steps S001, S002, S004, and S005 in Fig. 5, and therefore description thereof will be omitted.
[0071] First, the sound processing device 1 determines the front direction (S021). The front direction is, for example, the direction the user was facing immediately before the user's head started to move, that is, when the user was stationary. For example, when the user's head has been stationary for a predetermined period of time or more, the sound processing device 1 determines the current head direction as the front direction. The front direction may also be the direction relative to a monitor (not shown) provided in the sound processing device 1. For example, if the head tracking sensor 2 is provided with a camera, the camera may detect the monitor and determine the front direction. The front direction may also be arbitrarily set by the user via the user I / F 16.
[0072] In the second modification, when it is determined that the user's head is stationary (S002: YES), the sound processing device 1 further determines whether the user's head is facing forward based on the direction information (S022).
[0073] It is rare that the direction in which the user's head is facing perfectly coincides with the forward direction, so the sound processing device 1 may set an arbitrary threshold value for the forward direction, and determine that the user's head is facing forward when the direction in which the user's head is facing is within a range that does not exceed the threshold value.
[0074] If it is determined that the user's head is facing forward (S022: YES), the sound processing device 1 outputs the first sound signal (S023). That is, the user feels as if the sound image of the sound signal is localized inside his or her head.
[0075] If it is determined that the user's head faces in a direction other than the front (non-front facing) (S022: NO), the sound processing device 1 generates a second sound signal (S004).
[0076] Next, the sound processing device 1 outputs the second sound signal (S005) and ends the processing (END).
[0077] It should be noted that the operation of the sound processing device 1 is not limited to the above. For example, in the operation of the sound processing device 1, step S001 and step S021 may be interchanged.
[0078] In this way, the sound processing device 1 of Modification 2 switches between enabling and disabling the sound image localization process based on the motion information and direction information, which allows the user to have a new customer experience of hearing the second sound signal with a sense of realism when moving their head or not facing forward.
[0079] Variation 3 10 is a diagram showing the relationship between the user's head and the sound image localization positions of sound signals according to Modification 3. However, in the present invention, the relationship between the user's head and the sound image localization positions of sound signals is not limited to this example.
[0080] Fig. 10 shows the state in which the user's head, which was in a stationary state, rotates 90° to the left and then returns to a stationary state. The virtual sound source positions VS3-1, VS3-2, VS3-3, VS3-4, and VS3-5 shown in Fig. 10 represent the positions where the sound images of the sound signals are localized.
[0081] In Modification 3, the direction in which the user's head faces when the user's head is initially stationary (angle θ=0°) is defined as the first front direction. Also, the direction in which the user's head faces when the user's head rotates 90° to the left and returns to a stationary state (angle θ=90°) is defined as the second front direction.
[0082] The sound processing device 1 according to the third modification changes the front direction from the first front direction to the second front direction when the user's head returns from a non-stationary state to a stationary state again.
[0083] How the sound image of the sound signal changes when the front direction of the sound processing device 1 is changed from the first front direction to the second front direction will be described with reference to FIG.
[0084] When the user's head is stationary (angle θ = 0°), the user feels as if the sound image of the sound signal is localized at virtual sound source position VS3-1, which is the position of the user's head. When the user's head becomes stationary and a predetermined time has passed, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS3-2. When the user's head becomes stationary and a further predetermined time has passed, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS3-3. When the user's head becomes stationary and a further predetermined time has passed, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS3-4. Thereafter, when the user's head returns to a stationary state, the direction in which the sound image of the sound signal is localized changes from the first front direction to the second front direction. As a result, the user feels as if the sound image of the sound signal is localized at virtual sound source position VS3-5, which is the second front direction.
[0085] The further the direction of the user's head deviates from the front, the more the sound image seems to deviate from the front. However, with the configuration of variant example 3, when the user brings their head back to a still position, the user feels as if the sound image of the sound signal is again positioned in the front.
[0086] Fig. 11 is a flowchart showing the operation of the sound processing device 1 according to Modification 3. The processing of steps S001 to S005 in Fig. 11 is the same as the processing of steps S001 to S005 in Fig. 5, and therefore a description thereof will be omitted.
[0087] After outputting the second sound signal in step S005, the sound processing device 1 again determines whether the user's head is in a stationary state based on the motion information (S031). Note that the sound processing device 1 determines that the user's head is in a stationary state when the user's head maintains the stationary state for a certain period of time.
[0088] If it is determined that the head is not stationary (S031: NO), the sound processing device 1 repeats the process from step S004.
[0089] When it is determined that the head is stationary (S031: YES), the sound processing device 1 changes the setting so that the direction in which the user's head is facing becomes the front direction (S032). That is, the sound processing device 1 changes the front direction from the first front direction to the second front direction.
[0090] Next, the sound processing device 1 generates a second sound signal so that the sound image of the sound signal is localized in the second front direction (S033).
[0091] Next, the sound processing device 1 outputs the second sound signal (S034) and ends the processing (END).
[0092] In this way, the sound processing device 1 according to Modification 3 switches the front direction in response to the movement of the user's head, allowing the user to freely change the orientation of their head and body without worrying about the direction.
[0093] Variation 4 The sound processing device 1 of Modification 4 will be described with reference to Fig. 12(A) and (B). Fig. 12(A) and (B) are diagrams showing the relationship between the user's head and the sound image localization positions of sound signals according to Modification 4. However, in the present invention, the relationship between the user's head and the sound image localization positions of sound signals is not limited to this example.
[0094] The virtual sound source positions VS4-1 and VS4-2 shown in FIGS. 12(A) and 12(B) represent the positions where the sound images of the sound signals are localized.
[0095] The sound processing device 1 according to the fourth modification differs from the sound processing device 1 according to the first embodiment in that a head-related transfer function (hereinafter referred to as a second head-related transfer function) that realizes the localization of multiple near-distance sound images is convolved with the first sound signal. The near distance is the position where both ears of the user are located. The head-related transfer function that realizes far-distance sound image localization is the first head-related transfer function, and the far distance is a position farther than both ears of the user.
[0096] 12(A) shows a state in which the user's head is stationary. As shown in FIG. 12(A), when the user's head is stationary, the sound processing device 1 outputs only the first sound signal. At this time, the user feels as if the sound image of the sound signal is localized at the virtual sound source position VS4-1, which is the position of the user's head.
[0097] FIG. 12(B) shows a state in which the user's head is rotating to the left. When the user's head is moving, the sound processing device 1 performs sound image localization processing on the first sound signal. Specifically, a sound signal (hereinafter referred to as the second sound signal) is generated by convolving the L channel of the first sound signal with a second head related transfer function that localizes the sound to a position in front of the user and the R channel with a second head related transfer function that localizes the sound to a position behind the user's head. The second head related transfer function is changed depending on the direction of rotation of the user. In other words, the second sound signal is always localized to the position where the user's ears are located (virtual sound source position VS4-2) when the user's head is facing forward. In other words, the user feels as if the sound image of the sound signal is localized to the virtual sound source position VS4-2, regardless of the direction the user faces.
[0098] In this way, the sound processing device 1 according to Modification 4 achieves close-range localization of a sound image at the same position no matter which direction the user faces. Therefore, the user can properly listen to the first sound signal with high sound quality and the second sound signal with a sense of left and right localization without feeling uncomfortable due to changes in the sound image localization position of the sound signal.
[0099] Variation 5 The head-related transfer function is affected by the shape of the head and pinna, and therefore varies greatly from person to person. Furthermore, when a head-related transfer function set based on the shape of a model head is convolved, some users may not be able to obtain a sufficient sense of positioning.
[0100] Therefore, the sound processing device 1 according to the fifth modification changes the head-related transfer function to be convoluted with the first sound signal in accordance with the characteristics of the user's ears.
[0101] The sound processing device 1 according to the fifth modification stores in advance in the memory 12 a plurality of pieces of ear shape data and corresponding head-related transfer functions. The sound processing device 1 also receives input of an image of the ear from the user via the user I / F 16. Based on the acquired image of the user's ear, the sound processing device 1 reads out the optimal head-related transfer function for the user from the memory 12 and convolves it into the first sound signal.
[0102] This allows the user to hear the second sound signal convolved with a head-related transfer function that matches the shape of the user's ear, and provides a sufficient sense of localization.
[0103] Variation 6 The sound processing device 1 according to the sixth modification further includes a camera (not shown) that captures an image of the room the user is in. The memory 12 of the sound processing device 1 stores reverb data corresponding to different room sizes and acoustic environments.
[0104] Reverb is the sound that is reflected from a sound source on walls, floors, ceilings, etc.
[0105] The sound processing device 1 according to the sixth modification detects the size of a room based on an image of the room captured by a camera. Next, the sound processing device 1 reads reverb data corresponding to the size of the room from the memory 12, and performs reverb processing on the first sound signal. Note that the reverb processing is a process that simulates the reverberation (reflected sound) of a certain room by adding a predetermined delay time to the sound signal and generating level-adjusted pseudo-reflected sound.
[0106] The method of acquiring the reverb data is not limited to the above. For example, the sound processing device 1 may receive a selection of desired reverb data from the user via the user I / F 16.
[0107] In this way, the sound processing device 1 according to the sixth modification performs reverb processing on the first sound signal according to the size of the user's room or the user's preference. This allows the user to hear a more natural sound compared to when listening to only the first sound signal. In other words, the user can hear a more realistic sound.
[0108] Variation 7 A sound processing system 100A according to Modification 7 will be described with reference to Fig. 13. Fig. 13 is a block diagram showing the basic configuration of the sound processing system 100A according to Modification 7. Note that the same reference numerals are used to designate components common to the sound processing system 100, and descriptions thereof will be omitted.
[0109] The sound processing system 100A differs from the sound processing system 100 in that it further includes a cloud device 4 connected to a network 5.
[0110] The cloud device 4 includes a memory, a RAM, and a CPU (not shown). The memory, the RAM, and the CPU included in the cloud device 4 have the same functions as the memory 12, the RAM 13, and the CPU 14 included in the sound processing device 1.
[0111] The cloud device 4 according to the seventh modification executes a program in place of the sound processing device 1. Specifically, the CPU included in the cloud device 4 has functions such as sound signal acquisition processing, sensor information acquisition processing, determination processing, localization processing, and sound signal control processing, and generates a second sound signal based on the motion information and direction information and outputs the second sound signal to the headphones 3.
[0112] Finally, the description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope of the claims.
[0113] For example, the first to seventh modifications may be combined as appropriate. [Explanation of symbols]
[0114] 1: Sound processing device, 2: Head tracking sensor, 3: Headphones, 4: Cloud device, 5: Network, 11: Communication unit, 12, 32: Memory, 13, 33: RAM, 14, 31: CPU, 15, 34: Audio interface, 16: User interface, 35: Output unit, 100, 100A: Sound processing system, 141: Sound signal acquisition unit, 142: Sensor information acquisition unit, 143: Determination unit, 144: Localization processing unit, 145: Sound signal control unit, 351: DA converter, 352: Amplifier, 353L, 353R: Speaker unit, VS1-1, VS1-2, VS2-1 to VS2-4, VS3-1 to VS3-5, VS4-1, VS4-2: Virtual sound source positions
Claims
1. a sound processing device including an output bus; a sensor connected to the sound processing device and detecting a state of a user of the sound processing device; A system including the following, which executes a program: receiving a first sound signal; determining whether the user's head is in a stationary state based on the information detected by the sensor; outputting the first sound signal to the output bus when it is determined that the user's head is in a stationary state; when it is determined that the user's head is not stationary, a second sound signal obtained by localizing a sound image of the first sound signal at a predetermined position based on a head-related transfer function is output to the output bus. Sound processing system.
2. When it is determined that the user's head is in a stationary state and then determined that the user's head is in a non-stationary state, both the first sound signal and the second sound signal are output, and output levels of the first sound signal and the second sound signal are adjusted based on information acquired by the sensor and output to the output bus. The sound processing system of claim 1 .
3. when it is determined that the user's head is in a non-stationary state after it has been determined that the user's head is in a stationary state, both the first sound signal and the second sound signal are output to the output bus, and the output level of the first sound signal is gradually weakened and the output level of the second sound signal is gradually strengthened as time passes after it has been determined that the user's head is in the non-stationary state. The sound processing system of claim 2 .
4. If it is determined that the head of the user is in a stationary state, it is further determined whether the head of the user is facing forward; If it is determined that the speaker is not facing forward, the second sound signal is output to the output bus. The sound processing system of claim 1 .
5. If the user's head returns to a stationary state and a certain period of time has passed, it is determined that the user is facing forward. The sound processing system according to claim 1 or 2.
6. the head-related transfer functions include a first head-related transfer function that realizes localization of a far-distance sound image and a second head-related transfer function that realizes localization of a near-distance sound image; generating the second sound signal to which the first head-related transfer function or the second head-related transfer function is assigned based on the information acquired by the sensor; The sound processing system according to claim 1 or 2.
7. generating the second sound signal on either a user terminal or a terminal on the cloud; transmitting the generated second sound signal to headphones; The sound processing system according to claim 1 or 2.
8. The head-related transfer function is changed according to the characteristics of the user. The sound processing system according to claim 1 or 2.
9. Get a picture of the room Detecting the size of the room based on the image; imparting reverb to the first sound signal in accordance with the detected size of the room; The sound processing system according to claim 1 or 2.
10. A sound processing device including an output bus receiving a first sound signal; the sound processing device determines whether the user's head is in a stationary state based on information acquired by a sensor connected to the sound processing device; outputting the first sound signal to the output bus when it is determined that the user's head is in a stationary state; if it is determined that the user's head is not stationary, a second sound signal is generated by localizing a sound image of the first sound signal at a position determined based on a head-related transfer function, and the second sound signal is output to the output bus. Sound processing methods.
11. When it is determined that the user's head is in a stationary state and then determined that the user's head is in a non-stationary state, both the first sound signal and the second sound signal are output, and output levels of the first sound signal and the second sound signal are adjusted based on information acquired by the sensor and output to the output bus. The sound processing method according to claim 10.
12. when it is determined that the user's head is in a non-stationary state after it has been determined that the user's head is in a stationary state, both the first sound signal and the second sound signal are output to the output bus, and the output level of the first sound signal is gradually weakened and the output level of the second sound signal is gradually strengthened as time passes after it has been determined that the user's head is in the non-stationary state. The sound processing method according to claim 11.
13. If it is determined that the user's head is in a stationary state, it is further determined whether the user's head is facing forward; If it is determined that the speaker is not facing forward, the second sound signal is output to the output bus. The sound processing method according to claim 10.
14. If the user's head returns to a stationary state and a certain period of time has passed, it is determined that the user is facing forward. The sound processing method according to claim 10 or 11.
15. the head-related transfer functions include a first head-related transfer function that realizes localization of a far-distance sound image and a second head-related transfer function that realizes localization of a near-distance sound image; generating the second sound signal to which the first head-related transfer function or the second head-related transfer function is assigned based on the information acquired by the sensor; The sound processing method according to claim 10 or 11.
16. generating the second sound signal on either a user terminal or a terminal on the cloud; transmitting the generated second sound signal to headphones; The sound processing method according to claim 10 or 11.
17. The head-related transfer function is changed according to the characteristics of the user. The sound processing method according to claim 10 or 11.
18. Get a picture of the room Detecting the size of the room based on the image; imparting reverb to the first sound signal in accordance with the detected size of the room; The sound processing method according to claim 10 or 11.
Citation Information
Patent Citations
Acoustic processing device and method for distributing data for acoustic processing
JP2003330477A