Sound reproduction method, computer program, and sound reproduction device

JP7901197B2Active Publication Date: 2026-08-05PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-02-20
Publication Date
2026-08-05

Smart Images

  • Figure 0007901197000001
    Figure 0007901197000001
  • Figure 0007901197000002
    Figure 0007901197000002
  • Figure 0007901197000003
    Figure 0007901197000003
Patent Text Reader

Abstract

To provide an acoustic reproduction method capable of improving a sensation level of a sound reached to a rear side of a listener.SOLUTION: An acoustic reproduction method contains: a signal acquisition step of acquiring a first audio signal that indicates a first sound as a sound reached to a listener from a first range as a range of a predetermined angle and a second audio signal that indicates a second sound as a sound reached to the listener from a predetermined direction; an information acquisition step of acquiring direction information as information of a direction where a head part of the listener is directed; and a correction processing step of executing a correction processing to at least one part of the first and second audio signals that are acquired in the case where it is determined that at least one part of the first range and the predetermined direction is contained in a second range defined on the basis of the direction information to be acquired.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an acoustic reproduction method and the like.

Background Art

[0002] In Patent Document 1, a technique related to a stereophonic reproduction system that realizes immersive sound by outputting sound from a plurality of speakers arranged around a listener has been proposed.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, for a human (here, a listener who listens to sound), among the sounds reaching the human from the surroundings, the perception level of the sound reaching from behind the human is lower than that of the sound reaching from in front of the human.

[0005] Therefore, an object of the present disclosure is to provide an acoustic reproduction method and the like that improve the perception level of the sound reaching from behind the listener.

Means for Solving the Problems

[0006] A sound reproduction method according to one aspect of the present disclosure includes a signal acquisition step of acquiring a first audio signal indicating a first sound which is a sound reaching a listener from a first range which is a range of predetermined angles, and a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction; an information acquisition step of acquiring direction information which is information about the direction in which the listener's head is facing; and a correction processing step of applying correction processing to at least one of the acquired first audio signal and the acquired second audio signal when it is determined that at least a part of the first range and the predetermined direction are included in a second range determined based on the direction information, based on the acquired direction information.

[0007] A program relating to one aspect of this disclosure causes a computer to execute the above-described sound reproduction method.

[0008] An audio reproduction device according to one aspect of the present disclosure includes: a signal acquisition unit that acquires a first audio signal indicating a first sound which is a sound reaching a listener from a first range which is a range of predetermined angles, and a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction; an information acquisition unit that acquires direction information which is information about the direction the listener's head is facing; and a correction processing unit that, based on the acquired direction information, determines that at least a part of the first range and the predetermined direction is included in a second range determined based on the direction information, and applies correction processing to at least one of the acquired first audio signal and the acquired second audio signal.

[0009] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]

[0010] One aspect of the present disclosure, such as a sound reproduction method, can improve the perceived level of sound arriving from behind the listener. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a block diagram showing the functional configuration of an audio playback device according to Embodiment 1. [Figure 2] Figure 2 is a schematic diagram showing an example of how sound output from multiple speakers according to Embodiment 1 can be used. [Figure 3] Figure 3 is a flowchart showing an example of the operation of the sound reproduction device according to Embodiment 1. [Figure 4] Figure 4 is a schematic diagram illustrating an example of a decision made by the correction processing unit according to Embodiment 1. [Figure 5] Figure 5 is a schematic diagram illustrating another example of the decision made by the correction processing unit according to Embodiment 1. [Figure 6] Figure 6 is a schematic diagram illustrating another example of the decision made by the correction processing unit according to Embodiment 1. [Figure 7] Figure 7 illustrates an example of the correction process performed by the correction processing unit according to Embodiment 1. [Figure 8] Figure 8 illustrates another example of the correction processing performed by the correction processing unit according to Embodiment 1. [Figure 9] Figure 9 illustrates another example of the correction processing performed by the correction processing unit according to Embodiment 1. [Figure 10] Figure 10 is a schematic diagram showing an example of a correction process applied to the first audio signal according to Embodiment 1. [Figure 11] Figure 11 is a schematic diagram showing another example of the correction process applied to the first audio signal according to Embodiment 1. [Figure 12] Figure 12 is a block diagram showing the functional configuration of the sound reproduction device and sound acquisition device according to Embodiment 2. [Figure 13] Figure 13 is a schematic diagram illustrating sound collection by the sound collection device according to Embodiment 2. [Figure 14] FIG. 14 is a schematic diagram showing an example of correction processing applied to a plurality of first audio signals according to Embodiment 2.

Embodiments of the Invention

[0012] (Knowledge on which the present disclosure is based) Conventionally, there has been known a technique related to acoustic reproduction that realizes an immersive sound by outputting sounds represented by a plurality of different audio signals from a plurality of speakers arranged around a listener.

[0013] For example, a stereophonic reproduction system disclosed in Patent Document 1 includes a main speaker, a surround speaker, and a stereophonic reproduction device.

[0014] The main speaker amplifies the sound represented by the main audio signal at a position where the listener is arranged within the directivity angle, the surround speaker amplifies the sound represented by the surround audio signal toward the wall surface of the sound field space, and the stereophonic reproduction device amplifies each speaker.

[0015] In addition, this stereophonic reproduction device has signal adjustment means, delay time addition means, and output means. The signal adjustment means adjusts the frequency characteristics of the surround audio signal based on the propagation environment during amplification. The delay time addition means adds a delay time corresponding to the surround signal to the main audio signal. The output means outputs the main audio signal with the added delay time to the main speaker and the adjusted surround audio signal to the surround speaker.

[0016] According to such a stereophonic reproduction system, it is possible to create a sound field space that provides a high sense of immersion.

[0017] Incidentally, humans (in this case, listeners receiving sound) perceive sounds arriving from behind them at a lower level than sounds arriving from in front of them. For example, humans have a perceptual characteristic (more specifically, an auditory characteristic) that makes it difficult for them to perceive the location or direction of sounds arriving from behind them. This perceptual characteristic is derived from the shape and discrimination limit of the human ear.

[0018] Furthermore, when two types of sound (for example, a target sound and ambient sound) arrive from behind the listener, one sound (for example, the target sound) may be drowned out by the other sound (for example, ambient sound). In this case, the listener will have difficulty hearing the target sound, making it difficult to perceive the location or direction of the target sound arriving from behind them.

[0019] For example, in the 3D audio reproduction system disclosed in Patent Document 1, when the sound indicated by the main audio signal and the sound indicated by the surround audio signal arrive from behind the listener, the listener has difficulty perceiving the sound indicated by the main audio signal. Therefore, there is a need for an audio reproduction method that improves the perceived level of sound arriving from behind the listener.

[0020] Therefore, an audio reproduction method according to one aspect of the present disclosure includes: a signal acquisition step of acquiring a first audio signal indicating a first sound which is a sound reaching the listener from a first range which is a range of a predetermined angle, and a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction; an information acquisition step of acquiring direction information which is information about the direction the listener's head is facing; a correction processing step of applying a correction process to at least one of the acquired first audio signal and the acquired second audio signal, which is a process that makes the intensity of the second audio signal stronger than the intensity of the first audio signal, when the range behind the direction the listener's head is facing is defined as the front and the second range is defined as the second range based on the acquired direction information; and a mixing processing step of mixing at least one of the first audio signal and the second audio signal that has undergone the correction process and outputting it to an output channel.

[0021] As a result, the intensity of the second audio signal indicating the second sound increases when the first range and a predetermined direction are included in the second range. Therefore, the listener can more easily hear the second sound arriving from behind (i.e., from behind the listener) when the direction the listener's head is facing is considered forward. In other words, an acoustic reproduction method is realized that can improve the perceived level of the second sound arriving from behind the listener.

[0022] For example, if the first sound is an ambient sound and the second sound is the target sound, this method can prevent the target sound from being drowned out by the ambient sound. In other words, it enables an acoustic reproduction method that can improve the perceived level of the target sound arriving from behind the listener.

[0023] For example, the first range is the range behind the reference bearing determined by the position of the output channel.

[0024] This makes it easier for the listener to hear the second tone, which is coming from behind them, even if the first tone is reaching the listener from a range behind the reference direction.

[0025] For example, the correction process is a process of correcting at least one of the gain of the acquired first audio signal and the gain of the acquired second audio signal.

[0026] This allows for gain correction of at least one of the first audio signal representing the first tone and the second audio signal representing the second tone, making it easier for the listener to hear the second tone, which is arriving from behind them.

[0027] For example, the correction process is at least one of the following: a process to reduce the gain of the acquired first audio signal, and a process to increase the gain of the acquired second audio signal.

[0028] As a result, at least one of the following processes is applied: a process that reduces the gain of the first audio signal representing the first tone, and a process that increases the gain of the second audio signal representing the second tone. This makes it easier for the listener to hear the second tone, which is arriving from behind them.

[0029] For example, the correction process is a process of correcting at least one of the frequency components based on the acquired first audio signal and the frequency components based on the acquired second audio signal.

[0030] This allows for correction of at least one of the frequency components based on the first audio signal representing the first tone and the frequency components based on the second audio signal representing the second tone, making it easier for the listener to hear the second tone arriving from behind them.

[0031] For example, the correction process is a process that reduces the spectrum of frequency components based on the acquired first audio signal to be smaller than the spectrum of frequency components based on the acquired second audio signal.

[0032] This reduces the intensity of the frequency components in the spectrum based on the first audio signal representing the first tone, making it easier for the listener to hear the second tone arriving from behind them.

[0033] For example, the correction processing step is performed based on the positional relationship between the second range and the predetermined orientation, and the correction processing is a process that corrects at least one of the gain of the acquired first audio signal and the gain of the acquired second audio signal, or a process that corrects at least one of the frequency characteristics based on the acquired first audio signal and the frequency characteristics based on the acquired second audio signal.

[0034] This allows for correction processing based on the positional relationship between the second range and a predetermined direction, making it easier for the listener to hear the second sound arriving from behind them.

[0035] For example, when the second range is divided into a right rear range, which is the range to the right rear of the listener, a left rear range, which is the range to the left rear, and a central rear range, which is the range between the right rear range and the left rear range, the correction processing step determines that the predetermined direction is included in the right rear range or the left rear range, and performs the correction processing which is a process to decrease the gain of the acquired first audio signal or a process to increase the gain of the acquired second audio signal. If the predetermined direction is included in the central rear range, the correction processing which is a process to decrease the gain of the acquired first audio signal and a process to increase the gain of the acquired second audio signal is performed.

[0036] As a result, when a predetermined direction is included in the central rear range, a correction process is applied so that the intensity of the second audio signal indicating the second tone is stronger than the intensity of the first audio signal indicating the first tone, compared to when the predetermined direction is included in the right rear range or left rear range. Therefore, the listener will be able to hear the second tone arriving from behind them more easily.

[0037] For example, the signal acquisition step acquires a plurality of first audio signals and second audio signals representing a plurality of first sounds, and classification information which is information classifying the plurality of first audio signals based on the frequency characteristics of each of the plurality of first audio signals. The correction processing step performs the correction processing based on the acquired orientation information and classification information, and each of the plurality of first sounds is a sound picked up from each of the plurality of first ranges.

[0038] This allows the correction processing step to apply correction processing to each group into which multiple first audio signals have been classified. Therefore, the processing load of the correction processing step can be reduced.

[0039] For example, an audio reproduction method according to one aspect of the present disclosure includes: a signal acquisition step of acquiring a plurality of first audio signals indicating a plurality of first sounds which are a plurality of sounds reaching a listener from a plurality of first ranges which are a plurality of predetermined angular ranges, and a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction; an information acquisition step of acquiring direction information which is information about the direction the listener's head is facing; a correction processing step of applying a correction process to at least one of the acquired plurality of first audio signals and the acquired second audio signals, which is a process that makes the intensity of the second audio signal stronger than the intensity of the plurality of first audio signals, when the range behind the direction the listener's head is facing is defined as the front and the second range is defined as the second range based on the acquired direction information; and a mixing processing step of mixing at least one of the plurality of first audio signals and the acquired second audio signals that has undergone the correction process and outputting it to an output channel, wherein each of the plurality of first sounds is a sound picked up from each of the plurality of first ranges.

[0040] As a result, the intensity of the second audio signal indicating the second sound increases when the first range and a predetermined direction are included in the second range. Therefore, the listener can more easily hear the second sound arriving from behind (i.e., from behind the listener) when the direction the listener's head is facing is considered forward. In other words, an acoustic reproduction method is realized that can improve the perceived level of the second sound arriving from behind the listener.

[0041] Furthermore, the correction processing step can be performed on each group into which the multiple first audio signals are classified. This reduces the processing load of the correction processing step.

[0042] For example, a program relating to one aspect of this disclosure may be a program that causes a computer to execute the above-described sound reproduction method.

[0043] This allows the computer to perform the above sound playback method according to the program.

[0044] For example, an audio playback device according to one aspect of the present disclosure includes: a signal acquisition unit that acquires a first audio signal indicating a first sound which is a sound reaching the listener from a first range which is a range of a predetermined angle, and a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction; an information acquisition unit that acquires direction information which is information about the direction the listener's head is facing; a correction processing unit that, when the direction the listener's head is facing is defined as the front and the rear range as the second range, determines based on the acquired direction information that the first range and the predetermined direction are included in the second range, and applies a correction process to at least one of the acquired first audio signal and the acquired second audio signal, which is a process that makes the intensity of the second audio signal stronger than the intensity of the first audio signal; and a mixing processing unit that mixes at least one of the first audio signal and the second audio signal that has undergone the correction process and outputs it to an output channel.

[0045] As a result, the intensity of the second audio signal indicating the second sound increases when the first range and a predetermined direction are included in the second range. Therefore, the listener can more easily hear the second sound arriving from behind (i.e., from behind the listener) when the direction the listener's head is facing is considered forward. In other words, an acoustic reproduction device is realized that can improve the perceived level of the second sound arriving from behind the listener.

[0046] For example, if the first sound is an ambient sound and the second sound is the target sound, it is possible to suppress the target sound from being drowned out by the ambient sound. In other words, an acoustic reproduction device can be realized that can improve the perceived level of the target sound arriving from behind the listener.

[0047] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0048] The embodiments will be described in detail below with reference to the drawings.

[0049] The embodiments described below are all general or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit the scope of the claims.

[0050] Furthermore, in the following explanation, elements may be assigned ordinal numbers such as 1st, 2nd, and 3rd. These ordinal numbers are assigned to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly assigned, or removed as appropriate.

[0051] Furthermore, each figure is a schematic diagram and not necessarily a strictly accurate representation. Therefore, the scale and other aspects may not necessarily be consistent across all figures. In each figure, substantially identical components are given the same reference numerals, and redundant explanations are omitted or simplified.

[0052] (Embodiment 1) [composition] First, the configuration of the sound reproduction device 100 according to Embodiment 1 will be described. Figure 1 is a block diagram showing the functional configuration of the sound reproduction device 100 according to this embodiment. Figure 2 is a schematic diagram showing an example of how sound output from multiple speakers 1, 2, 3, 4 and 5 according to this embodiment is used.

[0053] The sound reproduction device 100 according to this embodiment processes multiple acquired audio signals and outputs them to multiple speakers 1, 2, 3, 4, and 5 shown in Figures 1 and 2, thereby allowing a listener L to hear the sounds represented by the multiple audio signals. More specifically, the sound reproduction device 100 is a stereophonic sound reproduction device that allows a listener L to hear stereophonic sound.

[0054] Furthermore, the sound reproduction device 100 processes multiple audio signals acquired based on the directional information output by the head sensor 300. The directional information is the direction in which the listener L's head is facing. The direction in which the listener L's head is facing is also the direction in which the listener L's face is facing. Note that "directional" here refers to, for example, a direction.

[0055] The head sensor 300 is a device that senses the direction the listener L's head is facing. The head sensor 300 is preferably a device that senses 6DOF (Degrees of Freedom) information of the listener L's head. For example, the head sensor 300 is a device that is attached to the listener L's head and may be an inertial measurement unit (IMU), accelerometer, gyroscope, magnetic sensor, or a combination thereof.

[0056] As shown in Figure 2, in this embodiment, multiple speakers (five in this case) 1, 2, 3, 4, and 5 are arranged to surround the listener L. In Figure 2, to explain direction, 0 o'clock, 3 o'clock, 6 o'clock, and 9 o'clock are shown, corresponding to the time indicated on the clock face. The white arrows indicate the direction towards which the listener L's head is pointing, and the direction towards which the listener L's head, located at the center (also called the origin) of the clock face, is pointing is the direction of 0 o'clock. Hereafter, the direction connecting the listener L and 0 o'clock may be referred to as the "direction of 0 o'clock," and the same applies to the other times indicated on the clock face.

[0057] In this embodiment, the five speakers 1, 2, 3, 4, and 5 consist of a center speaker, a front right speaker, a rear right speaker, a rear left speaker, and a front left speaker. Speaker 1, which is the center speaker, is positioned at the 12 o'clock position.

[0058] Each of the five speakers 1, 2, 3, 4, and 5 is a sound amplification device that outputs the sound indicated by multiple audio signals output from the sound reproduction device 100.

[0059] As shown in Figure 1, the sound reproduction device 100 includes a first signal processing unit 110, a first decoding unit 121, a second decoding unit 122, a first correction processing unit 131, a second correction processing unit 132, an information acquisition unit 140, and a mixing processing unit 150.

[0060] The first signal processing unit 110 is a processing unit that acquires multiple audio signals. The first signal processing unit 110 may acquire multiple audio signals by receiving multiple audio signals transmitted by other components not shown in Figure 2, or it may acquire multiple audio signals stored in a storage device not shown in Figure 2. The multiple audio signals acquired by the first signal processing unit 110 are signals that include a first audio signal and a second audio signal.

[0061] Here, we will explain the first audio signal and the second audio signal.

[0062] The first audio signal is a signal indicating the first sound, which is the sound reaching the listener L from a predetermined angular range, the first range D1. For example, the first range D1 is the range behind the reference direction determined by the positions of the five output channels, speakers 1, 2, 3, 4, and 5. In this embodiment, the reference direction is the direction from the listener L toward speaker 1, which is the center speaker, and is not limited to, for example, the 12 o'clock direction. Behind the 12 o'clock direction, which is the reference direction, is the 6 o'clock direction, and the first range D1 should include the 6 o'clock direction, which is behind the reference direction. Also, the first range D1 is not limited to the range from the 3 o'clock direction to the 9 o'clock direction (i.e., an angular range of 180°). Since the reference direction is constant regardless of the direction the listener L's head is facing, the first range D1 is also constant regardless of the direction the listener L's head is facing.

[0063] The first sound is a sound that reaches the listener L from all or part of the first range D1, which has such an extent, and is a so-called ambient sound or noise. The first sound is also sometimes called ambient sound. In this embodiment, the first sound is an ambient sound that reaches the listener L from all of the first range D1. Here, the first sound is a sound that reaches the listener L from the entire area marked with dots in Figure 2.

[0064] The second audio signal is a signal indicating a second tone, which is the sound reaching the listener L from a predetermined direction.

[0065] The second sound is, for example, a sound whose sound image is localized at the black dot shown in Figure 2. The second sound may also reach the listener L from a narrower range than the first sound. The second sound is, for example, a so-called target sound, which is a sound primarily heard by the listener L. Furthermore, a target sound can be defined as any sound other than ambient noise.

[0066] Furthermore, as shown in Figure 2, in this embodiment, the predetermined direction is the 5 o'clock direction, and the arrow indicates that the second tone reaches the listener L from the predetermined direction. Also, the predetermined direction is constant regardless of the direction the listener L's head is facing.

[0067] Let me explain the first signal processing unit 110 again.

[0068] Furthermore, the first signal processing unit 110 performs a process to separate multiple audio signals into a first audio signal and a second audio signal. The first signal processing unit 110 outputs the separated first audio signal to the first decoding unit 121 and the separated second audio signal to the second decoding unit 122. In this embodiment, the first signal processing unit 110 is a demultiplexer as an example, but is not limited to this.

[0069] In this embodiment, it is preferable that the multiple audio signals acquired by the first signal processing unit 110 are encoded using a method such as MPEG-H 3D Audio (ISO / IEC 23008-3) (hereinafter referred to as MPEG-H 3D Audio). In other words, the first signal processing unit 110 acquires multiple audio signals, which are encoded bitstreams.

[0070] The first decoding unit 121 and the second decoding unit 122, which are examples of signal acquisition units, acquire multiple audio signals. Specifically, the first decoding unit 121 acquires and decodes the first audio signal separated by the first signal processing unit 110. The second decoding unit 122 acquires and decodes the second audio signal separated by the first signal processing unit 110. The first decoding unit 121 and the second decoding unit 122 perform decoding processing based on the above-mentioned MPEG-H 3D Audio, etc.

[0071] The first decoding unit 121 outputs the decoded first audio signal to the first correction processing unit 131, and the second decoding unit 122 outputs the decoded second audio signal to the second correction processing unit 132.

[0072] Furthermore, the first decoding unit 121 outputs first information to the information acquisition unit 140, which is information indicating the first range D1 included in the first audio signal. The second decoding unit 122 outputs second information to the information acquisition unit 140, which is information indicating a predetermined direction, which is the direction in which the second sound included in the second audio signal reaches the listener L.

[0073] The information acquisition unit 140 is a processing unit that acquires orientation information output from the head sensor 300. The information acquisition unit 140 also acquires the first information output by the first decoding unit 121 and the second information output by the second decoding unit 122. The information acquisition unit 140 outputs the acquired orientation information, the first information, and the second information to the first correction processing unit 131 and the second correction processing unit 132.

[0074] The first correction processing unit 131 and the second correction processing unit 132 are examples of correction processing units. A correction processing unit is a processing unit that applies correction processing to at least one of the first audio signal and the second audio signal.

[0075] The first correction processing unit 131 acquires the first audio signal acquired by the first decoding unit 121, and the direction information, first information, and second information acquired by the information acquisition unit 140. The second correction processing unit 132 acquires the second audio signal acquired by the second decoding unit 122, and the direction information, first information, and second information acquired by the information acquisition unit 140.

[0076] The correction processing units (first correction processing unit 131 and second correction processing unit 132) perform correction processing on at least one of the first audio signal and the second audio signal when predetermined conditions (described later in Figures 3 to 6) are met, based on the acquired orientation information. More specifically, the first correction processing unit 131 performs correction processing on the first audio signal, and the second correction processing unit 132 performs correction processing on the second audio signal.

[0077] If the first audio signal and the second audio signal have been corrected, the first correction processing unit 131 outputs the corrected first audio signal to the mixing processing unit 150, and the second correction processing unit 132 outputs the corrected second audio signal to the mixing processing unit 150.

[0078] Furthermore, if the first audio signal is corrected, the first correction processing unit 131 outputs the corrected first audio signal to the mixing processing unit 150, and the second correction processing unit 132 outputs the uncorrected second audio signal to the mixing processing unit 150.

[0079] Furthermore, if the second audio signal is corrected, the first correction processing unit 131 outputs the first audio signal that has not been corrected to the mixing processing unit 150, and the second correction processing unit 132 outputs the second audio signal that has been corrected to the mixing processing unit 150.

[0080] The mixing processing unit 150 is a processing unit that mixes at least one of the first audio signal and the second audio signal, which have been corrected by the correction processing unit, and outputs them to a plurality of speakers 1, 2, 3, 4, and 5, which are output channels.

[0081] More specifically, if the first audio signal and the second audio signal have been corrected, the mixing processing unit 150 mixes the corrected first audio signal and the second audio signal and outputs the result. If the first audio signal has been corrected, the mixing processing unit 150 mixes the corrected first audio signal and the uncorrected second audio signal and outputs the result. If the second audio signal has been corrected, the mixing processing unit 150 mixes the uncorrected first audio signal and the corrected second audio signal and outputs the result.

[0082] As another example, if headphones placed near the ears of the listener L are used as the output channel, rather than multiple speakers 1, 2, 3, 4, and 5 placed around the listener L, the mixing processing unit 150 performs the following processing. In this case, when the mixing processing unit 150 mixes the first audio signal and the second audio signal, it performs a process of convolving the head-related transfer function and outputs the result.

[0083] [Example of operation] The following describes an example of the operation of the sound reproduction method performed by the sound reproduction device 100. Figure 3 is a flowchart of an example of the operation of the sound reproduction device 100 according to this embodiment.

[0084] The first signal processing unit 110 acquires multiple audio signals (S10).

[0085] The first signal processing unit 110 separates the multiple audio signals acquired by the first signal processing unit 110 into a first audio signal and a second audio signal (S20).

[0086] The first decoding unit 121 and the second decoding unit 122 acquire the separated first audio signal and second audio signal, respectively (S30). Step S30 is a signal acquisition step. More specifically, the first decoding unit 121 acquires the first audio signal, and the second decoding unit 122 acquires the second audio signal. Furthermore, the first decoding unit 121 decodes the first audio signal, and the second decoding unit 122 decodes the second audio signal.

[0087] Here, the information acquisition unit 140 acquires directional information output by the head sensor 300 (S40). Step S40 is an information acquisition step. The information acquisition unit 140 also acquires first information indicating a first range D1 included in the first audio signal representing the first sound, and second information indicating a predetermined direction which is the direction in which the second sound reaches the listener L.

[0088] Furthermore, the information acquisition unit 140 outputs the acquired azimuth information, first information, and second information to the first correction processing unit 131 and the second correction processing unit 132 (i.e., the correction processing unit).

[0089] The correction processing unit acquires the first audio signal, the second audio signal, azimuth information, first information, and second information. Furthermore, the correction processing unit determines, based on the azimuth information, whether the first range D1 and a predetermined azimuth are included in the second range D2 (S50). More specifically, the correction processing unit makes the above determination based on the acquired azimuth information, first information, and second information.

[0090] Here, the judgment made by the correction processing unit and the second range D2 will be explained using Figures 4 to 6.

[0091] Figures 4 to 6 are schematic diagrams illustrating an example of a decision made by the correction processing unit according to this embodiment. More specifically, in Figures 4 and 5, the correction processing unit determines that the first range D1 and the predetermined direction are included in the second range D2, while in Figure 6, the correction processing unit determines that the first range D1 and the predetermined direction are not included in the second range D2. Furthermore, Figures 4, 5, and 6 show, in that order, how the direction towards which the listener L's head is facing changes clockwise.

[0092] The second range D2, as shown in Figures 4 to 6, is the range behind the listener L when the direction their head is facing is considered forward. In other words, the second range D2 is the range behind the listener L. Furthermore, the second range D2 is the range centered on the direction directly opposite to the direction the listener L's head is facing. For example, as shown in Figure 4, when the direction the listener L's head is facing is 12 o'clock, the second range D2 is the range from 4 o'clock to 8 o'clock, centered on 6 o'clock, which is the direction directly opposite to 12 o'clock (i.e., a range of 120° in terms of angle). However, the second range D2 is not limited to this. Also, the second range D2 is determined based on the direction information acquired by the information acquisition unit 140. As shown in Figures 4 to 6, when the direction in which the listener L's head is facing changes, the second range D2 changes accordingly, but as described above, the first range D1 and the predetermined direction do not change.

[0093] In other words, the correction processing unit determines whether the first range D1 and the predetermined direction are included in the second range D2, which is the range behind the listener L determined based on the direction information. The specific positional relationship between the first range D1, the predetermined direction, and the second range D2 is explained below.

[0094] First, as shown in Figures 4 and 5, we will explain the case where the correction processing unit determines that both the first range D1 and the predetermined direction are included in the second range D2 (Yes in step S50).

[0095] As shown in Figure 4, if the direction towards which the listener L's head is facing is 12 o'clock, then the second range D2 is the range from 4 o'clock to 8 o'clock. Furthermore, the first range D1 for the first sound, which is an ambient sound, is the range from 3 o'clock to 9 o'clock, and the predetermined direction for the second sound, which is the target sound, is 5 o'clock. In other words, the predetermined direction is included in part of the first range D1, and part of the first range D1 is included in the second range D2. At this time, the correction processing unit determines that both the first range D1 and the predetermined direction are included in the second range D2. Moreover, the first and second sounds are sounds that reach the listener L from the second range D2 (behind the listener L).

[0096] Furthermore, the same is true when the direction in which the listener L's head is facing, as shown in Figure 5, moves more clockwise than in the case shown in Figure 4.

[0097] In the cases shown in Figures 4 and 5, the correction processing unit applies correction processing to at least one of the first audio signal and the second audio signal. Here, as an example, the correction processing unit applies correction processing to both the first audio signal and the second audio signal (S60). More specifically, the first correction processing unit 131 applies correction processing to the first audio signal, and the second correction processing unit 132 applies correction processing to the second audio signal. Step S60 is a correction processing step.

[0098] Furthermore, the correction process performed by the correction processing unit is a process that makes the intensity of the second audio signal stronger than the intensity of the first audio signal. "Making the audio signal intensity stronger" means, for example, that the volume or sound pressure of the sound represented by the audio signal becomes stronger. The correction process will be explained in detail in the first to third examples below.

[0099] The first correction processing unit 131 outputs the first audio signal, which has undergone correction processing, to the mixing processing unit 150, and the second correction processing unit 132 outputs the second audio signal, which has undergone correction processing, to the mixing processing unit 150.

[0100] The mixing processing unit 150 mixes the first audio signal and the second audio signal, which have been corrected by the correction processing unit, and outputs them to the output channels, which are a plurality of speakers 1, 2, 3, 4, and 5 (S70). Step S70 is a mixing processing step.

[0101] Next, as shown in Figure 6, we will explain the case where the correction processing unit determines that the first range D1 and the predetermined direction are not included in the second range D2 (No in step S50).

[0102] If the direction towards which the listener L's head is facing is 2 o'clock, as shown in Figure 6, then the second range D2 is the range from 6 o'clock to 10 o'clock. Also, the first range D1 and the predetermined direction do not change from those in Figures 4 and 5. In this case, the correction processing unit determines that the predetermined direction is not included in the second range D2. More specifically, the correction processing unit determines that at least one of the first range D1 and the predetermined direction is not included in the second range D2.

[0103] In the case shown in Figure 6, the correction processing unit does not apply correction processing to the first audio signal and the second audio signal (S80). The first correction processing unit 131 outputs the first audio signal, which has not been corrected, to the mixing processing unit 150, and the second correction processing unit 132 outputs the second audio signal, which has not been corrected, to the mixing processing unit 150.

[0104] The mixing processing unit 150 mixes the first audio signal and the second audio signal, which have not been corrected by the correction processing unit, and outputs them to the output channels, which are a plurality of speakers 1, 2, 3, 4, and 5 (S90).

[0105] Thus, in this embodiment, when the correction processing unit determines that the first range D1 and a predetermined direction are included in the second range D2, the correction processing unit applies a correction process to at least one of the first audio signal and the second audio signal. This correction process is a process that makes the intensity of the second audio signal stronger than the intensity of the first audio signal.

[0106] As a result, the intensity of the second audio signal indicating the second tone increases when the first range D1 and a predetermined direction are included in the second range D2. Therefore, the listener L can more easily hear the second tone that reaches the listener L from behind (i.e., from behind the listener L) when the direction the listener L's head is facing is considered forward. In other words, an acoustic reproduction device 100 and acoustic reproduction method are realized that can improve the perceived level of the second tone arriving from behind the listener L.

[0107] For example, if the first sound is an ambient sound and the second sound is the target sound, it is possible to suppress the target sound from being drowned out by the ambient sound. In other words, an acoustic reproduction device 100 is realized that can improve the perceived level of the target sound arriving from behind the listener L.

[0108] Furthermore, the first range D1 is the range behind the reference direction determined by the positions of the five speakers 1, 2, 3, 4, and 5.

[0109] This makes it easier for listener L to hear the second tone, which reaches listener L from behind, even if the first tone reaches listener L from a range behind the reference direction.

[0110] Here, we will describe the first to third examples of correction processing performed by the correction processing unit.

[0111] <Example 1> In the first example, the correction process is a process of correcting at least one of the gain of the first audio signal acquired by the first decoding unit 121 and the gain of the second audio signal acquired by the second decoding unit 122. More specifically, the correction process is at least one of a process of decreasing the gain of the first audio signal and a process of increasing the gain of the second audio signal.

[0112] Figure 7 illustrates an example of the correction processing performed by the correction processing unit according to this embodiment. More specifically, Figure 7(a) shows the relationship between the time and amplitude of the first audio signal and the second audio signal before the correction processing is applied. Note that the first range D1 and the multiple speakers 1, 2, 3, 4 and 5 are omitted in Figure 7, and the same applies to Figures 8 and 9, which will be described later.

[0113] Figure 7(b) shows an example where no correction processing is applied to the first audio signal and the second audio signal. The positional relationship between the first range D1, the predetermined direction, and the second range D2 shown in Figure 7(b) corresponds to Figure 6, meaning that Figure 7(b) shows the case where step S50 shown in Figure 3 is No. In this case, the correction processing unit does not apply correction processing to the first audio signal and the second audio signal.

[0114] Figure 7(c) shows an example in which the first audio signal and the second audio signal have been corrected. The positional relationship between the first range D1, the predetermined direction, and the second range D2 shown in Figure 7(c) corresponds to Figure 4, meaning that Figure 7(c) shows the case where step S50 shown in Figure 3 is Yes.

[0115] In this case, the correction processing unit performs at least one of the following correction processes: reducing the gain of the first audio signal and increasing the gain of the second audio signal. Here, the correction processing unit performs both correction processes: reducing the gain of the first audio signal and increasing the gain of the second audio signal. As the gains of the first and second audio signals are corrected in this way, the amplitudes of the first and second audio signals are corrected, as shown in Figure 7. In other words, the correction processing unit performs both processes: reducing the amplitude of the first audio signal representing the first tone and increasing the amplitude of the second audio signal representing the second tone. Therefore, the listener L can hear the second tone more easily.

[0116] In the first example, the correction process is a process that corrects the gain of at least one of the first audio signal and the second audio signal. As a result, the amplitude of at least one of the first audio signal representing the first sound and the second audio signal representing the second sound is corrected, making it easier for the listener L to hear the second sound.

[0117] More specifically, the correction process involves at least one of the following: reducing the gain of the first audio signal representing the first tone, and increasing the gain of the second audio signal representing the second tone. This makes it easier for the listener L to hear the second tone.

[0118] <Example 2> In the second example, the correction process is a process of correcting at least one of the frequency components based on the first audio signal acquired by the first decoding unit 121 and the frequency components based on the second audio signal acquired by the second decoding unit 122. More specifically, the correction process is a process of reducing the spectrum of the frequency components based on the first audio signal so that it is smaller than the spectrum of the frequency components based on the second audio signal. Here, as an example, the correction process is a process of subtracting the spectrum of the frequency components based on the second audio signal from the spectrum of the frequency components based on the first audio signal.

[0119] Figure 8 illustrates another example of the correction processing performed by the correction processing unit according to this embodiment. More specifically, Figure 8(a) shows the spectrum of the frequency components based on the first audio signal and the second audio signal before the correction processing is performed. The spectrum of the frequency components is obtained, for example, by performing a Fourier transform on the first audio signal and the second audio signal.

[0120] Figure 8(b) shows an example where no correction processing is applied to the first audio signal and the second audio signal. The positional relationship between the first range D1, the predetermined direction, and the second range D2 shown in Figure 8(b) corresponds to Figure 6, meaning that Figure 8(b) shows the case where step S50 shown in Figure 3 is No. In this case, the correction processing unit does not apply correction processing to the first audio signal and the second audio signal.

[0121] Figure 8(c) shows an example where the first audio signal has been corrected. The positional relationship between the first range D1, the predetermined direction, and the second range D2 shown in Figure 8(c) corresponds to Figure 4, meaning that Figure 8(c) shows the case where step S50 shown in Figure 3 is Yes.

[0122] In this case, the correction processing unit (more specifically, the first correction processing unit 131) subtracts the spectrum of the frequency components based on the second audio signal from the spectrum of the frequency components based on the first audio signal. As a result, as shown in Figure 8(c), the intensity of the spectrum of the frequency components based on the first audio signal representing the first tone decreases. On the other hand, since no correction processing is applied to the second audio signal, the intensity of the spectrum of the frequency components based on the second audio signal representing the second tone remains constant. In other words, the intensity of the spectrum of some of the frequency components based on the first audio signal decreases, while the intensity of the second audio signal remains constant. Therefore, the listener L can hear the second tone more easily.

[0123] In the second example, the correction process corrects at least one of the frequency components based on the first audio signal representing the first tone and the frequency components based on the second audio signal representing the second tone. This makes it easier for the listener L to hear the second tone.

[0124] Furthermore, the correction process reduces the spectrum of frequency components based on the first audio signal so that it is smaller than the spectrum of frequency components based on the second audio signal. Here, the correction process is the process of subtracting the spectrum of frequency components based on the second audio signal from the spectrum of frequency components based on the first audio signal. As a result, the intensity of some of the spectrum of frequency components based on the first audio signal representing the first sound is reduced, making it easier for the listener L to hear the second sound.

[0125] Furthermore, the correction process may be a process that reduces the spectrum of frequency components based on the first audio signal to be smaller than the spectrum of frequency components based on the second audio signal by a predetermined ratio. For example, the correction process may be applied so that the peak intensity of the spectrum of frequency components based on the second audio signal is less than or equal to a predetermined ratio to the peak intensity of the spectrum of frequency components based on the first audio signal.

[0126] <Example 3> In the third example, the correction processing unit performs a correction process based on the positional relationship between the second range D2 and a predetermined direction. In this case, the correction process is a process that corrects at least one of the gains of the first audio signal and the second audio signal, or a process that corrects at least one of the frequency characteristics based on the first audio signal and the frequency characteristics based on the second audio signal. Here, the correction process is a process that corrects at least one of the gains of the first audio signal and the second audio signal.

[0127] Figure 9 illustrates another example of the correction processing performed by the correction processing unit according to this embodiment. More specifically, Figure 9(a) shows the relationship between the time and amplitude of the first audio signal and the second audio signal before the correction processing is performed. Figures 9(b) and (c) show examples in which at least one of the gains of the first audio signal and the second audio signal is corrected. In Figure 9(c), an example is shown in which the second sound is a sound that reaches the listener L from the 7 o'clock direction.

[0128] Furthermore, in the third example, the second range D2 is divided as follows: As shown in Figures 9(b) and (c), the second range D2 is divided into the right rear range D21, which is the area to the right rear of the listener L; the left rear range D23, which is the area to the left rear of the listener L; and the central rear range D22, which is the area between the right rear range D21 and the left rear range D23. It is desirable that the central rear range D22 includes the direction directly behind the listener L.

[0129] Figure 9(b) shows an example where the correction processing unit determines that a predetermined direction (in this case, the 5 o'clock direction) is included in the right rear range D21. At this time, the correction processing unit performs a correction process, which is either a process to decrease the gain of the first audio signal or a process to increase the gain of the second audio signal. Here, the correction processing unit (more specifically, the second correction processing unit 132) performs a correction process that increases the gain of the second audio signal.

[0130] This makes it easier for listener L to hear the second tone.

[0131] Although not shown in the diagram, the same correction process is also applied in cases where the correction processing unit determines that a predetermined direction is included in the left rear range D23.

[0132] Furthermore, Figure 9(c) shows an example in which the correction processing unit determines that a predetermined direction (in this case, the 7 o'clock direction) is included in the central rear range D22. At this time, the correction processing unit performs correction processing, which involves reducing the gain of the first audio signal and increasing the gain of the second audio signal. Here, the first correction processing unit 131 performs correction processing to reduce the gain of the first audio signal, and the second correction processing unit 132 performs correction processing to increase the gain of the second audio signal. As a result, the amplitude of the first audio signal is reduced and the amplitude of the second audio signal is increased.

[0133] As a result, listener L is more likely to hear the second tone compared to the example shown in Figure 9(b).

[0134] As mentioned above, humans have a low perception level of sounds that reach them from behind. Furthermore, the closer the direction from which the sound is coming is to directly behind the person, the lower the perception level of that sound becomes.

[0135] Therefore, a correction process is performed as shown in the third example. In other words, a correction process is performed based on the positional relationship between the second range D2 and a predetermined direction. More specifically, when the predetermined direction is included in the central rear range D22, which includes the direction directly behind the listener L, the following correction process is performed. In this case, compared to when the predetermined direction is included in the right rear range D21, a correction process is performed so that the intensity of the second audio signal indicating the second tone is stronger than the intensity of the first audio signal indicating the first tone. Consequently, the listener L will be able to hear the second tone more easily.

[0136] [Details of the correction process] Furthermore, the details of how the correction processing unit applies correction processing to the first audio signal representing the first tone will be explained using Figures 10 and 11.

[0137] Figure 10 is a schematic diagram showing an example of the correction process applied to the first audio signal according to this embodiment. Figure 11 is a schematic diagram showing another example of the correction process applied to the first audio signal according to this embodiment. In Figures 10 and 11, as in Figure 2, the direction towards which the listener L's head is facing is the 12 o'clock direction.

[0138] In the first to third examples described above, the correction processing unit may apply correction processing to the first audio signal that represents a portion of the first sound, as shown below.

[0139] For example, as shown in Figure 10, the correction processing unit applies correction processing to the first audio signal that represents the sound reaching the listener L from the entire range of the second range D2 of the first tone. The sound reaching the listener L from the entire range of the second range D2 of the first tone is the sound reaching the listener L from the entire area marked with light dots in Figure 10. The other sounds of the first tone are the sounds reaching the listener L from the entire area marked with dark dots in Figure 10.

[0140] In this case, the correction processing unit performs a correction process, for example, which is a process of reducing the gain of the first audio signal that represents the sound reaching the listener L from the entire range of the second range D2 of the first sound.

[0141] Furthermore, as shown in Figure 11, for example, the correction processing unit applies correction processing to the first audio signal of the first sound that represents the sound reaching the listener L from around a predetermined direction in which the second sound reaches the listener L. The area around the predetermined direction is, as shown in Figure 11, for example, a range D11 of an angle of about 30° centered on the predetermined direction, but is not limited to this.

[0142] Furthermore, among the first tones, the sound that reaches the listener L from the vicinity of the predetermined direction is the sound that reaches the listener L from the entire area marked with light dots in Figure 11. The other sounds among the first tones are the sounds that reach the listener L from the entire area marked with dark dots in Figure 11.

[0143] In this case, the correction processing unit performs a correction process, for example, which is a process of reducing the gain of the first audio signal that represents the sound reaching the listener L from around a predetermined direction in which the second of the first tones reaches the listener L.

[0144] Thus, correction processing may be applied to the first audio signal that represents only a portion of the first sound. This eliminates the need to apply correction processing to the entire first audio signal, thereby reducing the processing load on the first correction processing unit 131 that corrects the first audio signal.

[0145] The same processing may also be applied to the first audio signal that represents all of the first tones.

[0146] (Embodiment 2) Next, the sound reproduction device 100a according to Embodiment 2 will be described.

[0147] Figure 12 is a block diagram showing the functional configuration of the sound reproduction device 100a and the sound acquisition device 200 according to this embodiment.

[0148] In this embodiment, the sound picked up by the sound pickup device 500 is output from multiple speakers 1, 2, 3, 4, and 5 via the sound acquisition device 200 and the sound reproduction device 100a. More specifically, the sound acquisition device 200 acquires multiple audio signals based on the sound picked up by the sound pickup device 500 and outputs them to the sound reproduction device 100a. The sound reproduction device 100a acquires the multiple audio signals output by the sound acquisition device 200 and outputs them to multiple speakers 1, 2, 3, 4, and 5.

[0149] The sound-collecting device 500 is a device that collects sound that reaches the sound-collecting device 500, and is, for example, a microphone. The sound-collecting device 500 may be directional. Therefore, the sound-collecting device 500 can collect sound from a specific direction. The sound-collecting device 500 converts the collected sound with an A / D converter and outputs it as an audio signal to the sound acquisition device 200. Multiple sound-collecting devices 500 may be provided.

[0150] The sound collection device 500 will be explained in more detail using Figure 13.

[0151] Figure 13 is a schematic diagram illustrating sound collection by the sound collection device 500 according to this embodiment.

[0152] In Figure 13, as in Figure 2, 0 o'clock, 3 o'clock, 6 o'clock, and 9 o'clock are shown to illustrate direction, corresponding to the time indicated on the clock face. The sound-collecting device 500 is located at the center (also called the origin) of the clock face and collects sound reaching the sound-collecting device 500. Hereafter, the direction connecting the sound-collecting device 500 and 0 o'clock may be referred to as the "direction of 0 o'clock," and the same applies to the other times indicated on the clock face.

[0153] The sound-collecting device 500 collects multiple first sounds and second sounds.

[0154] Here, the sound-collecting device 500 collects four first tones as multiple first tones. For identification purposes, these will be labeled as first tone A, first tone B-1, first tone B-2, and first tone B-3, as shown in Figure 13.

[0155] Since the sound-collecting device 500 can collect sound from a specific direction, as an example, as shown in Figure 13, the area around the sound-collecting device 500 is divided into four sections, and sound is collected in each of these sections. In this case, the area around the sound-collecting device 500 is divided into four sections: from 12 o'clock to 3 o'clock, from 3 o'clock to 6 o'clock, from 6 o'clock to 9 o'clock, and from 9 o'clock to 12 o'clock.

[0156] In this embodiment, each of the multiple first sounds is a sound that reaches the sound-collecting device 500 from a first range D1 which is a predetermined range of angles; in other words, it is a sound collected by the sound-collecting device 500 from each of the multiple first ranges D1. The first range D1 corresponds to any of the four ranges.

[0157] Specifically, as shown in Figure 13, the first tone A is the sound that reaches the sound-collecting device 500 from the first range D1, which is the range from the 12 o'clock direction to the 3 o'clock direction. In other words, the first tone A is the sound collected from the said first range D1. Similarly, the first tone B-1, first tone B-2, and first tone B-3 are the sounds that reach the sound-collecting device 500 from the first range D1, which is the range from the 3 o'clock direction to the 6 o'clock direction, from the 6 o'clock direction to the 9 o'clock direction, and from the 9 o'clock direction to the 12 o'clock direction, respectively. In other words, each of the first tone B-1, first tone B-2, and first tone B-3 is the sound collected from each of the three said first ranges D1. Note that the first tone B-1, first tone B-2, and first tone B-3 may be collectively referred to as the first tone B.

[0158] Furthermore, in this case, the first tone A is the sound that reaches listener L from the entire shaded area in Figure 13. Similarly, the first tones B-1, B-2, and B-3 are the sounds that reach listener L from the entire dotted area in Figure 13. The same applies to Figure 14.

[0159] The second tone is the sound that reaches the sound-collecting device 500 from a predetermined direction (in this case, the 5 o'clock direction). The second tone, like the multiple first tones, may also be collected in divided sections.

[0160] Next, the relationship between the sound picked up by the sound pickup device 500 and the sounds output from the multiple speakers 1, 2, 3, 4, and 5 will be explained. The multiple speakers 1, 2, 3, 4, and 5 output sound in a manner that reproduces the sound picked up by the sound pickup device 500. In other words, in this embodiment, since both the listener L and the sound pickup device 500 are positioned at the origin, the second sound reaching the sound pickup device 500 from a predetermined direction is heard by the listener L as sound reaching the listener L from the predetermined direction. Similarly, the first sound A reaching the sound pickup device 500 from the first range D1 (the range from the 12 o'clock direction to the 3 o'clock direction) is heard by the listener L as sound reaching the listener L from the first range D1.

[0161] The sound acquisition device 500 outputs multiple audio signals to the sound acquisition device 200. These multiple audio signals include multiple first audio signals representing multiple first tones and second audio signals representing second tones. Furthermore, the multiple first audio signals include a first audio signal representing first tone A and a first audio signal representing first tone B. More specifically, the first audio signal representing first tone B includes three first audio signals representing first tone B-1, first tone B-2, and first tone B-3, respectively.

[0162] The sound acquisition device 200 acquires multiple audio signals output by the sound collection device 500. The sound acquisition device 200 may also acquire classification information at this time.

[0163] Classification information refers to information that classifies multiple first audio signals based on the respective frequency characteristics of each first audio signal. In other words, in classification information, multiple first audio signals are classified into different groups based on their respective frequency characteristics.

[0164] In this embodiment, the first sound A and the first sound B are different types of sounds and have different frequency characteristics. Therefore, the first audio signal representing the first sound A and the first audio signal representing the first sound B are classified into different groups.

[0165] In other words, the first audio signal representing the first tone A is classified into one group, while the three first audio signals representing the first tone B-1, first tone B-2, and first tone B-3 are each classified into another group.

[0166] Furthermore, instead of the sound acquisition device 200 acquiring classification information, the sound acquisition device 200 may generate classification information based on the multiple audio signals it has acquired. In other words, the classification information may be generated by a processing unit included in the sound acquisition device 200, which is not shown in Figure 13.

[0167] Next, the components of the sound acquisition device 200 will be described. As shown in Figure 12, the sound acquisition device 200 is a device that comprises an encoding unit (a plurality of first encoding units 221 and second encoding units 222) and a second signal processing unit 210.

[0168] The encoding unit (a plurality of first encoding units 221 and second encoding units 222) acquires a plurality of audio signals output by the sound acquisition device 500 and classification information. After acquiring the plurality of audio signals, the encoding unit encodes them. More specifically, the plurality of first encoding units 221 acquire and encode a plurality of first audio signals, and the second encoding unit 222 acquires and encodes a second audio signal. The plurality of first encoding units 221 and second encoding units 222 perform encoding processing based on the above-mentioned MPEG-H 3D Audio, etc.

[0169] Here, each of the multiple first encoding units 221 is preferably associated one-to-one with each of the multiple first audio signals classified into different groups indicated by the classification information. Each of the multiple first encoding units 221 encodes each of the multiple first audio signals it is associated with. For example, the classification information indicates two groups (a group in which the first audio signal representing the first tone A is classified, and a group in which the first audio signal representing the first tone B is classified). Therefore, here, two first encoding units 221 are provided, with one of the two first encoding units 221 encoding the first audio signal representing the first tone A, and the other of the two first encoding units 221 encoding the first audio signal representing the first tone B. If the sound acquisition device 200 is equipped with one first encoding unit 221, that one first encoding unit 221 acquires and encodes the multiple first audio signals.

[0170] The encoding unit outputs the encoded first audio signals, the encoded second audio signals, and classification information to the second signal processing unit 210.

[0171] The second signal processing unit 210 acquires a plurality of encoded first audio signals and a plurality of encoded second audio signals, along with classification information. The second signal processing unit 210 combines the plurality of encoded first audio signals and the plurality of encoded second audio signals to form a plurality of encoded audio signals. A plurality of encoded audio signals is a so-called multiplexed plurality of audio signals. In this embodiment, the second signal processing unit 210 is a multiplexer as an example, but is not limited to this.

[0172] The second signal processing unit 210 outputs multiple audio signals, which are encoded bitstreams, and classification information to the sound playback device 100a (more specifically, the first signal processing unit 110).

[0173] The following description of the processing performed by the sound reproduction device 100a mainly focuses on the differences from Embodiment 1. Note that in this embodiment, the sound reproduction device 100a differs from Embodiment 1 in that it includes a plurality of first decoding units 121.

[0174] The first signal processing unit 110 acquires the output audio signals and classification information, and processes the audio signals to separate them into multiple first audio signals and multiple second audio signals. The first signal processing unit 110 outputs the separated multiple first audio signals and classification information to multiple first decoding units 121, and the separated second audio signals and classification information to second decoding units 122.

[0175] Multiple first decoding units 121 acquire and decode multiple first audio signals separated by the first signal processing unit 110.

[0176] Here, each of the multiple first decoding units 121 is preferably associated one-to-one with each of the multiple first audio signals classified into different groups indicated by the classification information. Each of the multiple first decoding units 121 decodes each of the associated multiple first audio signals. Similar to the first encoding unit 221 described above, here two first decoding units 121 are provided, with one of the two first decoding units 121 decoding the first audio signal representing first tone A, and the other of the two first decoding units 121 decoding the first audio signal representing first tone B. Note that if the sound reproduction device 100a is equipped with one first decoding unit 121, that one first decoding unit 121 acquires and decodes multiple first audio signals.

[0177] Multiple first decoding units 121 output the decoded multiple first audio signals and classification information to the first correction processing unit 131. The second decoding unit 122 outputs the decoded second audio signal and classification information to the second correction processing unit 132.

[0178] Furthermore, the first correction processing unit 131 acquires multiple first audio signals and classification information acquired by multiple first decoding units 121, as well as directional information, first information, and second information acquired by the information acquisition unit 140.

[0179] Similarly, the second correction processing unit 132 acquires the second audio signal and classification information acquired by the second decoding unit 122, and the orientation information, first information, and second information acquired by the information acquisition unit 140.

[0180] Furthermore, the first information according to this embodiment includes information indicating one first range D1 relating to the first tone A and three first range D1 relating to the first tone B, which are included in a plurality of first audio signals.

[0181] Next, the correction processing performed by the correction processing unit will be explained using Figure 14. Figure 14 is a schematic diagram showing an example of the correction processing performed on a plurality of first audio signals according to this embodiment. Figure 14(a) shows an example before the correction processing is performed, and Figure 14(b) shows an example after the correction processing is performed.

[0182] In this embodiment, the correction processing unit performs correction processing based on azimuth information and classification information. Here, we will describe the case in which the correction processing unit determines that one of the multiple first ranges D1 and a predetermined azimuth are included in the second range D2. In this case, the correction processing unit performs correction processing on at least one of the first audio signal and second audio signal that represent one first sound reaching the listener L from the said first range D1. More specifically, the correction processing unit performs correction processing on at least one of all first audio signals and second audio signals classified into the same group as the said first audio signal based on the classification information.

[0183] For example, in Figure 14, the correction processing unit determines that the first range D1 (the range from the 3 o'clock direction to the 6 o'clock direction) and a predetermined direction (the 5 o'clock direction) are included in the second range D2 (the range from the 4 o'clock direction to the 8 o'clock direction). The sound that reaches the listener L from the first range D1 is the first tone B-1. All first audio signals classified in the same group as the first audio signal representing the first tone B-1 are the three first audio signals representing the first tone B-1, the first tone B-2, and the first tone B-3, respectively.

[0184] In other words, the correction processing unit applies correction processing to at least one of the three first audio signals (in other words, the first audio signal representing the first tone B) that represent the first tone B-1, first tone B-2, and first tone B-3 respectively, and the second audio signal.

[0185] This allows the correction processing unit to perform correction processing on each group into which multiple first audio signals are classified. Here, the correction processing unit can perform correction processing on three first audio signals, each representing the first tone B-1, first tone B-2, and first tone B-3, together. Therefore, the processing load on the correction processing unit can be reduced.

[0186] (Other embodiments) The sound reproduction apparatus and sound reproduction method according to the embodiments of this disclosure have been described above based on embodiments, but this disclosure is not limited to these embodiments. For example, other embodiments realized by arbitrarily combining the components described herein, or by excluding some of the components, may also be considered embodiments of this disclosure. Furthermore, modifications obtained by applying various modifications to the above embodiments that a person skilled in the art could conceive of without departing from the spirit of this disclosure, that is, the meaning indicated by the wording in the claims, are also included in this disclosure.

[0187] Furthermore, the following forms may also be included within the scope of one or more aspects of this disclosure.

[0188] (1) Some of the components constituting the above-described sound reproduction device may be a computer system consisting of a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. A computer program is stored in the RAM or hard disk unit. The microprocessor achieves its function by operating in accordance with the computer program. Here, the computer program is composed of a combination of multiple instruction codes that indicate commands to the computer in order to achieve a predetermined function.

[0189] (2) Some of the components constituting the above-described sound reproduction device and sound reproduction method may be composed of a single system LSI (Large Scale Integration). The system LSI is a highly functional LSI manufactured by integrating multiple components onto a single chip, and specifically, it is a computer system comprising a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. The system LSI achieves its function by operating the microprocessor in accordance with the computer program.

[0190] (3) Some of the components constituting the above-described sound reproduction device may consist of detachable IC cards or standalone modules attached to each device. The IC card or module is a computer system consisting of a microprocessor, ROM, RAM, etc. The IC card or module may include the above-described multi-functional LSI. The microprocessor operates according to a computer program, thereby enabling the IC card or module to perform its function. The IC card or module may be tamper-resistant.

[0191] (4) Furthermore, some of the components constituting the above-described sound reproduction device may be the computer program or the digital signal recorded on a recording medium that can be read by a computer, such as a flexible disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray® Disc), semiconductor memory, etc. Alternatively, it may be the digital signal recorded on these recording media.

[0192] Furthermore, some of the components constituting the above-mentioned sound reproduction device may transmit the computer program or the digital signal via telecommunications lines, wireless or wired communication lines, networks such as the Internet, data broadcasting, etc.

[0193] (5) The disclosure may also be the methods described above. Alternatively, it may be a computer program that implements these methods using a computer, or a digital signal consisting of the computer program.

[0194] (6) The Disclosure may also provide a computer system comprising a microprocessor and memory, wherein the memory stores the computer program, and the microprocessor operates in accordance with the computer program.

[0195] (7) Alternatively, the program or the digital signal may be carried out by another independent computer system by recording it on the recording medium and transferring it, or by transferring the program or the digital signal via the network or the like.

[0196] (8) The above embodiments and the above modified examples may be combined.

[0197] Furthermore, although not shown in Figure 2, etc., images synchronized with sounds output from multiple speakers 1, 2, 3, 4, and 5 may be presented to the listener L. In this case, for example, a display device such as a liquid crystal panel or an organic EL (Electro-Luminescence) panel may be provided around the listener L, and the images may be presented on such a display device. Alternatively, the images may be presented to the listener L by wearing a head-mounted display or the like.

[0198] In the above embodiment, as shown in Figure 2, five speakers 1, 2, 3, 4, and 5 are provided, but the system is not limited to this. For example, a 5.1ch surround system may be used, which includes the five speakers 1, 2, 3, 4, and 5 and a speaker corresponding to a subwoofer. Alternatively, a multi-channel surround system with two speakers may be used, but the system is not limited to these. [Industrial applicability]

[0199] This disclosure is applicable to sound reproduction devices and sound reproduction methods, and is particularly applicable to stereophonic sound reproduction systems. [Explanation of Symbols]

[0200] 1, 2, 3, 4, 5 Speakers 100, 100a Sound reproduction device 110 First Signal Processing Unit 121 First Decoding Section 122 Second Decoding Section 131 First Correction Processing Unit 132 Second Correction Processing Unit 140 Information Acquisition Department 150 Mixing Processing Unit 200 Acoustic Acquisition Device 210 Second Signal Processing Unit 221 1st encoding section 222 Second encoding section 300 head sensors 500 Sound Collection Device D1 First Range D2 Second Range D11 Range D21 Right rear range D22 Central rear range D23 Left rear range L listener

Claims

1. A signal acquisition step involves acquiring a plurality of first audio signals that represent a first sound, which is a sound reaching the listener from a first range, which is a predetermined range of angles, and classification information, which is information that classifies the plurality of first audio signals. When one of the multiple first tones is determined to be a trigger tone, the process includes a correction step of applying correction processing to the first audio signal representing the determined trigger tone and one or more first audio signals representing one or more first tones classified in the same group as the trigger tone, based on the acquired classification information. Sound reproduction method.

2. Includes an information acquisition step of acquiring directional information which is information about the direction the listener's head is facing, When the direction the listener's head is facing is defined as the front, and the range behind it is defined as the second range, The correction processing step determines, based on the acquired directional information, that one of the multiple first tones has a first range that falls within the second range, and then determines that one of the first tones is the trigger tone. The sound reproduction method according to claim 1.

3. The signal acquisition step acquires a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction, The correction processing step determines, based on the acquired direction information, that one of the multiple first sounds has a first range and a predetermined direction included in the second range, and then determines that one of the first sounds is the trigger sound. The sound reproduction method according to claim 2.

4. The first sound is an ambient sound. The sound reproduction method according to claim 1.

5. The classification information is information obtained by classifying the plurality of first audio signals based on the frequency characteristics of each of the plurality of first audio signals. The sound reproduction method according to claim 1.

6. The signal acquisition step acquires a second audio signal indicating a second sound which is a sound reaching the listener from a predetermined direction, The correction processing step involves performing a correction process that makes the intensity of the acquired second audio signal stronger than the intensity of the first audio signal indicating the trigger sound and the intensity of one or more first audio signals. The sound reproduction method according to claim 1.

7. The range indicated by the second range changes according to the change in the direction the listener's head is facing. The sound reproduction method according to claim 2.

8. The correction processing step involves performing the correction processing, which is a process of weakening the intensity of the first audio signal indicating the trigger sound and the intensity of one or more of the first audio signals. The sound reproduction method according to claim 6.

9. The correction process is a process of correcting the gain of the first audio signal indicating the trigger sound and the gain of one or more of the first audio signals. The sound reproduction method according to claim 6.

10. The correction process is a process of reducing the gain of the first audio signal indicating the trigger sound and the gain of one or more of the first audio signals. The sound reproduction method according to claim 6.

11. The correction process is a process of correcting the frequency components based on the first audio signal indicating the trigger sound and one or more frequency components based on the first audio signal. The sound reproduction method according to claim 6.

12. The correction process is a process that reduces the spectrum of frequency components based on the first audio signal indicating the trigger sound and the spectrum of frequency components based on one or more of the first audio signals to be smaller than the spectrum of frequency components based on the acquired second audio signal. The sound reproduction method according to claim 6.

13. A computer program for causing a computer to execute the sound reproduction method described in any one of claims 1 to 12.

14. A signal acquisition unit that acquires a plurality of first audio signals representing a first sound that reaches the listener from a first range which is a predetermined range of angles, and classification information which is information that classifies the plurality of first audio signals. The system includes a correction processing unit that, when one of the multiple first tones is determined to be a trigger tone, applies correction processing to the first audio signal representing the determined trigger tone and one or more first audio signals representing one or more first tones classified in the same group as the trigger tone, based on the acquired classification information, among the multiple acquired first audio signals. Sound reproduction device.