Sound reproduction method, sound reproduction apparatus, and recording medium
By using anchor sound as a reference position for the second sound image in the sound reproduction device, combined with head-related transfer function and microphone-collected ambient sound, the problem of inaccurate sound image localization in virtual reality and extended reality is solved, improving sound quality and user immersion.
Patent Information
- Application Number
- CN202180020831.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-15
- Filing Date
- 2021-03-11
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2041-03-11
AI Technical Summary
In existing technologies for virtual reality and extended reality, sound localization struggles to accurately perceive height direction, and the use of marker sounds leads to a decline in sound quality and negatively impacts the user's immersive experience.
By using the anchor sound as the reference position for the second sound image in the sound reproduction device, and combining the head correlation transfer function and the microphone to collect surrounding sounds, the first sound image and the anchor sound are located, ensuring that the user can correctly perceive the sound image in relative positional relationship.
It improves the accuracy of audio-visual cues, suppresses the degradation of sound quality, reduces interference with the user's immersion, and enhances the accuracy of audio-visual perception in the height direction.
Smart Images

Figure CN115336290B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for sound reproduction, a device for sound reproduction, and a recording medium. Background Technology
[0002] Previously, there were known techniques for sound reproduction that allowed users to perceive stereo sound by placing sound images at desired locations in three-dimensional space (e.g., see Patent Document 1 and Non-Patent Document 1).
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2017-92732
[0006] Non-patent literature
[0007] Non-Patent Literature 1: Kazunori Ito, Yoshimichi Yonezawa, and Kenichi Kido, “Improvement of Display Efficiency Based on Audio-Visual Location Control for Hearing Information Transmission – Improvement of Display Efficiency Based on Marker Tone Addition,” Journal of the Japan Society for Audio Science, Vol. 42, No. 9 (1986), pp. 708-715. Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] The purpose of this disclosure is to provide a sound reproduction method, sound reproduction device, and recording medium for improving audio-visual cues.
[0010] Methods used to solve problems
[0011] A sound reproduction method according to a technical solution of this disclosure includes: a step of positioning a first sound image at a first position in an object space where the user is located; and a step of positioning a second sound image representing an anchor sound at a second position in the object space, wherein the anchor sound is used to represent a reference position.
[0012] The recording medium of one of the technical solutions of this disclosure is a non-transitory recording medium that can be read by a computer and contains a program for enabling a computer to execute the above-described sound reproduction method.
[0013] An audio reproduction device according to a technical solution of the present disclosure includes: a decoding unit that decodes an encoded audio signal that enables a user to perceive a first audio image; a first positioning unit that positions the first audio image to a first position in an object space where the user is located, according to the decoded audio signal; and a second positioning unit that positions a second audio image representing an anchor sound to a second position in the object space, wherein the anchor sound is used to represent a reference position.
[0014] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0015] Invention Effects
[0016] The sound reproduction method, recording medium, and sound reproduction device disclosed herein can improve audio-visual cues. Attached Figure Description
[0017] Figure 1 This is a block diagram showing a structural example of the sound reproduction device according to Embodiment 1.
[0018] Figure 2A This is an explanatory diagram schematically showing the object space of the sound reproduction device according to Embodiment 1.
[0019] Figure 2B This is a flowchart illustrating an example of an audio reproduction method of the audio reproduction apparatus according to Embodiment 1.
[0020] Figure 3 This is a block diagram showing a structural example of the sound reproduction device according to Embodiment 2.
[0021] Figure 4A This is a flowchart illustrating an example of an audio reproduction method of the audio reproduction device according to Embodiment 2.
[0022] Figure 4B This is a flowchart illustrating a processing example in which the second position is adaptively determined in the sound reproduction device of Embodiment 2.
[0023] Figure 5 This is a block diagram showing a modified example of the sound reproduction device according to Embodiment 2.
[0024] Figure 6 This is a diagram illustrating an example of the hardware structure of the sound reproduction device according to embodiments 1 and 2. Detailed Implementation
[0025] (Based on the understanding that forms the basis of this disclosure)
[0026] The inventors of this disclosure have discovered the following problems with the prior art described in the "Background Art" section.
[0027] Patent Document 1 proposes an auditory support system that assists a user's hearing by reproducing the three-dimensional sound environment observed by the user in object space. Based on the location of the sound source and facial posture, the auditory support system of Patent Document 1 uses a head-related transfer function from the location of the sound source to each ear of the user in object space to synthesize the sound signals used to reproduce the signals from the separated sounds to each ear of the user. Furthermore, the auditory support system adjusts the volume of each frequency band according to hearing attenuation characteristics. Thus, the auditory support system achieves auditory support without dissonance, and by separating the various sounds in the environment, it can selectively control the sounds needed and unwanted by the user.
[0028] However, according to Patent Document 1, there are the following problems. Although Patent Document 1 operates on frequency characteristics, it only uses the head-related transfer function for sound localization. Regarding the height direction, the user has difficulty accurately perceiving the sound image location. In other words, compared to the left-right direction based on the user's head or ears, it is difficult to accurately perceive the sound image in the vertical direction, i.e., the height direction.
[0029] Non-Patent Document 1, as a method for assisting the visually impaired, proposes a technique for transmitting images containing characters via auditory perception. The audio-visual display device of Non-Patent Document 1 establishes a correspondence between the position of synthesized sounds and the position of pixels, depicting the displayed image by scanning the space perceived by both ears with point-like sound images that vary over time. Furthermore, the audio-visual display device of Non-Patent Document 1 adds point-like sound images (called marker sounds) as indicators of the positions of the sound images that do not merge with the display points onto the display surface, thereby clarifying the relative positional relationship with the display points and improving the auditory-based positioning accuracy of the display points. The marker sounds use white noise with good additive effects and are set at the center position in the left-right direction.
[0030] However, according to Patent Document 1, there is the following problem. Since the marked sound is noise for the point image displayed as a point, it leads to a decrease in sound quality in applications such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), which hinders the user's immersion.
[0031] Therefore, this disclosure provides an audio reproduction method, an audio reproduction device, and a recording medium for improving audio-visual cues.
[0032] Therefore, a sound reproduction method of a technical solution of this disclosure includes the steps of: positioning a first sound image to a first position in an object space where the user is located; and positioning a second sound image representing an anchor sound to a second position in the object space, wherein the anchor sound is used to represent a reference position.
[0033] This improves the sonic image cues for the first sound. Specifically, the first sound image can be perceived based on the relative positional relationship between the second sound image and the first sound image, which serves as the anchor sound, so that the sonic image cues for the first sound can be correctly provided even when the first sound image is located in the vertical direction.
[0034] For example, the sound reproduction method may also involve using a portion of the surrounding sound or reproduced sound of the object space as the source of the anchor sound in the step of locating the second sound image.
[0035] Therefore, by using a portion of the surrounding sound or the space from which the sound is reproduced as the source of the anchor sound, it is possible to suppress the degradation of sound quality. For example, it is possible to prevent the anchor sound from hindering the user's immersion.
[0036] For example, the sound reproduction method may also include the step of using a microphone to acquire the surrounding sound of the user coming from the direction of the second position in the object space, and using the acquired sound as the source of the anchor sound in the step of locating the second sound image.
[0037] Therefore, by using a portion of the surrounding sound space as the source of the anchor sound, it is possible to suppress the degradation of sound quality. For example, it is possible to prevent the anchor sound from hindering the user's immersion.
[0038] For example, the sound reproduction method may also include: a step of acquiring ambient sound that comes to the user in the object space using a microphone; a step of selectively acquiring sound that meets a specified condition from the acquired ambient sound; and a step of determining the position of the selectively acquired sound in the direction of the second position.
[0039] This allows for greater freedom in selecting the sound source as the anchor sound, and enables adaptive setting of the second position.
[0040] For example, the above-mentioned conditions may relate to at least one of the following: the direction of sound arrival, the time of sound, the intensity of sound, the frequency of sound, and the type of sound.
[0041] Therefore, it is possible to select an appropriate sound as the source of the anchor sound.
[0042] For example, the aforementioned conditions may include an angle range as a condition for indicating the direction of sound arrival, wherein the angle range indicates a direction that does not include the user's vertical direction, includes the front direction, and includes the horizontal direction.
[0043] Therefore, it is possible to select a sound that is close to the direction that is perceived more correctly, i.e., the horizontal direction, as the anchor sound.
[0044] For example, the aforementioned conditions could include a specified intensity range as a condition for representing the intensity of the sound.
[0045] Therefore, it is possible to select a sound of appropriate intensity as the anchor sound.
[0046] For example, the aforementioned conditions could include a specific frequency range as a condition for representing the frequency of sound.
[0047] Therefore, it is possible to select an easily perceptible sound of an appropriate frequency as the anchor sound.
[0048] For example, the above-mentioned conditions may include human voice or special sounds as conditions for representing the category of sound.
[0049] Therefore, it is possible to select an appropriate sound as the anchor sound.
[0050] For example, in the step of locating the second sound image, the intensity of the anchor sound can be adjusted according to the intensity of the first sound source.
[0051] Therefore, the volume of the anchor sound can be adjusted relative to the first sound source.
[0052] For example, the angle of elevation or depression of the second position relative to the user may be smaller than the specified angle.
[0053] Therefore, it is possible to select a sound that is close to the direction that is perceived more correctly, i.e., the horizontal direction, as the anchor sound.
[0054] Furthermore, the recording medium of one of the technical solutions disclosed herein is a non-temporary recording medium that can be read by a computer and contains a program for enabling a computer to execute the aforementioned sound reproduction method.
[0055] This improves the sonic image cues for the first sound. Specifically, since the first sound image can be perceived in relation to the relative position of the second sound image and the first sound image, which serves as the anchor sound, the sonic image cues for the first sound can be correctly provided even when the first sound image is located in the vertical direction.
[0056] An audio reproduction device according to a technical solution of the present disclosure includes: a decoding unit that decodes an encoded audio signal that enables a user to perceive a first audio image; a first positioning unit that positions the first audio image to a first position in an object space where the user is located, according to the decoded audio signal; and a second positioning unit that positions a second audio image representing an anchor sound to a second position in the object space, wherein the anchor sound is used to represent a reference position.
[0057] This improves the sonic image cues for the first sound. Specifically, since the first sound image can be perceived in relation to the relative position of the second sound image and the first sound image, which serves as the anchor sound, the sonic image cues for the first sound can be correctly provided even when the first sound image is located in the vertical direction.
[0058] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0059] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.
[0060] Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, the arrangement and connection of constituent elements, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit this disclosure.
[0061] (Implementation Method 1)
[0062] [Definitions of the term]
[0063] First, the definitions of some technical terms appearing in this disclosure will be explained.
[0064] The "encoded audio signal" contains audio objects that allow the user to perceive an audio image. The encoded audio signal may, for example, be a signal based on the MPEG H Audio standard. This audio signal contains multiple audio channels and audio objects representing the first audio image. The multiple audio channels may include, for example, a maximum of 64 or 128 audio channels.
[0065] A "sound object" is data representing a virtual sound image that is perceived by the user. Hereinafter, it is assumed that a sound object contains the sound of a first sound image and data representing its first position. Furthermore, the "sound" of sound signals, sound objects, etc., is not limited to sounds made by humans or animals; any audible sound is acceptable.
[0066] "Localization of sound image" refers to the process of convolving the head correlation transfer function (HRTF) corresponding to the left ear and the HRTF corresponding to the right ear with the sound signal in the user's object space, so that the user can perceive the sound image at a virtual location.
[0067] "Binaural signal" refers to the signal obtained by convolving the sound signal, which is the source of the sound image, with the HRTF corresponding to the left ear and the HRTF corresponding to the right ear, respectively.
[0068] "Object space" refers to the virtual or real three-dimensional space in which the user resides. For example, object space is the three-dimensional space perceived by the user in virtual reality (VR), extended reality (AR), and mixed reality (MR).
[0069] An "anchor sound" refers to a sound in object space that originates from a sound image used by the user to perceive a reference position. Hereinafter, the sound image emitting the anchor sound will be referred to as the second sound image. The second sound image, as the anchor sound, allows the first sound image to be perceived in relative position, thus enabling the user to more accurately perceive the position of the first sound image even when it is located in the vertical direction.
[0070] [structure]
[0071] Next, the structure of the sound reproduction device 100 of Embodiment 1 will be described. Figure 1 This is a block diagram showing a structural example of the sound reproduction device 100 according to Embodiment 1. Furthermore, Figure 2A This is an explanatory diagram schematically showing the object space 200 of the sound reproduction device 100 according to Embodiment 1. Figure 2A In the diagram, the front of the user's face is set as the Z-axis, the top as the Y-axis, and the right as the X-axis.
[0072] exist Figure 1 In this device, the sound reproduction apparatus 100 includes a decoding unit 101, a first positioning unit 102, a second positioning unit 103, a position estimation unit 104, an anchor direction estimation unit 105, an anchor sound generation unit 106, a mixer 107, and a head-mounted device 110. The head-mounted device 110 includes headphones 111, a head sensor 112, and a microphone 113. Additionally, in... Figure 1 The head of the user is schematically depicted within the head-mounted device 110.
[0073] The decoding unit 101 decodes the encoded audio signal. The encoded audio signal may be, for example, a signal based on the MPEG-H Audio standard.
[0074] The first positioning unit 102 positions the first sound image at a first position in the object space where the user is located, based on the position of the sound object contained in the decoded sound signal, the relative position of the user 99, and the direction of their head. A first binaural signal is output from the first positioning unit 102, positioning the first sound image at the first position. Figure 2A The diagram schematically illustrates the positioning of the first sound image 201 within the object space 200 where the user 99 is located. The first sound image 201 is defined at any position within the object space 200 by the sound object. As the first sound image 201... Figure 2A When positioned vertically (along the Y-axis), the user (99) has difficulty accurately perceiving the location compared to when positioned horizontally (along the X and Z axes). In particular, if the HRTF is not the user's own HRTF or if the headphone characteristics are not properly corrected, the user (99) cannot accurately perceive the location of the first sound image.
[0075] The second positioning unit 103 positions the second acoustic image, representing the anchor sound, at a second position in the object space, where the anchor sound represents the reference position. A second binaural signal is output from the second positioning unit 103, positioning the second acoustic image at the second position. At this time, the second positioning unit 103 controls the volume and frequency band of the second sound source to make it appropriate relative to the first sound source and other reproduced sounds. For example, it can control the frequency response of the second sound source to be flattened by reducing the peaks or valleys, or it can control the high-frequency domain of the signal to be emphasized. Figure 2A The diagram schematically illustrates the positioning of the second sound image 202 within the object space 200 where the user 99 is located. This second position can be a pre-set, fixed position or an adaptive position based on ambient sound or reproduced sound. For example, the second position could be a pre-set position of the user's face in the initial state (i.e., along the Z-axis), or it could be... Figure 2A That is, a pre-set position on the right side relative to the front of the user's face 99. Since the second sound image 202 is positioned, for example, in a direction close to the horizontal direction, that is, within a predetermined angle range from the horizontal direction, the anchor sound is perceived by the user 99 more accurately. The anchor sound allows the first sound image to be perceived in relative position, so even when the first sound image is located in the vertical direction, the user 99 can perceive the position of the first sound image more accurately. In addition, the positioning of the first sound image and the positioning of the second sound image can be simultaneous or not simultaneous. When they are not simultaneous, the shorter the time interval between the positioning of the first sound image and the positioning of the second sound image, the easier it is to perceive more accurately.
[0076] The position estimation unit 104 obtains the orientation information output from the head sensor 112 and estimates the direction of the user 99's head, i.e. the direction of the face.
[0077] The anchor direction estimation unit 105 estimates a new anchor direction, i.e., the direction of the new second position, based on the direction estimated by the position estimation unit 104, as the user 99 moves. The estimated direction of the second position is then communicated to the anchor sound generation unit 106.
[0078] In addition, the anchor direction can be fixed based on the object space, or it can be fixed according to the environment.
[0079] The anchor sound generation unit 106 selectively acquires sounds arriving from the direction of a new anchor sound predicted by the anchor direction estimation unit 105, from the omnidirectional ambient sounds collected by the microphone 113. Furthermore, the anchor sound generation unit 106 uses the selectively acquired sounds as the source of the anchor sound, adjusting the intensity (volume) and frequency characteristics to generate an appropriate anchor sound. The intensity and frequency characteristics of the anchor sound can also be adjusted based on the sound of the first sound image.
[0080] The mixer 107 mixes the first binaural signal from the first positioning unit 102 with the second binaural signal from the second positioning unit 103. The mixed sound signal includes a signal for the left ear and a signal for the right ear, and is output to the headphones 111.
[0081] The earphone 111 has a speaker for the left ear and a speaker for the right ear. The speaker for the left ear converts the signal for the left ear into sound, and the speaker for the right ear converts the signal for the right ear into sound. The earphone 111 can also be an earbud type that is inserted into the outer ear.
[0082] The head sensor 112 detects the direction in which the user 99's head is facing, i.e., the direction of their face, and outputs this as orientation information. The head sensor 112 can also be a sensor that detects the 6DOF (Degrees of Freedom) information of the user 99's head. The head sensor 112 can, for example, be composed of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetometer, or a combination thereof.
[0083] Microphone 113 collects ambient sound arriving at the user 99 in the object space and converts it into electrical signals. Microphone 113 may have, for example, a left microphone and a right microphone. The left microphone may be positioned near the left ear speaker, and the right microphone near the right ear speaker. Furthermore, microphone 113 may be a directional microphone capable of arbitrarily specifying the direction of sound collection, or it may have three microphones. Additionally, microphone 113 may collect sound reproduced by the earphone 111, either in addition to ambient sound or other ambient sounds, and convert it into electrical signals. When locating the second sound image, the second positioning unit 103 may use a portion of the reproduced sound as the source of the anchor sound, instead of ambient sound arriving at the user from the direction of the second position in the object space.
[0084] Furthermore, the head-mounted device 110 and the main body of the sound reproduction device 100 can be either separate or integrated. When the head-mounted device 110 and the main body of the sound reproduction device 100 are integrated, the head-mounted device 110 and the sound reproduction device 100 can also be wirelessly connected.
[0085] [action]
[0086] Next, a general description of the operation of the sound reproduction device 100 according to Embodiment 1 will be given.
[0087] Figure 2B This is a flowchart illustrating an example of the sound reproduction method of the sound reproduction apparatus 100 according to Embodiment 1. As shown in the figure, the sound reproduction apparatus 100 first decodes the encoded sound signal of the first sound image perceived by the user (S21). Next, the sound reproduction apparatus 100 positions the first sound image at a first position in the object space where the user is located, according to the decoded sound signal (S22). Specifically, the sound reproduction apparatus 100 generates a first binaural signal by convolving the HRTF corresponding to the left ear and the HRTF corresponding to the right ear with the sound signal of the first sound image, respectively. Furthermore, the sound reproduction apparatus 100 positions a second sound image, representing an anchor sound used to represent a reference position, at a second position in the object space (S23). Specifically, the sound reproduction apparatus 100 generates a second binaural signal by convolving the HRTF corresponding to the left ear and the HRTF corresponding to the right ear with the sound source signal of the anchor sound of the second sound image, respectively. The sound reproduction apparatus 100 periodically repeats steps S21 to S23. Alternatively, the audio reproduction device 100 may continuously decode the audio signal as a bitstream (S21) while periodically repeating steps S22 and S23.
[0088] By reproducing the first binaural signal used to locate the first sound image and the second binaural signal used to locate the second sound image through the headphones 111, the user 99 perceives the first sound image and the second sound image. At this time, the user 99 perceives the first sound image in relative position relative to the anchor sound from the second sound image, so even when the first sound image is located in the height direction, the position of the first sound image can be perceived more accurately.
[0089] Additionally, the source of the anchor sound from the second sound image can be a directional portion of the surrounding sound or a directional portion of the reproduced sound, but is not limited to these. It can also be a preset sound that does not feel discordant with the surrounding sound or the reproduced sound.
[0090] (Implementation Method 2)
[0091] Next, the sound reproduction device 100 of Embodiment 2 will be described.
[0092] In Embodiment 2, an example is described using a directional portion of the surrounding sounds arriving at the user in the object space as the source of the anchor sound. For example, the sound reproduction device 100 uses a microphone to acquire the surrounding sounds arriving at the user in the object space, selectively acquires sounds that meet predetermined conditions from the acquired surrounding sounds, and uses the selectively selected sounds as the source of the anchor sound in the step of locating the second sound image. As a result, the user can more accurately perceive the position of the first sound image based on its relative positional relationship with the anchor sound. Furthermore, since the anchor sound is part of the surrounding sounds, even if the user hears the anchor sound, there is almost no sense of dissonance. In this way, it is easy to suppress the anchor sound from hindering the user's immersion.
[0093] [structure]
[0094] Figure 3 This is a block diagram showing a structural example of the sound reproduction device according to Embodiment 2. The sound reproduction device 100 in this diagram is... Figure 1 In comparison, it differs in that it adds an ambient sound acquisition unit 301, a directional control unit 302, a first direction acquisition unit 303, an anchor direction estimation unit 304, and a first volume acquisition unit 305, and in that it has an anchor sound generation unit 106a instead of an anchor sound generation unit 106. Hereinafter, the differences will be explained in detail.
[0095] The ambient sound acquisition unit 301 acquires the ambient sound collected by the microphone 113. Figure 3The microphone 113 not only collects ambient sound from all directions, but also has directional sound collection capability under the control of the directional control unit 302. Here, it is assumed that the ambient sound acquisition unit 301 acquires ambient sound in the direction in which the second sound image should be located via the microphone 113.
[0096] The directional control unit 302 controls the directionality of sound collection by the microphone 113. Specifically, the directional control unit 302 controls the microphone 113 to be directional in the new anchor direction predicted by the anchor direction prediction unit 304. As a result, the sound collected by the microphone 113 is ambient sound coming from the direction of the new anchor direction, i.e., the new second position, predicted as the user 99 moves.
[0097] The first direction acquisition unit 303 acquires the direction and first position of the first sound image from the sound object decoded by the decoding unit 101.
[0098] The anchor direction estimation unit 304 estimates a new anchor direction, i.e. the direction of the new second position, based on the direction of the user 99’s face as estimated by the position estimation unit 104 and the direction of the first sound image obtained by the first direction acquisition unit 303, as the user 99 moves.
[0099] The first volume acquisition unit 305 acquires the first volume as the volume of the first sound image from the sound object decoded by the decoding unit 101.
[0100] The anchor sound generation unit 106a generates anchor sound using the ambient sound acquired by the ambient sound acquisition unit 301 as the sound source.
[0101] [action]
[0102] Next, the operation of the sound reproduction device 100 of Embodiment 2 will be explained.
[0103] Figure 4A This is a flowchart illustrating an example of an audio reproduction method of the audio reproduction apparatus 100 according to Embodiment 2. Figure 4A and Figure 2B The difference lies in the addition of steps S43 to S44. The following explanation will focus on this difference.
[0104] After locating the first sound image in step S22, the sound reproduction device 100 detects the orientation of the user 99's face (S43). The detection of the face orientation is performed by the head sensor 112 and the position estimation unit 104.
[0105] Furthermore, the sound reproduction device 100 infers the anchor direction based on the detected orientation of the face (S44). The inference of the anchor direction is performed by the anchor direction inference unit 304. That is, when there is movement of the user 99's head, the anchor direction inference unit 304 infers a new anchor direction, that is, the direction of the new second position. When there is no movement of the user 99's head, the direction that is the same as the current anchor direction is inferred as the new anchor direction.
[0106] Next, the sound reproduction device 100 generates anchor sound using ambient sound from the predicted anchor direction as the sound source (S45). The acquisition of ambient sound from the predicted anchor direction is performed by the directional control unit 302, the microphone 113, and the ambient sound acquisition unit 301. The process of generating anchor sound using this ambient sound as the sound source is performed by the anchor sound generation unit 106a.
[0107] Then, the sound reproduction device 100 positions the second sound image representing the anchor sound at the second position in the inferred anchor direction (S23).
[0108] according to Figure 4A The sound reproduction device 100 is able to position the second sound image by following the movement of the user's head 99.
[0109] Furthermore, the second position, which serves as the location of the second sound image, can be a pre-set position, or it can be adaptively determined based on the surrounding sound. Next, an example of processing that adaptively determines the second position based on the surrounding sound will be described.
[0110] Figure 4B This is a flowchart illustrating a processing example of adaptively determining the second position in the sound reproduction apparatus of Embodiment 2. The sound reproduction apparatus 100, for example, in... Figure 4A Execution before processing begins Figure 4B The processing, and then with Figure 4A The processing is executed repeatedly in parallel. Figure 4B In this process, the sound reproduction device 100 first uses a microphone to acquire the ambient sound arriving at the user 99 in the target space (S46). The ambient sound acquired can be omnidirectional or encompass the entire circumference, including the horizontal angular range. Then, the sound reproduction device 100 searches for directions that satisfy predetermined conditions from the acquired ambient sound (S47). For example, the sound reproduction device 100 selectively acquires sounds that satisfy predetermined conditions from the acquired ambient sound and determines the direction of arrival of those sounds as the direction that satisfies the predetermined conditions. Furthermore, the sound reproduction device 100 determines a second position so that the second position exists in the direction of the search result (S48).
[0111] Here, the specified conditions are explained. The specified conditions relate to at least one of the following: the direction of arrival of the sound, the timing of the sound, the intensity of the sound, the frequency of the sound, and the type of sound.
[0112] For example, the specified conditions include an angle range as a condition for indicating the direction of sound arrival. This angle range includes a direction that does not include the user's vertical direction, includes the front direction, and includes the horizontal direction. Therefore, it is possible to select a sound from a direction that is close to the direction perceived more accurately, i.e., the horizontal direction, as the anchor sound.
[0113] Furthermore, the specified conditions can also include a specified intensity range as a condition representing the intensity of the sound. This allows for the selection of a sound of appropriate intensity as the anchor sound.
[0114] Furthermore, the specified conditions can also include a specific frequency range as a condition for representing the frequency of the sound. This allows for the selection of an easily perceptible sound of an appropriate frequency as the anchor sound.
[0115] Furthermore, the specified conditions can also include human voices or special sounds as conditions representing the category of sound. Thus, a sound can select an appropriate sound as its anchor.
[0116] Furthermore, the specified conditions can also include a duration of at least a specified time for representing the time of the sound. This allows for the selection of sounds with temporal characteristics as anchor sounds. By ensuring that the source of the anchor sound meets the specified conditions, it is possible to generate appropriate anchor sounds that do not cause any sense of disharmony for the user.
[0117] according to Figure 4B The second position of the second sound image can be adaptively determined based on the surrounding sound. In addition, the anchor sound can use a portion of the directionality of the surrounding sound as a sound source.
[0118] Alternatively, the sound reproduction device 100 in each of the above embodiments can also replace the head-mounted device 110 and include an HMD (Head Mounted Display). In this case, the HMD can also include a display unit in addition to the headphones 111, head sensor 112, and microphone 113. Furthermore, it is also possible to configure the sound reproduction device 100 to be built into the HMD body.
[0119] Furthermore, in implementation method 2 Figure 3 The sound reproduction device can also be modified as follows. Figure 5 This is a block diagram illustrating a modified example of the sound reproduction device 100 according to Embodiment 2. In this modified example, a configuration example is shown where reproduced sound is used instead of ambient sound. Figure 5 The sound reproduction device 100 and Figure 3In contrast, it has a sound reproduction unit 401 instead of an ambient sound acquisition unit 301.
[0120] The reproduced sound acquisition unit 401 acquires the reproduced sound decoded by the decoding unit 101. The anchor sound generation unit 106a uses the reproduced sound acquired by the reproduced sound acquisition unit 401 as the sound source to generate an anchor sound. For example, Figure 5 The sound reproduction device 100 reproduces a sound signal including a first sound source and other sound channels. From the reproduced sounds contained in the reproduced sound signal, it selectively extracts sounds that meet predetermined conditions and uses these selectively extracted sounds as the sound source of the anchor sound. As a result, the user can more accurately perceive the position of the first sound image based on its relative position to the anchor sound. Furthermore, since the anchor sound is part of the reproduced sound, the user experiences almost no dissonance even when hearing the anchor sound. This makes it easy to suppress the anchor sound from interfering with the user's immersion.
[0121] (Other implementation methods)
[0122] The above description, based on embodiments, outlines the sound reproduction apparatus and method relating to the technical solutions of this disclosure. However, this disclosure is not limited to these embodiments. For example, other embodiments implemented by arbitrarily combining or removing certain constituent elements described in this specification may also be considered embodiments of this disclosure. Furthermore, variations derived from the above embodiments by applying various modifications conceived by those skilled in the art without departing from the spirit of this disclosure, i.e., the meaning of the language expressed in the claims, are also included in this disclosure.
[0123] Furthermore, the forms shown below may also be included within the scope of one or more technical solutions disclosed herein.
[0124] (1) A component of the aforementioned sound reproduction device may also be a computer system consisting of a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. The RAM or hard disk unit stores a computer program. Each device performs its function by operating the microprocessor according to the computer program. Here, the computer program is composed of multiple command codes representing instructions for a computer, designed to achieve a specified function.
[0125] Such an audio reproduction device 100 can also be, for example, an input... Figure 6 The hardware structure is shown. In Figure 6The audio playback device 100 includes an I / O unit 11, a display control unit 12, a memory 13, a processor 14, headphones 111, a head sensor 112, a microphone 113, and a display unit 114. A portion of the components constituting the audio playback device 100 of embodiments 1 to 3 achieve their functions by the processor 14 executing a program stored in the memory 13. The hardware structure shown in FIG7 could also be, for example, an HMD, a combination of a head-mounted device 110 and a tablet computer terminal, a combination of a head-mounted device 110 and a smartphone, or a combination of a head-mounted device 110 and an information processing device (e.g., a PC, a television).
[0126] (2) A portion of the constituent elements of the aforementioned sound reproduction device and sound reproduction method may also be constituted by a single system LSI (Large Scale Integration). A system LSI is a multifunctional LSI manufactured by integrating multiple components onto a single chip; specifically, it is a computer system comprising a microprocessor, ROM, RAM, etc. The computer program is stored in the RAM. The system LSI performs its function by having the microprocessor operate according to the computer program.
[0127] (3) A component of the aforementioned audio reproduction device may also be a module consisting of a detachable IC card or a single unit relative to each device. The aforementioned IC card or module is a computer system composed of a microprocessor, ROM, RAM, etc. The aforementioned IC card or module may also include the aforementioned multi-functional LSI. The IC card or module performs its function by operating according to a computer program via a microprocessor. The IC card or module may also be tamper-resistant.
[0128] (4) Furthermore, a component of the aforementioned sound reproduction device may also be a recording medium capable of being read by a computer, such as a floppy disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray Disc), semiconductor memory, etc. Alternatively, it may be a digital signal recorded on these recording media.
[0129] Furthermore, the computer program or digital signal that constitutes part of the aforementioned sound reproduction device can also be transmitted via electrical communication lines, wireless or wired communication lines, networks such as the Internet, data broadcasting, etc.
[0130] (5) This disclosure may also be the method shown above. In addition, it may be a computer program that implements these methods by a computer, or a digital signal composed of the computer program described above.
[0131] (6) In addition, this disclosure may also be a computer system having a microprocessor and a memory, wherein the memory stores the computer program and the microprocessor operates according to the computer program.
[0132] (7) Alternatively, the above-mentioned program or digital signal may be recorded in the above-mentioned recording medium and transferred, or the above-mentioned program or digital signal may be transferred via the above-mentioned network, etc., by an independent computer system.
[0133] (8) The above-described embodiments and the above-described variations can also be combined separately.
[0134] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by a microprocessor executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0135] Furthermore, this disclosure is not limited to the embodiments. Various modifications that can be conceived by those skilled in the art, or forms constructed by combining the constituent elements of different embodiments, may be included within the scope of one or more technical solutions, as long as they do not depart from the spirit of this disclosure.
[0136] Industrial usability
[0137] This disclosure can be used with sound reproduction devices and sound reproduction methods, such as stereo sound reproduction devices.
[0138] Label Explanation
[0139] 10 Ministry of Communications
[0140] 11 I / O Department
[0141] 12 Display Control Unit
[0142] 13 Memory
[0143] 14 processors
[0144] 99 users
[0145] 100 Sound Reproduction Device
[0146] 101 Decoding Department
[0147] 102 First Positioning Section
[0148] 103 Second Positioning Section
[0149] 104 Location Prediction Department
[0150] Anchor Direction Prediction Section 105, 304
[0151] Anchor sound generation units 106, 106a, and 106b
[0152] 107 Mixer
[0153] 110 Headset
[0154] 111 Headphones
[0155] 112 Head Sensor
[0156] 113 Microphone
[0157] 114 Display Section
[0158] 200 object space
[0159] 201 First Audio-Visual
[0160] 202 Second Audio-Visual
[0161] 301 Ambient Sound Acquisition Department
[0162] 302 Directional Control Unit
[0163] 303 First Direction Acquisition Section
[0164] 305 1st volume acquisition unit
[0165] 401 Sound Reproduction Acquisition Department
[0166] 402 Sound Source Investigation Department
[0167] 403 Sound Source Direction Acquisition Department
Claims
1. A sound reproduction method, wherein, comprising: a step of positioning a first sound image to a first position in an object space in which a user is present; and a step of positioning a second sound image indicating an anchor sound to a second position in the object space, the anchor sound being used to indicate a reference position, the anchor sound being used to perceive the first sound image in a relative positional relationship with the reference position as a reference.
2. The sound reproduction method according to claim 1, wherein in the step of positioning the second sound image, a part of a surrounding sound or a reproduced sound of the object space is used as a sound source of the anchor sound.
3. The sound reproduction method according to claim 1, wherein further comprising a step of acquiring, using a microphone, a surrounding sound in the object space coming from a direction of the second position to the user, the acquired sound being used as the sound source of the anchor sound in the step of positioning the second sound image.
4. The sound reproduction method according to claim 1, wherein further comprising: a step of acquiring, using a microphone, a surrounding sound in the object space coming to the user; a step of selectively acquiring a sound satisfying a prescribed condition from the acquired surrounding sound; and a step of determining a position in a direction of the selectively acquired sound as the second position.
5. The sound reproduction method according to claim 4, wherein the prescribed condition relates to at least one of a direction of arrival of the sound, a time of the sound, an intensity of the sound, a frequency of the sound, and a category of the sound.
6. The sound reproduction method according to claim 4, wherein the prescribed condition includes an angle range as a condition indicating a direction of arrival of the sound, the angle range indicating a direction not including a vertical direction of the user, including a front direction, and including a horizontal direction.
7. The sound reproduction method according to claim 4, wherein the prescribed condition includes a prescribed intensity range as a condition indicating an intensity of the sound.
8. The sound reproduction method according to claim 4, wherein the prescribed condition includes a prescribed frequency range as a condition indicating a frequency of the sound.
9. The sound reproduction method according to claim 4, wherein the prescribed condition includes a human voice or a special sound as a condition indicating a category of the sound.
10. The sound reproduction method according to claim 1, wherein in the step of positioning the second sound image, an intensity of the anchor sound is adjusted in accordance with an intensity of a first sound source.
11. The sound reproduction method according to claim 1, wherein the second position is determined in accordance with an orientation of a face of the user.
12. The sound reproduction method according to claim 1, wherein the second position is changed based on a change in the orientation of the face of the user.
13. The sound reproduction method according to any one of claims 1 to 12, wherein the second position is smaller than a prescribed angle with respect to a pitch angle or a roll angle of the user.
14. The sound reproduction method according to claim 1, wherein the anchor sound is positioned at the second position closer to a horizontal direction than the first position.
15. A recording medium having recorded thereon a program for causing a computer to execute the acoustic reproduction method according to claim 1.
16. A sound reproducing apparatus, wherein, provided with: a decoding section for decoding an encoded sound signal for making a user perceive a first sound image; a first positioning section for positioning the first sound image to a first position in an object space where the user is present, based on the decoded sound signal; and a second positioning section for positioning a second sound image representing an anchor sound to a second position in the object space, the anchor sound being used to represent a reference position, the anchor sound being used to perceive the first sound image in a relative positional relationship with reference to the reference position.
Citation Information
Patent Citations
Auditory supporting system and auditory supporting device
JP2017092732A
Dynamic focus for audio augmented reality (AR)
US10506362B1
System and method for user controllable auditory environment customization
US20150195641A1
System and Method For User Alerts During An Immersive Computer-Generated Reality Experience
US20190392830A1