Audio signal processing method, program, and audio signal processing device
The method analyzes and adjusts sound characteristics to improve discriminability of three-dimensional sound perception by addressing overlapping sounds, ensuring clear audio experience in virtual reality systems.
Patent Information
- Application Number
- JP2025071529
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-31
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-03
Smart Images

Figure 2025100877000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio signal processing device, as well as an audio signal processing method and program related to the audio signal processing device.
Background Art
[0002] Conventionally, there has been known a technique related to acoustic reproduction for causing a user to perceive three-dimensional sound by controlling the position of a sound image, which is a virtual sound source object, in a virtual three-dimensional space (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the other hand, when causing a user to perceive sound as three-dimensional sound in a three-dimensional sound field, there may be a case where sound that is difficult for the user to perceive is generated. In the information processing method in a conventional acoustic reproduction device or the like, there may be a case where appropriate processing is not performed on such difficult-to-perceive sound.
[0005] In view of the above, an object of the present disclosure is to provide an information processing method or the like that more appropriately causes a user to perceive three-dimensional sound.
Means for Solving the Problems
[0006] An information processing method according to an aspect of the present disclosure is an information processing method for generating an output sound signal for causing a user to perceive a predetermined sound as a sound arriving from an arrival direction on a three-dimensional sound field corresponding to the predetermined direction from sound information including information about the predetermined sound and information about the predetermined direction, the method including: a first analysis step of analyzing a type of the predetermined sound; a second analysis step of analyzing a type of external sound that is listened to by the user as the external sound; a third analysis step of analyzing an arrival direction of the external sound; a first determination step of determining whether or not the type of the predetermined sound and the type of the external sound match by comparing the analyzed type of the predetermined sound and the analyzed type of the external sound; a second determination step of determining whether or not the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound overlap by comparing the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound; and an adjustment step of performing at least one of (a) adjusting a sound pressure of at least one of the predetermined sound and the external sound and (b) adjusting an arrival direction of the predetermined sound based on determination results of the first determination step and the second determination step.
[0007] In addition, an acoustic playback device according to one aspect of the present disclosure is an acoustic playback device that generates and plays an output sound signal for causing a user to perceive a predetermined sound as a sound arriving from an arrival direction on a three-dimensional sound field corresponding to the predetermined direction from sound information including information about the predetermined sound and information about the predetermined direction, the acoustic playback device including: an acquisition unit that acquires the sound information; a first analysis unit that analyzes the type of the predetermined sound; a second analysis unit that analyzes the type of external sound that is listened to by the user as external sound from outside; a third analysis unit that analyzes the arrival direction of the external sound; a first determination unit that determines whether or not the type of the predetermined sound and the type of the external sound match by comparing the analyzed type of the predetermined sound and the analyzed type of the external sound; a second determination unit that determines whether or not the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound overlap by comparing the arrival direction of the predetermined sound and the analyzed arrival direction of the external sound; an adjustment unit that performs at least one of (a) adjusting the sound pressure of at least one of the predetermined sound and the external sound and (b) adjusting the arrival direction of the predetermined sound based on the determination results of the first determination step and the second determination step; and an output unit that outputs sound by the output sound signal generated by the adjustment.
[0008] Also, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the acoustic playback method described above.
[0009] Note that these general or specific aspects may be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
Advantages of the Invention
[0010] According to the present disclosure, it is possible to more appropriately cause a user to perceive three-dimensional sound.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0012] (Knowledge underlying the disclosure) Conventionally, there has been known a technique related to acoustic reproduction for causing a user to perceive three-dimensional sound by controlling the position of a sound image, which is a sound source object in the user's sense, in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) (see, for example, Patent Document 1). By localizing the sound image at a predetermined position in the virtual three-dimensional space, the user can perceive this sound as if it is a sound arriving from a direction parallel to the straight line connecting the predetermined position and the user (i.e., a predetermined direction). To localize the sound image at a predetermined position in the virtual three-dimensional space in this way, for example, calculation processing for generating an interaural time difference of sound and an interaural level difference (or sound pressure difference) of sound, which are perceived as three-dimensional sound, for the recorded sound is required.
[0013] As an example of such calculation processing, a process of convolving a head-related transfer function for causing a sound to be perceived as a sound arriving from a predetermined direction with a signal of the target sound is known. By performing the convolution process of this head-related transfer function with higher resolution, the sense of presence experienced by the user is improved. On the other hand, in such a sound listening environment, there is known a phenomenon in which the sound becomes difficult to hear due to the overlap of external sounds arriving from the outside and being listened to by the user 99. In particular, in a situation where there is a reproduced predetermined sound and an external sound of the same type and arriving from the same direction, it may be difficult to distinguish either the predetermined sound or the external sound.
[0014] In recent years, the development of technologies related to virtual reality (VR) has been actively carried out. In virtual reality, the main focus is that the position in a virtual three-dimensional space does not follow the user's movement, and the user can feel as if they are moving within the virtual space. In particular, in this virtual reality technology, attempts have been made to enhance the sense of presence by incorporating auditory elements into visual elements. For example, when a sound image is localized in front of the user, if the user turns to the right, the sound image moves to the left direction of the user, and if the user turns to the left, the sound image moves to the right direction of the user. Thus, it is necessary to move the localization position of the sound image in the virtual space in the direction opposite to the user's movement with respect to the user's movement. Such processing is performed by applying a stereophonic filter to the original sound information.
[0015] In view of the above, in the present disclosure, while using a stereophonic filter for making a user perceive sound from a predetermined direction within a three-dimensional sound field, more appropriate calculation processing is performed to improve the discriminability when a reproduced predetermined sound and an external sound arriving from the outside overlap. The purpose of the present disclosure is to provide an information processing method or the like for making a user perceive a three-dimensional sound by such appropriate calculation processing.
[0016] More specifically, an information processing method according to one aspect of the present disclosure is an information processing method for generating an output sound signal for causing a user to perceive a predetermined sound as a sound arriving from an arrival direction on a spherical sound field corresponding to a predetermined direction from sound information including information about a predetermined sound and information about a predetermined direction, the method including: a first analysis step of analyzing the type of the predetermined sound; a second analysis step of analyzing the type of external sound that is listened to by the user as external sound from outside; a third analysis step of analyzing the arrival direction of the external sound; a first determination step of determining whether or not the type of the predetermined sound and the type of the external sound match by comparing the analyzed type of the predetermined sound and the analyzed type of the external sound; a second determination step of determining whether or not the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound overlap by comparing the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound; and an adjustment step of performing at least one of (a) adjusting the sound pressure of at least one of the predetermined sound and the external sound, and (b) adjusting the arrival direction of the predetermined sound, based on the determination results of the first determination step and the second determination step.
[0017] According to such an information processing method, when at least one of the arrival directions of the external sound and the predetermined sound overlapping and the types of the external sound and the predetermined sound matching causes the external sound and the predetermined sound to affect each other and makes it difficult for the user to listen to either of them, by performing at least one of the adjustments (a) and (b), it is possible to facilitate the listening to at least one of the external sound and the predetermined sound and allow the user to more appropriately perceive a three-dimensional sound.
[0018] Also, for example, in the adjustment step, when it is determined in the determination in the first determination step that the type of the predetermined sound and the type of the external sound match, and it is determined in the determination in the second determination step that the arrival direction of the predetermined sound and the arrival direction of the external sound overlap, at least one of (a) and (b) may be performed.
[0019] According to this, when the arrival directions of the external sound and the predetermined sound overlap and the types of the external sound and the predetermined sound match, so that the external sound and the predetermined sound affect each other and it becomes difficult for the user to listen to either of them, by performing at least one of the adjustments (a) and (b), it is possible to facilitate the listening of at least one of the external sound and the predetermined sound and enable the user to more appropriately perceive a three-dimensional sound.
[0020] Also, for example, in the adjustment step, as (a), the sound pressure of the external sound may be attenuated by generating and superimposing a superimposed sound having a phase opposite to that of the external sound.
[0021] According to this, by superimposing the superimposed sound on the external sound and letting the user listen to it, the sound pressure of the external sound can be attenuated and the user can more appropriately perceive the predetermined sound.
[0022] Also, for example, in the adjustment step, as (b), the arrival direction of the predetermined sound may be varied by a preset angle in a direction away from the arrival direction of the external sound.
[0023] According to this, by preventing the arrival directions of the predetermined sound and the external sound from overlapping, it is possible to facilitate the listening of at least one of the external sound and the predetermined sound and enable the user to more appropriately perceive a three-dimensional sound.
[0024] Also, for example, in the adjustment step, as (b), information regarding the predetermined direction may be corrected so that the arrival direction of the predetermined sound is varied by a preset angle in a direction away from the arrival direction of the external sound.
[0025] According to this, the arrival directions of the predetermined sound and the external sound do not overlap, facilitating the listening of at least one of the external sound and the predetermined sound, and enabling the user to more appropriately perceive a three-dimensional sound. For this purpose, by correcting the information regarding the predetermined direction included in the sound information, the subsequent selected stereophonic filter can be made into a stereophonic filter for preventing the arrival directions of the predetermined sound and the external sound from overlapping. As a result, the listening of at least one of the external sound and the predetermined sound can be facilitated, and the user can more appropriately perceive a three-dimensional sound.
[0026] Also, for example, in the analysis of the type of the predetermined sound and the analysis of the type of the external sound, the sound to be analyzed is divided for each unit time in the time domain, and the divided sound is input into a machine learning model, thereby calculating the likelihood for each of a plurality of preset types, and outputting an analysis result indicating that the type of the input sound corresponds to the type with the highest calculated likelihood.
[0027] According to this, using the machine learning model, it is possible to output, as an analysis result, that the sound to be analyzed corresponds to the type with the highest likelihood among a plurality of preset types.
[0028] Also, for example, the types of the predetermined sound and the external sound may consist of two types: voice and non-voice.
[0029] According to this, based on whether the types of the external sound and the predetermined sound are either of the two types: voice and non-voice, it is possible to determine whether the types of the external sound and the predetermined sound match.
[0030] Also, for example, the determination of whether the arrival directions of the predetermined sound and the external sound overlap is performed based on whether the angular difference between the arrival direction of the predetermined sound and the arrival direction of the external sound is smaller than a threshold value. For the virtual boundary surface that divides the user's head into front and back, the first threshold value, which is the threshold value when the arrival directions of the predetermined sound and the external sound are on the rear side of the boundary surface, may be larger than the second threshold value, which is the threshold value when the arrival directions of the predetermined sound and the external sound are on the front side of the boundary surface.
[0031] According to this, on the rear side of the boundary surface where it is easy to perceive that the arrival directions overlap due to the large minimum discrimination angle of the arrival direction, it is possible to determine whether the arrival directions of the external sound and the predetermined sound overlap based on a criterion that is larger than that on the front side of the boundary surface.
[0032] Also, a program according to one aspect of the present disclosure is a program for causing a computer to execute the information processing method described above.
[0033] According to this, the same effects as the information processing method described above can be achieved using a computer.
[0034] Also, an acoustic playback device according to one aspect of the present disclosure is an acoustic playback device that generates and plays an output sound signal for causing a user to perceive a predetermined sound as a sound arriving from an arrival direction on a third-order sound field corresponding to a predetermined direction from sound information including information regarding the predetermined sound and information regarding the predetermined direction, the acoustic playback device including: an acquisition unit that acquires the sound information; a first analysis unit that analyzes the type of the predetermined sound; a second analysis unit that analyzes the type of an external sound that is listened to by the user as an external sound; a third analysis unit that analyzes the arrival direction of the external sound; a first determination unit that determines whether the type of the predetermined sound and the type of the external sound match by comparing the analyzed type of the predetermined sound and the analyzed type of the external sound; a second determination unit that determines whether the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound overlap by comparing the arrival direction of the predetermined sound and the arrival direction of the analyzed external sound; and an adjustment unit that performs at least one of (a) adjusting the sound pressure of at least one of the predetermined sound and the external sound, and (b) adjusting the arrival direction of the predetermined sound based on the determination results of the first determination step and the second determination step, and an output unit that outputs sound using the output sound signal generated by the adjustment.
[0035] According to this, the same effects as the information processing method described above can be achieved.
[0036] Furthermore, these general or specific aspects may be implemented in a system, apparatus, method, integrated circuit, computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.
[0037] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that all of the embodiments described below show general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components not described in the independent claims are described as optional components. Note that each figure is a schematic diagram and is not necessarily drawn precisely. Also, in each figure, the same reference numerals are given to substantially the same configurations, and duplicate descriptions may be omitted or simplified.
[0038] Also, in the following description, ordinal numbers such as first, second, and third may be attached to elements. These ordinal numbers are attached to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be appropriately interchanged, newly assigned, or removed.
[0039] (Embodiment) [Overview] First, an overview of the audio playback device according to the embodiment will be described. FIG. 1 is a schematic diagram showing a usage example of the audio playback device according to the embodiment. In FIG. 1, a user 99 who uses the audio playback device 100 is shown.
[0040] The audio playback device 100 shown in FIG. 1 is used simultaneously with the stereoscopic video playback device 200. By viewing stereoscopic images and stereoscopic sounds simultaneously, the images enhance the auditory sense of presence, and the sounds enhance the visual sense of presence, allowing the user to feel as if they are at the scene where the images and sounds were captured. For example, when an image (moving image) of a person speaking is being displayed, even if the sound image localization of the conversation sound is off from the person's mouth, it is known that the user 99 can perceive it as the conversation sound emitted from the person's mouth. In this way, the position of the sound image may be corrected by visual information, and the sense of presence may be enhanced by combining the image and the sound.
[0041] The stereoscopic video playback device 200 is an image display device worn on the head of the user 99. Therefore, the stereoscopic video playback device 200 moves integrally with the head of the user 99. For example, as shown in the figure, the stereoscopic video playback device 200 is a glasses-type device supported by the ears and nose of the user 99.
[0042] The stereoscopic video playback device 200 changes the image displayed according to the movement of the head of the user 99, so that the user 99 is perceived as moving their head within the three-dimensional image space. That is, when an object in the three-dimensional image space is located in front of the user 99, when the user 99 turns to the right, the object moves to the left direction of the user 99, and when the user 99 turns to the left, the object moves to the right direction of the user. In this way, the stereoscopic video playback device 200 moves the three-dimensional image space in the direction opposite to the movement of the user 99 with respect to the movement of the user 99.
[0043] The stereoscopic video playback device 200 displays two images with a disparity shift in the left and right eyes of the user 99 respectively. The user 99 can perceive the three-dimensional position of the object on the image based on the disparity shift of the displayed images. Note that when the user 99 uses it with their eyes closed, such as using the audio playback device 100 for playing healing sounds for sleep induction, etc., it is not necessary to use the stereoscopic video playback device 200 simultaneously. That is, the stereoscopic video playback device 200 is not an essential component of the present disclosure.
[0044] The audio reproduction device 100 is a sound presentation device worn on the head of the user 99. Therefore, the audio reproduction device 100 moves integrally with the head of the user 99. For example, the audio reproduction device 100 in the present embodiment is a so-called over-ear headphone type device. Note that the form of the audio reproduction device 100 is not particularly limited. For example, it may be two earplug type devices independently worn on the left and right ears of the user 99. These two devices communicate with each other to synchronously present the sound for the right ear and the sound for the left ear.
[0045] The audio reproduction device 100 changes the sound presented according to the movement of the head of the user 99, so as to make the user 99 perceive as if the head is moving in the three-dimensional sound field. For this reason, as described above, the audio reproduction device 100 moves the three-dimensional sound field in the direction opposite to the movement of the user with respect to the movement of the user 99.
[0046] Here, it is known that when the sound image presented to the user overlaps with the external sound arriving from the outside and listened to by the user, it becomes difficult for the user 99 to distinguish any of these sounds. The audio reproduction device 100 according to the present embodiment can make at least one of the sound image and the external sound be perceived by the user 99 by correcting the sound presented by information processing so as to avoid this phenomenon. That is, the audio reproduction device 100 operates to detect and eliminate the overlap between the sound image and the external sound, so as to make at least one of the sound image and the external sound be perceived by the user 99.
[0047] [Configuration] Next, with reference to FIG. 2, the configuration of the audio reproduction device 100 according to the present embodiment will be described. FIG. 2 is a block diagram showing the functional configuration of the audio reproduction device according to the embodiment.
[0048] As shown in FIG. 2, the audio reproduction device 100 according to the present embodiment includes a processing module 101, a communication module 102, a detector 103, and a driver 104.
[0049] The processing module 101 is an arithmetic unit for performing various signal processes in the acoustic playback device 100. The processing module 101 includes, for example, a processor and a memory, and various functions are exhibited by the program stored in the memory being executed by the processor.
[0050] The processing module 101 includes an acquisition unit 111, a filter selection unit 121, an output sound generation unit 131, and a signal output unit 141. Details of each functional unit included in the processing module 101 will be described below together with details of configurations other than the processing module 101.
[0051] The communication module 102 is an interface device for receiving input of sound information to the acoustic playback device 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device by wireless communication. More specifically, the communication module 102 receives a radio signal indicating sound information converted into a format for wireless communication using an antenna, and performs reconversion from the radio signal to sound information by the signal converter. Thereby, the acoustic playback device 100 acquires sound information from an external device by wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. In this way, the sound information is input to the processing module 101. Note that communication between the acoustic playback device 100 and an external device may be performed by wired communication.
[0052] The sound information acquired by the acoustic playback device 100 is encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3), for example. As an example, the encoded sound information includes information about a predetermined sound reproduced by the acoustic playback device 100 and information regarding a localization position when localizing the sound image of the sound at a predetermined position in a three-dimensional sound field (that is, making it perceived as sound arriving from a predetermined direction), that is, information regarding a predetermined direction. For example, the sound information includes information about a plurality of sounds including a first predetermined sound and a second predetermined sound, and the sound images are localized so that when each sound is reproduced, it is perceived as sound arriving from different directions in the three-dimensional sound field.
[0053] With this three-dimensional sound, for example, the sense of presence of content to be viewed, such as in combination with an image viewed using the three-dimensional video playback device 200, can be improved. Note that the sound information may include only information about a predetermined sound. In this case, information regarding a predetermined direction may be obtained separately. Also, as described above, the sound information includes first sound information regarding a first predetermined sound and second sound information regarding a second predetermined sound. However, a plurality of pieces of sound information including these separately may be obtained and reproduced simultaneously to localize sound images at different positions within the three-dimensional sound field. Thus, there is no particular limitation on the form of the input sound information, and the sound playback device 100 may be provided with an acquisition unit 111 corresponding to various forms of sound information.
[0054] Here, an example of the acquisition unit 111 will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. As shown in FIG. 3, the acquisition unit 111 in the present embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.
[0055] The encoded sound information input unit 112 is a processing unit to which the encoded (in other words, encoded) sound information acquired by the acquisition unit 111 is input. The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that generates information about a predetermined sound and information about a predetermined direction included in the sound information in a form used for subsequent processing by decoding (in other words, decoding) the sound information output from the encoded sound information input unit 112. The sensing information input unit 114 will be described below together with the function of the detector 103.
[0056] The detector 103 is a device for detecting the movement speed of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In the present embodiment, the detector 103 is built into the acoustic playback device 100. However, for example, it may be built into an external device such as a stereoscopic video playback device 200 that operates according to the movement of the head of the user 99 in the same manner as the acoustic playback device 100. In this case, the detector 103 may not be included in the acoustic playback device 100. Further, as the detector 103, an external imaging device or the like may be used to image the movement of the head of the user 99, and the movement of the user 99 may be detected by processing the captured image.
[0057] The detector 103 is, for example, integrally fixed to the housing of the acoustic playback device 100 and detects the speed of movement of the housing. Since the acoustic playback device 100 including the above housing moves integrally with the head of the user 99 after the user 99 wears it, the detector 103 can, as a result, detect the speed of movement of the head of the user 99.
[0058] The detector 103 may detect, for example, as the amount of movement of the head of the user 99, a rotation amount having at least one of three axes orthogonal to each other in a three-dimensional space as a rotation axis, or a displacement amount having at least one of the above three axes as a displacement direction. Further, the detector 103 may detect both the rotation amount and the displacement amount as the amount of movement of the head of the user 99.
[0059] The sensing information input unit 114 acquires the movement speed of the head of the user 99 from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of movement of the head of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 acquires at least one of the rotation speed and the displacement speed from the detector 103. The amount of movement of the head of the user 99 acquired here is used to determine the coordinates and orientation of the user 99 in the three-dimensional sound field. In the acoustic playback device 100, based on the determined coordinates and orientation of the user 99, the relative position of the sound image is determined and the sound is reproduced. Specifically, the filter selection unit 121 and the output sound generation unit 131 realize the above functions.
[0060] The filter selection unit 121 is a processing unit that determines, based on the determined coordinates and orientation of the user 99, from which direction in the three-dimensional sound field the user 99 should perceive a predetermined sound, and selects a stereophonic filter to be applied to the predetermined sound. The stereophonic filter is a function filter that convolves a specific head-related transfer function with the input predetermined sound so that the user 99 perceives the predetermined sound as coming from a predetermined direction based on the specific head-related transfer function. In other words, when a predetermined sound (or information regarding the predetermined sound) is input to the stereophonic filter, a sound pressure difference, a time difference, a phase difference, etc. are generated in the left and right sound signals of the predetermined sound, and a sound signal capable of reproducing a predetermined sound with a controlled arrival direction can be output.
[0061] A plurality of stereophonic filters that are candidates for selection are, for example, adjusted for each user 99 and prepared in advance. These plurality of stereophonic filters are each calculated and generated for each arrival direction and are stored in a storage device (not shown) or the like for storing the plurality of stereophonic filters.
[0062] Here, an example of the filter selection unit 121 will be described with reference to FIG. 4. FIG. 4 is a block diagram showing the functional configuration of the filter selection unit according to the embodiment. As shown in FIG. 4, the filter selection unit 121 in the present embodiment includes, for example, a first analysis unit 122, a second analysis unit 123, a third analysis unit 124, a first determination unit 125, a second determination unit 126, and an adjustment unit 127.
[0063] The first analysis unit 122 is a processing unit that analyzes the type of a predetermined sound included in the sound information. The first analysis unit 122 outputs, as an analysis result, information indicating which of a plurality of preset types the predetermined sound corresponds to.
[0064] Note that the type of the predetermined sound may be, for example, whether it is a voice of a human, that is, it may consist of two types, voice and other than voice, or it may be a type that does not require specific objects such as a first type, a second type,... classified by frequency characteristics from the sound source of the sound. Also, the number of types is not particularly limited, and it may be set according to the type of the predetermined sound included in the sound information and the type of external sound assumed from the environment in which the sound playback device 100 is used. This description regarding the type of the predetermined sound is similarly applicable to the type of the external sound.
[0065] The second analysis unit 123 is a processing unit that analyzes the type of external sound that arrives from outside the sound playback device 100 and is listened to by the user 99. The second analysis unit 123 outputs, as an analysis result, information indicating which of a plurality of preset types the predetermined sound corresponds to. Also, the analysis result of the type of the external sound by the second analysis unit 123 is used for comparison with the type of the above-described predetermined sound. Therefore, as the external sound, a sound that is assumed to make it difficult to listen to at least one of the predetermined sound and the external sound when they overlap is used, and the others may be deleted. For example, the sound pressure of the predetermined sound is determined in advance by the sound information and the set volume of the sound playback device 100 by the user 99. Therefore, a threshold for determining whether to use it as the external sound may be provided depending on whether it is within a sound pressure range that can sufficiently interfere with the reproduced predetermined sound.
[0066] The description of the analysis of the type of the predetermined sound by the first analysis unit 122 and the description of the analysis of the type of the external sound by the second analysis unit 123 will be further described later with reference to FIG. 7.
[0067] The third analysis unit 124 is a processing unit that analyzes the arrival direction of the external sound. The third analysis unit 124 acquires the external sound picked up by each of two or more sound collection devices as respective external sound information, identifies the same one external sound based on the external sound information among these two or more sound collection devices, and analyzes the arrival direction of the external sound by calculation based on the arrival time difference, sound pressure difference, phase difference, etc. The third analysis unit 124 outputs, as an analysis result, information on from which direction the external sound has arrived with respect to the user 99.
[0068] The first determination unit 125 is a processing unit that determines whether or not the types of the predetermined sound and the external sound match. For this purpose, the first determination unit 125 acquires the analysis results of the first analysis unit 122 and the second analysis unit 123. Based on these analysis results, the first determination unit 125 determines whether or not the arrival directions of the predetermined sound and the external sound match. The first determination unit 125 outputs, as a determination result, information indicating whether or not the types of the predetermined sound and the external sound match. When there are a plurality of predetermined sounds and external sounds respectively, the first determination unit 125 may perform determination for all combinations of the predetermined sound and the external sound, or may perform determination for all combinations of the predetermined sound and the external sound only within a predetermined range as seen from the user 99.
[0069] The second determination unit 126 is a processing unit that determines whether the arrival direction of a predetermined sound overlaps with the arrival direction of an external sound based on the analysis result of the third analysis unit 124. The second determination unit 126 calculates the arrival direction of the predetermined sound based on the predetermined direction included in the sound information and the coordinates and orientation of the user 99, and determines whether they overlap by comparing the calculated arrival direction of the predetermined sound with the arrival direction of the external sound. In the determination by the second determination unit 126, the arrival direction of the predetermined sound and the arrival direction of the external sound do not necessarily have to exactly match. For example, if it is found that when the arrival direction of the predetermined sound and the arrival direction of the external sound are within a certain angular range, they interfere with each other and make it difficult for the user 99 to distinguish them, a threshold value for such an angular range may be provided. Since this threshold value is affected by factors such as the sound pressure of the predetermined sound, the sound pressure of the external sound, and the minimum discrimination angle of the user 99, it may be set for each user 99, or may be set as fixed values such as 5 degrees, 10 degrees, 15 degrees, 20 degrees, etc. determined on average for a plurality of users 99.
[0070] The adjustment unit 127 is a processing unit that selects a stereophonic filter by performing adjustment to improve the discriminability of at least one of the predetermined sound and the external sound based on the determination result of the first determination unit 125 and the determination result of the second determination unit 126. Regarding which of the predetermined sound and the external sound the adjustment unit 127 improves the discriminability of, the user 99 can set it in advance. The adjustment unit 127 reads this set value and performs adjustment to improve the discriminability of at least one of the predetermined sound and the external sound according to the set value. The adjustment by the adjustment unit 127 will be described later together with the operation of the audio playback device 100.
[0071] The adjustment of the sound by the adjustment unit 127 is performed by changing the stereophonic filter from the stereophonic filter based on the predetermined direction on the original sound information to the stereophonic filter of the arrival direction of the sound for realizing the adjustment. That is, the adjustment of the sound by the adjustment unit 127 can also be regarded as the determination of the changed stereophonic filter. As a result, the changed stereophonic filter obtained by changing the stereophonic filter as the initial value is selected and output from the filter selection unit 121. The arrival direction of the sound in the output sound signal at this time is a direction different from the predetermined direction on the sound information.
[0072] Note that the stereophonic filter may be directly determined without setting the initial value of the stereophonic filter as described above. That is, the change of the stereophonic filter is an expression used for convenience of explanation, and directly selecting and outputting the stereophonic filter without using the initial value is also included in the present disclosure.
[0073] The output sound generation unit 131 is a processing unit that generates an output sound signal by inputting information about a predetermined sound included in sound information to the selected stereophonic filter using the stereophonic filter selected by the filter selection unit 121.
[0074] Here, an example of the output sound generation unit 131 will be described with reference to FIG. 5. FIG. 5 is a block diagram showing the functional configuration of the output sound generation unit according to the embodiment. As shown in FIG. 5, the output sound generation unit 131 in the present embodiment includes, for example, a filter processing unit 132. The filter processing unit 132 sequentially reads the filters continuously selected by the filter selection unit 121, and inputs information about the corresponding predetermined sound on the time axis, thereby continuously outputting a sound signal in which the arrival direction of the predetermined sound arriving on the three-dimensional sound field is controlled. In this way, the sound information segmented at each processing unit time on the time axis is output as a continuous sound signal (output sound signal) on the time axis.
[0075] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, generates a sound wave in the driver 104 based on the waveform signal, and presents sound to the user 99. The driver 104 has, for example, a diaphragm, a driving mechanism such as a magnet and a voice coil. The driver 104 operates the driving mechanism according to the waveform signal, and vibrates the diaphragm by the driving mechanism. In this way, the driver 104 generates a sound wave due to the vibration of the diaphragm according to the output sound signal, the sound wave propagates through the air and is transmitted to the ear of the user 99, and the user 99 perceives the sound.
[0076] [Operation] Next, with reference to FIGS. 6 and 7, the operation of the acoustic playback device 100 described above will be described. FIG. 6 is a flowchart showing the operation of the acoustic playback device according to the embodiment. Further, FIG. 7 is a flowchart showing the operations of the first analysis unit and the second analysis unit according to the embodiment. First, when the operation of the acoustic playback device 100 is started, the acquisition unit 111 acquires sound information via the communication module 102. The sound information is decoded by the decoding processing unit 113 into information regarding a predetermined sound and information regarding a predetermined direction, and filter selection is started.
[0077] In the filter selection unit 121, as an initial value, a stereophonic filter for reproducing a predetermined sound so as to have an arrival direction preset in the content is read from a storage device or the like.
[0078] The acoustic playback device 100 selects and applies a stereophonic filter so that a predetermined sound arrives from the arrival direction and performs sound playback. In parallel with the sound playback, the first analysis unit 122 analyzes the type of the predetermined sound being reproduced (S101) and continuously outputs the analysis result. The analysis of the type of the predetermined sound by the first analysis unit 122 is performed as shown in FIG. 7. First, the first analysis unit 122 divides the predetermined sound into predetermined processing unit times and generates divided data (S201). Next, the first analysis unit 122 inputs the divided data into a machine learning model such as a neural network constructed to cluster the divided data into classes associated with the types, and calculates the likelihood for each class (S202). As a result, the first analysis unit 122 outputs an analysis result indicating that the input divided data corresponds to the type corresponding to the class with the highest likelihood, assuming that the input divided data is of the type corresponding to the class with the highest likelihood (S203).
[0079] Returning to FIG. 6, the sound collection device for collecting external sound starts collecting external sound simultaneously with the start of the operation of the acoustic playback device 100, and sequentially outputs the external sound information to the second analysis unit 123. The second analysis unit 123 analyzes the type of the acquired external sound information (S102) and continuously outputs the analysis result in the same manner as the first analysis unit 122.
[0080] Further, the third analysis unit 124 analyzes the acquired external sound information for the arrival direction of the external sound and continuously outputs the analysis result. Since the analyses by these first analysis unit 122, second analysis unit 123, and third analysis unit 124 are performed in parallel, the order of steps S101 and S102 in the figure may be swapped.
[0081] Next, the first determination unit 125 determines whether or not the type of the predetermined sound matches the type of the external sound (S103). When the type of the predetermined sound matches the type of the external sound (Yes in S103), further, the second determination unit 126 determines whether or not the arrival direction of the predetermined sound overlaps with the arrival direction of the external sound (S104). When the arrival direction of the predetermined sound overlaps with the arrival direction of the external sound (Yes in S104), the adjustment unit 127 adjusts the stereo filter so that the sound discrimination is improved (S105). For example, the adjustment unit 127 determines the stereo filter to be changed to determine the stereo filter to be changed from the initial value stereo filter in which the predetermined direction and the arrival direction match to the stereo filter in which the predetermined direction and the arrival direction are different. On the other hand, when the type of the predetermined sound does not match the type of the external sound (No in S103), and when the arrival direction of the predetermined sound does not overlap with the arrival direction of the external sound (No in S104), the filter selection unit 121 ends the process and outputs the initial value stereo filter as the selected stereo filter.
[0082] Hereinafter, the determination (in other words, the change) of the stereo filter by the adjustment unit 127 will be described with reference to FIGS. 8 to 10. FIG. 8 is a first diagram for explaining the arrival direction of a predetermined sound by the selected stereo filter according to the embodiment. Further, FIG. 9 is a second diagram for explaining the arrival direction of a predetermined sound by the selected stereo filter according to the embodiment. Further, FIG. 10 is a third diagram for explaining the arrival direction of a predetermined sound by the selected stereo filter according to the embodiment. In FIGS. 8 to 10, the user 99 in an upright posture in the direction perpendicular to the paper surface is schematically shown as a circle with "U" attached, and this user 99 is standing upright in the direction perpendicular to the paper surface.
[0083] Further, in FIGS. 8 to 10, the positions where the predetermined sounds are localized are shown as black circles, and icons of virtual sound sources corresponding to the types of sounds are also shown.
[0084] As shown in FIG. 8, the position where the first predetermined sound is localized at a certain point in time is the first position S1. At the same point in time, the first external sound is arriving from the second position S2. For the first predetermined sound and the first external sound, the same speaker icon is attached, indicating that they are of the same type. Therefore, the determination result by the first determination unit 125 indicates a match in type. Also, the range with dot hatching in the figure (the front side in the figure) is a range that can be regarded as the arrival direction overlapping with the first predetermined sound, spreading around the arrival direction of the first predetermined sound. Since the arrival direction of the first external sound is within this range, it can be seen that the first predetermined sound and the first external sound overlap.
[0085] Therefore, the determination result by the second determination unit 126 indicates an overlap in arrival direction. As a result, in the example of FIG. 8, the stereophonic filter is changed so as to reduce the sound pressure of the first external sound and improve the discriminability of the first predetermined sound. For this purpose, the adjustment unit 127 changes the stereophonic filter so as to generate a signal of the inverse phase of the first external sound from the external sound information of the first external sound and superimpose it. As a result, the output sound signal obtained by inputting information regarding the predetermined sound into the stereophonic filter is a signal with a signal of the inverse phase of the first external sound added thereto, and by canceling out the arriving first external sound, the sound pressure of the first external sound is reduced.
[0086] Also, in FIG. 8, the dashed line extending left and right of the user 99 indicates a virtual boundary surface that divides the user 99's head into front and back. This boundary surface may be a surface along the user 99's external auditory canal, or a surface passing through the point at the rearmost end of the user 99's auricle, or simply a surface passing through the center of gravity of the user 99's head. It is known that there is a difference in the ease of hearing sound before and after such a boundary surface, that is, in front of and behind the user 99. Therefore, it is effective to make the characteristics of the change of the stereophonic filter different between the front side and the rear side with the boundary surface as the boundary.
[0087] In FIG. 8, the located position of the second predetermined sound at the same time point as above is the third position S3. At the same time point, the second external sound is arriving from the fourth position S4. For the second predetermined sound and the second external sound, the same speaker icon is attached, indicating that they are of the same type. Therefore, the determination result by the first determination unit 125 indicates a match in type. Also, the range hatched with dots in the figure (the rear side in the figure) is a range that can be regarded as a direction of arrival overlapping with the second predetermined sound, which spreads around the direction of arrival of the second predetermined sound. Since the direction of arrival of the second external sound is within this range, it can be seen that the second predetermined sound and the second external sound overlap. Therefore, the determination result by the second determination unit 126 indicates an overlap in the direction of arrival. As a result, in the example of FIG. 8, the stereophonic filter is changed so as to lower the sound pressure of the second external sound and improve the discriminability of the second predetermined sound.
[0088] Assume that the first predetermined sound and the second predetermined sound are the same sound differing only in the direction of arrival, and the first external sound and the second external sound are the same sound differing only in the direction of arrival. However, the range that can be regarded as the direction of arrival of the second predetermined sound and the second external sound overlapping on the rear side of the boundary surface is set larger than the range that can be regarded as the direction of arrival of the first predetermined sound and the first external sound overlapping on the front side of the boundary surface. In this way, a configuration corresponding to the width of the minimum discrimination angle with respect to the direction of arrival of the sound arriving from the rear side (that is, behind the user 99) may be provided as compared with the front side.
[0089] Also, as another example of the adjustment by the adjustment unit 127, as shown in FIG. 9, the arrival direction may be rotated to change the stereophonic filter so that the localization position of the first predetermined sound is the fifth position S1a. Here, the arrival direction of the first predetermined sound is rotationally varied in a direction away from the arrival direction of the first external sound until the hatched area does not overlap with the arrival direction of the external sound. In this example, the discriminability of both the first predetermined sound and the first external sound is improved and can be listened to by the user 99. Further, the adjustment unit 127 can also simply reduce the sound pressure of the first predetermined sound to improve the discriminability of the first external sound and make it audible.
[0090] Also, in the case shown in FIG. 10, the adjustment unit 127 may not particularly change the stereophonic filter. As shown in FIG. 10, for the first predetermined sound, the third external sound arrives from the sixth position S5, and the fourth external sound arrives from the seventh position S6. As shown in the figure, since the first predetermined sound and the third external sound are different types of sounds with different icons attached, they can be discriminated and listened to even if their arrival directions overlap. Also, although the first predetermined sound and the fourth external sound are the same type of sound with the same speaker icon attached, their arrival directions are sufficiently different, so they can be discriminated and listened to. Thus, when it is shown that they are of different types in the determination result of the first determination unit 125 and when it is shown that the arrival directions do not overlap in the determination result of the second determination unit 126, the adjustment unit 127 may not change the stereophonic filter.
[0091] However, when the arrival directions completely coincide even though the types of sounds are different, or when they do not overlap but are affected by each other due to sound pressure, etc., the stereophonic filter may be changed.
[0092] In this way, in the present embodiment, when it is difficult to distinguish a predetermined sound and an external sound from each other, for example, because the types of the predetermined sound and the external sound match and the arrival directions of the predetermined sound and the external sound overlap, at least one of (a) adjusting the sound pressure of at least one of the predetermined sound and the external sound, and (b) adjusting the arrival direction of the predetermined sound is performed. Thereby, the discriminability of at least one of the predetermined sound and the external sound can be improved, and it is possible to facilitate the listening of the one with improved discriminability. Therefore, it is possible to more appropriately allow the user 99 to perceive a stereoscopic sound.
[0093] (Other embodiments) As described above, the embodiments have been described. However, the present disclosure is not limited to the above-described embodiments.
[0094] For example, in the above-described embodiment, an example in which the sound does not follow the movement of the user's head has been described. However, the content of the present disclosure is also effective when the sound follows the movement of the user's head. That is, in the operation of making the user perceive a predetermined sound as a sound arriving from a first position that relatively moves with the movement of the user's head, when the types of the predetermined sound and the external sound match and the arrival directions overlap, etc., the stereophonic filter may be changed to improve the discriminability of at least one of them.
[0095] Further, for example, the acoustic playback device described in the above embodiment may be realized as a single device including all the components, or may be realized by allocating each function to a plurality of devices and having these plurality of devices cooperate with each other. In the latter case, an information processing device such as a smartphone, a tablet terminal, or a PC may be used as the device corresponding to the processing module.
[0096] As a configuration different from the description of the above embodiment, for example, the decoding processing unit can also select the changed stereo filter by correcting the original sound information. Specifically, the decoding processing unit in this example is a processing unit that generates information regarding a predetermined direction included in the sound information and corrects the original sound information. After the decoding processing unit performs the same operations as the first analysis unit, the second analysis unit, the third analysis unit, the first determination unit, and the second determination unit, if necessary, the information regarding the predetermined direction is corrected so that the arrival direction of the predetermined sound varies by an angle preset in a direction farther from the arrival direction of the external sound. As a result, based on the corrected information regarding the predetermined direction output from the decoding processing unit, only the stereo filter that defines the arrival direction from which the predetermined sound arrives is selected, and the changed stereo filter in the above embodiment is applied.
[0097] In this way, the information processing method and the like disclosed in the present application may also be realized by correcting the information regarding the predetermined direction in the original sound information. The decoding processing unit as described above can realize an acoustic playback device that can achieve the same effect as the present disclosure by simply replacing and inserting it with a processing unit that performs the decoding processing of a conventional stereo playback device.
[0098] Also, the acoustic playback device of the present disclosure can be realized as an acoustic playback device that is connected to a playback device equipped with only a driver and outputs an output sound signal using the stereo filter selected based on the acquired sound information only to the playback device. In this case, the acoustic processing device may be realized as hardware equipped with a dedicated circuit or as software for causing a general-purpose processor to execute specific processing.
[0099] Also, in the above embodiment, the processing executed by a specific processing unit may be executed by another processing unit. Also, the order of a plurality of processes may be changed, or a plurality of processes may be executed in parallel.
[0100] In addition, in the above-described embodiments, each component may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0101] Also, each component may be realized by hardware. For example, each component may be a circuit (or an integrated circuit). These circuits may constitute one circuit as a whole, or may be separate circuits respectively. Further, these circuits may be general-purpose circuits or dedicated circuits respectively.
[0102] Also, the general or specific aspects of the present disclosure may be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM. Further, the general or specific aspects of the present disclosure may be realized by any combination of an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0103] For example, the present disclosure may be realized as a method for reproducing an audio signal executed by a computer, or may be realized as a program for causing a computer to execute the audio signal reproduction method. The present disclosure may be realized as a computer-readable non-transitory recording medium on which such a program is recorded.
[0104] In addition, forms obtained by applying various modifications conceivable by those skilled in the art to each embodiment, or forms realized by arbitrarily combining the components and functions in each embodiment without departing from the spirit of the present disclosure are also included in the present disclosure.
Industrial Applicability
[0105] The present disclosure is useful in acoustic reproduction such as allowing a user to perceive three-dimensional sound.
Explanation of Signs
[0106] 99 users 100 audio playback device 101 processing module 102 communication module 103 detector 104 driver 111 acquisition unit 112 encoded audio information input unit 113 decoding processing unit 114 sensing information input unit 121 filter selection unit 122 first analysis unit 123 second analysis unit 124 third analysis unit 125 first determination unit 126 second determination unit 127 adjustment unit 131 output audio generation unit 132 filter processing unit 141 signal output unit 200 stereoscopic video playback device S1 first position S1a fifth position S2 second position S3 third position S4 fourth position S5 sixth position S6 seventh position
Claims
1. A step of acquiring first sound information regarding a first sound and second sound information regarding a second sound, wherein each of the first sound information and the second sound information includes information regarding the respective arrival directions of the first sound and the second sound at the user's listening position in the three-dimensional sound field of the first sound and the second sound; a first determination step of determining whether the types of the first sound and the second sound match based on the first sound information and the second sound information; a second determination step of determining whether the arrival directions of the first sound and the second sound overlap; and a generation step of generating an output sound signal based on the determination results of the first determination step and the second determination step and the first sound information and the second sound information. An audio signal processing method.
2. The generation step includes an adjustment step of adjusting the sound pressure of at least one of the first sound and the second sound based on the determination results of the first determination step and the second determination step, and generates the output sound signal after the adjustment step performs (a). The audio signal processing method according to Claim 1.
3. The generation step includes an adjustment step of adjusting the arrival direction of the first sound based on the determination results of the first determination step and the second determination step, and generates the output sound signal after the adjustment step performs (b). The audio signal processing method according to Claim 1.
4. The generation step includes an adjustment step of performing at least one of (a) adjusting the sound pressure of at least one of the first sound and the second sound and (b) adjusting the arrival direction of the first sound based on the determination results of the first determination step and the second determination step, generates the output sound signal after the adjustment step performs at least one of (a) and (b), and in the adjustment step, when it is determined in the determination of the first determination step that the types of the first sound and the second sound match and it is determined in the determination of the second determination step that the arrival directions of the first sound and the second sound overlap, at least one of (a) and (b) is performed. The audio signal processing method according to Claim 1.
5. In the adjustment step, as (a), a superimposed sound having a phase opposite to that of the second sound is generated and superimposed to attenuate the sound pressure of the second sound. The voice signal processing method according to claim 2 or 4.
6. In the adjustment step, as the (b), the arrival direction of the first sound is varied by a preset angle in a direction away from the arrival direction of the second sound. The voice signal processing method according to claim 3 or 4.
7. In the adjustment step, as the (b), information regarding the arrival direction of the first sound is corrected so that the arrival direction of the first sound is varied by a preset angle in a direction away from the arrival direction of the second sound. The voice signal processing method according to claim 6.
8. A first analysis step of analyzing the type of the first sound, and a second analysis step of analyzing the type of the second sound, and including In the analysis of the type of the first sound and the analysis of the type of the second sound, the sound to be analyzed is divided for each unit time in the time domain, by inputting the divided sound into a machine learning model, likelihoods for each of a plurality of preset types are calculated, and an analysis result indicating that the type of the input sound corresponds to the type with the highest calculated likelihood is output. The voice signal processing method according to any one of claims 1 to 7.
9. A first analysis step of analyzing the type of the first sound, and a second analysis step of analyzing the type of the second sound, and including The type of the first sound and the type of the second sound consist of two types: voice and non-voice. The voice signal processing method according to any one of claims 1 to 8.
10. The determination as to whether or not the arrival directions of the first sound and the second sound overlap is made based on whether or not the angular difference between the arrival direction of the first sound and the arrival direction of the second sound is smaller than a threshold value. For a virtual boundary surface that divides the user's head into front and back, a first threshold value, which is the threshold value when the arrival directions of the first sound and the second sound are on the rear side of the boundary surface, is larger than a second threshold value, which is the threshold value when the arrival directions of the first sound and the second sound are on the front side of the boundary surface. The voice signal processing method according to any one of claims 1 to 9.
11. For causing a computer to execute the voice signal processing method according to any one of claims 1 to 10 Program.
12. An acquisition unit that acquires first sound information regarding a first sound and second sound information regarding a second sound, wherein each of the first sound information and the second sound information includes information regarding a respective arrival direction to a user's listening position in a three-dimensional sound field of the first sound and the second sound. A first determination unit that determines whether or not the types of the first sound and the second sound match based on the first sound information and the second sound information. A second determination unit that determines whether or not the arrival direction of the first sound and the arrival direction of the second sound overlap. A generation unit that generates an output sound signal based on the determination results of the first determination unit and the second determination unit, and the first sound information and the second sound information. An audio signal processing device.
Citation Information
Patent Citations
Signal processing device, signal processing method and program
JP2010183451A
Acoustic controller, electronic apparatus and acoustic control method
JP2015198297A
Voice generation program in virtual space, generation method of quadtree, and voice generation device
JP2020018620A