Acoustic processing device and acoustic processing method
By obtaining sound space information in the audio processing device and controlling the selection of reflected sound, the problem of low audio processing efficiency of mobile listeners in the virtual space is solved, and more efficient audio processing and better sound quality experience is achieved.
Patent Information
- Application Number
- CN202380071404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-11
- Filing Date
- 2023-10-06
- Publication Date
- 2025-05-13
AI Technical Summary
Existing audio processing technologies are difficult to effectively handle the sound effects required by listeners moving in virtual spaces, resulting in problems such as large computing volume, short battery life and poor sound quality.
By using circuits and memory in the audio processing device, sound space information is acquired, and whether or not the reflected sound generated in the sound space is selected based on this information is controlled, thereby reducing unnecessary sound line processing.
It realizes the reduction of the computational volume and battery consumption of the audio processing, and improves the sound quality and listener's audio experience.
Smart Images

Figure CN119998867A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio processing device and the like. Background Art
[0002] In recent years, products and services using ER (Extended Reality) (also known as XR), including VR (Virtual Reality), AR (Augmented Reality) and MR (Mixed Reality) have become popular. As a result, the importance of sound processing technology that provides immersive audio (Immersive Audio) to listeners by giving the sound emitted by virtual sound sources in virtual space or real space an acoustic effect corresponding to the environment of the space has increased.
[0003] In addition, the listener may also be expressed as a listener or a user. Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1 disclose technologies related to the sound processing device and the sound processing method of the present disclosure.
[0004] Prior art literature
[0005] Patent Literature
[0006] Patent Document 1: Japanese Patent No. 6288100
[0007] Patent Document 2: Japanese Patent Application Publication No. 2019-22049
[0008] Patent Document 3: International Publication No. 2021 / 180938
[0009] Non-patent literature
[0010] Non-patent literature 1: BCJ Moore, "An Outline of Auditory Psychology", Chengxin Bookstore, April 20, 1994, Chapter 6: Spatial Perception, p.225 Summary of the invention
[0011] Problems to be solved by the invention
[0012] For example, Patent Document 1 discloses a technology for performing signal processing on a target audio signal and presenting it to a listener. With the popularization of ER technology and the diversification of services using ER technology, there is a demand for audio processing corresponding to differences in, for example, the audio quality required by each service, the signal processing capability of the terminal used, and the sound quality that can be provided by the audio presentation device. In addition, in order to provide these, further improvements in the audio processing technology are required.
[0013] Here, the improvement of the sound processing technology refers to the change of the existing sound processing. For example, the improvement of the sound processing technology provides a process for giving a new sound effect, a reduction in the amount of sound processing, an improvement in the quality of the sound obtained by the sound processing, a reduction in the amount of data of information used in the implementation of the sound processing, or an facilitation of obtaining or generating information used in the implementation of the sound processing, etc. Alternatively, the improvement of the sound processing technology may provide a combination of any two or more of these.
[0014] In particular, these improvements are required in devices or services that allow listeners to move freely in a virtual space. However, the above-mentioned effects obtained by improving the sound processing technology are just examples. One or more technical solutions grasped based on the present disclosure may also be technical solutions conceived based on different viewpoints from the above, technical solutions that achieve different purposes from the above, or technical solutions that can obtain different effects from the above.
[0015] Means used to solve problems
[0016] An audio device according to a technical solution disclosed in the present invention comprises a circuit and a memory; the circuit uses the memory to obtain sound space information related to the sound space; based on the sound space information, the circuit obtains characteristics related to a first sound generated from a sound source in the sound space; based on the characteristics related to the first sound, the circuit controls whether to select a second sound generated in the sound space corresponding to the first sound.
[0017] Furthermore, these inclusive or specific technical solutions may also be implemented by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or by any combination of these.
[0018] Effects of the Invention
[0019] A technical solution of the present disclosure can provide, for example, processing for imparting new sound effects, reduction in the amount of processing for sound processing, improvement in the sound quality of sound obtained through sound processing, reduction in the amount of data of information used in the implementation of sound processing, or facilitation of obtaining or generating information used in the implementation of sound processing, etc. Alternatively, a technical solution of the present disclosure can provide any combination of these. As a result, a technical solution of the present disclosure provides sound processing suitable for the use environment of the listener, and can contribute to improving the sound experience of the listener.
[0020] In particular, the above-mentioned effect can be obtained in a device or service that allows the listener to move freely in a virtual space. However, the above-mentioned effect is only an example of the effect of various technical solutions grasped by the present disclosure. One or more technical solutions grasped by the present disclosure may also be technical solutions conceived based on different viewpoints from the above, technical solutions that achieve different purposes from the above, or technical solutions that can obtain different effects from the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a diagram showing an example of direct sound and reflected sound generated in an acoustic space.
[0022] Figure 2 It is a diagram showing an example of a stereophonic sound reproduction system according to an embodiment.
[0023] Figure 3A This is a block diagram showing a configuration example of an encoding device according to an embodiment.
[0024] Figure 3B This is a block diagram showing a configuration example of a decoding device according to an embodiment.
[0025] Figure 3C This is a block diagram showing another configuration example of the encoding device according to the embodiment.
[0026] Figure 3D This is a block diagram showing another configuration example of the decoding device according to the embodiment.
[0027] Figure 4A This is a block diagram showing an example configuration of a decoder according to an embodiment.
[0028] Figure 4B This is a block diagram showing another configuration example of a decoder according to an embodiment.
[0029] Figure 5 It is a diagram showing an example of the physical structure of the sound signal processing device according to the embodiment.
[0030] Figure 6 This is a diagram showing an example of the physical structure of the encoding device according to the embodiment.
[0031] Figure 7 This is a block diagram showing a configuration example of a rendering unit according to an embodiment.
[0032] Figure 8 This is a flowchart showing an operation example of the sound signal processing device according to the embodiment.
[0033] Fig. 9 This diagram shows the positional relationship between the listener and the obstacle object, which is relatively far away.
[0034] Fig.10This diagram shows the positional relationship between the listener and the obstacle object, which is relatively close.
[0035] Fig.11 This is a diagram showing the relationship between the time difference between direct sound and reflected sound and the threshold value.
[0036] Fig. 12A This is a diagram showing a part of an example of a method of setting threshold data.
[0037] Fig. 12B This is a diagram showing a part of an example of a method of setting threshold data.
[0038] Fig. 12C This is a diagram showing a part of an example of a method of setting threshold data.
[0039] Fig.13 This is a diagram showing an example of a method of setting a threshold value.
[0040] Fig.14 This is a flowchart showing an example of selection processing.
[0041] Fig.15 This is a diagram showing the relationship between the direction of direct sound, the direction of reflected sound, the time difference, and the threshold value.
[0042] Fig.16 It is a diagram showing the relationship among the angle difference, time difference and threshold value.
[0043] Fig.17 This is a block diagram showing another configuration example of the rendering unit.
[0044] Fig.18 This is a flowchart showing another example of the selection process.
[0045] Fig.19 This is a flowchart showing still another example of the selection process.
[0046] Fig. 20 This is a flowchart showing a first variation of the operation of the audio signal processing device according to the embodiment.
[0047] Fig.21 This is a flowchart showing a second variation of the operation of the audio signal processing device according to the embodiment.
[0048] Fig. 22 This is a diagram showing an example of the arrangement of avatars, sound source objects, and obstacle objects.
[0049] Fig.23 This is a flowchart showing still another example of the selection process.
[0050] Fig.24 This is a block diagram showing an example of a configuration for a rendering unit to perform pipeline processing.
[0051] Fig.25 This diagram shows the transmission and diffraction of sound. DETAILED DESCRIPTION
[0052] (Understanding that is the basis of the present disclosure)
[0053] Figure 1 This is a diagram showing an example of direct sound and reflected sound generated in a sound space. In the audio processing that expresses the characteristics of a virtual space with sound, in order to express the width of the space and the material of the wall, etc., and to accurately grasp the position of the sound source (localization of the sound image), it is effective to reproduce not only the direct sound but also the reflected sound.
[0054] For example, in Figure 1 When listening to sound in a rectangular room, six primary reflections corresponding to the six walls are generated for one sound source. The reproduction of these reflections becomes a clue to the proper understanding of space and sound image. Furthermore, for each reflection, a secondary reflection is generated on a surface other than the reflection surface that generated the reflection. These reflections also become effective clues in perception.
[0055] However, even if only the second reflection is considered, one direct sound and 36 (6+6×5) reflected sounds are generated for one sound source, so 37 sound rays are generated. A considerable amount of calculation is required to process these sound rays.
[0056] In addition, in recent years' application products based on the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, there will inevitably be multiple sound sources, so even larger amounts of computing power will be required.
[0057] In addition, listeners who listen to sounds in a virtual space use headphones or VR goggles. In order to provide stereo sound to such listeners, binaural processing is performed on each sound line to reproduce the direction and sense of distance of the sound by giving a sound pressure ratio and phase difference between the two ears. Therefore, if all the reflected sounds generated are to be reproduced, the amount of calculation is very large.
[0058] On the other hand, as batteries for VR goggles worn by listeners experiencing in a virtual space, small storage batteries are sometimes used for their convenience. In order to extend the battery life, the computational load required for the above-mentioned processing is preferably small. For this reason, it is desirable to reduce the number of sound rays generated on a scale of several hundred within a range that does not impair the localization of the sound and the grasp of the space.
[0059] In addition, in a system that reproduces sound, sometimes 6DoF (6 Degrees of Freedom) or the like is allowed for the position and orientation of the listener. In this case, the positional relationship between the listener, the sound source, and the object that reflects the sound cannot be determined unless it is reproduced (rendered). Therefore, the reflected sound cannot be determined unless it is reproduced. Therefore, it is difficult to predetermine the reflected sound of the processing object.
[0060] Therefore, appropriately selecting one or more reflected sounds to be processed or not to be processed from among a plurality of reflected sounds generated in the sound space during reproduction is advantageous for appropriately reducing the amount of calculation and the calculation load.
[0061] Therefore, an object of the present disclosure is to provide an audio processing device or the like that can appropriately control whether or not to select a sound generated in a sound space.
[0062] In addition, controlling whether to select a sound corresponds to determining whether to select a sound. In addition, selecting a sound may be selecting the sound as a processing target sound or selecting the sound as a non-processing target sound.
[0063] (Public Summary)
[0064] The sound processing device of the first technical solution mastered based on the present disclosure includes a circuit and a memory; the circuit uses the memory to obtain sound space information related to the sound space; based on the sound space information, it obtains characteristics related to the first sound generated from the sound source in the sound space; based on the characteristics related to the first sound, it controls whether to select the second sound generated in the sound space corresponding to the first sound.
[0065] The device of the above technical solution can appropriately control whether to select the second sound generated in the sound space corresponding to the first sound based on the characteristics related to the first sound generated in the sound space. In other words, it can appropriately control whether to select the sound generated in the sound space. Therefore, it is possible to appropriately reduce the amount of calculation and the calculation load.
[0066] The sound processing device according to the second technical aspect grasped based on the present disclosure may be the sound processing device according to the first technical aspect, wherein the first sound is direct sound and the second sound is reflected sound.
[0067] The device of the above technical solution can appropriately control whether to select the reflected sound based on the characteristics related to the direct sound.
[0068] The sound processing device related to the third technical solution mastered based on the present disclosure may also be that, in the sound processing device of the second technical solution, the characteristic related to the first sound is the volume ratio of the direct sound to the volume of the reflected sound; the circuit calculates the volume ratio based on the sound space information, and controls whether to select the reflected sound based on the volume ratio.
[0069] The device of the above technical solution can appropriately select the reflected sound that has a greater impact on the listener's perception based on the volume ratio between the volume of the direct sound and the volume of the reflected sound.
[0070] The sound processing device of the fourth technical solution grasped based on the present disclosure may be, in the sound processing device of the third technical solution, when reflected sound is selected, the circuit generates sounds that reach the two ears of the listener respectively by applying binaural processing to the reflected sound and the direct sound.
[0071] The device of the above technical solution can appropriately select the reflected sound that has a great influence on the listener's perception, and apply binaural processing to the selected reflected sound.
[0072] The sound processing device related to the fifth technical solution mastered based on the present disclosure can also be that in the sound processing device of the third or fourth technical solution, the circuit calculates the time difference between the end time of the direct sound and the arrival time of the reflected sound based on the sound space information, and controls whether to select the reflected sound based on the time difference and the volume ratio.
[0073] The device of the above technical solution can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the time difference between the end time of the direct sound and the arrival time of the reflected sound, and the volume ratio between the volume of the direct sound and the volume of the reflected sound. Therefore, the device of the above technical solution can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the back masking effect.
[0074] The sound processing device related to the sixth technical solution mastered based on the present disclosure may also be that, in the sound processing device of the fifth technical solution, the circuit selects the reflected sound when the volume ratio is above the threshold; the first threshold used as the threshold when the time difference is the first value is greater than the second threshold used as the threshold when the time difference is the second value larger than the first value.
[0075] The device of the above technical solution can increase the possibility of selecting the reflected sound with a large time difference between the end time of the direct sound and the arrival time of the reflected sound. Therefore, the device of the above technical solution can more appropriately select the reflected sound with a large degree of influence on the listener's perception.
[0076] The sound processing device of the seventh technical solution based on the present disclosure may also be a sound processing device of the third or fourth technical solution, wherein the circuit calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound based on the sound space information; and controls whether to select the reflected sound based on the time difference and the volume ratio.
[0077] The device of the above technical solution can more appropriately select the reflected sound with a greater impact on the listener's perception based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and the volume ratio between the volume of the direct sound and the volume of the reflected sound. Therefore, the device of the above technical solution can more appropriately select the reflected sound with a greater impact on the listener's perception based on the priority effect.
[0078] The sound processing device of the 8th technical solution based on the present disclosure may also be that, in the sound processing device of the 7th technical solution, the circuit selects the reflected sound when the volume ratio is above the threshold; the first threshold used as the threshold when the time difference is the first value is greater than the second threshold used as the threshold when the time difference is the second value larger than the first value.
[0079] The device of the above technical solution can increase the possibility of selecting the reflected sound with a large time difference between the arrival time of the direct sound and the arrival time of the reflected sound. Therefore, the device of the above technical solution can appropriately select the reflected sound with a large degree of influence on the listener's perception.
[0080] The sound processing device according to the ninth technical solution grasped based on the present disclosure may be the sound processing device according to the eighth technical solution, wherein the circuit adjusts the threshold value based on the arrival direction of the direct sound and the arrival direction of the reflected sound.
[0081] The device of the above technical solution can appropriately select the reflected sound that has a greater impact on the listener's perception based on the arrival direction of the direct sound and the arrival direction of the reflected sound.
[0082] The sound processing device according to the tenth technical solution grasped based on the present disclosure may be the sound processing device according to any one of the second to ninth technical solutions, wherein when the reflected sound is not selected, the circuit corrects the volume of the direct sound based on the volume of the reflected sound.
[0083] The device of the above technical solution can appropriately reduce the sense of disharmony caused by the lack of volume of the reflected sound due to the reflected sound not being selected with a small amount of calculation.
[0084] The sound processing device according to the eleventh technical solution grasped based on the present disclosure may be the sound processing device according to any one of the second to ninth technical solutions, wherein when the reflected sound is not selected, the circuit synthesizes the reflected sound with the direct sound.
[0085] The device of the above technical solution can more accurately reflect the characteristics of the reflected sound to the direct sound. Therefore, the device of the above technical solution can reduce the sense of disharmony caused by the lack of volume of the reflected sound due to the reflected sound not being selected.
[0086] The sound processing device of the twelfth technical solution grasped based on the present disclosure may be, in any one of the third to ninth technical solutions, the volume ratio is the volume ratio of the direct sound at the first moment and the reflected sound at the second moment different from the first moment.
[0087] When the time when the direct sound is perceived is different from the time when the reflected sound is perceived, the device of the above technical solution can appropriately select the reflected sound that has a greater impact on the listener's perception based on the volume ratio of the direct sound to the reflected sound at different times.
[0088] The sound processing device according to the 13th technical solution grasped based on the present disclosure may be the sound processing device according to the first or second technical solution, wherein the circuit sets a threshold value based on characteristics related to the first sound, and controls whether to select the second sound based on the threshold value.
[0089] The device of the above-mentioned aspect can appropriately control whether to select the second sound based on the threshold value set according to the characteristics related to the first sound.
[0090] The sound processing device related to the 14th technical solution based on the present disclosure may also be a sound processing device of any one of the 1st, 2nd and 13th technical solutions, wherein the characteristic related to the first sound is one or a combination of two or more of the volume of the sound source, the visibility of the sound source and the localization of the sound source.
[0091] The device of the above technical solution can appropriately control whether to select the second sound based on the volume of the sound source, the visibility of the sound source, or the localization of the sound source.
[0092] The sound processing device according to the fifteenth technical solution grasped based on the present disclosure may be the sound processing device according to any one of the first, second and thirteenth technical solutions, wherein the characteristic related to the first sound is a frequency characteristic of the first sound.
[0093] The device of the above-mentioned technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the frequency characteristics of the first sound.
[0094] The sound processing device according to the sixteenth technical solution grasped based on the present disclosure may be the sound processing device according to any one of the first, second and thirteenth technical solutions, wherein the characteristic related to the first sound is a characteristic indicating discontinuity of amplitude of the first sound.
[0095] The device of the above-mentioned solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristic indicating the discontinuity of the amplitude of the first sound.
[0096] The sound processing device related to the 17th technical solution grasped based on the present disclosure may also be, in the sound processing device of any one of the 1st, 2nd, 13th and 16th technical solutions, the characteristic related to the first sound is a characteristic representing the duration of the voiced part of the first sound or the duration of the unvoiced part of the first sound.
[0097] The device of the above-mentioned technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristics representing the duration of the voiced part of the first sound or the duration of the unvoiced part of the first sound.
[0098] The sound processing device related to the 18th technical solution mastered based on the present disclosure may also be, in the sound processing device of any one of the 1st, 2nd, 13th, 16th and 17th technical solutions, the characteristic related to the first sound is a characteristic that represents the duration of the voiced part of the first sound and the duration of the unvoiced part of the first sound in a time series.
[0099] The device of the above technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristics of the duration of the voiced part of the first sound and the duration of the unvoiced part of the first sound expressed in time series.
[0100] The sound processing device according to the 19th technical solution grasped based on the present disclosure may be the sound processing device according to any one of the 1st, 2nd, 13th and 15th technical solutions, wherein the characteristic related to the first sound is a characteristic indicating a change in the frequency characteristic of the first sound.
[0101] The device of the above-mentioned aspect can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristics indicating the change in the frequency characteristics of the first sound.
[0102] The sound processing device of the 20th technical solution grasped based on the present disclosure may also be, in the sound processing device of any one of the 1st, 2nd, 13th, 15th and 19th technical solutions, the characteristic related to the first sound is a characteristic representing the smoothness of the frequency characteristic of the first sound.
[0103] The device of the above-mentioned solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristics indicating the smoothness of the frequency characteristics of the first sound.
[0104] The sound processing device according to the 21st technical means grasped based on the present disclosure may be, in the sound processing device according to any one of the 1st, 2nd, and 13th to 20th technical means, wherein the characteristic related to the first sound is acquired from the bit stream.
[0105] The device of the above technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the information obtained from the bit stream.
[0106] The sound processing device related to the 22nd technical solution mastered based on the present disclosure can also be that in the sound processing device of any one of the 1st, 2nd and 13th to 21st technical solutions, the circuit calculates the characteristics related to the second sound; based on the characteristics related to the first sound and the characteristics related to the second sound, controls whether to select the second sound.
[0107] The device of the above technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the characteristics related to the first sound and the characteristics related to the second sound.
[0108] The audio processing device related to the 23rd technical solution mastered based on the present disclosure may also be that, in the audio processing device of the 22nd technical solution, the circuit obtains a threshold value of the volume corresponding to the boundary of whether the sound can be heard; based on the characteristics related to the first sound, the characteristics related to the second sound and the threshold value, controls whether the second sound is selected.
[0109] The device of the above-mentioned technical solution can appropriately control whether to select the second sound based on a threshold value corresponding to whether the second sound can be heard in addition to the characteristics related to the first sound and the characteristics related to the second sound.
[0110] The sound processing device according to the twenty-fourth technical solution grasped based on the present disclosure may be, in the sound processing device according to the twenty-third technical solution, wherein the characteristic related to the second sound is the volume of the second sound.
[0111] The device of the above technical solution can appropriately control whether to select the second sound based on the volume of the second sound.
[0112] The audio processing device of the 25th technical solution mastered based on the present disclosure may also be that, in the audio processing device of the first or second technical solution, the sound space information includes information about the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit controls whether to select each of the plurality of second sounds based on characteristics related to the first sound, and selects one or more processing object sounds to which binaural processing is applied from the first sound and the plurality of second sounds.
[0113] The device of the above technical solution can appropriately control whether to select each of the plurality of second sounds generated in the sound space corresponding to the first sound based on the characteristics related to the first sound generated in the sound space. In addition, the device of the above technical solution can appropriately select one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds.
[0114] The sound processing device related to the 26th technical solution mastered based on the present disclosure may also be, in the sound processing device of any one of the 1st to 25th technical solutions, the timing of obtaining the characteristics related to the first sound is at least one of when the sound space is created, when the processing of the sound space starts, and when the information update thread in the processing of the sound space is generated.
[0115] The device of the above-mentioned technical solution can appropriately select one or more processing target sounds to which binaural processing is applied based on the information acquired at the adaptive timing.
[0116] The sound processing device according to the 27th technical solution grasped based on the present disclosure may be the sound processing device according to any one of the 1st to 26th technical solutions, wherein the characteristics related to the first sound are periodically acquired after the processing of the sound space starts.
[0117] The device of the above-mentioned technical solution can appropriately select one or more processing target sounds to which binaural processing is applied based on the information obtained regularly.
[0118] The audio processing device of the 28th technical solution mastered based on the present disclosure may also be that, in the audio processing device of the 1st or 2nd technical solution, the characteristic related to the first sound is the volume of the first sound; the circuit calculates the evaluation value of the second sound based on the volume of the first sound, and controls whether to select the second sound based on the evaluation value.
[0119] The device of the above-mentioned aspect can appropriately control whether to select the second sound based on the evaluation value calculated for the second sound based on the volume of the first sound.
[0120] The sound processing device according to the 29th technical solution grasped based on the present disclosure may be the sound processing device according to the 28th technical solution, wherein the volume of the first sound has a transition.
[0121] The device of the above-mentioned aspect can appropriately control whether to select the second sound based on the evaluation value calculated from the volume with transition.
[0122] The sound processing device according to the 30th technical solution grasped based on the present disclosure may be the sound processing device according to the 28th or 29th technical solution, wherein the circuit calculates the evaluation value so that the second sound is more likely to be selected as the volume of the first sound is louder.
[0123] The device of the above-mentioned aspect can appropriately control whether to select the second sound based on the evaluation value set to a value such that the second sound is more likely to be selected as the volume of the first sound increases.
[0124] The audio processing device of the 31st technical solution mastered based on the present disclosure may also be that in the audio processing device of the first or second technical solution, the sound space information is scene information including information of the sound source in the sound space and information of the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit obtains the signal of the first sound; based on the scene information and the signal of the first sound, a plurality of second sounds are calculated; characteristics related to the first sound are obtained from the information of the sound source; based on the characteristics related to the first sound, whether to select each of the plurality of second sounds as a sound to which binaural processing is not applied is controlled, thereby selecting one or more second sounds to which binaural processing is not applied from the plurality of second sounds.
[0125] The device of the above-mentioned solution can appropriately select one or more second sounds to which binaural processing is not applied from a plurality of second sounds generated in the sound space in response to the first sound, based on the characteristics related to the first sound.
[0126] The sound processing device related to the 32nd technical solution grasped based on the present disclosure may also be that, in the sound processing device of the 31st technical solution, the scene information is updated based on the input information; and the characteristics related to the first sound are obtained corresponding to the update of the scene information.
[0127] The device of the above-mentioned technical solution can appropriately select one or more second sounds to which binaural processing is not applied based on the information obtained in response to the update of the scene information.
[0128] The sound processing device according to the thirty-third technical solution grasped based on the present disclosure may be, in the sound processing device according to the thirty-first or thirty-second technical solution, acquiring scene information and characteristics related to the first sound from metadata included in the bitstream.
[0129] The device of the above-mentioned technical solution can appropriately select one or more second sounds to which binaural processing is not applied based on information obtained from metadata included in the bitstream.
[0130] The sound processing method related to the 34th technical solution mastered based on the present disclosure includes: a step of obtaining sound space information related to the sound space; a step of obtaining characteristics related to the first sound generated from the sound source in the sound space based on the sound space information; and a step of controlling whether to select a second sound corresponding to the first sound generated in the sound space based on the characteristics related to the first sound.
[0131] The method of the above technical solution can achieve the same effect as the sound processing device described in the first technical solution.
[0132] The program related to the 35th technical solution grasped based on the present disclosure is a program for causing a computer to execute the sound processing method of the 34th technical solution.
[0133] The program of the above technical solution can achieve the same effect as the sound processing method of the 35th technical solution using a computer.
[0134] In addition, these inclusive or specific technical solutions can be implemented by systems, devices, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, or recording media.
[0135] Hereinafter, the sound processing device, the encoding device, the decoding device, and the stereo sound reproduction system of the present disclosure will be described in detail with reference to the accompanying drawings. The stereo sound reproduction system can also be expressed as a sound signal reproduction system.
[0136] In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, configuration positions and connection forms of constituent elements, steps and the order of steps, etc. shown in the following embodiments are examples and do not limit the main purpose of the technical solution grasped by this disclosure. In addition, regarding the constituent elements of the following embodiments, for example, constituent elements not included in the basic technical solution described in this disclosure or constituent elements not recorded in the independent claims representing the highest concept, they are set as arbitrary constituent elements for description.
[0137] (Implementation Method)
[0138] (Example of Stereo Sound Reproduction System)
[0139] Figure 2 is a diagram showing an example of a stereophonic sound reproduction system. Specifically, Figure 2 A stereo sound reproduction system 1000 is shown as an example of a system to which the sound processing or decoding processing of the present disclosure can be applied. Stereo sound is also expressed as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound presentation device 1002 .
[0140] The sound signal processing device 1001 is also represented as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio signal after the audio processing for the listener. The sound signal is not limited to speech, as long as it is audible. For example, the audio processing is a signal processing performed on the sound signal in order to reproduce one or more effects to which the sound is subjected during the period from the sound source to the listener.
[0141] The sound signal processing device 1001 performs sound processing based on the spatial information describing the causes of the above-mentioned effects. The spatial information is, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, and parameters related to the propagation of the sound. The sound signal processing device 1001 is, for example, a PC (Personal Computer), a smart phone, a tablet computer, or a game console.
[0142] The acoustically processed signal is presented to the listener from the audio presentation device 1002. The audio presentation device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The acoustically processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio presentation device 1002 via wireless or wired communication.
[0143] When the sound prompting device 1002 is composed of a plurality of devices such as a device for the right ear and a device for the left ear, the plurality of devices synchronously prompt the sound through communication between the plurality of devices or communication between the plurality of devices and the sound signal processing device 1001. The sound prompting device 1002 is, for example, a headset, an earplug, a head-mounted display, or a surround speaker composed of a plurality of fixed speakers worn on the head of the listener.
[0144] In addition, the stereo sound reproduction system 1000 can also be used in combination with an image prompting device or a stereoscopic image prompting device that visually provides an ER experience including AR / VR. For example, the space handled by the spatial information is a virtual space, and the positions of the sound source, listener, and object in the space are virtual positions of the virtual sound source, virtual listener, and virtual object in the virtual space. The space can also be expressed as a sound space. In addition, the spatial information can also be expressed as sound space information.
[0145] also, Figure 2 The system configuration example in which the sound signal processing device 1001 and the sound prompting device 1002 are different devices is shown. However, the stereo sound reproduction system 1000 to which the sound processing method or decoding method disclosed in the present invention can be applied is not limited to Figure 2 For example, the audio signal processing device 1001 may be included in the audio prompting device 1002, and the audio prompting device 1002 may perform both audio processing and audio prompting.
[0146] Furthermore, the audio signal processing device 1001 and the audio prompting device 1002 may share and implement the audio processing described in this disclosure. Furthermore, a server connected to the audio signal processing device 1001 or the audio prompting device 1002 via a network may implement part or all of the audio processing described in this disclosure.
[0147] Furthermore, the audio signal processing device 1001 may perform audio processing by decoding a bit stream generated by encoding at least a part of the data of the audio signal and the spatial information used for the audio processing. Therefore, the audio signal processing device 1001 may be expressed as a decoding device.
[0148] (Example of encoding device)
[0149] Figure 3A is a block diagram showing an example of the configuration of an encoding device. Specifically, Figure 3A The configuration of an encoding device 1100 is shown as an example of an encoding device of the present disclosure.
[0150] Input data 1101 is encoding target data including spatial information and / or a sound signal input to encoder 1102. The spatial information will be described in detail later.
[0151] The encoder 1102 encodes the input data 1101 to generate encoded data 1103. The encoded data 1103 is, for example, a bit stream generated by encoding processing.
[0152] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or may be other memory.
[0153] In the above description, a bit stream generated by encoding is listed as an example of the encoded data 1103 stored in the memory 1104, but the encoded data 1103 may be data other than a bit stream. For example, the encoding device 1100 may store converted data generated by converting a bit stream into a predetermined data format in the memory 1104. The converted data may be, for example, a file or a multiplexed stream corresponding to one or more bit streams.
[0154] Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format) etc. In addition, the encoded data 1103 may be in the form of a plurality of packets generated by dividing the above-mentioned bit stream or file.
[0155] For example, the bit stream generated by the encoder 1102 may be transformed into data different from the bit stream. In this case, the encoding device 1100 includes a transform unit (not shown), and the transform process may be performed by the transform unit or by a CPU (Central Processing Unit) as an example of a processor described later.
[0156] (Example of decoding device)
[0157] Figure 3B is a block diagram showing an example of the structure of a decoding device. Specifically, Figure 3B The configuration of a decoding device 1110 is shown as an example of a decoding device of the present disclosure.
[0158] The memory 1114 stores, for example, the same data as the coded data 1103 generated by the coding device 1100. The stored data is read out from the memory 1114 and input to the decoder 1112 as input data 1113. The input data 1113 is, for example, a bit stream to be decoded. The memory 1114 may be, for example, a hard disk or an SSD, or other memory.
[0159] In addition, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 directly to the decoder 1112, but may transform the data read and input the transformed data to the decoder 1112 as input data 1113. The data before transformation may be, for example, multiplexed data including one or more bit streams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF.
[0160] In addition, the data before conversion may be a plurality of packets generated by dividing the bit stream or file described above. Data different from the bit stream may be read from the memory 1114 and converted into the bit stream. In this case, the decoding device 1110 may also include a conversion unit (not shown) to perform the conversion process, or the conversion process may be performed by a CPU as an example of a processor described later.
[0161] The decoder 1112 decodes the input data 1113 and generates a sound signal 1111 representing a sound to be presented to the listener.
[0162] (Another example of an encoding device)
[0163] Figure 3C is a block diagram showing another example of the configuration of the encoding device. Specifically, Figure 3C 1 shows the structure of an encoding device 1120 which is another example of the encoding device of the present disclosure. Figure 3C In Figure 3A The same constituent elements as the constituent elements of Figure 3A The same reference numerals as those in the figure are used to represent the components, and descriptions of these components are omitted.
[0164] The coding device 1100 stores the coded data 1103 in the memory 1104. On the other hand, the coding device 1120 is different from the coding device 1100 in that it includes a transmission unit 1121 that transmits the coded data 1103 to the outside.
[0165] The transmission unit 1121 transmits a transmission signal 1122 generated based on the coded data 1103 or data converted from the coded data 1103 into another data format to another device or server. The data used to generate the transmission signal 1122 is, for example, the bit stream, multiplexed data, file, or packet described in the coding device 1100.
[0166] (Another example of a decoding device)
[0167] Figure 3D is a block diagram showing another example of the structure of a decoding device. Specifically, Figure 3D 1 shows the structure of a decoding device 1130 which is another example of the decoding device of the present disclosure. Figure 3D In Figure 3B The same constituent elements as the constituent elements of Figure 3B The same reference numerals as those in the figure are used to represent the components, and descriptions of these components are omitted.
[0168] The decoding device 1110 reads the input data 1113 from the memory 1114. On the other hand, the decoding device 1130 is different from the decoding device 1110 in that it includes a receiving unit 1131 that receives the input data 1113 from the outside.
[0169] The receiving unit 1131 receives the reception signal 1132 to obtain reception data, and outputs input data 1113 to be input to the decoder 1112. The reception data may be the same as the input data 1113 to be input to the decoder 1112, or may be data of a data format different from that of the input data 1113.
[0170] When the data format of the received data is different from the data format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the received data may be converted into the input data 1113 by a conversion unit (not shown) or a CPU of the decoding device 1130. The received data is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 1120.
[0171] (Decoder example)
[0172] Figure 4A is a block diagram showing an example of the configuration of a decoder. Specifically, Figure 4A Indicate as Figure 3B or Figure 3D The configuration of a decoder 1200 is an example of a decoder 1112 in FIG.
[0173] The input data 1113 is a coded bit stream, and includes coded audio data, which is a coded audio signal, and metadata used in audio processing.
[0174] The spatial information management unit 1201 obtains metadata included in the input data 1113 and analyzes the metadata. The metadata includes information describing elements that act on the sound and are arranged in the sound space. The spatial information management unit 1201 manages the spatial information used in the sound processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1203.
[0175] In addition, in the present disclosure, the information used in the sound processing is expressed as spatial information, but other expressions may be used. For example, the information used in the sound processing may be expressed as sound spatial information or scene information. In addition, when the information used in the sound processing changes over time, the spatial information input to the rendering unit 1203 may be information expressed as a spatial state, a sound spatial state, or a scene state.
[0176] The information managed by the space information management unit 1201 is not limited to the information included in the bitstream. For example, the input data 1113 may include data representing the characteristics and structure of the space obtained from software or a server providing VR or AR as data not included in the bitstream.
[0177] In addition, the input data 1113 may also include data indicating characteristics and positions of the listener or object, etc. In addition, the input data 1113 may also include information about the position of the listener obtained by a sensor included in the terminal including the decoding device (1110, 1130), and may also include information indicating the position of the terminal estimated based on the information obtained by the sensor.
[0178] In addition, the space in the above description can be a virtually formed space, i.e., a VR space, or a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR space. In addition, the virtual space can also be represented as a sound field or a sound space. In addition, the information indicating the position in the above description can be information such as coordinate values indicating the position in the space, information indicating the relative position relative to a predetermined reference position, or information indicating the movement or acceleration of the position in the space.
[0179] The audio data decoder 1202 decodes the encoded audio data included in the input data 1113 to obtain an audio signal.
[0180] The coded audio data obtained by the stereo sound reproduction system 1000 is, for example, a bit stream coded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3) or the like. MPEG-H 3D Audio is merely an example of a coding method that can be used when generating coded audio data included in a bit stream. The coded audio data may also be a bit stream coded in another coding method.
[0181] For example, the encoding method may be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method may be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0182] Alternatively, any encoding method other than the above may be used. For example, PCM (pulse code modulation) data may be a type of encoded sound data. In this case, for example, when the number of quantization bits of the PCM data is N, the decoding process may be a process of converting an N-bit binary number into a number format (e.g., floating point format) that can be processed by the rendering unit 1203.
[0183] The rendering unit 1203 acquires the sound signal and the spatial information, performs acoustic processing on the sound signal using the spatial information, and outputs the sound signal after the acoustic processing (sound signal 1111 ).
[0184] Figure 4B is a block diagram showing another example of the configuration of a decoder. Specifically, Figure 4B Indicate as Figure 3B or Figure 3D The structure of decoder 1210 is another example of decoder 1112 in FIG.
[0185] Figure 4B The input data 1113 does not include coded audio data but includes an uncoded audio signal. Figure 4A The input data 1113 includes a bit stream including metadata and an audio signal.
[0186] The spatial information management unit 1211 is related to Figure 4A Since the spatial information management unit 1201 is the same as that of the spatial information management unit 1201, the description thereof will be omitted.
[0187] The rendering unit 1213 is related to Figure 4A The rendering unit 1203 is the same as that of FIG. 1 , so the description is omitted.
[0188] In addition, the decoders 1112, 1200, and 1210 may be expressed as an audio processing unit that performs audio processing. In addition, the decoding devices 1110 and 1130 may be the audio signal processing device 1001, or may be expressed as an audio processing device.
[0189] (Physical Structure of Sound Signal Processing Device)
[0190] Figure 5 1 is a diagram showing an example of the physical structure of the sound signal processing device 1001. Figure 5 The sound signal processing device 1001 may also be Figure 3B The decoding device 1110 or Figure 3D Decoding device 1130. Figure 3B or Figure 3D The multiple components shown can also be Figure 5 The multiple components shown are installed. In addition, part of the components described here can also be equipped in the sound prompting device 1002.
[0191] Figure 5 The sound signal processing device 1001 includes a processor 1402 , a memory 1404 , a communication IF (Interface) 1403 , a sensor 1405 , and a speaker 1401 .
[0192] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The audio processing or decoding processing of the present disclosure may be implemented by executing a program stored in the memory 1404 by the CPU, DSP or GPU. In addition, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 may also be a dedicated circuit that performs signal processing of sound signals including the audio processing of the present disclosure.
[0193] The memory 1404 is composed of, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory). The memory 1404 may also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1404 may also be an internal memory built into a CPU or a GPU. In addition, the memory 1404 may also store spatial information managed by the spatial information management unit (1201, 1211). In addition, threshold data described later may also be stored.
[0194] The communication IF 1403 is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices via the communication IF 1403 to obtain a bit stream to be decoded. The obtained bit stream is stored in the memory 1404, for example.
[0195] The communication IF 1403 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and Wi-Fi (registered trademark), and may also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark).
[0196] The communication method is not limited to the wireless communication method described above, but may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0197] The sensor 1405 performs sensing for estimating the position and orientation of the listener. Specifically, the sensor 1405 estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, speed, angular velocity, and acceleration of a part or the whole of the body, and generates position / orientation information indicating the position and / or orientation of the listener.
[0198] Alternatively, a device external to the sound signal processing device 1001 may include the sensor 1405. The body part may be the head of the listener, etc. The position / orientation information may be information indicating the position and / or orientation of the listener in real space, or information indicating the displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time point. Alternatively, the position / orientation information may be information indicating the relative position and / or orientation to the stereophonic reproduction system 1000 or the external device including the sensor 1405.
[0199] The sensor 1405 is, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Light Detection And Ranging). The sensor 1405 may also capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. In addition, a device that performs position estimation using wireless in any frequency band such as millimeter waves may also be used as the sensor 1405.
[0200] In addition, the sound signal processing device 1001 may obtain the position information from an external device having the sensor 1405 via the communication IF 1403. In this case, the sound signal processing device 1001 may not include the sensor 1405. Here, the external device is, for example, Figure 2 The sensor 1405 may be a combination of various sensors such as a gyro sensor and an acceleration sensor, or the like.
[0201] For example, as the speed of movement of the listener's head, the sensor 1405 can detect the angular velocity of rotation about at least one of the three axes orthogonal to each other in the sound space, or can detect the acceleration of displacement in the direction of displacement about at least one of the three axes.
[0202] For example, as the amount of movement of the listener's head, the sensor 1405 can detect the amount of rotation with at least one of the three axes orthogonal to each other in the sound space as the rotation axis, or can detect the amount of displacement with at least one of the three axes as the displacement direction. Specifically, the sensor 1405 detects the position (x, y, z) and angle (yaw, pitch, roll) of 6DoF as the position of the listener. The sensor 1405 is composed of a combination of various sensors for motion detection, such as a gyro sensor and an acceleration sensor.
[0203] The sensor 1405 may be realized by a camera or a GPS (Global Positioning System) receiver for detecting the position of the listener. Position information obtained by performing self-position estimation using LiDAR or the like as the sensor 1405 may also be used. For example, when the stereo sound reproduction system 1000 is realized by a smartphone, the sensor 1405 is built into the smartphone.
[0204] The sensor 1405 may include a temperature sensor such as a thermocouple for detecting the temperature of the sound signal processing device 1001. The sensor 1405 may include a battery included in the sound signal processing device 1001 or a sensor for detecting the remaining amount of a battery connected to the sound signal processing device 1001.
[0205] The speaker 1401 has a driving mechanism such as a vibration plate, a magnet or a voice coil, and an amplifier, and presents the sound signal after the sound processing to the listener as a sound. The speaker 1401 operates the driving mechanism according to the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), and the driving mechanism vibrates the vibration plate. In this way, the vibration plate vibrates according to the sound signal to generate sound waves, which propagate in the air and are transmitted to the listener's ears, and the listener perceives the sound.
[0206] In addition, although the example in which the sound signal processing device 1001 includes the speaker 1401 and presents the sound signal after the audio processing via the speaker 1401 is shown here, the presentation mechanism of the sound signal is not limited to the above-mentioned configuration.
[0207] For example, the sound signal after the sound processing can also be output to the external sound prompt device 1002 connected through the communication module. The communication performed through the communication module can be either wired or wireless. In addition, as another example, the sound signal processing device 1001 has a terminal for outputting an analog signal of sound, and a cable such as an earplug is connected to the terminal to prompt the sound signal from the earplug.
[0208] In the above case, the sound prompt device 1002 may also be a headset, earplug, head-mounted display, neck speaker, or wearable speaker, etc., which is worn on the head or a part of the body of the listener. Alternatively, the sound prompt device 1002 may also be a surround speaker composed of a plurality of fixed speakers, etc. Furthermore, the sound prompt device 1002 may also reproduce a sound signal.
[0209] (Physical Structure of Encoding Device)
[0210] Figure 6 This is a diagram showing an example of the physical structure of an encoding device. Figure 6The encoding device 1500 may also be Figure 3A The encoding device 1100 or Figure 3C The encoding device 1120 can also Figure 3A or Figure 3C The multiple components shown are Figure 6 The multiple components shown are installed.
[0211] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503 and a communication IF 1502.
[0212] The processor 1501 is, for example, a CPU, a DSP, or a GPU. The encoding process of the present disclosure may also be implemented by executing a program stored in the memory 1503 by the CPU, DSP, or GPU. In addition, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 may also be a dedicated circuit that performs signal processing of a sound signal including the encoding process of the present disclosure.
[0213] The memory 1503 is composed of, for example, a RAM or a ROM. The memory 1503 may also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1503 may also be an internal memory embedded in a CPU or a GPU.
[0214] The communication IF 1502 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WIGI (registered trademark). The encoding device 1500 communicates with another communication device via the communication IF 1502, for example, and transmits an encoded bit stream.
[0215] Communication IF1502 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and Wi-Fi (registered trademark), and may also be LTE, NR, or Wi-Fi (registered trademark). In addition, the communication method is not limited to the wireless communication method. The communication method may also be a wired communication method such as Ethernet (registered trademark), USB, or HDMI (registered trademark).
[0216] (Composition of the rendering unit)
[0217] Figure 7 is a block diagram showing an example of the configuration of a rendering unit. Specifically, Figure 7 Representation and Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213 .
[0218] The rendering unit 1300 is composed of an analyzing unit 1301 , a selecting unit 1302 , and a synthesizing unit 1303 , and performs acoustic processing on audio data included in an input signal and outputs the resultant signal.
[0219] The input signal is composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bit stream composed of sound data and metadata (control information), and in this case, the metadata may also include spatial information.
[0220] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000, and is composed of information related to objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound emitting objects that do not emit sound. The sound source object can also be simply expressed as a sound source.
[0221] A non-sound-generating object acts as an obstacle object that reflects the sound emitted by a sound source object, but a sound source object may also act as an obstacle object that reflects the sound emitted by another sound source object. An obstacle object may also be represented as a reflecting object.
[0222] The information given to both the sound source object and the non-sound emitting object includes position information, shape information, and a volume attenuation rate when the object reflects sound.
[0223] The position information is represented by the coordinate values of three axes, such as the X-axis, the Y-axis, and the Z-axis, in the Euclidean space, but it is not necessarily three-dimensional information. For example, the position information may be two-dimensional information represented by the coordinate values of two axes, the X-axis and the Y-axis. The position information of the object is determined by the representative position of the shape represented by the grid or voxel.
[0224] The shape information may also include information about the material of the surface.
[0225] The attenuation rate can be expressed by a real number greater than 0 and less than 1, or by a negative decibel value. Since the volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value, but for example, in order to present a sense of horror in an unreal space, an attenuation rate greater than 1, that is, a positive decibel value, can be deliberately set.
[0226] In addition, the attenuation rate may be set to a different value for each frequency band constituting the plurality of frequency bands, or may be set to a value independently for each frequency band. In addition, when the attenuation rate is set for each type of material on the surface of the object, the corresponding attenuation rate value may be used based on information related to the material on the surface.
[0227] In addition, the spatial information may also include information indicating whether the object is a living being and information indicating whether the object is a moving body. If the object is a moving body, the position indicated by the position information may also move over time. In this case, the information of the changed position or the amount of change is transmitted to the rendering unit 1300.
[0228] The information related to the sound source object includes not only the information given to the sound source object and the non-sound-generating object, but also the sound data. The sound data is data indicating information related to the frequency and strength of the sound, and is data expressing the sound perceived by the listener.
[0229] The audio data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal needs to be decoded at least before it reaches the synthesizer 1303, so the rendering unit 1300 may also include a decoding unit not shown. Alternatively, the signal may also be decoded by the audio data decoder 1202.
[0230] The information on the sound source object may include, for example, information on the direction of the sound source object (that is, information on the directionality of the sound emitted by the sound source object).
[0231] The information related to the orientation of the sound source object (orientation information) is typically expressed by yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted and the orientation information of the sound source object may be expressed by azimuth (yaw) and elevation (pitch). The orientation information of the sound source object may also change over time and, if changed, is transmitted to the rendering unit 1300.
[0232] Information related to the listener is information related to the position and orientation of the listener in the sound space. Information related to the position (position information) is represented by the position of the XYZ axis of the Euclidean space, but it is not necessarily three-dimensional information and can also be two-dimensional information. Information related to the orientation of the listener (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted and the orientation information of the listener can be represented by azimuth (yaw) and elevation (pitch).
[0233] The position information and orientation information of the listener may also change over time, and if changed, the information is transmitted to the rendering unit 1300 .
[0234] The sensor information is information including the amount of rotation or displacement detected by the sensor 1405 worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering unit 1300, and the rendering unit 1300 updates the information on the position and orientation of the listener based on the sensor information. The sensor information may also include, for example, location information obtained by the portable terminal through self-position estimation using GPS, camera, LiDAR, etc.
[0235] In addition, instead of the sensor 1405, information obtained from the outside via the communication module may be detected as sensor information. Information indicating the temperature of the sound signal processing device 1001 and information indicating the remaining amount of the battery may be obtained from the sensor 1405. In addition, the computing resources (CPU capacity, memory resources, PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 may be obtained in real time.
[0236] The analyzing unit 1301 analyzes the sound signal included in the input signal and the space information received from the space information management unit (1201, 1211), and detects information required for generating direct sound and reflected sound, and information required for selecting whether to generate reflected sound.
[0237] Information required for generating direct sound and reflected sound includes, for example, values related to the path of the direct sound and the reflected sound to reach the listening position, the time required to reach the position, and the volume when the sound reaches the listening position.
[0238] The information required for selecting the reflected sound to be output is information indicating the relationship between the direct sound and the reflected sound, such as a value related to the time difference between the direct sound and the reflected sound, and a value related to the volume ratio of the direct sound and the reflected sound at the listening position.
[0239] In addition, when the volume is expressed by the unit of decibel on the logarithmic axis (when the volume is expressed in the decibel area), the volume ratio of the two signals is of course expressed by the difference in decibel values. Specifically, the volume ratio of the two signals can be the difference when the amplitude values of each signal are expressed in the decibel area. This value can also be calculated based on energy value or power value. In addition, this difference can be called the difference in gain or simply the gain difference in the decibel area.
[0240] That is, the volume ratio in the present disclosure is essentially the ratio of the amplitudes of the signals, so it can also be expressed as Sound volume ratio, Volume ratio, Amplitude ratio, Sound level ratio, Sound intensity ratio or Gain ratio, etc. In addition, when the unit of volume is decibel, the volume ratio in the present disclosure can of course be changed to volume difference.
[0241] In the present disclosure, "volume ratio" typically refers to the gain difference when the volume of two sounds is expressed in decibel units. In the example of the implementation, the threshold data is also typically specified by the gain difference expressed in the decibel area. However, the volume ratio is not limited to the gain difference in the decibel area. When using a volume ratio expressed outside the decibel area, the threshold data specified in the decibel area can also be converted into the unit of the calculated volume ratio for use. Alternatively, the threshold data specified in each unit can also be stored in the memory in advance.
[0242] That is, even if a ratio of energy values or power values is used instead of the volume ratio, it is obvious that the algorithm in the present disclosure can be applied to solving the problem of the present disclosure.
[0243] The time difference between the direct sound and the reflected sound is, for example, the time difference between the arrival time (arrival time) of the direct sound and the arrival time (arrival time) of the reflected sound. The time difference between the direct sound and the reflected sound may also be the time difference between the time when the direct sound and the reflected sound reach the listening position, the time difference between the time when the direct sound ends and the time when the reflected sound reaches the listening position. The calculation method of these values will be described later.
[0244] The selection unit 1302 selects whether to generate reflected sound using the information calculated by the analysis unit 1301 and the threshold data. In other words, the selection unit 1302 determines whether to select the reflected sound as the generation target reflected sound. In other words, the selection unit 1302 selects which reflected sound to generate among the plurality of reflected sounds.
[0245] The threshold data is represented as a boundary (threshold) between whether the reflected sound is perceived or not perceived in a graph having a time difference value between the direct sound and the reflected sound on the horizontal axis and a volume ratio of the direct sound to the reflected sound on the vertical axis, for example. The threshold data may be represented by an approximate expression having the time difference value between the direct sound and the reflected sound as a variable, or by an array having the time difference value between the direct sound and the reflected sound as an index and corresponding threshold values.
[0246] For example, the selection unit 1302 selects to generate reflected sound when the volume ratio of the direct sound arrival volume to the reflected sound arrival volume ratio among the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is greater than the threshold value set by referring to the threshold data.
[0247] In other words, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is the difference in time required for the direct sound and the reflected sound to arrive at the listening position, respectively. In addition, the time difference between the time point when the sound of the direct sound ends and the time point when the reflected sound arrives at the listening position can also be used as the time difference between the direct sound and the reflected sound. In this case, threshold data different from the threshold data set by using the time difference between the arrival time of the direct sound and the arrival time of the reflected sound as a reference can also be used, or common threshold data can be used.
[0248] The threshold data may be acquired from the memory 1404 of the audio signal processing device 1001 or from an external storage device via a communication module. The method of storing the threshold data and the method of setting the threshold will be described later.
[0249] The synthesizing unit 1303 synthesizes the audio signal of the direct sound and the audio signal of the reflected sound selected and generated by the selecting unit 1302 .
[0250] Specifically, the synthesizing unit 1303 processes the input sound signal to generate the direct sound based on the information of the arrival time of the direct sound and the volume of the direct sound when it arrives calculated by the analyzing unit 1301. In addition, the synthesizing unit 1303 processes the input sound signal to generate the reflected sound based on the information of the arrival time of the reflected sound and the volume of the reflected sound when it arrives selected by the selecting unit 1302. Then, the synthesizing unit 1303 synthesizes the generated direct sound and reflected sound and outputs them.
[0251] (Rendering unit actions)
[0252] Figure 8 1 is a flowchart showing an example of the operation of the audio signal processing device 1001. Figure 8 , the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001 is shown.
[0253] In the analysis of the input signal ( Figure 8 In S101 of FIG. 1 , the analyzing unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct sound and reflected sound that can be generated in the sound space. The reflected sound detected here is a reflected sound candidate selected by the selecting unit 1302 as the reflected sound to be finally generated by the synthesizing unit 1303. In addition, the analyzing unit 1301 analyzes the input signal and calculates the information required for generating the direct sound and the reflected sound and the information required for selecting the generated target reflected sound.
[0254] First, the characteristics of the direct sound and the reflected sound are calculated. Specifically, the arrival time and volume of the direct sound and the reflected sound when they reach the listener are calculated. When there are multiple objects in the sound space as reflection objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0255] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is the path connecting the position information S (xs, ys, zs) of the sound source object and the position information A (xa, ya, za) of the listener. The direct sound arrival time (td) is the value obtained by dividing the length of the path connecting the position information S (xs, ys, zs) and the position information A (xa, ya, za) by the speed of sound (approximately 340 m / sec).
[0256] For example, the length of the path (X) is calculated by (xs-xa)^2+(ys-ya)^2+(zs-za)^2)^0.5. The volume attenuates inversely proportional to the distance. Therefore, when the volume in the location information S(xs, ys, zs) of the sound source object is N and the unit distance is U, the volume (ld) when the direct sound arrives is calculated by ld=N*U / X.
[0257] The reflected sound arrival time (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is a path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za).
[0258] In addition, the position of the sound image of the reflected sound can be derived using, for example, the "mirror method" or the "ray tracing method", or any other method for deriving the sound image position. The mirror method is a method of assuming that the reflected wave on the wall surface in the room has a mirror image at a position symmetrical to the wall surface and the sound source, and assuming that the sound wave is radiated from the position of the mirror image to simulate the sound image. The ray tracing method is a method of simulating an image (sound image) observed at a certain point by tracing a wave propagating in a straight line such as a light ray or a sound line.
[0259] Fig. 9 This diagram shows the positional relationship between the listener and the obstacle object, which is relatively far away. Fig.10 is a diagram showing the positional relationship between the listener and the obstacle object. Fig. 9 and Fig.10 Each shows an example of forming a sound image of reflected sound at a position symmetrical to the sound source position across a wall. By finding the position of the sound image of the reflected sound on the xyz axis based on such a relationship, the arrival time of the reflected sound can be found in the same way as the method of calculating the arrival time of the direct sound.
[0260] The arrival time of the reflected sound (tr) is the value obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za) by the speed of sound (about 340 m / sec). The volume decays inversely with the distance. Therefore, when the volume at the sound source position is N, the unit distance is U, and the attenuation rate of the volume in the reflection is G, the volume (lr) when the reflected sound arrives is calculated by lr = N*G*U / Y.
[0261] As described above, the attenuation rate G can be expressed by a real number greater than 0 and less than 1, or by a negative decibel value. In this case, the volume of the entire signal is attenuated by an amount corresponding to G. In addition, the attenuation rate can also be set for each frequency band constituting a plurality of frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate to each frequency component of the signal. In addition, in order to reduce the amount of calculation, the analysis unit 1301 can also use a representative value or an average value of a plurality of attenuation rates of a plurality of frequency bands as the overall attenuation rate, so that the volume of the entire signal is attenuated accordingly.
[0262] Next, the analysis unit 1301 calculates the volume ratio (L), which is the ratio of the volume (ld) of the direct sound when it arrives to the volume (lr) of the reflected sound when it arrives, and the time difference (T) between the direct sound and the reflected sound, which are required for selecting the generated target reflected sound.
[0263] The volume ratio (L), which is the ratio of the volume (ld) when the direct sound arrives to the above lr, is obtained by, for example, L = (N*G*U / Y) / (N*U / X) = G*X / Y. Since the obtained value is the ratio of the volume, the values of N and U can be any preset values.
[0264] The time difference (T) between the direct sound and the reflected sound may be, for example, the time difference between the direct sound and the reflected sound when they reach the listening position. For example, the time difference (T) between the direct sound and the reflected sound when they reach the listening position is obtained by T=tr-td.
[0265] In addition, the time difference (T) may also be the difference between the time when the direct sound and the reflected sound reach the listening position. In addition, the time difference (T) may also be the time difference between the time when the direct sound ends and the time when the reflected sound reaches the listening position. That is, the time difference (T) may also be the time difference between the time when the direct sound ends and the time when the reflected sound starts at the listening position.
[0266] Next, in the selection process of reflected sound ( Figure 8In S102 of FIG. 1 , the selection unit 1302 selects whether to generate the reflected sound calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether to select the reflected sound as the generated object reflected sound. In the case where there are multiple reflected sounds, the selection unit 1302 selects whether to generate each reflected sound. The result of the selection unit 1302 selecting whether to generate each reflected sound may be that more than one generated object reflected sound is selected from the multiple reflected sounds, or none of the generated object reflected sounds is selected.
[0267] In addition, the selection unit 1302 is not limited to the generation process, and can also select the reflected sound of the object to which other processes are applied. For example, the selection unit 1302 can also select the reflected sound of the object to which binaural processing is applied. In addition, the selection unit 1302 basically selects only one or more reflected sounds of the processing object. However, the selection unit 1302 can also select only one or more reflected sounds that are not the processing object. In addition, the processing can also be applied to one or more reflected sounds that are not selected.
[0268] For example, the selection of the reflected sound is performed based on the volume ratio (L) and the time difference (T) calculated by the analysis unit 1301. By performing the selection process based on the time difference (T) between the direct sound and the reflected sound, compared with the case where the selection process is performed based only on the volume difference between the direct sound and the reflected sound, the reflected sound that has a large impact on the listener's perception can be more appropriately selected.
[0269] Specifically, the selection of whether to generate reflected sound is performed, for example, by comparing the volume ratio of the direct sound and the reflected sound corresponding to the time difference between the direct sound and the reflected sound with a preset threshold. The threshold is set with reference to the threshold data. The threshold data is an indicator indicating the boundary of whether the reflected sound of the direct sound is perceived by the listener, and is defined by the ratio of the volume (Id) of the direct sound at the time of arrival to the volume (lr) of the reflected sound at the time of arrival.
[0270] In addition, the threshold value corresponds to a value represented by a numerical value set corresponding to the time difference (T). The threshold value data corresponds to the relationship between the time difference (T) and the threshold value, and corresponds to table data or a relational expression used to determine or calculate the threshold value under the time difference (T). The form and type of the threshold value data are not limited to the table data or the relational expression.
[0271] Fig.11 is a graph showing the relationship between the time difference between direct sound and reflected sound and the threshold value. Fig.11 Alternatively, the threshold data of the volume ratio pre-set for each value of the time difference between the direct sound and the reflected sound may be referred to. Fig.11 The threshold data shown are threshold data obtained by interpolation or extrapolation.
[0272] Furthermore, the threshold value of the volume ratio at the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. Furthermore, the selection unit 1302 determines whether to select the reflected sound as the generated target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than the threshold value.
[0273] By performing selection processing using threshold data of a volume ratio pre-set for each value of the time difference between direct sound and reflected sound, selection processing that takes into account post-masking or priority effects can be achieved. The types, formats, storage methods, and setting methods of threshold data will be described in detail later.
[0274] Next, in the generation process of direct sound and reflected sound ( Figure 8 In S103), the synthesizing unit 1303 generates and synthesizes the sound signal of the direct sound and the sound signal of the reflected sound selected by the selecting unit 1302 as the object reflected sound.
[0275] The sound signal of the direct sound is generated by applying the arrival time (td) and the volume (ld) at the time of arrival calculated by the analysis unit 1301 to the sound data of the sound source object included in the input information. Specifically, the sound data is delayed by the amount of the arrival time (td) and multiplied by the volume (ld) at the time of arrival. The process of delaying the sound data is a process of moving the position of the sound data forward or backward on the time axis. For example, the process of delaying the sound data without deteriorating the sound quality as disclosed in Patent Document 2 may also be applied.
[0276] The sound signal of the reflected sound is generated by applying the arrival time (tr) and the arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object, similarly to the direct sound.
[0277] However, the arrival volume (lr) in the generation of reflected sound is different from the arrival volume of direct sound, and is a value to which the attenuation rate G of the volume in the reflection is applied. G can be an attenuation rate applied to the entire frequency band. Alternatively, the reflectivity can be specified for each specified frequency band in order to reflect the bias of the frequency components generated by the reflection. In this case, the processing of applying the arrival volume (lr) can also be implemented as a processing of multiplying the attenuation rate for each frequency band, that is, a frequency equalizer processing.
[0278] In the above example, the path lengths of the direct sound and the reflected sound candidates are calculated when they reach the listener. Furthermore, the arrival time and the volume at the time of arrival are calculated based on the path lengths. And based on the time difference and volume ratio, the reflected sound candidate is selected.
[0279] In addition, as another example, the selection process can also be performed based on the path lengths when the direct sound and the reflected sound reach the listener, respectively, and the calculation of the arrival time and volume of the direct sound and the reflected sound, as well as the calculation of the time difference and the volume ratio, can be omitted. In this case, a threshold value corresponding to the path length difference can also be pre-set for the path length ratio. In addition, the selection process can also be performed based on whether the calculated path length ratio is above the threshold value corresponding to the calculated path length difference. In this way, the selection process can be performed based on the path length difference corresponding to the time difference while reducing the amount of calculation.
[0280] Furthermore, in addition to the path length difference, the value of a parameter indicating the propagation speed of sound or the value of a parameter affecting the parameter of the propagation speed of sound may be used.
[0281] (Select the details of the process)
[0282] The details of the selection process of whether to generate reflected sound will be described.
[0283] The selection of the reflected sound is performed by comparing a threshold value of the volume ratio, which is the ratio of the volume of the direct sound when it arrives to the volume of the reflected sound when the time difference (T) between the direct sound and the reflected sound is set, with the volume ratio (L) calculated by the analysis unit 1301. For example, the threshold value of the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced among the threshold values of the volume ratio pre-set for each value of the time difference between the direct sound and the reflected sound. And, depending on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value, it is determined whether the reflected sound is selected as the generated target reflected sound.
[0284] The time difference (T) may be, for example, the difference between the time when the direct sound and the reflected sound arrive at the listening position, the time difference between the time required for the direct sound and the reflected sound to arrive at the listening position, and the time difference between the time when the direct sound ends and the time when the reflected sound arrives at the listening position. Here, the end time of the direct sound may also be obtained by, for example, adding the duration of the direct sound to the arrival time of the direct sound.
[0285] Threshold data can also be determined by the auditory nerve function or the cognitive function of the brain, more specifically, by the priority effect described later, the temporal masking phenomenon described later, or a combination thereof, based on the listener's perception to detect the minimum time difference between two sounds. The specific value can be derived from the research results of known temporal masking effects, priority effects, or echo detection limits, or can be obtained through a listening experiment based on the application to the virtual space.
[0286] Fig. 12A , Fig. 12B and Fig. 12CFIG. 2 is a diagram showing an example of a method for setting threshold data. Fig. 12A , Fig. 12B and Fig. 12C As shown, the threshold data is represented by the boundary (threshold) where the reflected sound is perceived or not perceived in a graph having the time difference between the direct sound and the reflected sound on the horizontal axis and the volume ratio of the direct sound to the reflected sound on the vertical axis.
[0287] The threshold data may also be expressed by an approximate expression having the time difference between the direct sound and the reflected sound as a variable. Fig.11 The index of the time difference between the direct sound and the reflected sound and the arrangement of the threshold values corresponding to the index are stored in the area of the memory 1404.
[0288] In addition, after parsing ( Figure 8 In the case where multiple reflected sounds are generated by the S101 of the present invention, the selection process may be performed on all the reflected sounds, or the selection process may be performed only on the reflected sounds with high evaluation values based on the evaluation values derived for each reflected sound by a pre-set evaluation method. Here, the evaluation value of the reflected sound corresponds to the perceived importance of the reflected sound. In addition, a high evaluation value corresponds to a large evaluation value, and these expressions may also be replaced with each other.
[0289] The selection unit 1302 may calculate the evaluation value of the reflected sound by a pre-set evaluation method corresponding to the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.
[0290] Specifically, the louder the sound source volume, the higher the evaluation value. In addition, in order to make visual localization consistent with acoustic localization, the evaluation value may be high when the sound source object or reflection object (obstacle object) can be visually seen by the listener or when the localization of the sound source object is high.
[0291] In addition, the opening of the arrival angle of direct sound and reflected sound and the difference in arrival time of direct sound and reflected sound have a great influence on the grasp of space. Therefore, the evaluation value may be high when the opening of the arrival angle of direct sound and reflected sound is large or when the difference in arrival time of direct sound and reflected sound is large.
[0292] The above selection process can be interpreted as a process of selecting reflected sound according to the properties of direct sound. For example, in the process of selecting reflected sound according to the properties of direct sound, the threshold used in the selection of reflected sound is set or adjusted according to the properties of direct sound. Alternatively, the evaluation value used in the selection of reflected sound is calculated based on one or more of the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflection object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.
[0293] Furthermore, in the process of selecting the reflected sound according to the properties of the direct sound, it is not limited to the process of setting or adjusting the threshold according to the properties of the direct sound and the process of calculating the evaluation value used in the selection of the reflected sound to be processed, and other processes may be performed. Furthermore, in the case of performing the process of setting or adjusting the threshold according to the properties of the direct sound or the process of calculating the evaluation value used in the selection of the reflected sound to be processed, part of the process may be changed or a new process may be added.
[0294] In addition, setting the threshold value may include adjusting the threshold value and changing the threshold value.
[0295] (Threshold setting method)
[0296] The threshold data used in the selection process may be set by referring to a known value of an echo detection limit based on a precedence effect or a masking threshold based on a post-masking effect, for example.
[0297] The precedence effect is a phenomenon in which when two sounds are heard from two places, the source of the sound is recognized as the place heard first in time. If two short sounds merge and are heard as one sound, the position (localization position) of the overall sound is generally determined by the position of the first sound. The echo detection limit is a phenomenon that occurs through the precedence effect, and is the minimum time difference at which the listener's perception detects the deviation of two sounds.
[0298] exist Fig. 12C In Example 2, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the detectable reflected sound to the direct sound, specifically, the threshold value of whether the reflected sound arriving with the delay time can be detected.
[0299] Fig.13 This is a diagram showing an example of a method of setting a threshold value. Fig.13 The horizontal axis in corresponds to the arrival time of the reflected sound, specifically, corresponds to the time difference (T) between the direct sound and the reflected sound. Fig.13 The vertical axis in corresponds to the volume of the reflected sound. Specifically, Fig.13 The vertical axis in may correspond to the volume of the reflected sound set relatively to the volume of the direct sound (volume ratio), or may correspond to the volume of the reflected sound absolutely determined without depending on the volume of the direct sound.
[0300] For example, in Fig. 9 When the listener is far away from the obstacle, the arrival time of the reflected sound becomes later, as shown in Fig.13 As shown in C, the threshold is set low. Fig. 9On the other hand, in the case of Fig.10 When the listener is close to the obstacle, the arrival time of the reflected sound is Fig. 9 The situation becomes earlier, such as Fig.13 As shown in B, the threshold is set high. Fig.10 In this case, no reflected sound is generated.
[0301] Alternatively, the threshold data may be stored in the memory 1404 , retrieved from the memory 1404 when a selection process is performed, and used for the selection process.
[0302] Fig.14 1301. First, the selection unit 1302 specifies the reflected sound detected by the analysis unit 1301 (S201). Next, the selection unit 1302 detects the volume ratio (L) between the direct sound and the reflected sound and the time difference (T) between the direct sound and the reflected sound (S202 and S203).
[0303] The time difference (T) may be, for example, the time difference between the time required for the direct sound and the reflected sound to reach the listening position, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and the time difference between the time point when the direct sound ends and the time point when the reflected sound reaches the listening position. Here, an example based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is described.
[0304] Specifically, the selection unit 1302 calculates the difference between the length of the path of the direct sound and the length of the path of the reflected sound based on the position information of the sound source object and the listener and the position information and shape information of the obstacle object. Furthermore, the selection unit 1302 detects the time difference (T) between the time when the direct sound reaches the listener position and the time when the reflected sound reaches the listener position by dividing the length by the speed of sound.
[0305] The volume reaching the listener is attenuated in proportion to the distance from the listener relative to the volume of the sound source (inversely proportional to the distance). Therefore, the volume of the direct sound is obtained by dividing the volume of the sound source by the length of the path of the direct sound. The volume of the reflected sound is obtained by dividing the volume of the sound source by the length of the path of the reflected sound and multiplying it by the attenuation rate given to the virtual obstacle object. The selection unit 1302 detects the volume ratio by calculating the ratio of their volumes.
[0306] Furthermore, the selection unit 1302 determines a threshold value corresponding to the time difference (T) using the threshold value data (S204). Next, the selection unit 1302 determines whether the detected volume ratio (L) is equal to or greater than the threshold value (S205).
[0307] When the volume ratio (L) is greater than the threshold value ("Yes" in S205), the selection unit 1302 selects the reflected sound as the reflected sound to be generated (S206). When the volume ratio (L) is less than the threshold value ("No" in S205), the selection unit 1302 does not select the reflected sound as the reflected sound to be generated (S207). That is, in this case, the selection unit 1302 determines the reflected sound as the reflected sound not to be generated.
[0308] Then, the selection unit 1302 determines whether there is an unspecified reflected sound (S208). If there is an unspecified reflected sound ("Yes" in S208), the selection unit 1302 repeats the above-mentioned process (S201 to S207). If there is no unspecified reflected sound ("No" in S208), the selection unit 1302 ends the process.
[0309] This selection process may be executed for all reflected sounds generated in the analysis process, or may be executed only for the reflected sounds having high evaluation values as described above.
[0310] (Details of the Threshold Storage Method)
[0311] The threshold data related to the present embodiment is stored in the memory 1404 of the sound signal processing device 1001. The form and type of the stored threshold data can be any form and any type. In the case of storing multiple forms and multiple types of thresholds, in the selection process, it is possible to determine which form and which type of threshold is used for the selection process of the reflected sound. The method for determining which threshold data is used for the selection process will be described later.
[0312] In addition, multiple forms and types of threshold data may be combined and stored. The combined threshold data may be read from the spatial information management unit (1201, 1211) to set the threshold used in the selection process. In addition, the threshold data stored in the memory 1404 may be stored in the spatial information management unit (1201, 1211).
[0313] Threshold data can also be stored as thresholds at different time differences, for example, to depict Fig. 12C The threshold lines shown in [Example 1] and [Example 2].
[0314] Alternatively, threshold data can be Fig.11 As shown, the threshold value and the time difference (T) are stored as table data corresponding to each other. That is, the threshold value data can also be stored as table data with the time difference (T) as an index. Of course, Fig.11 The threshold value shown is an example and is not limited to Fig.11In addition, instead of storing the threshold value itself, the threshold value may be approximated by a function having the time difference (T) as a variable, and the coefficient of the function may be stored. In addition, a plurality of approximate expressions may be combined and stored.
[0315] In the memory 1404, information related to a relational expression representing the relationship between the time difference (T) and the threshold value may also be stored. That is, an expression having the time difference (T) as a variable may also be stored. The threshold value of each time difference (T) may also be approximated by a straight line or a curve, and parameters representing the geometric shape of the straight line or the curve may be stored. For example, when the geometric shape is a straight line, the starting point and slope used to represent the straight line may also be stored.
[0316] In addition, the type and form of the threshold data may be set and stored for each property of the direct sound. In addition, a parameter for adjusting the threshold according to the property of the direct sound and for selecting a process may be stored. The process of adjusting the threshold according to the property of the direct sound and for selecting a process will be described later as a modified example of the threshold setting method.
[0317] As an example of combining and storing a plurality of threshold data, it is also possible to Fig. 12C As shown in [Example 3] of , the larger value of the masking threshold and the echo detection limit threshold is stored for each time difference (T). Fig. 12C As shown in [Example 4] of FIG. 1 , the larger value of the minimum volume reproduced in the virtual space and the echo detection limit threshold is stored for each time difference (T).
[0318] The combination of the plurality of types of threshold data is not limited thereto. For example, the maximum value information may be stored for each time difference (T) among the plurality of threshold data.
[0319] In the above, the information on the threshold value has a time item as a one-dimensional index. The information on the threshold value may also have a two-dimensional or three-dimensional index further including a variable related to the arrival direction.
[0320] Fig.15 is a graph showing the relationship between the direction of direct sound, the direction of reflected sound, the time difference, and the threshold. Fig.15 As shown, a threshold value calculated in advance according to the relationship between the direction of direct sound (θ), the direction of reflected sound (γ), the time difference (T), and the volume ratio (L) may be stored.
[0321] The direction of direct sound (θ) corresponds to the angle of the arrival direction of direct sound relative to the listener. The direction of reflected sound (γ) corresponds to the angle of the arrival direction of reflected sound relative to the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of direct sound and the arrival time of reflected sound to the listening position. The volume ratio (L) corresponds to the volume ratio of the volume of direct sound at the time of arrival to the volume of reflected sound at the time of arrival.
[0322] certainly, Fig.15 The threshold value shown is an example and is not limited to Fig.15 In addition, Fig.15 , the threshold value is mainly exemplified when the angle (θ) of the arrival direction of the direct sound is 0 degrees. However, the threshold value is also stored in the memory 1404 when the arrival direction (θ) of the direct sound is other than 0 degrees.
[0323] In addition, in the above, the threshold is stored as an arrangement having the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound as independent variables or indices. However, the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound may not be used as independent variables.
[0324] For example, the angle difference between the angle of the arrival direction of the direct sound (θ) and the angle of the arrival direction of the reflected sound (γ) may be used. The angle difference corresponds to the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound, and may also be expressed as the arrival angles of the direct sound and the reflected sound.
[0325] Fig.16 is a graph showing the relationship between the angle difference, the time difference and the threshold. Fig.16 As shown in the example, the threshold value calculated in advance is stored using the angle difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound as a variable. Fig.16 The threshold value shown is an example and is not limited to Fig.16 Example.
[0326] exist Fig.16 In the example of , the number of variables used in deriving the threshold value can be reduced. Therefore, the number of threshold values stored in the memory 1404 can be reduced. Therefore, the amount of data stored in the memory 1404 can be reduced.
[0327] In addition, when using the angle difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound, the threshold data may be stored in a two-dimensional arrangement. In addition, in the selection process, a three-dimensional arrangement may be used to calculate the difference between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound.
[0328] A method of selecting reflected sound using a threshold value according to the arrival direction will be described later.
[0329] (First Modification of the Threshold Setting Method)
[0330] exist Fig. 12A , Fig. 12B and Fig. 12C In the example of , multiple forms and multiple types of thresholds may also be stored in the spatial information management unit (1201, 1211). In addition, it is also possible to determine which form and which type of threshold among the multiple forms and multiple types of thresholds is used for the selection process of the reflected sound. Specifically, it is also possible to Fig. 12C As shown in Example 3, the highest threshold is used at the time difference (T) corresponding to the arrival time of the reflected sound.
[0331] Furthermore, the masking threshold, the threshold of the echo detection limit, and the threshold indicating the minimum volume reproduced in the virtual space may be stored as shown in Example 4. Furthermore, the highest threshold may be used at the time difference (T) corresponding to the arrival time of the reflected sound.
[0332] (Second Modification of Threshold Setting Method)
[0333] As another example of a method of setting a threshold value, a method of setting a threshold value based on the properties of direct sound will be described.
[0334] Fig.17 Yes means Figure 7 13 is a block diagram of another configuration example of the rendering unit 1300 shown. Fig.17 The rendering unit 1300 and Figure 7 The rendering unit 1300 of FIG. 1 is different from the rendering unit 1300 of FIG. 1 in that it includes a threshold value adjustment unit 1304. The description other than the threshold value adjustment unit 1304 is the same as that of FIG. Figure 7 The contents described in are the same, so they are omitted.
[0335] The threshold adjustment unit 1304 selects a threshold to be used by the selection unit 1302 from the threshold data based on the information indicating the property of the audio signal. Alternatively, the threshold adjustment unit 1304 may adjust the threshold included in the threshold data based on the information indicating the property of the audio signal.
[0336] The information indicating the properties of the sound signal may also be included in the input signal. Furthermore, the threshold adjustment unit 1304 may also obtain the information indicating the properties of the sound signal from the input signal. Alternatively, the analysis unit 1301 may analyze the sound signal included in the received input signal, derive the properties of the sound signal, and output the information indicating the properties of the sound signal to the threshold adjustment unit 1304.
[0337] The information indicating the properties of the audio signal may be acquired before starting the rendering process or may be acquired at any time during the rendering process.
[0338] In addition, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, and another communication device may have the function of the threshold adjustment unit 1304. In this case, the analysis unit 1301 or the selection unit 1302 may obtain information indicating the property of the sound signal, threshold data corresponding to the property, or information for adjusting the threshold data according to the property from the other communication device via the communication IF 1403.
[0339] Fig.18 This is a flowchart showing another example of the selection process. Fig.19 is a flowchart showing yet another example of the selection process. Fig.18 and Fig.19 In , the threshold is set according to the nature of the direct sound. Fig.18 In , the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the properties of the sound signal. Fig.19 In the example, the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the property of the sound signal.
[0340] The following describes the operation of each example. Fig.14 Description of the common processing in the examples is omitted.
[0341] First, explain Fig.18 Here, the threshold data is pre-stored in the memory 1404 for each property of the direct sound. Thus, a plurality of threshold data corresponding to a plurality of properties are pre-stored in the memory 1404. And the threshold adjustment unit 1304 determines the threshold data used in the selection process of the reflected sound from the plurality of threshold data.
[0342] For example, the threshold adjustment unit 1304 obtains the property of the direct sound based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the property of the direct sound associated with the input signal. Next, the threshold adjustment unit 1304 determines a threshold corresponding to the time difference (T) and the property of the direct sound (S212).
[0343] In addition, if Fig.19 As shown, the threshold adjustment unit 1304 may adjust the threshold determined by the selection unit 1302 based on the properties of the direct sound (S221).
[0344] In any case, the input signal may include information indicating the properties of the audio signal, information for adjusting the threshold value according to the properties of the audio signal, or both. The threshold value adjustment unit 1304 may adjust the threshold value using either or both of them.
[0345] Furthermore, the information indicating the property of the sound signal, the information for adjusting the threshold value, or both of them may be transmitted via an input signal different from the input signal including the sound signal. In this case, the input signal including the sound signal may include information associated with the input signal different from the input signal, and the information associated with the input signal different from the input signal may be stored in the memory 1404 together with the information about the threshold value.
[0346] exist Fig.18 and Fig.19 In the example of , the threshold used in the selection of the reflected sound is set according to the properties of the direct sound, that is, the properties of the sound signal. Fig.18 In this way, the threshold data set in advance for each property can be used, or Fig.19 In this way, the threshold value is adjusted according to the properties of the sound signal. In addition, the parameters of the threshold value data can also be adjusted according to the properties of the sound signal.
[0347] Furthermore, the operation performed by the threshold adjustment unit 1304 may be performed by the analysis unit 1301 or the selection unit 1302. For example, the analysis unit 1301 may obtain the properties of the audio signal. Alternatively, the selection unit 1302 may set the threshold according to the properties of the audio signal.
[0348] Next, the relationship between the properties of the audio signal and the threshold value will be described.
[0349] If the time interval between two short sounds that reach the listener's ears is sufficiently short, they are heard as one sound. This phenomenon is called the precedence effect. It is known that the precedence effect occurs only for discontinuous, i.e., transient sounds (Non-Patent Document 1). Therefore, when the sound signal represents a stationary sound, the echo detection limit can also be set lower than when the sound signal represents a non-stationary sound.
[0350] That is, according to the characteristics of such a priority effect, for example, when the direct sound is a stable sound, the threshold value is set to be smaller. In addition, the higher the stability, the smaller the threshold value can be set.
[0351] An example of processing when the nature of the sound signal is stationary is described. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the stationarity based on the amount of change in the frequency component of the sound signal over time. For example, when the amount of change is small, it is determined that the stationarity is high. On the contrary, when the amount of change is large, it is determined that the stationarity is low. As a result of the determination, a flag indicating the level of stationarity can be set, and a parameter indicating stationarity can be set according to the amount of change.
[0352] Next, the threshold adjustment unit 1304 may adjust the threshold data or threshold based on information indicating the stationarity such as a flag or parameter indicating the stationarity of the sound signal, and set the adjusted threshold data or threshold as the threshold data or threshold used in the selection unit 1302 .
[0353] Alternatively, parameters for setting threshold data based on information indicating the stationarity of direct sound may be pre-stored in the memory 1404. In this case, the threshold adjustment unit 1304 may determine the stationarity of the sound signal and set the threshold data used for selecting the reflected sound based on the information indicating the stationarity and the parameters.
[0354] Alternatively, a plurality of parameters of the threshold data may be stored in advance in the memory 1404 in correspondence with a plurality of patterns of stationarity of the direct sound. In this case, the threshold adjustment unit 1304 may determine the stationarity of the sound signal, select the parameters of the threshold data based on the pattern of stationarity of the direct sound, and set the threshold data used for selecting the reflected sound based on the parameters of the threshold data.
[0355] Furthermore, the stationarity of the audio signal may be determined based on the amount of change in the frequency component of the audio signal each time the audio signal is input.
[0356] Alternatively, the stability of the sound signal may be determined based on information indicating the stability that is pre-associated with the sound signal. That is, the information indicating the stability of the sound signal may be pre-associated with the sound signal and stored in the memory 1404. The analysis unit 1301 may also obtain the information indicating the stability that is associated with the sound signal each time the sound signal is input. Furthermore, the threshold adjustment unit 1304 may adjust the threshold based on the information indicating the stability that is associated with the sound signal.
[0357] As another example of setting the threshold value according to the properties of the sound signal, when the sound signal represents a short sound (click, etc.), the application range of the echo detection limit may be set shorter than when the sound signal represents a long sound. This process is based on the characteristics of the priority effect.
[0358] It is known that due to the precedence effect, two short sounds that reach the listener's ears consecutively are heard as one sound if the time interval between them is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is about 5 ms in the case of a click, and sometimes 40 ms in the case of a complex sound such as a human voice or music (Non-Patent Document 1).
[0359] According to the characteristics of such a priority effect, for example, in the case of a sound with a shorter duration of direct sound, a threshold value of a shorter duration is set. In addition, the shorter the duration of direct sound, the shorter the threshold value of the shorter duration is set.
[0360] Setting a shorter time length threshold means setting a threshold corresponding to the echo detection limit based on the characteristic of the priority effect in a range where the time difference (T) between the direct sound and the reflected sound is small. Outside this range, the threshold corresponding to the echo detection limit based on the characteristic of the priority effect is not set. That is, outside this range, the threshold is small. Therefore, setting a shorter time length threshold for a shorter sound can correspond to setting a smaller threshold for a shorter sound.
[0361] As another example of setting the threshold value according to the properties of the direct sound, when the direct sound is intermittent sound (such as speech), the threshold value may be set lower than when the direct sound is continuous sound (such as music).
[0362] For example, when the direct sound corresponds to speech, there are repeated voiced and silent parts, and as a masking effect, only the back-masking effect occurs in the silent part. On the other hand, when the direct sound is a continuous sound such as music content, two masking effects occur, namely the back-masking effect and the simultaneous masking effect based on the sound generated at this time. Therefore, the comprehensive masking effect is higher in the case of music, etc. than in the case of speech, etc.
[0363] According to the characteristics of the masking effect as described above, the threshold value may be set higher in the case of music, etc. than in the case of speech, etc. On the contrary, the threshold value may be set lower in the case of speech, etc. than in the case of music, etc. That is, the threshold value may be set lower when there are many discontinuous parts in the direct sound.
[0364] By setting the threshold used in the selection of the reflected sound according to the property of the direct sound, the reflected sound required for auditory sense can be appropriately selected, and the auditory characteristics can be effectively reflected in the stereo sound reproduction system 1000. The process of detecting the property of the direct sound, the process of determining the threshold according to the property, and the process of adjusting the threshold according to the property can be performed during the rendering process or before starting the rendering process.
[0365] For example, these processes can be performed when the virtual space is created (when the software is created), when the virtual space processing starts (when the software is started or when rendering starts), or when the information update thread occurs periodically during the virtual space processing. In addition, when the virtual space is created, it can be the timing of constructing the virtual space before the start of the sound processing, it can be the time when the virtual space information (spatial information) is obtained, or it can be the time when the software is obtained.
[0366] (Third Modification of Threshold Setting Method)
[0367] As another example of a method for setting a threshold, the threshold may be set based on the computing resources (CPU capacity, memory resources, PC performance, or remaining battery level, etc.) used to process the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources, and sets the threshold to a high level when the amount of computing resources is small. As a result, the volume of more reflected sounds is smaller than the threshold, so the reflected sounds that are binaurally processed can be reduced, and the amount of computing can be reduced.
[0368] Alternatively, when the signal processing is performed by a battery-powered device such as a smartphone or VR goggles, it is desirable to prioritize the processing for a long time and save computing resources. In such a case, the threshold value may be set high without detecting the amount of computing resources or the remaining amount.
[0369] (Fourth Modification of Threshold Setting Method)
[0370] As another example of a method of setting the threshold, the audio signal processing device 1001 or the audio presentation device 1002 may include a threshold setting unit (not shown) so that the administrator of the virtual space or the listener may set the threshold.
[0371] For example, a listener wearing the sound prompting device 1002 may select an "energy saving mode" with less reflected sound from the listening object and less computation, or a "high performance mode" with more reflected sound from the listening object and more computation. Alternatively, an administrator who manages the stereo sound reproduction system 1000 or a producer of stereo sound content may select a mode. In addition, instead of a mode, a threshold value or threshold value data may be directly selected.
[0372] (First Modification Example of Operation of Rendering Unit)
[0373] Fig. 20 1 is a flowchart showing a first variation of the operation of the audio signal processing device 1001. Fig. 20 2 shows the processing mainly performed by the rendering unit 1300 of the audio signal processing device 1001. In this modification, a volume compensation process is added to the operation of the rendering unit 1300.
[0374] For example, the analysis unit 1301 obtains data (input signal) (S301). Then, the analysis unit 1301 analyzes the data (S302). Then, the selection unit 1302 determines whether to select the reflected sound based on the analysis result (S303). Then, the synthesis unit 1303 performs volume compensation processing based on the reflected sound that is not selected (S304). Then, the synthesis unit 1303 performs acoustic processing of the direct sound and the reflected sound (S305). And, the synthesis unit 1303 outputs the direct sound and the reflected sound as audio (S306).
[0375] Among the above-mentioned processes (S301 to S306), processes other than the volume compensation process (S304) are common to the above-mentioned other examples, and therefore their descriptions are omitted.
[0376] The volume compensation process is performed in response to the reflected sound that is not selected in the selection process. For example, by not selecting the reflected sound in the selection process, a lack of volume occurs. The volume compensation process can suppress the sense of disharmony that accompanies such a lack of volume. As examples of methods for compensating for the sense of volume, the following two methods are disclosed. Either method can be used.
[0377] First, a method of compensating for the sense of volume by increasing the volume of direct sound is described. The synthesizing unit 1303 generates direct sound by increasing the volume of direct sound by an amount corresponding to the volume of the reflected sound that is not selected. Thus, the sense of volume lost due to the non-generation of the reflected sound is compensated.
[0378] When increasing the volume, the synthesizer 1303 may increase the volume for each frequency component according to the frequency characteristics of the reflected sound. In order to perform such processing, an attenuation rate of the volume attenuated by the reflection object may be given for each specified frequency band. Thus, the frequency characteristics of the reflected sound can be derived.
[0379] Next, a method for compensating the volume perception by synthesizing the reflected sound into the direct sound is described. In this method, the synthesizer 1303 adds the non-selected reflected sound to the direct sound to generate the direct sound, thereby compensating the volume perception caused by not generating the reflected sound. The generated direct sound reflects the volume (amplitude), frequency, delay, etc. of the non-selected reflected sound.
[0380] In the case of the method of increasing the volume of the direct sound, the amount of calculation for the compensation process is very small, but only the volume is compensated. In the case of the method of synthesizing the reflected sound into the direct sound, the amount of calculation for the compensation process is larger than the method of increasing the volume of the direct sound, but the characteristics of the reflected sound are compensated more accurately.
[0381] In either case, no reflected sound is generated but only direct sound is generated, so the overall amount of calculation is reduced. In particular, the amount of calculation required for binaural processing including the convolution HRTF processing is reduced, so the overall amount of calculation is greatly reduced. The reason is that the amount of calculation required for binaural processing is much greater than the amount of calculation required for the above-mentioned compensation processing.
[0382] Furthermore, if the reason why the reflected sound is not selected is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume is not lost, the reflected sound may be simply removed without performing compensation processing.
[0383] (Second Modification Example of Operation of Rendering Unit)
[0384] Fig.21 1 is a flowchart showing a second variation of the operation of the audio signal processing device 1001. Fig.21 , mainly shows the processing performed by the rendering unit 1300 of the audio signal processing device 1001. In this modification, the left-right volume difference adjustment processing is added to the operation of the rendering unit 1300.
[0385] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the direction of arrival of the sound (S402). Next, the selection unit 1302 adjusts the difference in volume of the sound perceived by the left and right ears (S403). In addition, the selection unit 1302 adjusts the difference (delay) in the arrival time of the sound perceived by the left and right ears (S404). The selection unit 1302 determines whether to select the reflected sound based on the information of the adjusted sound (S405).
[0386] In the above-mentioned processing (S401 to S405), the processing other than the left and right volume difference adjustment (S403) and the delay adjustment (S404) is the same as that of the other examples described above, and therefore the description thereof is omitted.
[0387] Fig. 22 is a diagram showing an example of the configuration of an avatar, a sound source object, and an obstacle object. For example, when the front direction of the listener is 0 degrees, Fig. 22 As shown, if the polarity (for example, positive or negative) of the arrival direction (θ) of the direct sound and the arrival direction (γ) of the reflected sound are different, the volume difference generated between the two ears is corrected.
[0388] Specifically, when the polarities of θ and γ are different, the ears that mainly (first) perceive the sound in the direct sound and the reflected sound are different. In this case, the selection unit 1302 adjusts the volume of the direct sound to match the position of the ear that mainly perceives the reflected sound as the left and right volume difference adjustment (S403). For example, the selection unit 1302 multiplies the volume of the direct sound when it reaches the listener by (1.0-0.3sin(θ)) (0≤θ≤180), thereby attenuating the volume of the direct sound when it reaches the listener.
[0389] The selection unit 1302 determines whether to select the reflected sound by calculating the volume ratio of the volume of the direct sound and the volume of the reflected sound corrected as described above, and comparing the calculated volume ratio with the threshold value. As a result, the volume difference generated between the two ears is corrected, the volume of the direct sound that affects the reflected sound is more accurately derived, and the determination of whether to select the reflected sound is more accurately performed.
[0390] In addition to the left-right volume difference adjustment (S403), the selection unit 1302 may also delay the arrival time of the direct sound as a delay adjustment (S404) to match the position of the ear that perceives the reflected sound. Specifically, the selection unit 1302 may delay the arrival time of the direct sound by adding (a(sinθ+θ) / c) ms (a is the radius of the head, and c is the speed of sound) to the arrival time of the direct sound.
[0391] (Third Modification Example of Operation of Rendering Unit)
[0392] A method of setting a threshold value according to the arrival direction is described.
[0393] Fig.23 is a flowchart showing yet another example of selection processing. Fig.14 The common processing in the examples is omitted. Fig.23 In the example of , the selection unit 1302 selects the reflected sound using a threshold value corresponding to the arrival direction.
[0394] Specifically, the selection unit 1302 calculates the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) determined by using the orientation of the avatar as a reference, based on the direct sound arrival path (pd), the reflected sound arrival path (pr) and the orientation information D of the avatar calculated by the analysis unit 1301. That is, the selection unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. The orientation information D of the avatar may also be included in the input signal.
[0395] The selection unit 1302 uses three indices including the time difference (T) in addition to the direct sound arrival direction (θ) and the reflected sound arrival direction (γ), and selects the sound according to the following formula: Fig.15The three-dimensional arrangement shown is used to determine the threshold value used in the selection process (S232).
[0396] As an example, Fig. 22 The method of setting the threshold used in the selection process when an avatar, a sound source object, and an obstacle object are arranged is shown.
[0397] The position information of the avatar, the sound source object, and the obstacle object, and the avatar's orientation information D are obtained from the input signal. Using these position information and orientation information D, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound (γ) are calculated when the avatar's orientation is set to 0 degrees. Fig. 22 In the case of , the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees).
[0398] Next, refer to Fig.15 The threshold data is stored in a three-dimensional arrangement as shown, and the threshold is determined from the arrangement area corresponding to the values of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. In the case where there is no index corresponding to the calculated values of (θ), (γ), and (T), the threshold corresponding to the closest index may be determined.
[0399] As another method, the threshold value may be determined by performing interpolation, extrapolation, or the like based on one or more threshold values corresponding to one or more indices close to the calculated values of (θ), (γ), and (T). For example, the threshold value corresponding to (20 degrees, 265 degrees, T) may be determined based on four threshold values corresponding to four indices, namely, (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).
[0400] The selection process based on the difference between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound will be described.
[0401] For example, you can also pre-make and set Fig.16 The angle difference (Φ) between the arrival direction (θ) of the direct sound and the arrival direction (γ) of the reflected sound and the time difference (T) are arranged as two-dimensional indexes to have threshold data. In this case, the angle difference (Φ) and the time difference (T) are referred to in the selection process. Alternatively, the angle difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound can be calculated in the selection process, and the calculated angle difference (Φ) can be used to determine the threshold.
[0402] Alternatively, threshold data may be set that has a combination of the angle difference (Φ), the arrival direction of the direct sound (θ) and the time difference (T), or a combination of the angle difference (Φ), the arrival direction of the reflected sound (γ) and the time difference (T) arranged as an index.
[0403] Alternatively, you can set Fig.15 The values of (θ), (γ), and (T) are shown as threshold value data arranged as three-dimensional indices.
[0404] (Fourth Modification Example of Operation of Rendering Unit)
[0405] The processing performed by the above-mentioned analyzing unit 1301, selecting unit 1302, and synthesizing unit 1303 may be performed as pipeline processing as described in Patent Document 3, for example.
[0406] Fig.24 13 is a block diagram showing a configuration example for the rendering unit 1300 to perform pipeline processing.
[0407] Fig.24 The rendering unit 1300 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may also be Figure 7 The rendering unit 1300 shown in FIG. 1 may be composed of multiple components, or may be composed of Figure 5 The audio signal processing device 1001 shown in the figure is constituted by at least a part of the multiple components.
[0408] Pipeline processing means dividing the processing for providing an acoustic effect into a plurality of processes and executing the plurality of processes in sequence one by one. In each of the plurality of processes, for example, signal processing of a sound signal or generation of parameters used in the signal processing is executed.
[0409] The rendering unit 1300 may also perform reverberation processing, initial reflection processing, distance attenuation processing, binaural processing, etc. as pipeline processing. However, these processing are examples, and pipeline processing may include processing other than these, or may not include some processing. For example, pipeline processing may also include diffraction processing and occlusion processing. In addition, for example, reverberation processing may be omitted when it is not necessary.
[0410] In addition, each process may be represented as a stage. In addition, the result of each process and the generated sound signal such as reflected sound may be represented as a rendering item. The multiple stages in the pipeline processing and their order are not limited to Fig.24 Example shown.
[0411] Here, the parameters used in the selection process (arrival path, arrival time, and volume ratio related to direct sound and reflected sound) may also be calculated in one of the multiple stages used to generate the rendering item. That is, the parameters used in the selection of the reflected sound are calculated in a part of the pipeline processing used to generate the rendering item. In addition, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted or performed outside the rendering unit 1300.
[0412] The following describes the reverberation process, initial reflection process, distance attenuation process, selection process, generation process, and binaural process that can be included as stages in the pipeline process. In each stage, metadata included in the input signal can also be analyzed to calculate parameters used in the generation of reflected sound.
[0413] In the reverberation processing, the reverberation processing unit 1311 generates a sound signal representing a reverberation sound or a parameter used in generating a sound signal. The reverberation sound refers to the sound that reaches the listener as a reverberation after the direct sound. As an example, the reverberation sound is the sound that reaches the listener after being reflected more times (e.g., several dozen times) than the initial reflected sound at a relatively late stage (e.g., from the arrival of the direct sound to about one hundred and several dozen ms) after the initial reflected sound described later reaches the listener.
[0414] The reverberation processing unit 1311 refers to the audio signal and space information included in the input signal, and calculates the reverberation sound using a predetermined function prepared in advance as a function for generating the reverberation sound.
[0415] The reverberation processing unit 1311 may also generate reverberation sound by applying a known reverberation generation method to the sound signal included in the input signal. An example of a known reverberation generation method is the Schroeder method, but the known reverberation generation method is not limited to the Schroeder method. In addition, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by the spatial information in the application of the known reverberation generation method. Thus, the reverberation processing unit 1311 can calculate parameters for generating reverberation sound.
[0416] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating initial reflection sound based on the spatial information. The initial reflection sound is the reflection sound that reaches the listener after one or more reflections at a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct sound) after the direct sound reaches the listener from the sound source object.
[0417] The initial reflection processing unit 1312 calculates the path of the reflected sound from the sound source object to the listener by reflecting from the reflection object, for example, with reference to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflection object such as a structure, and the reflectivity of the reflection object may be used in the calculation of the path.
[0418] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct sound. Information on the path may be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected sound, or may be used as a parameter for the selection unit 1314 to select the reflected sound.
[0419] In the distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct sound and the reflected sound reaching the listener based on the length of the path of the direct sound and the reflected sound. The volume of the direct sound and the reflected sound reaching the listener is attenuated in proportion to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct sound by dividing the volume of the sound source by the length of the path of the direct sound, and can calculate the volume of the reflected sound by dividing the volume of the sound source by the length of the path of the reflected sound.
[0420] In the selection process, the selection unit 1314 selects the generation of the object reflected sound based on the parameters calculated before the selection process. In the selection of the generation of the object reflected sound, a certain selection method of the present disclosure may be used.
[0421] The selection process can be performed on all reflected sounds, or it can be performed only on reflected sounds with high evaluation values based on the evaluation process as described above. That is, for reflected sounds with low evaluation values, they are determined to be not selected without performing the selection process. For example, for reflected sounds with very low volume, it can also be regarded as having a low evaluation value of the reflected sound and determined to be not selected.
[0422] Furthermore, for example, the selection process may be performed on all reflected sounds. Also, the evaluation value of the reflected sound selected in the selection process may be determined, and the reflected sound having a low evaluation value may be re-determined as not selected.
[0423] In the generation process, the generator 1315 generates direct sound and reflected sound. For example, the generator 1315 generates direct sound based on the arrival time and volume of the direct sound according to the sound signal included in the input signal. In addition, the generator 1315 generates reflected sound based on the arrival time and volume of the reflected sound according to the sound signal included in the input signal for the reflected sound selected in the selection process.
[0424] In binaural processing, the binaural processing unit 1316 performs signal processing so that the sound signal of the direct sound is perceived as sound reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing so that the reflected sound selected by the selection unit 1314 is perceived as sound reaching the listener from the reflection object.
[0425] For example, the binaural processing unit 1316 performs processing using the HRIR DB based on the position and orientation of the listener in the sound space so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.
[0426] In addition, HRIR (Head-Related Impulse Responses) is the response characteristic when an impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from the frequency domain to the time domain through Fourier transform. The head-related transfer function expresses the change of the sound generated by the surrounding objects including the auricle, the head and shoulders as a transfer function. HRIR DB is a database containing such information.
[0427] In addition, the position and orientation of the listener in the sound space are, for example, the position and orientation of the virtual listener in the virtual sound space. Alternatively, the position and orientation of the virtual listener in the virtual sound space may change in accordance with the movement of the listener's head. In addition, the position and orientation of the virtual listener in the virtual sound space may also be determined based on information obtained from the sensor 1405.
[0428] The program, spatial information, HRIR DB, threshold data, other parameters, and the like used in the above-mentioned processing are acquired from the memory 1404 included in the sound signal processing device 1001 or from outside the sound signal processing device 1001 .
[0429] In addition, pipeline processing may include other processing. Furthermore, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in the pipeline processing. For example, the rendering unit 1300 may include a diffraction processing unit and an occlusion processing unit.
[0430] The diffraction processing unit performs processing for generating a sound signal representing a sound including diffracted sound caused by an obstacle object between a listener and a sound source object in a three-dimensional sound field (space). The diffracted sound is a sound that, when an obstacle object exists between the sound source object and the listener, bypasses the obstacle object and reaches the listener from the sound source object.
[0431] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object to the listener around the obstacle object, and generates the diffracted sound based on the path. In calculating the path, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), and the shape and size of the obstacle object may also be used.
[0432] When a sound source object exists on the opposite side of the obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on the spatial information and information such as the material of the obstacle object.
[0433] (Example of sound source object)
[0434] In the above description, the position information given to the sound source object indicates a "point" in the virtual space as the position of the sound source object. That is, in the above description, the sound source is defined as a "point sound source".
[0435] On the other hand, the sound source in the virtual space can also be defined as an object with length, size, shape, etc., that is, a sound source that is defined as a non-point sound source and extends in space. In this case, the distance between the listener and the sound source and the direction of arrival of the sound are uncertain. Therefore, the reflected sound caused by such a sound source does not need to be analyzed by the analysis unit 1301, or is limited to being selected by the selection unit 1302 regardless of the analysis result. In this way, it is possible to avoid the degradation of sound quality that may occur due to not selecting the reflected sound.
[0436] Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied assuming that the sound is generated from the representative point. In this case, the threshold may be adjusted based on spatial extension information of the sound source.
[0437] (Examples of direct sound and reflected sound)
[0438] For example, direct sound is sound that is not reflected by a reflective object, and reflected sound is sound that is reflected by a reflective object. Direct sound can also be sound that reaches the listener from the sound source without being reflected by a reflective object, and reflected sound can also be sound that reaches the listener from the sound source after being reflected by a reflective object.
[0439] Furthermore, direct sound and reflected sound are not limited to sounds reaching the listener, but may be sounds before reaching the listener. For example, direct sound may be sound output from a sound source, or in other words, sound of the sound source.
[0440] Fig.25 It is a diagram showing the transmission and diffraction of sound. Fig.25 As shown in FIG. 1 , sometimes the direct sound does not reach the listener because there is an obstacle object between the sound source object and the listener. In this case, the sound emitted from the sound source object, transmitted through the obstacle object and reached the listener can also be regarded as the direct sound. In addition, the sound emitted from the sound source object, diffracted through the obstacle object and reached the listener can also be regarded as the reflected sound.
[0441] In addition, the two sounds compared in the selection process are not limited to the direct sound and the reflected sound based on the sound emitted by a sound source. For example, the sound selection can also be performed by comparing two reflected sounds based on the sound emitted by a sound source. In this case, the direct sound in the present disclosure can also be replaced by the sound that reaches the listener first, and the reflected sound in the present disclosure can also be replaced by the sound that reaches the listener later.
[0442] (Bitstream Structure Example)
[0443] The bitstream includes, for example, an audio signal and metadata. The audio signal is audio data that represents the sound, and indicates information related to the frequency and strength of the sound, etc. In addition, the metadata includes spatial information related to the sound space, that is, the sound field space.
[0444] For example, spatial information is information about the space where a listener who listens to a sound based on a sound signal is located. Specifically, spatial information is information related to a predetermined position (localization position) for localizing a sound image in a sound space (e.g., a three-dimensional sound field), that is, for allowing a listener to perceive a sound coming from a direction corresponding to the predetermined position. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
[0445] The sound source object information is information about a sound source object that generates sound based on a sound signal. That is, the sound source object information is information about an object (sound source object) that reproduces a sound signal, and is information about a virtual sound source object configured in a virtual sound space. Here, the virtual sound space may also correspond to a real space where an object that generates sound is configured, and the sound source object in the virtual sound space may also correspond to an object that generates sound in the real space.
[0446] The sound source object information may also indicate the position of the sound source object arranged in the sound space, the direction of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living thing, whether the sound source object is a moving body, etc. For example, a sound signal is associated with one or more sound source objects indicated by the sound source object information.
[0447] The bit stream has a data structure composed of, for example, metadata (control information) and an audio signal.
[0448] The audio signal and metadata may be included in one bit stream or in multiple bit streams. In addition, the audio signal and metadata may be included in one file or in multiple files.
[0449] The bit stream may exist for each sound source or for each playback time. When the bit stream exists for each playback time, a plurality of bit streams may be processed in parallel at the same time.
[0450] The metadata may be assigned to each bit stream, or may be assigned to multiple bit streams together as information for controlling the multiple bit streams. In this case, the metadata may be shared by multiple bit streams. In addition, the metadata may be assigned for each playback time.
[0451] When there are multiple bitstreams or multiple files, information indicating related bitstreams or related files may be included in more than one bitstream or more than one file. Alternatively, information indicating related bitstreams or related files may be included in all bitstreams or all files.
[0452] Here, the associated bitstream or associated file refers to, for example, a bitstream or file that may be used simultaneously during audio processing. In addition, a bitstream or file in which information indicating the associated bitstream or associated file is recorded together may also be included.
[0453] Here, the information indicating the associated bitstream or the associated file may be, for example, an identifier indicating the associated bitstream or the associated file. In addition, the information indicating the associated bitstream or the associated file may be, for example, a file name, a URL (Uniform Resource Locator) or a URI (Uniform Resource Identifier) indicating the associated bitstream or the associated file.
[0454] In this case, the acquisition unit may also determine and acquire the associated bitstream or associated file based on the information indicating the associated bitstream or associated file. In addition, the information indicating the associated bitstream or associated file may be included in a bitstream or file, and the information indicating the associated bitstream or associated file may be included in another bitstream or other file.
[0455] Here, the file including information indicating the associated bitstream or the associated file may be, for example, a control file such as a declaration file used for content distribution.
[0456] In addition, all or part of the metadata may be obtained from outside the bitstream of the audio signal. For example, metadata for controlling the audio or metadata for controlling the video may be obtained from outside the bitstream, or both metadata may be obtained from outside the bitstream.
[0457] Furthermore, metadata for controlling images may be included in the bit stream obtained by the stereoscopic sound reproduction system 1000. In this case, the stereoscopic sound reproduction system 1000 may output metadata for controlling images to a display device that displays images or a stereoscopic image reproduction device that reproduces stereoscopic images.
[0458] (Example of information included in metadata)
[0459] Metadata may be information used to describe a scene represented by a sound space. Here, a scene is a term that refers to a collection of all elements of three-dimensional images and sound events in a sound space modeled by a sound signal reproduction system using metadata.
[0460] That is, metadata may include not only information for controlling audio processing but also information for controlling video processing. Metadata may include only one of the information for controlling audio processing and the information for controlling video processing, or both.
[0461] The stereo sound reproduction system 1000 generates virtual sound effects by performing sound processing on sound signals using metadata included in a bitstream and interactive listener position information obtained by addition. The sound effects may include initial reflection processing, obstacle processing, diffraction processing, occlusion processing, and reverberation processing, and other sound processing may be performed using metadata. For example, sound effects such as distance attenuation effect, localization, or Doppler effect may be added.
[0462] Furthermore, information on switching on / off of all or part of the additional acoustic effects or priority information on a plurality of processes for the acoustic effects may be added to the metadata.
[0463] In addition, as an example, metadata includes information related to the sound space including the sound source object and the obstacle object, and information related to the localization position used to localize the sound image at a specified position in the sound space (i.e., to make the listener perceive the sound coming from a specified direction).
[0464] Here, an obstacle object is an object that may affect the sound perceived by the listener, such as blocking or reflecting the sound, during the period from the sound emitted by the sound source object to the sound reaching the listener. In addition to stationary objects, obstacle objects may also include moving objects such as animals or machinery. Animals may also be humans, etc.
[0465] Furthermore, when there are multiple sound source objects in the sound space, other sound source objects may become obstacle objects for any sound source object. That is, objects that do not make sound, such as building materials or inanimate objects, i.e., non-sound-generating objects, and sound source objects that make sound may all become obstacle objects.
[0466] The metadata includes all or part of information indicating the shape of the sound space, the shape and position of obstacle objects in the sound space, the shape and position of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
[0467] The sound space may be a closed space or an open space. In addition, the metadata may also include information indicating the reflectivity of an obstacle object that can reflect sound in the sound space. For example, a floor, a wall, or a ceiling that constitutes a boundary of the sound space may also be an obstacle object.
[0468] The reflectivity is the energy ratio of the reflected sound to the incident sound, and can also be set for each frequency band of the sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, for example, uniformly set parameters such as the attenuation rate, diffracted sound, and initial reflected sound can also be used.
[0469] The metadata may also include information other than reflectivity as a parameter related to an obstacle object or a sound source object. For example, the metadata may also include information related to the material of the object as a parameter related to both the sound source object and the non-sound-emitting object. Specifically, the metadata may also include information such as diffusivity, transmittance, and sound absorption coefficient.
[0470] The information about the sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information indicating the sound source area in the object. The reproduction conditions may also be set to be, for example, whether the sound is continuously flowing or the sound is triggered by an event. The sound source area in the object may be set based on the relative relationship between the position of the listener and the position of the object, or may be set using the object as a reference.
[0471] For example, when the sound source area is set based on the relative relationship between the listener's position and the object's position, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.
[0472] Furthermore, when the sound source area is set using the object as a reference, it is possible to use the object as a reference to fix which sound is emitted from which area of the object. For example, when the listener is looking at the object from the front, the listener can perceive high tones from the right side of the object and low tones from the left side of the object. Also, when the listener is looking at the object from the back, the listener can perceive low tones from the right side of the object and high tones from the left side of the object.
[0473] Metadata related to the space may include time until initial reflected sound, reverberation time, and the ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the listener can perceive only direct sound.
[0474] (Replenish)
[0475] In addition, the aspects grasped based on the present disclosure are not limited to the embodiments, and various modifications can be made and implemented.
[0476] For example, the processing executed by a specific component in the embodiment may be executed by another component instead of the specific component. In addition, the order of a plurality of processing may be changed, or a plurality of processing may be executed in parallel.
[0477] In addition, the ordinal numbers such as 1 and 2 used in the description may be replaced, removed, or newly assigned as appropriate. These ordinal numbers do not necessarily correspond to a meaningful order, but may be used to identify elements.
[0478] In addition, for example, in the comparison of the threshold, above the threshold and greater than the threshold can be replaced with each other. Similarly, below the threshold and less than the threshold can be replaced with each other. In addition, for example, time and moment can be replaced with each other.
[0479] Furthermore, in the process of selecting one or more processing target sounds from a plurality of sounds, if there is no sound satisfying the condition, none of the sounds may be selected as the processing target sounds. That is, in the process of selecting one or more processing target sounds from a plurality of sounds, the case of not selecting the processing target sound may also be included.
[0480] Furthermore, for example, an expression such as at least one of the first element, the second element, and the third element may correspond to the first element, the second element, the third element, or any combination thereof.
[0481] In addition, for example, the embodiments describe the case where the form grasped based on the present disclosure is implemented as an audio processing device, an encoding device, or a decoding device. However, the form grasped based on the present disclosure is not limited to these, and can also be implemented as software for executing the audio processing method, encoding method, or decoding method.
[0482] For example, a program for executing the above-mentioned sound processing method, encoding method, or decoding method may be stored in advance in the ROM, and the CPU may operate according to the program.
[0483] In addition, a program for executing the above-mentioned sound processing method, encoding method or decoding method may be stored in a computer-readable recording medium. Furthermore, the computer may record the program stored in the recording medium into the RAM of the computer and operate according to the program.
[0484] Furthermore, each of the above-mentioned components can be typically implemented as an integrated circuit (LSI) having input terminals and output terminals. They can be formed into one chip individually, or can be formed into one chip in a manner that includes all or part of the components of the implementation mode. LSI can also be expressed as IC, system LSI, super LSI or ultra-large-scale LSI according to the difference in integration.
[0485] In addition, it is not limited to LSI, and a dedicated circuit or a general-purpose processor can also be used. In addition, an FPGA that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connection or setting of the circuit unit inside the LSI can also be used. Furthermore, if a technology for integrated circuitization that replaces LSI appears due to the progress of semiconductor technology or other derived technologies, then of course, this technology can also be used to integrate the components. It may be an application of biotechnology, etc.
[0486] In addition, the FPGA or CPU may download all or part of the software for implementing the sound processing method, encoding method or decoding method described in the present disclosure through wireless communication or wired communication. Furthermore, the whole or part of the software for updating may be downloaded through wireless communication or wired communication. Furthermore, the digital signal processing described in the present disclosure may be performed by storing the downloaded software in a memory by the FPGA or CPU and operating based on the stored software.
[0487] At this time, the device with FPGA or CPU etc. can also be connected to the signal processing device wirelessly or by wire, or can be connected to the signal processing server via a network. Furthermore, the device and the signal processing device or the signal processing server can also perform the audio processing method, encoding method or decoding method described in the present disclosure.
[0488] For example, the sound processing device, encoding device, or decoding device of the present disclosure may also include an FPGA or a CPU, etc. Furthermore, the sound processing device, encoding device, or decoding device may also include an interface for obtaining software for operating the FPGA or the CPU, etc. from the outside, and a memory for storing the obtained software. Furthermore, the FPGA or the CPU, etc. may also perform the signal processing described in the present disclosure by operating based on the stored software.
[0489] Alternatively, the server may provide software related to the audio processing, encoding processing, or decoding processing of the present disclosure. Furthermore, the terminal or device may operate as the audio processing device, encoding device, or decoding device described in the present disclosure by installing the software. Alternatively, the terminal or device may be connected to the server via a network to install the software.
[0490] In addition, another device different from the terminal or device may be connected to the server via a network to obtain data for software installation, and the software may be installed in the terminal or device by providing the software installation data to the terminal or device. In addition, an example of software may also be VR software or AR software for causing the terminal or device to execute the audio processing method described in the embodiment.
[0491] In addition, in the above-mentioned embodiments, each component may be formed by dedicated hardware, or implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.
[0492] In the above, the device and the like of one or more forms are described based on the embodiments, but the forms grasped by the present disclosure are not limited to the embodiments. As long as it does not depart from the main purpose of the present disclosure, the forms obtained by applying various modifications to the embodiments that can be thought of by those skilled in the art, and the forms constructed by combining the constituent elements in different modified examples are also included in the scope of one or more forms.
[0493] (Note)
[0494] The following techniques are disclosed through the description of the above embodiments.
[0495] (Technology 1) An audio processing device comprising a circuit and a memory; the circuit uses the memory to obtain sound space information related to a sound space; based on the sound space information, obtains characteristics related to a first sound generated from a sound source in the sound space; based on the characteristics related to the first sound, controls whether to select a second sound generated in the sound space corresponding to the first sound.
[0496] (Technique 2) In the sound processing device according to Technique 1, the first sound is direct sound, and the second sound is reflected sound.
[0497] (Technology 3) In the sound processing device as described in Technology 2, the characteristic related to the first sound is the volume ratio of the direct sound to the volume of the reflected sound; the circuit calculates the volume ratio based on the sound space information; and controls whether to select the reflected sound based on the volume ratio.
[0498] (Technique 4) In the sound processing device according to Technique 3, when the reflected sound is selected, the circuit generates sounds that reach both ears of the listener by applying binaural processing to the reflected sound and the direct sound.
[0499] (Technology 5) In the audio processing device as described in Technology 3 or 4, the circuit calculates the time difference between the end time of the direct sound and the arrival time of the reflected sound based on the sound space information; and controls whether to select the reflected sound based on the time difference and the volume ratio.
[0500] (Technique 6) In the sound processing device as described in Technique 5, the circuit selects the reflected sound when the volume ratio is above a threshold value; and the first threshold value used as the threshold value when the time difference is a first value is greater than the second threshold value used as the threshold value when the time difference is a second value larger than the first value.
[0501] (Technique 7) In the audio processing device as described in Technique 3 or 4, the circuit calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound based on the sound space information; and controls whether to select the reflected sound based on the time difference and the volume ratio.
[0502] (Technique 8) In the sound processing device as described in Technique 7, the circuit selects the reflected sound when the volume ratio is above a threshold value; and the first threshold value used as the threshold value when the time difference is a first value is greater than the second threshold value used as the threshold value when the time difference is a second value larger than the first value.
[0503] (Technique 9) In the sound processing device according to Technique 8, the circuit adjusts the threshold value based on the arrival direction of the direct sound and the arrival direction of the reflected sound.
[0504] (Technique 10) In the sound processing device according to any one of Techniques 2 to 9, when the reflected sound is not selected, the circuit corrects the volume of the direct sound based on the volume of the reflected sound.
[0505] (Technique 11) In the sound processing device according to any one of Techniques 2 to 9, when the reflected sound is not selected, the circuit synthesizes the reflected sound with the direct sound.
[0506] (Technique 12) In the sound processing device according to any one of Techniques 3 to 9, the volume ratio is a volume ratio of the volume of the direct sound at a first time and the volume of the reflected sound at a second time different from the first time.
[0507] (Technique 13) In the sound processing device according to Technique 1 or 2, the circuit sets a threshold value based on characteristics related to the first sound, and controls whether to select the second sound based on the threshold value.
[0508] (Technique 14) In the sound processing device as described in any one of Techniques 1, 2 and 13, the characteristic related to the first sound is one or a combination of two or more of the volume of the sound source, the visibility of the sound source and the localization of the sound source.
[0509] (Technique 15) In the sound processing device according to any one of Techniques 1, 2, and 13, the characteristic related to the first sound is a frequency characteristic of the first sound.
[0510] (Technique 16) In the sound processing device according to any one of Techniques 1, 2, and 13, the characteristic related to the first sound is a characteristic indicating discontinuity of amplitude of the first sound.
[0511] (Technique 17) In the sound processing device according to any one of Techniques 1, 2, 13 and 16, the characteristic related to the first sound is a characteristic indicating a duration of a voiced part of the first sound or a duration of a silent part of the first sound.
[0512] (Technique 18) In the sound processing device described in any one of Techniques 1, 2, 13, 16 and 17, the characteristic related to the first sound is a characteristic in which the duration of the voiced part of the first sound and the duration of the unvoiced part of the first sound are expressed in a time series.
[0513] (Technique 19) In the sound processing device according to any one of Techniques 1, 2, 13, and 15, the characteristic related to the first sound is a characteristic indicating a change in a frequency characteristic of the first sound.
[0514] (Technique 20) In the sound processing device according to any one of Techniques 1, 2, 13, 15 and 19, the characteristic related to the first sound is a characteristic indicating the smoothness of the frequency characteristic of the first sound.
[0515] (Technique 21) The sound processing device according to any one of Techniques 1, 2, and 13 to 20, wherein the characteristic related to the first sound is obtained from the bit stream.
[0516] (Technique 22) In the sound processing device described in any one of Techniques 1, 2, and 13 to 21, the circuit calculates characteristics related to the second sound; and controls whether to select the second sound based on the characteristics related to the first sound and the characteristics related to the second sound.
[0517] (Technology 23) In the sound processing device as described in Technology 22, the circuit obtains a threshold value indicating the volume corresponding to the boundary of whether the sound can be heard; based on the characteristics related to the first sound, the characteristics related to the second sound and the threshold, controls whether the second sound is selected.
[0518] (Technique 24) In the sound processing device according to Technique 23, the characteristic related to the second sound is the volume of the second sound.
[0519] (Technique 25) In the sound processing device as described in Technique 1 or 2, the sound space information includes information on the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; and the circuit controls whether to select each of the plurality of second sounds based on characteristics related to the first sound, thereby selecting one or more processing object sounds to which binaural processing is applied from the first sound and the plurality of second sounds.
[0520] (Technique 26) In the sound processing device described in any one of Techniques 1 to 25, the timing for obtaining the characteristics related to the first sound is at least one of when the sound space is created, when the processing of the sound space starts, and when an information update thread occurs during the processing of the sound space.
[0521] (Technique 27) The sound processing device according to any one of Techniques 1 to 26, wherein the characteristics related to the first sound are periodically acquired after the processing of the sound space starts.
[0522] (Technique 28) In the audio processing device as described in Technique 1 or 2, the characteristic related to the first sound is the volume of the first sound; the circuit calculates an evaluation value of the second sound based on the volume of the first sound; and based on the evaluation value, controls whether to select the second sound.
[0523] (Technique 29) In the sound processing device according to Technique 28, the volume of the first sound has a transition.
[0524] (Technique 30) In the sound processing device according to Technique 28 or 29, the circuit calculates the evaluation value so that the second sound is more likely to be selected as the volume of the first sound is louder.
[0525] (Technique 31) In the sound processing device as described in Technique 1 or 2, the sound space information is scene information including information about the sound source in the sound space and information about the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit obtains a signal of the first sound; calculates the plurality of second sounds based on the scene information and the signal of the first sound; obtains characteristics related to the first sound from the information about the sound source; and controls whether to select each of the plurality of second sounds as a sound to which binaural processing is not applied based on the characteristics related to the first sound, thereby selecting one or more second sounds to which the binaural processing is not applied from the plurality of second sounds.
[0526] (Technique 32) In the sound processing device according to Technique 31, the scene information is updated based on input information; and the characteristics related to the first sound are obtained according to the update of the scene information.
[0527] (Technique 33) The sound processing device as described in Technique 31 or 32 obtains the scene information and the characteristics related to the first sound from metadata included in the bit stream.
[0528] (Technique 34) A sound processing method, comprising: a step of obtaining sound space information related to a sound space; a step of obtaining characteristics related to a first sound generated from a sound source in the sound space based on the sound space information; and a step of controlling whether to select a second sound generated in the sound space corresponding to the first sound based on the characteristics related to the first sound.
[0529] (Technique 35) A program for causing a computer to execute the sound processing method described in Technique 34.
[0530] Industrial Applicability
[0531] The present disclosure includes, for example, aspects that can be applied to an audio processing device, an encoding device, a decoding device, or a terminal or device including any of these devices.
[0532] Description of symbols
[0533] 1000 stereo sound reproduction system
[0534] 1001 Sound signal processing device (sound processing device)
[0535] 1002 Sound prompt device
[0536] 1100, 1120, 1500 encoding device
[0537] 1101, 1113 Input data
[0538] 1102 Encoder
[0539] 1103 Encoded Data
[0540] 1104, 1114, 1404, 1503 memory
[0541] 1110, 1130 Decoding device
[0542] 1111 Sound signal
[0543] 1112, 1200, 1210 decoders
[0544] 1121 Sending Department
[0545] 1122 Send signal
[0546] 1131 Receiving Department
[0547] 1132 Receiving signal
[0548] 1201, 1211 Space Information Management Department
[0549] 1202 Sound Data Decoder
[0550] 1203, 1213, 1300 Rendering Department
[0551] 1301 Analysis Department
[0552] 1302, 1314 Selection Department
[0553] 1303 Synthesis Department
[0554] 1304 Threshold Adjustment Unit
[0555] 1311 Reverberation Processing Unit
[0556] 1312 Initial reflection processing unit
[0557] 1313 Distance Attenuation Processing Unit
[0558] 1315 Generation Department
[0559] 1316 Binaural Processing Department
[0560] 1401 Speaker
[0561] 1402, 1501 processors
[0562] 1403, 1502 communication IF
[0563] 1405 Sensor
Claims
1. A sound processing device, wherein: A circuit and a memory are provided, The circuit uses the memory, Acquiring sound space information related to the sound space, acquiring characteristics related to a first sound generated from a sound source in the sound space based on the sound space information, Based on the characteristics related to the first sound, whether to select the second sound generated in the sound space corresponding to the first sound is controlled.
2. The sound processing device according to claim 1, wherein: The first sound is a direct sound. The second sound is a reflected sound.
3. The sound processing device according to claim 2, wherein: The characteristic related to the first sound is a volume ratio of the volume of the direct sound to the volume of the reflected sound, The circuit, Based on the sound space information, calculating the volume ratio, Based on the volume ratio, whether to select the reflected sound is controlled.
4. The sound processing device according to claim 3, wherein: When the reflected sound is selected, the circuit generates sounds that reach both ears of the listener by applying binaural processing to the reflected sound and the direct sound.
5. The sound processing device according to claim 3 or 4, wherein: The circuit, Based on the sound space information, the time difference between the end time of the direct sound and the arrival time of the reflected sound is calculated, Based on the time difference and the volume ratio, whether to select the reflected sound is controlled.
6. The sound processing device according to claim 5, wherein: The circuit selects the reflected sound when the volume ratio is greater than a threshold value, A first threshold value used as the threshold value when the time difference is a first value is larger than a second threshold value used as the threshold value when the time difference is a second value larger than the first value.
7. The sound processing device according to claim 3 or 4, wherein: The circuit, Based on the sound space information, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is calculated, Based on the time difference and the volume ratio, whether to select the reflected sound is controlled.
8. The sound processing device according to claim 7, wherein: The circuit selects the reflected sound when the volume ratio is greater than a threshold value, A first threshold value used as the threshold value when the time difference is a first value is larger than a second threshold value used as the threshold value when the time difference is a second value larger than the first value.
9. The sound processing device according to claim 8, wherein: The circuit adjusts the threshold based on an arrival direction of the direct sound and an arrival direction of the reflected sound.
10. The sound processing device according to any one of claims 2 to 4, wherein: When the reflected sound is not selected, the circuit corrects the volume of the direct sound based on the volume of the reflected sound.
11. The sound processing device according to any one of claims 2 to 4, wherein: In a case where the reflected sound is not selected, the circuit synthesizes the reflected sound into the direct sound.
12. The sound processing device according to claim 3 or 4, wherein: The volume ratio is a volume ratio between the volume of the direct sound at a first time and the volume of the reflected sound at a second time different from the first time.
13. The sound processing device according to claim 1, wherein: The circuit sets a threshold value based on characteristics related to the first sound, and controls whether to select the second sound based on the threshold value.
14. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is one or a combination of two or more of the volume of the sound source, the visibility of the sound source, and the localization of the sound source.
15. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a frequency characteristic of the first sound.
16. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a characteristic indicating the discontinuity of the amplitude of the first sound.
17. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a characteristic indicating a duration of a voiced part of the first sound or a duration of a silent part of the first sound.
18. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a characteristic that represents the duration of the voiced part of the first sound and the duration of the unvoiced part of the first sound in a time series.
19. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a characteristic indicating a change in a frequency characteristic of the first sound.
20. The sound processing device according to claim 1, wherein: The characteristic related to the first sound is a characteristic indicating the smoothness of the frequency characteristic of the first sound.
21. The sound processing device according to any one of claims 1, 2, and 13 to 20, wherein: A characteristic related to the first sound is obtained from the bit stream.
22. The sound processing device according to any one of claims 1, 2, and 13 to 20, wherein: The circuit, calculating a characteristic related to the second sound, Whether or not to select the second sound is controlled based on the characteristics related to the first sound and the characteristics related to the second sound.
23. The sound processing device according to claim 22, wherein: The circuit, A threshold value indicating the volume corresponding to the boundary of whether or not the sound can be heard is obtained. Whether or not to select the second sound is controlled based on the characteristic related to the first sound, the characteristic related to the second sound, and the threshold.
24. The sound processing device according to claim 23, wherein: The characteristic related to the second sound is the volume of the second sound.
25. The sound processing device according to claim 1 or 2, wherein: The sound space information includes information about the position of the listener in the sound space, The second sound is a plurality of second sounds each generated in the sound space in response to the first sound. The circuit controls whether to select each of the plurality of second sounds based on the characteristics related to the first sound, thereby selecting one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds.
26. The sound processing device according to claim 1 or 2, wherein: The timing for acquiring the characteristic related to the first sound is at least one of when the sound space is created, when processing of the sound space starts, and when an information update thread occurs during processing of the sound space.
27. The sound processing device according to claim 1 or 2, wherein: After the processing of the sound space is started, the characteristics related to the first sound are periodically acquired.
28. The sound processing device according to claim 1, wherein: the characteristic related to the first sound is the volume of the first sound, The circuit, calculating an evaluation value of the second sound based on the volume of the first sound, Based on the evaluation value, whether or not to select the second sound is controlled.
29. The sound processing device according to claim 28, wherein: The volume of the first sound has a transition.
30. The sound processing device according to claim 28 or 29, wherein: The circuit calculates the evaluation value so that the second sound is more likely to be selected as the volume of the first sound is louder.
31. The sound processing device according to claim 1, wherein: The sound space information is scene information including information about the sound source in the sound space and information about the position of the listener in the sound space. The second sound is a plurality of second sounds each generated in the sound space in response to the first sound. The circuit, obtaining the first sound signal, calculating the plurality of second sounds based on the scene information and the signal of the first sound, acquiring characteristics related to the first sound from the information of the sound source, Whether or not to select each of the plurality of second sounds as a sound to which binaural processing is not applied is controlled based on the characteristics related to the first sound, thereby selecting one or more second sounds to which the binaural processing is not applied from among the plurality of second sounds.
32. The sound processing device according to claim 31, wherein: The scene information is updated based on the input information, The characteristics related to the first sound are acquired according to the update of the scene information.
33. The sound processing device according to claim 31 or 32, wherein: The scene information and characteristics related to the first sound are acquired from metadata included in the bitstream.
34. A sound processing method, wherein: include: The step of obtaining sound space information related to the sound space; A step of acquiring characteristics related to a first sound generated from a sound source in the sound space based on the sound space information; as well as A step of controlling whether to select a second sound generated in the sound space corresponding to the first sound based on the characteristics related to the first sound.
35. A program for causing a computer to execute the sound processing method according to claim 34.
Citation Information
Patent Citations
Signal processor
JP2019022049A
Apparatus and method for rendering a sound scene using pipeline stages
WO2021180938A1