Acoustic processing device, threshold determination device, and acoustic processing method

By controlling the selection of reflected sounds through an audio processing device, and based on sound spatial information and volume ratio, the problems of high computational load and poor sound quality in virtual reality and augmented reality are solved, achieving efficient audio processing and extended battery life.

CN121925867APending Publication Date: 2026-04-24PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2024-10-03
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In virtual reality and augmented reality environments, existing audio processing technologies struggle to effectively manage large amounts of reflected sound, resulting in high computational demands, short battery life, and poor sound quality.

Method used

The sound processing device controls whether to select reflected sound, and based on the sound spatial information and volume ratio, appropriately reduces the amount of computation and the computational load, and applies binaural processing to generate sound.

Benefits of technology

It improves the efficiency of audio processing, reduces computational load and battery consumption, and enhances the audio quality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925867A_ABST
    Figure CN121925867A_ABST
Patent Text Reader

Abstract

A sound processing device (1001) is provided with a circuit (1402) and a memory (1404), and the circuit (1402) uses the memory (1404) to acquire sound space information relating to a sound space, acquires, on the basis of the sound space information, a volume ratio between the volume of direct sound generated from a sound source in the sound space and the volume of reflected sound generated in the sound space corresponding to the direct sound, and transmits the acquired volume ratio to the sound space. Information relating to the additional processing is acquired, and whether or not to select the reflected sound is controlled on the basis of the information relating to the additional processing and the volume ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to sound processing devices, etc. Background Technology

[0002] In recent years, goods and services utilizing ER (Extended Reality), including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), have become increasingly popular. Consequently, the importance of audio processing technologies—which imbue virtual sound sources with sound effects corresponding to the environment of that space, thus providing listeners with immersive audio in virtual or real spaces—has grown.

[0003] Additionally, the listener can also be a listener or a user. Furthermore, technologies related to the sound processing apparatus and sound processing method disclosed herein are shown in Patent Document 1, Patent Document 2, Patent Document 3, Patent Document 4, Non-Patent Document 1, and Non-Patent Document 2.

[0004] Existing technical documents Patent documents Patent Document 1: Japanese Patent No. 6288100 Patent Document 2: Japanese Patent Application Publication No. 2019-22049 Patent Document 3: International Publication No. 2021 / 180938 Patent Document 4: International Publication No. 2023 / 083780 Non-patent literature Non-Patent Literature 1: BCJ Moore, *An Outline of Auditory Psychology*, Chengxin Bookstore, April 20, 1994, Chapter 6: Spatial Perception, p. 225 Non-patent literature 2: Fabian Brinkmann et al., "The FABIAN head-related transfer function database", 2020 / 1 / 22, https: / / sofacoustics.org / data / database / tu-berlin / FABIAN_Documentation.pdf Summary of the Invention

[0005] The problem that the invention aims to solve For example, Patent Document 1 discloses a technique for processing an audio signal and providing a prompt to the listener. With the increasing prevalence of audio processing (ER) technology and the diversification of services using ER technology, there are growing demands for audio processing tailored to the specific audio quality required by each service, the signal processing capabilities of the terminal used, and the audio quality that the audio prompting device can provide. Furthermore, to meet these demands, further improvements in audio processing technology are required.

[0006] Here, improvements to sound processing technology refer to changes to existing sound processing methods. For example, improvements to sound processing technology may provide processing that imparts new sound effects, reduce the amount of processing required, improve the quality of sound obtained through sound processing, reduce the amount of data used in the implementation of sound processing, or facilitate the acquisition or generation of information used in the implementation of sound processing. Alternatively, improvements to sound processing technology may provide any combination of two or more of these.

[0007] In particular, these improvements are required in devices or services where listeners can move freely within virtual space. However, the aforementioned effects achieved through improvements in sound processing technology are merely examples. One or more solutions based on this disclosure may also be solutions conceived from different viewpoints, solutions achieving different objectives, or solutions yielding different effects than those described above.

[0008] Methods used to solve problems An audio processing apparatus based on a solution disclosed herein includes a circuit and a memory. The circuit uses the memory to obtain sound space information related to a sound space, and based on the sound space information, obtains a volume ratio between the volume of a direct sound generated from a sound source in the sound space and the volume of a reflected sound generated in the sound space corresponding to the direct sound, obtains information related to additional processing, and controls whether to select the reflected sound based on the information related to additional processing and the volume ratio.

[0009] In addition, these inclusive or specific solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of these.

[0010] Invention Effects One solution disclosed herein can provide, for example, processing that imparts new sound effects, reduction of the processing load of sound processing, improvement of the sound quality obtained through sound processing, reduction of the amount of data used in the implementation of sound processing, or facilitation of the acquisition or generation of information used in the implementation of sound processing. Alternatively, one solution disclosed herein can provide any combination of these. As a result, one solution disclosed herein provides sound processing suitable for the listener's usage environment and can contribute to improving the listener's audio experience.

[0011] In particular, the aforementioned effects can be achieved in devices or services that allow listeners to move freely within virtual space. However, the aforementioned effects are merely one example of the effects of various solutions based on this disclosure. One or more solutions based on this disclosure may also be solutions conceived from different viewpoints, solutions achieving different purposes, or solutions achieving different effects than those described above. Attached Figure Description

[0012] Figure 1 This is a diagram illustrating an example of direct and reflected tones generated in sound space.

[0013] Figure 2 This is a diagram illustrating an example of a stereo sound reproduction system implemented in this way.

[0014] Figure 3A This is a block diagram illustrating an example of the configuration of an encoding device in an embodiment.

[0015] Figure 3B This is a block diagram illustrating an example of the configuration of a decoding device in an embodiment.

[0016] Figure 3C This is a block diagram illustrating another configuration example of the encoding device in the implementation method.

[0017] Figure 3D This is a block diagram illustrating another configuration example of the decoding device in the embodiment.

[0018] Figure 4A This is a block diagram illustrating an example of the configuration of a decoder in an implementation method.

[0019] Figure 4B This is a block diagram illustrating another configuration example of the decoder in an implementation scheme.

[0020] Figure 5 This is a diagram illustrating an example of the physical configuration of a sound signal processing device according to an embodiment.

[0021] Figure 6 This is a diagram illustrating an example of the physical configuration of the encoding device in an embodiment.

[0022] Figure 7This is a block diagram illustrating an example of the configuration of the rendering unit in an implementation method.

[0023] Figure 8 This is a flowchart illustrating an example of the operation of the sound signal processing device in an embodiment.

[0024] Figure 9 It is a diagram that shows the relative positions of the listener and the obstacle.

[0025] Figure 10 It is a diagram that shows the relative positional relationship between the listener and the obstacle object.

[0026] Figure 11 It is a graph showing the relationship between the time difference and threshold of direct and reflected tones.

[0027] Figure 12A This is a diagram that represents part of an example of how threshold data is set.

[0028] Figure 12B This is a diagram that represents part of an example of how threshold data is set.

[0029] Figure 12C This is a diagram that represents part of an example of how threshold data is set.

[0030] Figure 13 This is a diagram illustrating an example of how a threshold is set.

[0031] Figure 14 This is a flowchart representing an example of a selection process.

[0032] Figure 15 This is a diagram representing an audio experiment involving preceding and following sounds related to the priority effect.

[0033] Figure 16 It is a graph representing the experimental parameters in the horizontal direction.

[0034] Figure 17 It is a graph representing the experimental parameters in the vertical plane direction.

[0035] Figure 18 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold.

[0036] Figure 19 It is a graph showing the relationship between angle difference, time difference, and threshold.

[0037] Figure 20 This is a block diagram representing another example of the rendering unit.

[0038] Figure 21 This is a flowchart representing another example of the selection process.

[0039] Figure 22 This is another flowchart representing a selection process.

[0040] Figure 23 This is a flowchart illustrating a first variation of the operation of the sound signal processing apparatus according to the embodiment.

[0041] Figure 24 This is a flowchart illustrating a second variation of the operation of the sound signal processing apparatus according to the embodiment.

[0042] Figure 25 This is a diagram showing a configuration example of avatars, sound source objects, and obstacle objects.

[0043] Figure 26 This is another flowchart representing a selection process.

[0044] Figure 27 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit.

[0045] Figure 28 It is a diagram showing the transmission and diffraction of sound.

[0046] Figure 29 It is a graph representing the threshold.

[0047] Figure 30 This is a diagram representing the first application example of the threshold.

[0048] Figure 31 This is a diagram representing the second application example of the threshold.

[0049] Figure 32 This is a diagram representing the third application example of the threshold.

[0050] Figure 33 This is a diagram representing the fourth application example of the threshold.

[0051] Figure 34 This is a flowchart illustrating the first basic operation example of the sound signal processing device in the embodiment.

[0052] Figure 35 This is a flowchart illustrating a second basic operation example of the sound signal processing device in the embodiment.

[0053] Figure 36 This is a flowchart illustrating the third basic operation example of the sound signal processing device in the embodiment. Detailed Implementation

[0054] (The understanding that forms the basis of this disclosure) Figure 1This diagram illustrates an example of direct and reflected sounds generated in a sound space. In sound processing that uses sound to represent the characteristics of a virtual space, in order to represent the breadth of the space and the material of the walls, as well as to accurately determine the location of the sound source (sound image localization), it is effective to reproduce not only direct sounds but also reflected sounds.

[0055] For example, in such Figure 1 When listening to sound within a rectangular room, a sound source produces six primary reflections corresponding to the six walls. The reproduction of these reflections provides clues for a proper understanding of the space and the sound image. Furthermore, for each reflection, secondary reflections are produced on surfaces other than the one that produced the reflection. These reflections also serve as perceptually valid clues.

[0056] However, even considering only the case of secondary reflection, a sound source will produce 1 direct tone and 36 (6+6×5) reflected tones, resulting in 37 vocal lines. Processing these vocal lines requires a considerable amount of computation.

[0057] Furthermore, in recent years, applications related to the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, have inevitably involved multiple sound sources, thus requiring a much larger amount of computation.

[0058] Furthermore, listeners in virtual spaces use headphones or VR goggles to receive audio. To provide stereo sound to such listeners, binaural processing is performed on each sound ray, assigning sound pressure levels and phase differences between the two ears to reproduce the direction of arrival and the sense of distance. Therefore, the computational load is extremely high if all the reflected sounds are to be reproduced.

[0059] On the other hand, small rechargeable batteries are sometimes used as batteries for VR goggles worn by listeners experiencing virtual space, due to their convenience. To extend battery life, the computational load required for the processing described above should ideally be low. Therefore, it is desirable to reduce the number of sound lines generated on a scale of several hundred without compromising sound localization and spatial accuracy.

[0060] Furthermore, in systems that reproduce sound, sometimes the listener's position and orientation are allowed 6 degrees of freedom (6DoF). In this case, the positional relationship between the listener, the sound source, and the object reflecting the sound cannot be determined if it is not during reproduction (rendering). Therefore, the reflected sound also cannot be determined if it is not during reproduction. Thus, it is difficult to predetermine the reflected sound of the object being processed.

[0061] Therefore, appropriately selecting one or more reflected sounds of the object to be processed or the object not to be processed from among the multiple reflected sounds generated in the sound space during reproduction is beneficial to the appropriate reduction of computational load.

[0062] Therefore, the purpose of this disclosure is to provide an audio processing device, etc., that can appropriately control whether to select the sound generated in the sound space.

[0063] Additionally, the control over whether to select a sound corresponds to the determination of whether to select a sound. Furthermore, selecting a sound can be either choosing a sound as the object of processing or choosing a sound as a non-processing object.

[0064] (Public Summary) The audio processing apparatus based on the first embodiment disclosed herein includes a circuit and a memory; the circuit uses the memory to obtain sound space information related to the sound space; based on the sound space information, it obtains characteristics related to a first sound generated from a sound source in the sound space; based on the characteristics related to the first sound, it controls whether to select a second sound generated in the sound space corresponding to the first sound.

[0065] The device described above can appropriately control whether to select a second sound corresponding to the first sound in the sound space based on characteristics related to the first sound generated in the sound space. That is, it can appropriately control whether to select a sound generated in the sound space. Therefore, it can appropriately reduce the computational load and computational complexity.

[0066] The sound processing apparatus based on the second embodiment disclosed herein can also be, in the sound processing apparatus of the first embodiment, where the first sound is a direct sound and the second sound is a reflected sound.

[0067] The device described above can appropriately control whether to select reflected sound based on characteristics related to direct sound.

[0068] The audio processing apparatus based on the third embodiment disclosed herein can also be, in the audio processing apparatus of the second embodiment, where the characteristic related to the first sound is the volume ratio of the direct sound volume to the reflected sound volume; the circuit calculates the volume ratio based on the sound space information, and controls whether to select the reflected sound based on the volume ratio.

[0069] The device described above can appropriately select the reflected sound that has a greater impact on the listener's perception based on the volume ratio of the direct sound to the reflected sound.

[0070] The sound processing apparatus based on the fourth embodiment disclosed herein can also be, in the sound processing apparatus of the third embodiment, in the case of selected reflected sound, the circuit generates sound that reaches the listener's two ears respectively by applying binaural processing to the reflected sound and the direct sound.

[0071] The device described above can appropriately select reflected sounds that have a significant impact on the listener's perception, and apply binaural processing to the selected reflected sounds.

[0072] The audio processing apparatus based on the fifth embodiment disclosed herein may also be, in the audio processing apparatus of the third or fourth embodiment, wherein the circuit calculates the time difference between the end time of the direct tone and the arrival time of the reflected tone based on the sound space information, and controls whether to select the reflected tone based on the time difference and the volume ratio.

[0073] The device described above can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the time difference between the end of the direct sound and the arrival of the reflected sound, as well as the volume ratio of the direct sound to the reflected sound. Therefore, the device described above can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the aftermasking effect.

[0074] The audio processing apparatus based on the sixth embodiment disclosed herein may also be, in the audio processing apparatus of the fifth embodiment, wherein the circuit selects reflected sound when the volume ratio is above a threshold; and the first threshold used as the threshold when the time difference is a first value is greater than the second threshold used as the threshold when the time difference is a second value greater than the first value.

[0075] The device described above increases the likelihood that reflected sounds with a large time difference between the end of the direct tone and the arrival of the reflected sound will be selected. Therefore, the device described above can more appropriately select reflected sounds that have a greater impact on the listener's perception.

[0076] The audio processing apparatus based on the seventh embodiment disclosed herein may also be, in the audio processing apparatus of the third or fourth embodiment, wherein the circuit calculates the time difference between the arrival time of the direct tone and the arrival time of the reflected tone based on the sound spatial information; and controls whether to select the reflected tone based on the time difference and the volume ratio.

[0077] The device described above can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, as well as the volume ratio of the direct sound to the reflected sound. Therefore, the device described above can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the priority effect.

[0078] The audio processing apparatus based on the present disclosure regarding the eighth embodiment may also be, in the audio processing apparatus of the seventh embodiment, wherein the circuit selects reflected sound when the volume ratio is above a threshold; and the first threshold used as the threshold when the time difference is a first value is greater than the second threshold used as the threshold when the time difference is a second value greater than the first value.

[0079] The device described above increases the likelihood that reflected sounds with a large time difference between the arrival time of the direct sound and the arrival time of the reflected sound will be selected. Therefore, the device described above can appropriately select reflected sounds that have a significant impact on the listener's perception.

[0080] The audio processing apparatus based on the ninth embodiment disclosed herein can also be, in the audio processing apparatus of the eighth embodiment, wherein the circuit adjusts the threshold based on the direction of arrival of the direct sound and the direction of arrival of the reflected sound.

[0081] The device described above can appropriately select the reflected sound that has a greater impact on the listener's perception based on the direction of arrival of the direct sound and the direction of arrival of the reflected sound.

[0082] The audio processing apparatus based on the 10th embodiment disclosed herein may also be, in any of the 2nd to 9th embodiments, in the absence of selecting a reflected tone, the circuit corrects the volume of the direct tone based on the volume of the reflected tone.

[0083] The device described above can appropriately reduce the unpleasantness caused by the lack of volume of reflected sound due to the absence of selected reflected sound with less computation.

[0084] The audio processing apparatus based on the 11th embodiment disclosed herein may also be, in any of the 2nd to 9th embodiments, synthesize the reflected sound into the direct sound in the absence of selecting the reflected sound.

[0085] The device described above can more accurately reflect the characteristics of reflected sound into the direct sound. Therefore, the device described above can reduce the unpleasant feeling caused by the lack of volume of reflected sound due to the inability to select the reflected sound.

[0086] The sound processing apparatus based on the 12th embodiment disclosed herein may also be, in any of the 3rd to 9th embodiments, a volume ratio that is the volume ratio of the direct sound at the first moment to the reflected sound at the second moment, which is different from the first moment.

[0087] When the time when the direct sound is perceived is different from the time when the reflected sound is perceived, the device of the above scheme can appropriately select the reflected sound that has a greater impact on the listener's perception based on the volume ratio of the direct sound to the reflected sound at different times.

[0088] The audio processing apparatus based on the 13th embodiment disclosed herein may also be, in the audio processing apparatus of the 1st or 2nd embodiment, wherein the circuit sets a threshold based on the characteristics related to the first sound, and controls whether to select the second sound based on the threshold.

[0089] The device described above can appropriately control whether to select the second sound based on a threshold set according to the characteristics related to the first sound.

[0090] The sound processing apparatus based on the 14th embodiment disclosed herein may also be, in any of the 1st, 2nd and 13th embodiments, the characteristic related to the first sound is one or a combination of two or more of the following: the volume of the sound source, the visuality of the sound source and the localization of the sound source.

[0091] The device described above can appropriately control whether to select a second sound based on the volume of the sound source, the visual nature of the sound source, or the localization of the sound source.

[0092] The sound processing apparatus based on the 15th embodiment disclosed herein may also be, in any of the 1st, 2nd and 13th embodiments, the characteristic related to the first sound is the frequency characteristic of the first sound.

[0093] The device described above can appropriately control whether to select a second sound corresponding to the first sound based on the frequency characteristics of the first sound.

[0094] The sound processing apparatus based on the 16th embodiment disclosed herein may also be, in any of the 1st, 2nd and 13th embodiments, a characteristic related to the first sound that represents the discontinuity of the amplitude of the first sound.

[0095] The device described above can appropriately control whether to select a second sound corresponding to the first sound based on the characteristic representing the discontinuity of the amplitude of the first sound.

[0096] The sound processing apparatus based on the 17th embodiment disclosed herein may also be, in any of the 1st, 2nd, 13th and 16th embodiments, a characteristic related to the first sound that represents the duration of the audible portion of the first sound or the duration of the silent portion of the first sound.

[0097] The apparatus described above can appropriately control whether to select a second sound corresponding to the first sound based on the characteristics of the duration of the audible part of the first sound or the duration of the silent part of the first sound.

[0098] The sound processing apparatus based on the 18th embodiment disclosed herein may also be, in any of the 1st, 2nd, 13th, 16th and 17th embodiments, a characteristic related to the first sound that is to represent the duration of the sound portion of the first sound and the duration of the silent portion of the first sound in a time sequence.

[0099] The apparatus described above can appropriately control whether to select a second sound corresponding to the first sound, based on the characteristic that the duration of the audible part of the first sound and the duration of the silent part of the first sound are represented in a time sequence.

[0100] The sound processing apparatus based on the present disclosure regarding the 19th embodiment may also be, in any of the 1st, 2nd, 13th and 15th embodiments, a characteristic related to the first sound that represents a change in the frequency characteristics of the first sound.

[0101] The device described above can appropriately control whether to select a second sound corresponding to the first sound based on the characteristics representing the variation of the frequency characteristics of the first sound.

[0102] The sound processing apparatus based on the present disclosure regarding the 20th embodiment may also be, in any of the 1st, 2nd, 13th, 15th and 19th embodiments, a characteristic related to the first sound that represents the stability of the frequency response of the first sound.

[0103] The device described above can appropriately control whether to select a second sound corresponding to the first sound based on the stability of the frequency characteristics representing the first sound.

[0104] The audio processing apparatus based on the present disclosure regarding the 21st embodiment may also be an audio processing apparatus in any of the 1st, 2nd, and 13th to 20th embodiments, in which characteristics related to the first sound are obtained from the bitstream.

[0105] The apparatus described above can appropriately control whether to select a second sound corresponding to the first sound based on information obtained from the bit stream.

[0106] The audio processing apparatus based on the present disclosure regarding the 22nd embodiment may also be an audio processing apparatus in any of the 1st, 2nd, and 13th to 21st embodiments in which the circuit calculates characteristics related to the second sound; and controls whether to select the second sound based on the characteristics related to the first sound and the characteristics related to the second sound.

[0107] The device described above can appropriately control whether to select a second sound corresponding to the first sound based on characteristics related to the first sound and characteristics related to the second sound.

[0108] The audio processing apparatus based on the present disclosure of the 23rd embodiment may also be, in the audio processing apparatus of the 22nd embodiment, wherein the circuit obtains a threshold representing the volume corresponding to the boundary of whether a sound can be heard; and controls whether to select the second sound based on the characteristics related to the first sound, the characteristics related to the second sound, and the threshold.

[0109] The device described above can, in addition to the characteristics related to the first sound and the characteristics related to the second sound, also appropriately control whether to select the second sound based on the threshold corresponding to whether it can be heard.

[0110] The sound processing apparatus based on the present disclosure regarding the 24th embodiment may also be, in the sound processing apparatus of the 23rd embodiment, a characteristic related to the second sound being the volume of the second sound.

[0111] The device described above can appropriately control whether to select the second sound based on the volume of the second sound.

[0112] The audio processing apparatus based on the 25th embodiment disclosed herein may also be, in the audio processing apparatus of the 1st or 2nd embodiment, the sound space information includes information about the location of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit controls whether to select each of the plurality of second sounds based on characteristics related to the first sound, and selects one or more processing target sounds from the plurality of second sounds to be applied to binaural processing.

[0113] The device described above is able to appropriately select one or more target sounds for processing from a plurality of second sounds generated in the sound space corresponding to the first sound, based on characteristics related to the first sound generated in the sound space.

[0114] The audio processing apparatus based on the present disclosure regarding the 26th embodiment may also be, in the audio processing apparatus of the 25th embodiment, at least one of the following: when the sound space is created, when the sound space processing begins, and when the information update thread in the sound space processing is generated.

[0115] The device described above can appropriately select one or more target sounds to be processed by binaural processing based on information obtained at adaptive timing.

[0116] The sound processing apparatus based on the present disclosure regarding the 27th embodiment may also be, in the sound processing apparatus of the 26th embodiment, periodically acquire characteristics related to the first sound after the processing of the sound space begins.

[0117] The device described above can appropriately select one or more target sounds to be processed by binaural processing based on information obtained periodically.

[0118] The audio processing apparatus based on the 28th embodiment disclosed herein may also be, in the audio processing apparatus of the 1st or 2nd embodiment, the characteristic related to the first sound is the volume of the first sound; the circuit calculates the evaluation value of the second sound based on the volume of the first sound, and controls whether to select the second sound based on the evaluation value.

[0119] The device described above can appropriately control whether to select the second sound based on an evaluation value calculated for the second sound according to the volume of the first sound.

[0120] The sound processing apparatus based on the present disclosure regarding the 29th embodiment may also be, in the sound processing apparatus of the 28th embodiment, where the volume of the first sound has a change.

[0121] The device described above can appropriately control whether to select a second sound based on an evaluation value calculated according to the volume change.

[0122] The audio processing apparatus based on the present disclosure regarding the 30th embodiment may also be, in the audio processing apparatus of the 28th or 29th embodiment, where the circuit calculates an evaluation value such that the greater the volume of the first sound, the easier it is to select the second sound.

[0123] The device described above can appropriately control whether to select the second sound based on an evaluation value that is set so that the larger the volume of the first sound, the easier it is to select the second sound.

[0124] The audio processing apparatus based on the 31st embodiment disclosed herein may also be, in the audio processing apparatus of the 1st or 2nd embodiment, the sound space information is scene information including information about sound sources in the sound space and information about the location of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit acquires the signal of the first sound; calculates a plurality of second sounds based on the scene information and the signal of the first sound; acquires characteristics related to the first sound from the information about the sound sources; controls whether to select each of the plurality of second sounds as a sound that is not subject to binaural processing based on the characteristics related to the first sound, thereby selecting one or more second sounds from the plurality of second sounds that are not subject to binaural processing.

[0125] The device described above is able to appropriately select one or more second sounds that are not subject to binaural processing from a plurality of second sounds generated in the sound space corresponding to the first sound, based on characteristics related to the first sound.

[0126] The audio processing apparatus based on the present disclosure of the 32nd embodiment can also be, in the audio processing apparatus of the 31st embodiment, where scene information is updated based on input information; and corresponding to the update of scene information, characteristics related to the first sound are obtained.

[0127] The device described above can appropriately select one or more second sounds that are not processed by binaural processing, based on information obtained from updates corresponding to scene information.

[0128] The audio processing apparatus based on the 33rd scheme disclosed herein may also be, in the audio processing apparatus of the 31st or 32nd scheme, obtain scene information and characteristics related to the first sound from the metadata contained in the bitstream.

[0129] The apparatus described above can appropriately select one or more second sounds that are not processed by binaural processing, based on information obtained from the metadata contained in the bitstream.

[0130] The sound processing method based on the 34th embodiment disclosed herein includes: a step of obtaining sound space information related to a sound space; a step of obtaining characteristics related to a first sound generated from a sound source in the sound space based on the sound space information; and a step of controlling whether to select a second sound generated in the sound space corresponding to the first sound based on the characteristics related to the first sound.

[0131] The method described above can achieve the same effect as the audio processing device described in the first solution.

[0132] The program based on the 35th scheme disclosed herein is a program used to enable a computer to execute the audio processing method of the 34th scheme.

[0133] The above-described program can achieve the same sound processing effect as the 34th scheme using a computer.

[0134] The audio processing apparatus based on the 36th embodiment disclosed herein may also include, in any of the 3rd to 9th and 12th embodiments, an additional unit that applies additional processing to the second sound; an additional increment / decrease acquisition unit that acquires the amplitude increment / decrease in the additional unit; and a volume ratio adjustment unit that adjusts the volume ratio based on the increment / decrease acquired by the additional increment / decrease acquisition unit.

[0135] The device described above can appropriately adjust the volume ratio based on the amplitude increased or decreased through additional processing.

[0136] The audio processing apparatus based on the 37th embodiment disclosed herein may also include, in any of the 5th to 9th embodiments, an additional unit that applies additional processing to the second sound; an additional delay amount acquisition unit that acquires the delay increase or decrease amount in the additional unit; and a time difference adjustment unit that adjusts the time difference based on the increase or decrease amount acquired by the additional delay amount acquisition unit.

[0137] The device described above can appropriately adjust the time difference based on the delay increased or decreased through additional processing.

[0138] The sound processing apparatus based on the 38th embodiment disclosed herein may also be, in the sound processing apparatus of the 36th or 37th embodiment, an additional processing which is the processing not applied to the first sound.

[0139] The apparatus described above can apply additional processing to the second sound that is not applied to the first sound. Therefore, the apparatus described above can apply additional processing to the second sound that is suitable for the second sound.

[0140] The audio processing apparatus based on the 39th embodiment disclosed herein may also be, in the audio processing apparatus of the 36th or 37th embodiment, an additional unit may be a filter unit that emphasizes the diffusion of sound in terms of hearing.

[0141] The device described above can apply additional processing to the second sound by using a filter unit that emphasizes the diffusion of sound in an auditory sense.

[0142] The audio processing apparatus based on the 40th embodiment disclosed herein may also include a circuit and a memory. The circuit uses the memory to obtain sound space information related to the sound space. Based on the sound space information, it obtains the volume ratio of the volume of the direct sound generated from the sound source in the sound space to the volume of the reflected sound generated in the sound space corresponding to the direct sound. It obtains information related to additional processing and controls whether to select the reflected sound based on the information related to additional processing and the volume ratio.

[0143] The apparatus described above can appropriately control whether to select a reflected sound based on information related to additional processing and the volume ratio of the direct sound to the reflected sound. That is, it can appropriately control whether to select a reflected sound that corresponds to the direct sound generated from the sound source in the sound space. Therefore, it can appropriately reduce the computational load and computational complexity.

[0144] The audio processing apparatus based on the 41st embodiment disclosed herein may also be, in the audio processing apparatus of the 40th embodiment, wherein the circuit adjusts the volume ratio according to information related to additional processing, and controls whether to select reflected sound based on the adjusted volume ratio.

[0145] The device described above can appropriately control whether to select reflected sound based on the volume ratio, which reflects the effect of additional processing.

[0146] The audio processing apparatus based on the 42nd embodiment disclosed herein may also be, in the audio processing apparatus of the 40th or 41st embodiment, wherein the circuit calculates the time when the sound emitted from the sound source arrives as a reflected sound, and controls whether to select the reflected sound based on the time, information related to additional processing, and volume ratio.

[0147] The device described above can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the arrival time of the reflected sound, information related to additional processing, and the volume ratio of the direct sound to the reflected sound.

[0148] The audio processing apparatus based on the 43rd embodiment disclosed herein may also be an audio processing apparatus in any of the 40th to 42nd embodiments, in which the circuit calculates the time difference between the arrival time of the direct tone and the arrival time of the reflected tone based on the sound space information, and controls whether to select the reflected tone based on the time difference, information related to additional processing, and the volume ratio.

[0149] The device described above can more appropriately select a reflected tone that has a greater impact on the listener's perception based on the time difference between the arrival time of the direct tone and the arrival time of the reflected tone, information related to additional processing, and the volume ratio of the direct tone to the reflected tone. Therefore, the device described above can more appropriately select a reflected tone that has a greater impact on the listener's perception based on a priority effect.

[0150] The audio processing apparatus based on the 44th embodiment disclosed herein may also be an audio processing apparatus in any of the 40th to 42nd embodiments, in which the circuit calculates the time difference between the end time of the direct tone and the arrival time of the reflected tone based on the sound space information, and controls whether to select the reflected tone based on the time difference, information related to additional processing, and the volume ratio.

[0151] The device described above can more appropriately select a reflected tone that has a greater impact on the listener's perception, based on the time difference between the end of the direct tone and the arrival of the reflected tone, information related to additional processing, and the volume ratio of the direct tone to the reflected tone. Therefore, the device described above can more appropriately select a reflected tone that has a greater impact on the listener's perception, based on the back-shielding effect.

[0152] The audio processing apparatus based on the 45th embodiment disclosed herein may also be an audio processing apparatus in any of the 40th to 44th embodiments, wherein the additional processing is the processing of the amplitude increase or decrease accompanying the reflection sound, and the information related to the additional processing is information indicating the amplitude increase or decrease in the processing accompanying the amplitude increase or decrease of the reflection sound, and the circuit adjusts the volume ratio according to the amplitude increase or decrease, and controls whether to select the reflection sound based on the adjusted volume ratio.

[0153] The device described above can appropriately adjust the volume ratio based on the amplitude increased or decreased through additional processing. Furthermore, the device described above can appropriately control whether to select reflected sound based on the volume ratio adjusted through additional processing.

[0154] The audio processing apparatus based on the 46th embodiment disclosed herein may also be an audio processing apparatus of the 43rd or 44th embodiment, wherein the additional processing is a processing of the delay increase or decrease accompanying the reflected sound, and the information related to the additional processing is information indicating the delay increase or decrease amount in the processing of the delay increase or decrease accompanying the reflected sound, the circuit adjusts the time difference according to the delay increase or decrease amount, and controls whether to select the reflected sound based on the adjusted time difference and the volume ratio.

[0155] The apparatus described above can appropriately adjust the time difference based on the delay increased or decreased through additional processing. Furthermore, the apparatus described above can appropriately control whether to select reflected sound based on the time difference adjusted through additional processing.

[0156] The sound processing apparatus based on the 47th embodiment disclosed herein may also be, in any of the 40th to 46th embodiments, an additional processing which is a processing not applied to the direct sound.

[0157] The apparatus described above can apply additional processing to reflected sounds without applying additional processing to direct sounds. That is, the apparatus described above can apply additional processing to reflected sounds that is not applied to direct sounds. Therefore, the apparatus described above can apply additional processing to reflected sounds suitable for reflected sounds.

[0158] The audio processing apparatus based on the 48th embodiment disclosed herein may also be an additional processing method in any of the 40th to 47th embodiments, wherein the additional processing is a filtering process that emphasizes the diffusion of sound in terms of auditory perception.

[0159] The device described above can appropriately control whether to select reflected sound based on information related to filtering processing that emphasizes the diffusion of sound in auditory perception, and the volume ratio of direct sound to reflected sound.

[0160] The audio processing apparatus based on the 49th embodiment disclosed herein may also be, in the audio processing apparatus of the 41st or 45th embodiment, not adjust the volume ratio when the information related to additional processing indicates that additional processing is not performed.

[0161] The device described above can appropriately control whether to adjust the volume ratio depending on whether additional processing is performed.

[0162] The audio processing apparatus based on the 50th embodiment disclosed herein may also be, in the audio processing apparatus of the 46th embodiment, not adjust the time difference when the information related to additional processing indicates that no additional processing is performed.

[0163] The device described above can appropriately control whether to adjust the time difference based on whether additional processing is performed.

[0164] The sound processing apparatus based on the 51st embodiment disclosed herein may also include a circuit and a memory. The circuit uses the memory to obtain sound space information related to the sound space. Based on the sound space information, it determines whether the sound source in the sound space is a point sound source. If the sound source is determined to be a point sound source, (i) based on the sound space information, it obtains the volume ratio of the volume of the direct sound generated from the sound source to the volume of the reflected sound generated in the sound space corresponding to the direct sound, (ii) based on the sound space information, it calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and (iii) based on the volume ratio and the time difference, it controls whether to select the reflected sound. If the sound source is determined not to be a point sound source, the reflected sound is selected regardless of the volume ratio and the time difference.

[0165] Even if the distance between the listener and the sound source is not determined, making it difficult to properly calculate the volume ratio and time difference between the direct and reflected sounds, the device described above can select the reflected sound regardless of the volume ratio and time difference. Therefore, the device described above can suppress the sound quality degradation that may occur when the sound source is not a point sound source.

[0166] The threshold determination device based on the 52nd scheme disclosed herein includes a circuit and a memory. The circuit uses the memory to obtain the time difference between the arrival time of the previous tone and the arrival time of the subsequent tone generated corresponding to the previous tone, determines a threshold based on the time difference, and outputs the threshold. The threshold is a threshold of the volume ratio of the subsequent tone to the volume of the previous tone, corresponding to the boundary of whether the subsequent tone is subject to the priority effect brought by the previous tone. In determining the threshold, in the region of the first axis with time difference and the second axis with volume ratio, on the boundary line determined as the boundary line, the volume ratio corresponding to the time difference is determined as the threshold. The domain of the boundary line relative to the first axis is the time range in which the priority effect is generated.

[0167] The apparatus described above can appropriately determine the threshold corresponding to the boundary of the priority effect based on the relationship between the time difference between the preceding and following sounds and the volume ratio between the preceding and following sounds. Furthermore, the apparatus described above can efficiently determine the threshold for the time range in which the priority effect occurs.

[0168] The threshold determination device based on the 53rd scheme disclosed herein can also be, in the threshold determination device of the 52nd scheme, when the first axis is the horizontal axis and the second axis is the vertical axis, the boundary line in the domain is a curve with the right shoulder descending and the top convex.

[0169] The apparatus described above can appropriately determine the threshold based on the boundary characteristic that a larger time difference results in a smaller volume ratio, and that the decrease in volume ratio is proportional to the increase in time difference. Therefore, the apparatus described above can provide an appropriate threshold for determining whether a subsequent sound is subject to a priority effect.

[0170] The threshold determination device based on the 54th scheme disclosed herein can also be, in the threshold determination device of the 52nd or 53rd scheme, where the boundary line is defined by a mathematical formula with time difference as a variable, and the threshold is calculated based on the mathematical formula.

[0171] The apparatus described above can efficiently determine the threshold based on mathematical formulas. Therefore, the apparatus described above can efficiently provide the threshold.

[0172] The sound processing method based on the 55th embodiment disclosed herein includes: a step of obtaining sound space information related to the sound space; a step of obtaining, based on the sound space information, a volume ratio of the volume of a direct sound generated from a sound source in the sound space to the volume of a reflected sound generated in the sound space corresponding to the direct sound; a step of obtaining information related to additional processing; and a step of controlling whether to select a reflected sound based on the information related to additional processing and the volume ratio.

[0173] The method described above can achieve the same effect as the audio processing device described in Scheme 40.

[0174] The program based on the 56th scheme disclosed herein is a program for causing a computer to execute the audio processing method of the 55th scheme.

[0175] The above-described program can achieve the same effect as the audio processing method in Scheme 55 using a computer.

[0176] Furthermore, these inclusive or specific solutions can be implemented by systems, devices, methods, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, or recording media.

[0177] The audio processing apparatus, encoding apparatus, decoding apparatus, and stereo audio reproduction system of this disclosure will now be described in detail with reference to the accompanying drawings. The stereo audio reproduction system can also be described as a sound signal reproduction system.

[0178] Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement and connection schemes of constituent elements, steps, and order of steps shown in the following embodiments are examples and are not intended to limit the scope of the solutions based on this disclosure. In addition, constituent elements in the following embodiments that are not included in the basic solutions described in this disclosure or that are not described in the independent claims representing the highest-level concept are described as arbitrary constituent elements.

[0179] (Implementation Method) (Example of a stereo sound reproduction system) Figure 2 This diagram illustrates an example of a stereo sound reproduction system. Specifically, Figure 2 This describes a stereo sound reproduction system 1000 as an example of a system capable of applying the sound processing or decoding techniques disclosed herein. Stereo sound is also referred to as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound prompting device 1002.

[0180] The sound signal processing device 1001 is also manifested as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio-processed sound signal for the listener. The sound signal is not limited to speech; any audible sound is acceptable. Audio processing, for example, is signal processing performed on the sound signal to reproduce the effects that the sound undergoes from the sound source to the listener.

[0181] The sound signal processing device 1001 performs sound processing based on spatial information describing the reasons for the aforementioned effects. Spatial information includes, for example, information indicating the location of the sound source, the listener, and surrounding objects; information indicating the shape of the space; and parameters related to sound propagation. The sound signal processing device 1001 may be, for example, a PC (Personal Computer), a smartphone, a tablet computer, or a game console.

[0182] The processed audio signal is presented to the listener by the audio prompt device 1002. The audio prompt device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio prompt device 1002 via wireless or wired communication.

[0183] When the sound prompting device 1002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices simultaneously provide prompts through communication between the multiple devices or through communication between each of the multiple devices and the sound signal processing device 1001. The sound prompting device 1002 may be, for example, a headset, earplugs, head-mounted display, or a surround sound system composed of multiple fixed speakers, worn on the listener's head.

[0184] Furthermore, the stereo sound reproduction system 1000 can also be used in combination with image prompting devices or stereoscopic image prompting devices that visually provide ER experiences, including AR / VR. For example, the space processed by spatial information is a virtual space, where the positions of sound sources, listeners, and objects are virtual positions of virtual sound sources, virtual listeners, and virtual objects within the virtual space. This space can also be represented as a sound space. Additionally, spatial information can also be represented as sound spatial information.

[0185] also, Figure 2 This is an example of a system configuration where the sound signal processing device 1001 and the sound prompting device 1002 are different devices, but the stereo sound reproduction system 1000, which can apply the sound processing method or decoding method of this disclosure, is not limited to this. Figure 2 The configuration can be as follows. For example, the sound signal processing device 1001 can be included in the sound prompting device 1002, and the sound prompting device 1002 performs both sound processing and sound prompting.

[0186] Alternatively, the sound signal processing device 1001 and the sound prompting device 1002 may share the implementation of the sound processing described in this disclosure. Furthermore, a portion or all of the sound processing described in this disclosure may be implemented via a server connected to the sound signal processing device 1001 or the sound prompting device 1002 through a network.

[0187] Furthermore, the audio signal processing apparatus 1001 can also perform audio processing by decoding a bitstream generated by encoding at least a portion of the audio signal and spatial information data used for audio processing. Therefore, the audio signal processing apparatus 1001 can also be exemplified as a decoding device.

[0188] (Example of an encoding device) Figure 3A This is a block diagram illustrating an example of the configuration of an encoding device. Specifically, Figure 3A The configuration of encoding device 1100, which is an example of the encoding device of this disclosure, is shown.

[0189] Input data 1101 is encoded object data containing spatial information and / or audio signals input to encoder 1102. Details regarding the spatial information will be explained later.

[0190] Encoder 1102 encodes the input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated through encoding processing.

[0191] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or other types of memory.

[0192] Furthermore, in the above description, the bitstream generated through encoding processing was listed as an example of encoded data 1103 stored in memory 1104, but the encoded data 1103 can also be data other than a bitstream. For example, the encoding device 1100 may also store transformed data generated by converting the bitstream into a specified data format in memory 1104. The transformed data may, for example, be a file or multiplexed stream corresponding to more than one bitstream.

[0193] Here, the file is a file with a file format such as ISOBMFF (ISO Base Media File Format). Furthermore, the encoded data 1103 can also be in the form of multiple packets generated by splitting the aforementioned bitstream or file.

[0194] For example, the bitstream generated by encoder 1102 can be transformed into data different from the bitstream. In this case, encoding device 1100 has a transformation unit (not shown), which can perform the transformation processing either by the transformation unit or by a CPU (Central Processing Unit), which is an example of a processor described later.

[0195] (Example of a decoding device) Figure 3B This is a block diagram illustrating an example of the configuration of a decoding device. Specifically, Figure 3B The configuration of decoding device 1110, which is an example of the decoding device of this disclosure, is shown.

[0196] Memory 1114 stores, for example, the same data as the encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and input as input data 1113 into decoder 1112. Input data 1113 is, for example, a bitstream intended for decoding. Memory 1114 can be, for example, a hard disk or SSD, or other type of storage.

[0197] Alternatively, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 into the decoder 1112 as is, but instead transform the read data and input the transformed data as input data 1113 into the decoder 1112. The data before transformation may be, for example, multiplexed data containing more than one bitstream. Here, the multiplexed data may also be a file with a file format such as ISOBMFF.

[0198] Furthermore, the data before conversion can also be multiple packets generated by splitting the aforementioned bitstream or file. Alternatively, data different from the bitstream can be read from memory 1114 and converted into a bitstream. In this case, the decoding device 1110 may also include a conversion unit (not shown), and the conversion processing may be performed by the conversion unit, or by a CPU, as an example of a processor described later.

[0199] Decoder 1112 decodes the input data 1113 and generates an audio signal 1111 representing a prompt to the listener.

[0200] (Another example of an encoding device) Figure 3C This is a block diagram illustrating another configuration example of an encoding device. Specifically, Figure 3C This illustrates the configuration of encoding device 1120, which is another example of the encoding device disclosed herein. Figure 3C In China, for the sake of Figure 3A The same constituent elements are assigned to the same constituent elements as the constituent elements. Figure 3A The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.

[0201] Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 differs from encoding device 1100 in that it has a transmitting unit 1121 that transmits encoded data 1103 to the outside.

[0202] The transmitting unit 1121 transmits a transmission signal 1122 generated based on encoded data 1103 or data converted from encoded data 1103 into other data formats to other devices or servers. The data used in generating the transmission signal 1122 may be, for example, a bit stream, multiplexed data, file, or packet as described in the encoding device 1100.

[0203] (Another example of a decoding device) Figure 3D This is a block diagram illustrating another configuration example of a decoding device. Specifically, Figure 3D This illustrates the configuration of decoding device 1130, another example of a decoding device disclosed herein. Figure 3D In China, for the sake of Figure 3B The same constituent elements are assigned to the same constituent elements as Figure 3B The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.

[0204] Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 differs from decoding device 1110 in that it has a receiving unit 1131 that receives input data 1113 from the outside.

[0205] The receiving unit 1131 receives the received signal 1132 to obtain received data, and outputs the input data 1113 input to the decoder 1112. The received data can be the same as the input data 1113 input to the decoder 1112, or it can be data in a different format than the input data 1113.

[0206] If the format of the received data differs from the format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the receiving data may be converted into the input data 1113 by a conversion unit (not shown) or CPU of the decoding device 1130. The received data may be, for example, a bit stream, multiplexed data, a file, or a packet as described in the encoding device 1120.

[0207] (Example of a decoder) Figure 4A This is a block diagram representing an example of the decoder's structure. Specifically, Figure 4A Indicates as Figure 3B or Figure 3D The decoder 1200 is an example of the decoder 1112 in the example.

[0208] Input data 1113 is the encoded bitstream, which contains encoded audio data as the encoded audio signal and metadata used in audio processing.

[0209] The Spatial Information Management Unit 1201 acquires the metadata contained in the input data 1113 and parses the metadata. The metadata contains information describing the elements that act on sound and are configured in the sound space. The Spatial Information Management Unit 1201 manages the spatial information used in sound processing obtained by parsing the metadata and provides the spatial information to the Rendering Unit 1203.

[0210] Furthermore, in this disclosure, the information used in sound processing is represented as spatial information, but other representations may also be used. For example, the information used in sound processing may be represented as sound spatial information or scene information. In addition, when the information used in sound processing changes over time, the spatial information input to the rendering unit 1203 may also be represented as information such as spatial state, sound spatial state, or scene state.

[0211] Furthermore, spatial information can be managed on a per-sound-space or per-scene basis. For example, when multiple distinct rooms are represented as virtual spaces, each room can be managed as a separate scene. Additionally, even the same space can be managed as different scenes depending on the context it presents.

[0212] Therefore, it is also possible to manage multiple spatial information for multiple sound spaces or multiple scenes. In the management of multiple spatial information, each spatial information can be assigned an identifier to identify its own spatial information.

[0213] Spatial information data can also be included in a bitstream, which is one example of input data 1113. Alternatively, the bitstream may contain an identifier for the spatial information, and the spatial information data may be obtained from an information source other than the bitstream. Specifically, if the bitstream only includes an identifier for the spatial information, the identifier for the spatial information can be used during rendering to obtain spatial information data stored in the device's memory or an external server as input data 1113.

[0214] Furthermore, the information managed by the Spatial Information Management Department 1201 is not limited to the information contained in the bitstream. For example, the input data 1113 may include data on the characteristics and structure of the representation space obtained from software or servers providing VR or AR, as data not included in the bitstream.

[0215] Furthermore, the input data 1113 may also include data representing the characteristics and location of the listener or object. Additionally, the input data 1113 may include information about the listener's location obtained by sensors possessed by the terminal, including the decoding devices (1110, 1130), and may also include information representing the terminal's location inferred based on the information obtained from the sensors.

[0216] That is, the spatial information management unit 1201 can also communicate with external systems or servers to obtain spatial information and the location of listeners. In addition, the spatial information management unit 1201 can obtain clock synchronization information from external systems and perform clock synchronization processing with the rendering unit 1203.

[0217] Furthermore, the space described above can be a virtually formed space, i.e., VR space, or a real space or a virtual space corresponding to a real space, i.e., AR space or MR space. Additionally, the virtual space can also be represented as a sound field or acoustic space. Furthermore, the information indicating position described above can be information such as coordinate values ​​representing the position within the space, information representing the relative position with respect to a defined reference position, or information representing the motion or acceleration of the position within the space.

[0218] The audio data decoder 1202 decodes the encoded audio data contained in the input data 1113 to obtain the audio signal.

[0219] The encoded audio data acquired by the stereo sound reproduction system 1000 is, for example, a bitstream encoded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3). Furthermore, MPEG-H 3D Audio is merely one example of an encoding method that can be used when generating encoded audio data contained within a bitstream. The encoded audio data can also be a bitstream encoded in other encoding methods.

[0220] For example, the encoding method can also be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method can be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).

[0221] Alternatively, any encoding method other than those described above can be used. For example, PCM (pulse code modulation) data can be used as encoded audio data. In this case, for example, if the number of quantization bits of the PCM data is N, the decoding process can be a process of converting the N-bit binary number into a number form (e.g., floating-point form) that the rendering unit 1203 can process.

[0222] The rendering unit 1203 acquires the sound signal and spatial information, uses the spatial information to perform sound processing on the sound signal, and outputs the sound processed sound signal (sound signal 1111).

[0223] Before rendering begins, the Spatial Information Management Department 1201 reads the metadata of the input signal, detects the rendering items such as objects and sounds specified by the spatial information, and sends them to the Rendering Department 1203. After rendering begins, the Spatial Information Management Department 1201 monitors the temporal changes in the spatial information and the listener's location, updates and manages the spatial information, and sends the updated spatial information to the Rendering Department 1203.

[0224] The rendering unit 1203 generates and outputs an audio signal with added audio processing based on the audio signal contained in the input data 1113 and the spatial information received from the spatial information management unit 1201.

[0225] Spatial information update processing and sound signal output processing with added audio processing can be performed in the same thread. Alternatively, the spatial information management unit 1201 and the rendering unit 1203 can assign processing to separate independent threads. When performing spatial information update processing and sound signal output processing with added audio processing in different threads, the spatial information management unit 1201 and the rendering unit 1203 can set the thread start frequency separately, or they can perform processing in parallel.

[0226] When the spatial information management unit 1201 and the rendering unit 1203 perform processing through different independent threads, computing resources can be allocated to the rendering unit 1203 first. As a result, it is possible to safely perform sound output processing that does not allow even the slightest delay, such as a delay of 0.02 msec (one sample) which would produce noise like a pop.

[0227] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared to the output processing of audio signals, the updating of spatial information is a low-frequency process (e.g., updating the orientation of the listener's face), and therefore does not need to be instantaneous like the output processing of audio signals. Therefore, even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality.

[0228] Spatial information updates can be performed periodically at preset times or intervals, or when preset conditions are met. Furthermore, spatial information updates can be performed manually by the listener or the administrator of the sound space, or triggered by changes in external systems.

[0229] For example, the listener can operate the controller to update the spatial information when the listener's avatar's standing position is momentarily distorted or when it moves forward or backward. Alternatively, the spatial information can be updated when the virtual space administrator implements a performance that suddenly changes the environment of the venue. In these cases, the thread used to update the spatial information managed by the Spatial Information Management Department 1201 can be started either periodically or as a single interrupt.

[0230] Figure 4B This is a block diagram representing another example of a decoder. Specifically, Figure 4B Indicates as Figure 3B or Figure 3D Another example of decoder 1112 is the construction of decoder 1210.

[0231] Figure 4B The point is that input data 1113 does not contain encoded audio data but rather unencoded audio signals. Figure 4ADifferent. Input data 1113 includes a bitstream containing metadata and an audio signal.

[0232] Spatial Information Management Department 1211 due to its relationship with Figure 4A The same applies to the Spatial Information Management Department 1201, so the explanation is omitted.

[0233] Rendering Department 1213 due to its relationship with Figure 4A The rendering unit 1203 is the same, so the description is omitted.

[0234] Alternatively, decoders 1112, 1200, and 1210 can also be represented as audio processing units that perform audio processing. Furthermore, decoders 1110 and 1130 can also be audio signal processing devices 1001, or can be represented as audio processing devices.

[0235] (Physical structure of a sound signal processing device) Figure 5 This diagram illustrates an example of the physical configuration of the sound signal processing device 1001. Additionally, Figure 5 The sound signal processing device 1001 can also be Figure 3B Decoding device 1110 or Figure 3D Decoding device 1130. Figure 3B or Figure 3D The multiple constituent elements shown can also be obtained through Figure 5 The various components shown are installed. Furthermore, a portion of the components described herein can also be incorporated into the sound prompting device 1002.

[0236] Figure 5 The sound signal processing device 1001 includes a processor 1402, a memory 1404, a communication interface 1403, a sensor 1405, and a speaker 1401.

[0237] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The audio processing or decoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1404. Furthermore, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 can also be a dedicated circuit that performs signal processing of sound signals, including the audio processing of this disclosure.

[0238] The memory 1404 may be composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory 1404 may also include magnetic recording media such as a hard disk or semiconductor memory such as an SSD. Furthermore, the memory 1404 may be internal memory built into the CPU or GPU. In addition, the memory 1404 may store spatial information managed by the spatial information management unit (1201, 1211). Furthermore, it may also store threshold data, which will be described later.

[0239] The communication IF1403 is, for example, a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices, for example, via the communication IF1403, to obtain the bitstream of the decoded object. The obtained bitstream is stored, for example, in the memory 1404.

[0240] The communication IF1403, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark), etc.

[0241] Furthermore, the communication method is not limited to the wireless communication methods mentioned above. The communication method can also be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).

[0242] Sensor 1405 performs sensing to infer the listener's position and orientation. Specifically, sensor 1405 infers the listener's position and / or orientation based on the detection results of one or more of the following: position, orientation, movement, velocity, angular velocity, and acceleration of a part or the whole of the body, and generates position / or orientation information representing the listener's position and / or orientation.

[0243] Alternatively, the sensor 1405 may be an external device of the sound signal processing device 1001. A part of the body may also be the listener's head, etc. The position / or orientation information may be information indicating the listener's position and / or orientation in real space, or it may be information indicating the displacement of the listener's position and / or orientation based on a predetermined point in time. Furthermore, the position / or orientation information may also be information indicating the relative position and / or orientation to the stereo sound reproduction system 1000 or the external device equipped with the sensor 1405.

[0244] Sensor 1405 can be, for example, a camera or a ranging device such as LiDAR (Light Detection and Ranging). Sensor 1405 can also capture images of the listener's head movements and detect these movements by processing the captured images. Furthermore, a device that uses wireless communication in any frequency band, such as millimeter waves, to perform position estimation can also be used as sensor 1405.

[0245] Furthermore, the sound signal processing device 1001 can also obtain location information from an external device equipped with sensor 1405 via communication IF 1403. In this case, the sound signal processing device 1001 may also not include sensor 1405. Here, the external device is, for example, a... Figure 2 The sound prompt device 1002 described herein may be a stereoscopic image reproduction device worn on the head of the listener. In this case, the sensor 1405 is configured by combining various sensors such as a gyroscope sensor and an accelerometer sensor.

[0246] For example, as the speed of the listener's head movement, sensor 1405 can detect the angular velocity of rotation about at least one of three mutually orthogonal axes in the sound space, and can also detect the acceleration of displacement about at least one of the three axes.

[0247] For example, as a measure of head movement in the listener, sensor 1405 can detect rotational motion about at least one of three mutually orthogonal axes in the sound space, and displacement about at least one of these three axes. Specifically, sensor 1405 detects 6DoF position (x, y, z) and angle (yaw, pitch, roll) as the listener's position. Sensor 1405 is constructed by combining various sensors used for motion detection, such as gyroscopes and accelerometers.

[0248] Alternatively, sensor 1405 can be implemented using a camera or a GPS (Global Positioning System) receiver, which detects the listener's location. Location information obtained by inferring location using LiDAR or similar devices as sensor 1405 can also be used. For example, in the case where the stereo sound reproduction system 1000 is implemented in a smartphone, sensor 1405 can be built into the smartphone.

[0249] Furthermore, sensor 1405 may also include a temperature sensor such as a thermocouple that detects the temperature of the sound signal processing device 1001. Additionally, sensor 1405 may also include a battery included in the sound signal processing device 1001, or a sensor that detects the remaining amount of the battery connected to the sound signal processing device 1001.

[0250] The loudspeaker 1401 includes, for example, a drive mechanism and an amplifier, such as a diaphragm, a magnet, or a voice coil, which transmits the processed sound signal as a sound prompt to the listener. The loudspeaker 1401 activates the drive mechanism based on the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), causing the diaphragm to vibrate. The diaphragm, vibrating in response to the sound signal, generates sound waves, which propagate through the air and reach the listener's ear, allowing the listener to perceive the sound.

[0251] Additionally, an example is given here of a sound signal processing device 1001 that includes a speaker 1401 and prompts the sound signal after sound processing via the speaker 1401, but the sound signal prompting mechanism is not limited to the above configuration.

[0252] For example, the processed sound signal can also be output to an external sound prompting device 1002 connected via a communication module. Communication via the communication module can be either wired or wireless. Furthermore, as another example, the sound signal processing device 1001 has a terminal for outputting an analog sound signal, to which a cable for an earphone or similar device is connected, and a prompting sound signal is received from the earphone or similar device.

[0253] In the above-described cases, the sound prompting device 1002 may also be a headset, earphone, head-mounted display, neck speaker, or wearable speaker worn on the head or part of the listener's body. Alternatively, the sound prompting device 1002 may also be a surround sound system consisting of multiple fixed speakers. Furthermore, the sound prompting device 1002 can also reproduce sound signals.

[0254] (Physical structure of the encoding device) Figure 6 This is a diagram illustrating an example of the physical configuration of an encoding device. Figure 6 The encoding device 1500 can also be Figure 3A Encoding device 1100 or Figure 3C The encoding device 1120 can also Figure 3A or Figure 3C The multiple constituent elements shown are composed of Figure 6 The installation of the multiple components shown.

[0255] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503, and a communication IF 1502.

[0256] The processor 1501 is, for example, a CPU, DSP, or GPU. The encoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1503. Furthermore, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 can also be a dedicated circuit that performs signal processing on an audio signal, including the encoding processing of this disclosure.

[0257] The memory 1503 may be composed of, for example, RAM or ROM. The memory 1503 may also include magnetic recording media, such as a hard disk, or semiconductor memory, such as an SSD. Furthermore, the memory 1503 may also be internal memory embedded in a CPU or GPU.

[0258] The communication IF1502 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with other communication devices, for example, via the communication IF1502, and transmits the encoded bit stream.

[0259] The communication IF1502, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE, NR, or Wi-Fi (registered trademark), etc. Furthermore, the communication method is not limited to wireless communication. The communication method can also be wired communication methods such as Ethernet (registered trademark), USB, or HDMI (registered trademark).

[0260] (The composition of the rendering department) Figure 7 This is a block diagram illustrating an example of the structure of the rendering unit. Specifically, Figure 7 Indicates and Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213.

[0261] The rendering unit 1300 is composed of a resolution unit 1301, a selection unit 1302, and a synthesis unit 1303. It performs additional audio processing on the audio data contained in the input signal and outputs it.

[0262] The input signal may consist of spatial information, sensor information, and sound data. Alternatively, the input signal may contain a bitstream of sound data and metadata (control information), in which case spatial information may also be included in the metadata.

[0263] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000. It consists of information related to the objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound-emitting objects that do not emit sound. Sound source objects can also be simply represented as sound sources.

[0264] Non-sound-producing objects can act as obstacles reflecting the sound emitted by a sound source object, but there are also cases where a sound source object acts as an obstacle reflecting the sound emitted by other sound source objects. Obstacle objects can also be represented as reflecting objects.

[0265] As information shared by both the sound source object and the non-sound-producing object, it includes location information, shape information, and the attenuation rate of the volume when the object reflects the sound.

[0266] Position information is represented by coordinate values ​​along three axes in Euclidean space, such as the X, Y, and Z axes, but it doesn't necessarily have to be three-dimensional. For example, position information can also be two-dimensional, represented by coordinate values ​​along only the X and Y axes. The position information of an object is determined by the representative position of its shape, represented by a mesh or voxels.

[0267] Shape information can also include information related to the material of the surface.

[0268] The attenuation rate can be represented by a real number above 0 and below 1, or by a negative decibel value. Since volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value. However, for example, to present the sense of horror in an unreal space, the attenuation rate above 1, i.e., a positive decibel value, can be deliberately set.

[0269] Furthermore, the attenuation rate can be set to a different value for each of the multiple frequency bands, or it can be set independently for each frequency band. Additionally, when setting the attenuation rate for each type of material on the object's surface, the corresponding attenuation rate value can be used based on information related to the surface material.

[0270] Furthermore, spatial information can also include information indicating whether an object is a living organism or whether it is a moving object. If the object is a moving object, the position indicated by the location information can also change over time. In this case, information about the changed position or the amount of change is transmitted to the rendering unit 1300.

[0271] Information related to the sound source object includes not only the information shared by the sound source object and the non-sound-producing object, but also the sound data and the information needed to project the sound data into the sound space. Sound data represents information related to the frequency and intensity of the sound, and is data that represents the sound perceived by the listener.

[0272] Audio data is typically a PCM signal, but it can also be data compressed using encoding methods such as MP3. In this case, decoding is required at least before the signal reaches the synthesis unit 1303, so the rendering unit 1300 may also include a decoding unit (not shown). Alternatively, the signal can also be decoded by the audio data decoder 1202.

[0273] For a sound source object, one sound data point or multiple sound data points can be set. Additionally, recognition information can be assigned to each sound data point, and information related to the sound source object can also include this recognition information.

[0274] Information needed to project sound data into the sound space may include, for example, information about the reference volume used as a reference when reproducing sound data, information about the properties (also called characteristics) of the sound data, information about the location of the sound source object, and information about the orientation of the sound source object (i.e., information about the directivity of the sound emitted by the sound source object).

[0275] Reference volume information can be, for example, the effective value of the amplitude of the sound data at the sound source location when the sound data is radiated into the sound space, or it can be represented as a decibel (dB) value in floating point.

[0276] For example, a reference volume of 0 dB could mean that the volume of the signal level represented by the audio data is not increased or decreased, and sound is radiated into the sound space at the original volume from the position indicated by the information related to the location of the sound source object. Alternatively, a reference volume of -6 dB could mean that the volume of the signal level represented by the audio data is reduced to approximately half, and sound is radiated into the sound space from the position indicated by the information related to the location of the sound source object.

[0277] The reference volume information can be assigned to each sound data point or to multiple sound data points at once.

[0278] Information representing the properties of sound data can be, for example, information related to the volume of the sound source, and information representing the time-series variation of the sound source's volume.

[0279] For example, in a virtual conference room where the sound space is a speaker and the sound source is the speaker, the volume changes intermittently over a short period of time. That is, the spoken and silent parts alternate. Conversely, in a concert hall where the sound space is a performance hall and the sound source is the performer, the volume is maintained for a certain duration. Furthermore, in a battlefield where the sound space is an explosive device and the sound source is an explosive device, the volume of the explosion only increases momentarily, then remains silent or low.

[0280] In this way, the volume information of the sound source includes not only the loudness of the sound, but also information about changes in loudness. This information can also be used to represent the properties of the sound data.

[0281] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can be represented by data showing the duration of a sound interval. Transition information can be represented by time-series data showing the duration of both the sound interval and the duration of the silent interval. Transition information can also be represented by data that lists multiple groups of time series, showing the duration of which the amplitude of the sound signal can be considered stable (or approximately constant), and the amplitude values ​​of the signal during those periods.

[0282] Transition information can also be represented by data showing the duration during which the frequency characteristics of a sound signal can be considered stable. Transition information can also be represented by data obtained by listing multiple groups of time series, showing the duration during which the frequency characteristics of a sound signal can be considered stable, and the groups of frequency characteristics during that period. Transition information can also be represented, for example, by data showing the approximate shape of a spectrogram.

[0283] Furthermore, the volume used as the reference for the aforementioned frequency characteristics can also be the aforementioned reference volume. Information about the reference volume and information representing the properties of the sound data can be used for calculating the volume of direct or reflected sounds perceived by the listener, or for selecting whether or not to allow the listener to perceive them. Other examples and methods of using information representing the properties of the sound data will be described later.

[0284] Information related to the sound source object may include, for example, information related to the orientation of the sound source object (i.e., information related to the directionality of the sound emitted by the sound source object).

[0285] Information related to the orientation of a sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the roll rotation can be omitted, and the orientation information of the sound source object can be represented by azimuth (yaw) and pitch (pitch). The orientation information of the sound source object can also change over time, and any changes are transmitted to the rendering unit 1300.

[0286] Information relevant to the listener relates to their position and orientation within the sound space. Position information is represented by the XYZ axes in Euclidean space, but it doesn't necessarily have to be three-dimensional; it can also be two-dimensional. Orientation information is typically represented by yaw, pitch, and roll. Alternatively, the roll can be omitted, and the listener's orientation information can be represented by azimuth (yaw) and pitch (pitch).

[0287] The listener's location and orientation information can also change over time, and if changes occur, they will be transmitted to the rendering unit 1300.

[0288] The sensor information includes information such as the amount of rotation or displacement detected by the sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may also include, for example, position information obtained by the portable terminal using GPS, a camera, or LiDAR for self-position estimation.

[0289] Alternatively, instead of sensor 1405, information obtained from an external source via a communication module can be used as sensor information for detection. Information indicating the temperature of the sound signal processing device 1001 and the remaining battery level can also be obtained from sensor 1405. Furthermore, the computing resources (CPU capacity, memory resources, or PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 can be obtained in real time.

[0290] The analysis unit 1301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management unit (1201, 1211), and detects the information required for the generation of direct sound and reflected sound, as well as the information required for the selection of whether to generate reflected sound.

[0291] The information required for the generation of direct and reflected tones includes, for example, values ​​related to the path of each direct and reflected tone to the listening position, the time required to reach it, and the volume at the time of arrival.

[0292] The information needed for selecting the output reflected tone is information representing the relationship between the direct tone and the reflected tone, such as values ​​related to the time difference between the direct tone and the reflected tone, and values ​​related to the volume ratio of the direct tone to the reflected tone at the listening position.

[0293] Furthermore, when volume is expressed in decibels on a logarithmic axis (in the case of representing volume in the decibel range), the volume ratio of two signals is naturally represented by the difference in decibel values. Specifically, the volume ratio of two signals can be the difference between the amplitude values ​​of each signal expressed in the decibel range. This value can also be calculated based on energy values ​​or power values, etc. Moreover, this difference can be referred to as the gain difference or simply the gain difference in the decibel range.

[0294] That is, the volume ratio in this disclosure is essentially the ratio of signal amplitudes, so it can also be expressed as Soundvolume ratio, Volume ratio, Amplitude ratio, Soundlevel ratio, Sound intensity ratio, or Gain ratio, etc. Furthermore, when the unit of volume is decibels, the volume ratio in this disclosure can of course also be referred to as volume difference.

[0295] In this disclosure, "volume ratio" typically refers to the gain difference between two sounds expressed in decibels. In examples of embodiments, the threshold data is also typically defined by the gain difference expressed in decibels. However, the volume ratio is not limited to the gain difference in decibels. When using a volume ratio expressed outside the decibel range, the threshold data defined in the decibel range can be converted to the units of the calculated volume ratio for use. Alternatively, the threshold data defined in each unit can be stored in memory in advance.

[0296] That is, even if a ratio such as energy value or power value is used instead of volume ratio, it is obvious that the algorithm in this disclosure can be applied to the solution of the problem of this disclosure.

[0297] The time difference between direct and reflected tones can be, for example, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone. Alternatively, it can be the time difference between the arrival times of the direct and reflected tones at the listening position, the difference in time required for each to reach the listening position, or the time difference between the end of the direct tone's sound production and the arrival time of the reflected tone at the listening position. The calculation methods for these values ​​will be described later.

[0298] The selection unit 1302 uses the information calculated by the analysis unit 1301 and the threshold data to select whether to generate a reflected sound. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. In other words, the selection unit 1302 selects which reflected sound among multiple reflected sounds to generate.

[0299] Threshold data, for example, can be represented in a graph where the horizontal axis represents the time difference between the direct and reflected sounds, and the vertical axis represents the volume ratio of the direct and reflected sounds. This represents the boundary (threshold) at which the reflected sound is perceived or not. Threshold data can also be represented by an approximation using the time difference between the direct and reflected sounds as a variable, or by an arrangement of values ​​indexed by the time difference between the direct and reflected sounds and corresponding thresholds.

[0300] For example, if the ratio of the volume of the direct tone to the volume of the reflected tone is greater than a threshold set by the reference threshold data, the selection unit 1302 selects to generate the reflected tone.

[0301] The time difference between the arrival time of the direct tone and the arrival time of the reflected tone is, in other words, the difference in time required for the direct tone and the reflected tone to reach the listening position respectively. Alternatively, the time difference between the end of the direct tone's emission and the arrival time of the reflected tone at the listening position can also be used as the time difference between the direct tone and the reflected tone. In this case, different threshold data can be used than the threshold data set based on the time difference between the arrival times of the direct tone and the reflected tone, or a common threshold data can be used.

[0302] The threshold data can be obtained from the memory 1404 of the audio signal processing device 1001, or from an external storage device via the communication module. The method for storing the threshold data and the method for setting the threshold will be described later.

[0303] The synthesis unit 1303 synthesizes the direct sound signal with the reflected sound signal selected and generated by the selection unit 1302.

[0304] Specifically, the synthesis unit 1303 processes the input sound signal to generate a direct tone based on the information about the arrival time and volume of the direct tone calculated by the analysis unit 1301. Furthermore, the synthesis unit 1303 processes the input sound signal to generate a reflected tone based on the information about the arrival time and volume of the reflected tone selected by the selection unit 1302. Then, the synthesis unit 1303 synthesizes the generated direct tone and reflected tone and outputs it.

[0305] (The actions of the rendering department) Figure 8This is a flowchart illustrating an example of the operation of the sound signal processing device 1001. Figure 8 The text indicates that the processing is mainly performed by the rendering unit 1300 of the sound signal processing device 1001.

[0306] In the analysis and processing of input signals ( Figure 8 In step S101, the analysis unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct tones and reflected tones that can be generated in the sound space. The reflected tones detected here are candidate reflected tones selected by the selection unit 1302 as the final reflected tones to be generated by the synthesis unit 1303. Furthermore, the analysis unit 1301 analyzes the input signal and calculates the information needed for generating direct tones and reflected tones, as well as the information needed for selecting the target reflected tones.

[0307] First, the characteristics of direct and reflected tones are calculated. Specifically, the arrival time and volume of the direct and reflected tones when they reach the listener are calculated. If multiple objects exist in the sound space as reflection targets, the characteristics of the reflected tones are calculated for each object separately.

[0308] The direct arrival time (td) is calculated based on the direct arrival path (pd). The direct arrival path (pd) is the path connecting the location information S(xs, ys, zs) of the sound source object to the location information A(xa, ya, za) of the listener. The direct arrival time (td) is the value obtained by dividing the length of the path connecting the location information S(xs, ys, zs) and the location information A(xa, ya, za) by the speed of sound (approximately 340 m / s).

[0309] For example, the path length (X) is calculated using (xs-xa)^2 + (ys-ya)^2 + (zs-za)^2)^0.5. Volume decreases inversely with distance. Therefore, given the volume as N in the location information S(xs, ys, zs) of the sound source object and the unit distance as U, the volume (ld) when the direct sound arrives is calculated using ld = N. Find U / X.

[0310] The volume N at the sound source location can also be the reference volume described earlier.

[0311] The arrival time (tr) of the reflected sound is calculated based on the arrival path (pr). The arrival path (pr) is the path that connects the position of the sound image of the reflected sound with the position information A (xa, ya, za).

[0312] Furthermore, the location of the sound image of reflected sound can be derived using methods such as the "mirror method" or the "ray tracing method," or any other method. The mirror method assumes that the reflected wave on the wall of a room has a mirror image at a position symmetrical to the sound source relative to the wall, and simulates the sound image by assuming that sound waves are emitted from that mirror image location. The ray tracing method simulates the image (sound image) observed at a certain point by tracing waves that propagate in straight lines, such as light rays or sound rays.

[0313] Figure 9 It is a diagram that shows the relative positions of the listener and the obstacle. Figure 10 This is a diagram showing the relatively close positional relationship between the listener and the obstacle object. That is, Figure 9 and Figure 10 These examples illustrate sound images formed at positions symmetrical to the sound source location, separated by a wall. By determining the position of the sound image of the reflected sound on the x, y, and z axes based on this relationship, the arrival time of the reflected sound can be calculated in the same way as the arrival time of the direct sound.

[0314] The arrival time (tr) of the reflected sound is obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound to the position information A (xa, ya, za) by the speed of sound (approximately 340 m / s). Volume decreases inversely with distance. Therefore, given a volume of N at the sound source location, a unit distance of U, and a volume attenuation rate of G in the reflection, the volume (lr) upon arrival of the reflected sound decreases by lr = N. G Find U / Y.

[0315] As explained above, the attenuation rate G can be represented by a real number greater than or equal to 0 and less than 1, or by a negative decibel value. In this case, the overall volume attenuation of the signal corresponds to the amount of G. Furthermore, the attenuation rate can also be set for each of the multiple frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate for each frequency component of the signal. Additionally, to reduce computational complexity, the analysis unit 1301 can use representative values ​​or average values ​​of multiple attenuation rates from multiple frequency bands as the overall attenuation rate, thereby correspondingly attenuating the overall volume of the signal.

[0316] Next, the analysis unit 1301 calculates the volume ratio (L) required in the selection of the reflected sound of the generated object, namely the ratio of the volume (ld) when the direct sound arrives to the volume (lr) when the reflected sound arrives, and the time difference (T) between the direct sound and the reflected sound.

[0317] The ratio of the volume (ld) of the direct tone to the volume (lr) mentioned above is the volume ratio (L). For example, L = (N G U / Y) / (N) U / X) = G X / Y is calculated. Since the calculated value is a volume ratio, the values ​​of N and U can be any pre-set values.

[0318] Furthermore, the volume ratio (L) can also be adjusted in the following situations. For example, to remove unwanted noise or emphasize specific components, the audio signal is sometimes filtered to allow or attenuate components of a specific frequency band, or to change their phase. Filtering processes can be, for example, low-pass filtering, high-pass filtering, band-pass filtering, or full-pass filtering, or other filtering processes.

[0319] The reflected sound selected without considering the filtering process described above may not be suitable as the reflected sound for generation after applying filtering. That is, due to the newly assigned sound effect, the reflected sound may not be properly selected, leading to a deterioration in quality. Therefore, the volume ratio (L) can be adjusted to match the appropriate selection of the reflected sound with the newly assigned sound effect.

[0320] Specifically, when applying filtering to an audio signal, the parameters of the set filter coefficients can be obtained, and the volume ratio (L) can be adjusted based on these parameters. By adjusting the volume ratio (L) according to the parameters of the filter coefficients, the reflected sound of the output object can be appropriately selected, taking into account the changes in the volume ratio caused by the filtering process.

[0321] Furthermore, as explained above, filtering can be performed to allow or attenuate sounds in a specific frequency band, or to change the phase, or other processing can be performed. For example, filtering can be performed to allow sounds in all frequency bands to pass through, or it can be performed to change only the phase.

[0322] Specifically, the rendering unit 1300 may also have one or more of the following: an additional unit (additional unit), an additional increment / decrement acquisition unit (additional increment / decrement acquisition unit), a volume ratio adjustment unit (volume ratio adjustment unit), an additional delay acquisition unit (additional delay acquisition unit), and a time difference adjustment unit (time difference adjustment unit).

[0323] For example, the additional unit is an example of the filtering process described above, which is a filtering unit (filtering section) that performs filtering processing on reflected sounds to emphasize the phase deviation in order to emphasize the diffusion of sound or the naturalness of reflected sounds.

[0324] The additional increment / decrement acquisition unit acquires the amplitude increment / decrement in the additional unit. The volume ratio adjustment unit adjusts the volume ratio based on the increment / decrement acquired by the additional increment / decrement acquisition unit. The additional delay acquisition unit acquires the delay increment / decrement in the additional unit. The time difference adjustment unit adjusts the time difference based on the increment / decrement acquired by the additional delay acquisition unit.

[0325] These units can also be installed via processor 1402 and memory 1404, etc.

[0326] For example, sometimes a filter (diffusion filter) is applied to a reflected sound to emphasize the phase deviation of that reflected sound. When a diffusion filter that causes a large phase deviation is applied to a reflected sound, the amplitude of the reflected sound is sometimes amplified. Therefore, a parameter can be obtained to control the degree to which the diffusion filter causes a phase deviation, and the volume ratio (L) can be adjusted by taking into account the increase or decrease of the corresponding parameter.

[0327] Alternatively, for example, if the user can set a parameter that increases or decreases the amplitude of the filter coefficients of the diffusion filter, the set parameter can be obtained and the volume ratio (L) can be multiplied by the amplitude increase or decrease specified by the parameter.

[0328] Furthermore, the aforementioned diffusion filter can be called a dispersion filter or diffusion filter, or a decorrelation filter that causes a phase deviation in the sound signal. Additionally, the aforementioned diffusion filter can also be used for purposes other than emphasizing sound diffusion. For example, the aforementioned diffusion filter can also be a filter that amplifies amplitude or delay.

[0329] Alternatively, it can be determined whether filtering is applied to the audio signal, and the volume ratio (L) can be adjusted if filtering is applied. In determining whether filtering is applied to the audio signal, for example, the spatial information management unit (1201, 1211) or the additional unit can parse the metadata contained in the bitstream, and determine whether filtering is applied if the information contained in the metadata indicates that filtering is effective.

[0330] Alternatively, the volume ratio (L) can be adjusted based on the parameters of the filter coefficients applied to the audio signal.

[0331] Furthermore, the determination of whether to apply filtering and whether to adjust the volume ratio (L) can be switched according to each sound category. Sound categories can be considered, for example, direct sound, initial reflection, nth reflection, reverberation, and diffraction. That is, it is also possible to make a determination such as applying filtering to initial reflections and not applying filtering to diffractions.

[0332] The time difference (T) between the direct tone and the reflected tone can also be the time difference between the time required for the direct tone and the reflected tone to reach the listening position, respectively. For example, the time difference (T) between the time required for the direct tone and the reflected tone to reach the listening position is calculated using T = tr - td.

[0333] Furthermore, the time difference (T) can also be the difference between the arrival times of the direct tone and the reflected tone at the listening position. Alternatively, the time difference (T) can also be the time difference between the end of the direct tone's speech and the arrival time of the reflected tone at the listening position. In other words, the time difference (T) can also be the time difference between the end of the direct tone and the beginning of the reflected tone at the listening position.

[0334] Furthermore, the time difference (T) can also be adjusted in the following situations. For example, to remove unwanted noise or emphasize specific components, sometimes the audio signal is filtered to allow or attenuate components of a specific frequency band, or to change their phase. Filtering processes can be, for example, low-pass filtering, high-pass filtering, band-pass filtering, or full-pass filtering, or other filtering processes.

[0335] The reflected sound selected without considering the filtering process described above may not be suitable as the reflected sound for generation after applying filtering. That is, due to the newly assigned sound effect, the reflected sound may not be properly selected, leading to a deterioration in quality. Therefore, the volume ratio (L) can be adjusted to match the appropriate selection of the reflected sound with the newly assigned sound effect.

[0336] Specifically, when applying filtering to an audio signal, the parameters of the set filter coefficients can be obtained, and the time difference (T) can be adjusted based on these parameters. By adjusting the time difference (T) according to the filter coefficient parameters, the reflected sound of the output object can be appropriately selected, taking into account the time difference changed by the filtering process.

[0337] Furthermore, as explained above, filtering can be performed to allow or attenuate sounds in a specific frequency band, or to change the phase, or other processing can be performed. For example, filtering can be performed to allow sounds in all frequency bands to pass through, or it can be performed to change only the phase.

[0338] Specifically, the rendering unit 1300 may also have one or more of the following: an additional unit (additional unit), an additional increment / decrement acquisition unit (additional increment / decrement acquisition unit), a volume ratio adjustment unit (volume ratio adjustment unit), an additional delay acquisition unit (additional delay acquisition unit), and a time difference adjustment unit (time difference adjustment unit).

[0339] For example, the additional unit is an example of the filtering process described above, which is a filtering unit (filtering section) that performs filtering processing on reflected sounds to emphasize the phase deviation in order to emphasize the diffusion of sound or the naturalness of reflected sounds.

[0340] The additional increment / decrement acquisition unit acquires the amplitude increment / decrement in the additional unit. The volume ratio adjustment unit adjusts the volume ratio based on the increment / decrement acquired by the additional increment / decrement acquisition unit. The additional delay acquisition unit acquires the delay increment / decrement in the additional unit. The time difference adjustment unit adjusts the time difference based on the increment / decrement acquired by the additional delay acquisition unit.

[0341] These units can also be installed via processor 1402 and memory 1404, etc.

[0342] For example, sometimes a filter (diffusion filter) is applied to the reflected sound to emphasize the phase deviation of the reflected sound. When a diffusion filter that causes a large phase deviation is applied to the reflected sound, the arrival time (time taken for the reflected sound to arrive) may increase or decrease. Therefore, it is also possible to obtain parameters that control the diffusion time of the diffusion filter (filter length, number of filter coefficients, number of taps), and adjust the time difference (T) by taking into account the increase or decrease in the arrival time corresponding to these parameters.

[0343] Alternatively, for example, if the user can set a parameter to increase or decrease the filter length of the diffusion filter, the set parameter can be obtained, and the increase or decrease in the filter length specified by the parameter can be added to the time difference (T).

[0344] Furthermore, the aforementioned diffusion filter can be called a dispersion filter or diffusion filter, or a decorrelation filter that causes a phase deviation in the sound signal. Additionally, the aforementioned diffusion filter can also be used for purposes other than emphasizing sound diffusion.

[0345] Alternatively, it can be determined whether filtering is applied to the audio signal, and if filtering is applied, the time difference (T) can be adjusted. In determining whether filtering is applied to the audio signal, for example, the spatial information management unit (1201, 1211) or the additional unit can parse the metadata contained in the bitstream, and if the information contained in the metadata indicates that filtering is effective, it can be determined that filtering is applied.

[0346] Alternatively, the time difference (T) can be adjusted based on the parameters of the filter coefficients applied to the audio signal.

[0347] Furthermore, the determination of whether to apply filtering and whether to adjust the time difference (T) can be switched according to each category of sound. For example, sound categories could include direct sound, initial reflection, nth reflection, reverberation, and diffraction. That is, it's also possible to make a determination such as applying filtering to initial reflections and not applying filtering to diffractions.

[0348] The metadata of the input signal can contain various parameters as information related to filtering processing, which can be used to select the processing method.

[0349] For example, as described in Patent Document 4, it may include a flag for switching the filtering process on and off, a flag for switching the filtering process on and off for initial reflections, a flag for switching the filtering process on and off for diffracted sounds, a parameter indicating the duration of the filtering process, a parameter indicating the gain that is increased or decreased during the filtering process, or a parameter indicating the angle information of spatial diffusion, etc. The parameter indicating the duration of the filtering process may also indicate the number of filter coefficients (number of taps).

[0350] The additional unit can also parse the metadata contained in the input signal to determine whether to perform filtering processing. In addition, the additional increment / decrement acquisition unit or the additional delay acquisition unit can also parse the metadata contained in the input signal to obtain the parameters of the set filtering processing, thereby obtaining the amplitude increment / decrement or delay increment / decrement in the additional unit.

[0351] The filtering process described above is also an additional processing method. This additional processing can be applied only to the reflected tones and the direct tones, or only to the direct tones, or it can be applied to both the reflected tones and the direct tones.

[0352] Next, in the selective processing of reflected sounds ( Figure 8 In step S102, the selection unit 1302 selects whether to generate the reflected tone calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether to select the reflected tone as the target reflected tone for generation. If multiple reflected tones exist, the selection unit 1302 selects whether to generate each reflected tone. In other words, the selection unit 1302 selects one or more target reflected tones for generation from the multiple reflected tones.

[0353] Furthermore, the selection unit 1302 is not limited to generation processing; it can also select reflected sounds that are the application targets of other processing. For example, the selection unit 1302 can also select reflected sounds that are the application targets of binaural processing. In addition, the selection unit 1302 generally selects only one or more reflected sounds that are the processing targets. However, the selection unit 1302 can also select only one or more reflected sounds that are not the processing targets. Furthermore, processing can be applied to one or more reflected sounds that are not selected.

[0354] For example, the selection of reflected tones is based on the volume ratio (L) and time difference (T) calculated by the analysis unit 1301. By performing selection processing based on the time difference (T) between the direct tone and the reflected tone, compared to the case where selection processing is based solely on the volume difference between the direct tone and the reflected tone, it is possible to more appropriately select reflected tones that have a greater impact on the listener's perception.

[0355] Specifically, the decision to generate a reflected tone is made, for example, by comparing the volume ratio of the direct tone to the reflected tone, corresponding to the time difference between the direct tone and the reflected tone, with a pre-set threshold. The threshold is set with reference to threshold data. The threshold data is an indicator representing the boundary at which the reflected tone of a direct tone is perceived by the listener, defined by the ratio of the volume (Id) of the arriving direct tone to the volume (lr) of the arriving reflected tone.

[0356] Furthermore, a threshold corresponds to a value expressed as a numerical value set in relation to a time difference (T). Threshold data corresponds to the relationship between the time difference (T) and the threshold, and to tabular data or formulas used to determine or calculate the threshold under the time difference (T). The form and type of threshold data are not limited to tabular data or formulas.

[0357] Figure 11 This is a graph showing the relationship between the time difference of the direct tone and the reflected tone and the threshold. For example, you can also refer to... Figure 11 The threshold data shown is for the volume ratio preset according to each value of the time difference between the direct tone and the reflected tone. Alternatively, one can refer to the data from... Figure 11 The threshold data shown is obtained through interpolation or extrapolation.

[0358] Furthermore, the threshold for the volume ratio under the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. The selection unit 1302 then decides whether to select the reflected sound as the target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than this threshold.

[0359] By using threshold data of volume ratios preset according to each value of the time difference between direct and reflected tones, selection processing that takes into account backmasking or priority effects can be achieved. Detailed explanations of the types, formats, storage methods, and setting methods of the threshold data will be provided later.

[0360] Next, in the generation and processing of direct and reflected tones ( Figure 8 In S103, the synthesis unit 1303 generates a direct sound signal and a reflected sound signal selected by the selection unit 1302 as the target reflected sound, and synthesizes them.

[0361] The direct sound signal is generated by applying the arrival time (td) and arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object contained in the input information. Specifically, the sound data is delayed by the arrival time (td) and multiplied by the arrival volume (ld). The sound data delay process is a process of shifting the position of the sound data back and forth on the time axis. For example, the sound data delay process disclosed in Patent Document 2 without degrading the sound quality can also be applied.

[0362] The sound signal of the reflected sound is generated in the same way as the direct sound by applying the arrival time (tr) and the volume (ld) at the time of arrival calculated by the analysis unit 1301 to the sound data of the sound source object.

[0363] However, the arrival volume (lr) in the generation of reflected tones differs from that of direct tones; it is the value of the attenuation rate G applied to the volume of the reflected tones. G can be an attenuation rate applied across the entire frequency band. Alternatively, it can be a reflectance rate specified for each defined frequency band to reflect the bias of the frequency components generated by reflection. In this case, the application of the arrival volume (lr) can also be implemented as a process of multiplying by the attenuation rate for each frequency band, i.e., the process of a frequency equalizer.

[0364] In the example above, the path lengths of the direct and reflected tone candidates as they reach the listener are calculated. Then, the arrival time and volume are calculated based on each path length. Finally, the reflected tone candidates are selected based on their time difference and volume ratio.

[0365] Alternatively, as another example, selection can be performed based on the path lengths of the direct and reflected sounds as they reach the listener, omitting the calculations of arrival time and volume, as well as time difference and volume ratio. In this case, a threshold corresponding to the path length difference can be pre-set for the path length ratio. Furthermore, selection can be performed based on whether the calculated path length ratio is above the threshold corresponding to the calculated path length difference. Thus, selection can be performed based on the path length difference corresponding to the time difference while reducing computational load.

[0366] In addition to the path length difference, the value of a parameter representing the speed of sound propagation, or the value of a parameter that affects the speed of sound propagation, can also be used.

[0367] (Select the detailed processing option) The details of the selection process for whether or not to generate reflected sound are explained.

[0368] The selection of the reflected sound is performed by comparing a threshold value for the volume ratio of the direct sound to the reflected sound at a given time difference (T), i.e., a volume ratio threshold, with a volume ratio (L) calculated by the analysis unit 1301. For example, among the volume ratio thresholds preset for each value of the time difference between the direct sound and the reflected sound, the threshold value for the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced. Furthermore, whether the reflected sound is selected as the target reflected sound is determined based on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value.

[0369] The time difference (T) can be, for example, the difference in the time when the direct tone and the reflected tone arrive at the listening position, the time difference in the time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the time when the direct tone ends and the time when the reflected tone arrives at the listening position. Here, the end time of the direct tone can also be obtained, for example, by adding the duration of the direct tone to the time of its arrival.

[0370] Regarding threshold data, it can also be determined through auditory nerve action or cognitive function of the brain. More specifically, it can be determined based on the minimum time difference between two sounds that the listener can perceive and detect, through the prioritization effect, the time-dependent masking phenomenon, or a combination thereof (described later). Specific values ​​can be derived from known research findings on time-dependent masking effects, prioritization effects, or echo detection limits, or they can be determined through audiovisual experiments applied to this virtual space.

[0371] Figure 12A , Figure 12B and Figure 12C This is a diagram illustrating an example of how threshold data is set. For example... Figure 12A , Figure 12B and Figure 12C As shown, the threshold data is represented by the boundary (threshold) at which the reflected sound is perceived or not in the graph where the horizontal axis represents the time difference between the direct sound and the reflected sound, and the vertical axis represents the volume ratio between the direct sound and the reflected sound.

[0372] Threshold data can also be approximated by using the time difference between the direct tone and the reflected tone as a variable. Furthermore, threshold data can also be used as... Figure 11 The indexes of the time differences between the direct and reflected tones, and the corresponding thresholds, are stored in the region of memory 1404.

[0373] In addition, in the Figure 12CIn Example 4, when the height of the line parallel to the horizontal axis (the minimum audible limit) is used as the threshold, the comparison is not between the volume ratio (L) of the reflected sound and the direct sound, but rather between the volume of the reflected sound itself and the threshold. This is because the threshold represents the volume boundary between what can be perceived by the listener, and is used to determine sounds with volumes less than this threshold as unreproducible sounds. In other words, the threshold corresponding to the minimum audible limit is not a threshold for the ratio of the volume of the reflected sound to the volume of the direct sound.

[0374] When using the minimum audible limit as the threshold, the threshold is constant and independent of the time difference (T), so the time difference (T) can be disregarded.

[0375] Additionally, in the parsing process ( Figure 8 In the case where multiple reflected sounds are generated in S101, selection processing can be performed on all reflected sounds, or selection processing can be performed only on reflected sounds with high evaluation values ​​based on the evaluation values ​​derived from each reflected sound using a pre-set evaluation method. Here, the evaluation value of a reflected sound corresponds to its perceived importance. Furthermore, a high evaluation value corresponds to a large evaluation value, and these interpretations can be interchanged.

[0376] The selection unit 1302 may, for example, calculate the evaluation value of the reflected sound using a pre-set evaluation method corresponding to the volume of the sound source, the visuality of the sound source, the locality of the sound source, the visuality of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.

[0377] Specifically, a higher volume of the sound source can result in a higher evaluation value. Furthermore, to ensure consistency between visual and acoustic localization, a higher evaluation value can also be achieved when the sound source object or its reflection (obstacle) is visible to the listener, or when the sound source object has high localization.

[0378] Furthermore, the opening of the angles of arrival of the direct and reflected tones, as well as the difference in their arrival times, significantly impact spatial perception. Therefore, a higher evaluation value can be achieved when the openings of the angles of arrival of the direct and reflected tones are large, or when the difference in their arrival times is significant.

[0379] The volume information of an audio source can also represent a reference volume determined by each content, a change in volume over time, or both.

[0380] For example, in a virtual conference room where the virtual space is a virtual meeting room and the direct audio is conversational sound, the volume changes intermittently over a short period of time. That is, the audible and silent parts alternate. Conversely, in a concert hall where the virtual space is a concert hall and the direct audio is a musical performance, the volume is maintained for a certain duration. Finally, in a battlefield where the virtual space is a battlefield and the direct audio is an explosion, the volume increases only momentarily, then remains silent or low.

[0381] In this way, the volume information of the sound source not only includes information about the reference volume that corresponds to the volume setting when the sound is radiated into the virtual space, but also information about the change in sound volume.

[0382] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can be represented by data showing the duration of a sound interval. Transition information can be represented by time-series data showing the duration of both the sound interval and the duration of the silent interval. Transition information can also be represented by data that lists multiple groups of time series, showing the duration of which the amplitude of the sound signal can be considered stable (or approximately constant), and the amplitude values ​​of the signal during those periods.

[0383] Information about the transition can also be represented by data showing that the frequency characteristics of a sound signal can be considered to be stable over a certain duration. Information about the transition can also be represented by data obtained by listing multiple groups of time series, showing the stable duration of the frequency characteristics of a sound signal and the groups of frequency characteristics during that period.

[0384] Furthermore, there has been extensive effort to use the temporal transformation of the frequency characteristics of a signal for audio processing in virtual space (Patent Document 1, etc.). Given this prior art, the aforementioned group could of course also be a group of time lengths with constant frequency characteristics and such frequency characteristics.

[0385] Geometric relationships can also refer to the relationships between the positions of a sound source, a listener, and a reflecting object within a virtual space. Based on these relationships, the path lengths of both direct and reflected sounds can be calculated geometrically. Therefore, by utilizing the inverse relationship between volume and distance, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.

[0386] The reflection coefficient of the reflecting object can be used in the calculation of the reference volume of the reflected sound. Alternatively, a typical value that is commonly used can be used as the reflection coefficient. On the other hand, in special cases, such as when the reflecting object is covered by sound-absorbing material, a specially assigned reflection coefficient can be used as the reflection coefficient of the reflecting object.

[0387] Reflected sound can also be evaluated by its volume. The volume of the reflected sound can be determined based on the geometric relationship between the direct sound and the reflected sound as described above, as well as the index assigned to the reflecting object. The volume can also be compared to a predetermined threshold to evaluate the reflected sound.

[0388] Furthermore, information representing the temporal change in the volume of the sound source can also be reflected in the evaluation. For example, if the information representing the temporal change in the volume of the sound source represents the duration of the sound interval, the evaluation value of the reflected sound can remain unchanged when the time is within the sound interval. On the other hand, when the time is outside the sound interval, even if the reference volume of the reflected sound exceeds a threshold, the evaluation value of the reflected sound can be reduced or reduced to zero.

[0389] Alternatively, information representing the temporal changes in the volume of a sound source can also be data obtained by listing multiple groups in a time series, where the amplitude of the sound signal is considered to be approximately constant for a duration and the amplitude value of the signal during that period. In this case, the processing of the reflected sound can be evaluated by changing the reference volume of the reflected sound in conjunction with the changes in the amplitude values ​​in the data.

[0390] Alternatively, the volume information representing the direct tone can be obtained by using both reference volume information and volume information that has changed over time. For example, after calculating an evaluation value based on the reference volume information, the evaluation value can be corrected using the volume information to be changed.

[0391] In evaluating reflected sounds, all of the above methods can be performed, or only some of them can be performed. For example, reflected sounds can be evaluated using multiple evaluation methods, or they can be evaluated using only one evaluation method.

[0392] When evaluating reflected sounds using multiple evaluation methods, the decision on whether to select a reflected sound can be based on the overall evaluation value determined by the multiple evaluation methods, or it can be based on the individual evaluation values ​​of the multiple evaluation methods.

[0393] When the sound signal processing device 1001 decides whether to select a reflected sound based on multiple evaluation methods, it may also select the sound if all evaluation results based on multiple evaluation methods indicate that the sound is selected. Alternatively, the sound signal processing device 1001 may select the sound if any one of the evaluation results based on multiple evaluation methods indicates that the sound is selected.

[0394] Alternatively, priorities can be set for the first to third evaluation methods. Furthermore, the sound signal processing device 1001 can also ultimately determine that sound is not selected if the first evaluation method determines that sound is not selected, regardless of the determination results in the second and third evaluation methods. Additionally, the sound signal processing device 1001 can also ultimately determine that sound is selected if either the second or third evaluation method determines that sound is not selected, but the other method determines that sound is selected.

[0395] Furthermore, the selection process and the evaluation process can be performed independently, or only one of them can be performed. Alternatively, the evaluation process can be performed only on the reflected sounds that are determined to be selected in the selection process, and the selection of the reflected sounds can be determined again in the evaluation process. Or, the evaluation process can be performed only on the reflected sounds that are determined not to be selected in the selection process, and the selection of the reflected sounds can be determined again in the evaluation process.

[0396] The selection process described above can be interpreted as selecting reflected sounds based on the properties of the direct sound. For example, in selecting reflected sounds based on the properties of the direct sound, a threshold used in the selection of reflected sounds is set or adjusted according to the properties of the direct sound. Alternatively, an evaluation value used in the selection of reflected sounds can be calculated based on one or more of the following: the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.

[0397] Furthermore, the process of selecting reflected tones based on the properties of direct tones is not limited to setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values ​​used in selecting reflected tones of the processed object; other processes may also be performed. Moreover, when performing processes such as setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values ​​used in selecting reflected tones of the processed object, a portion of the process may be modified, or new processes may be added.

[0398] In addition, setting thresholds can also include adjusting thresholds and changing thresholds.

[0399] (Method for setting the threshold) The threshold data used in the selection process can be set by referring to known values ​​of echo detection limits based on priority effects or masking thresholds based on backmasking effects.

[0400] The priority effect refers to the phenomenon where, when two sounds are heard from different locations, the listener perceives which sound source was heard first in time. If two short sounds merge and sound like one, the overall sound's perceived location (locality) is largely determined by the location of the initial sound. The echo detection limit, a phenomenon occurring through the priority effect, is the minimum time difference that allows a listener to perceive the discrepancy between two sounds.

[0401] exist Figure 12C In Example 2, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the detectable reflected sound to the direct sound, specifically, the threshold for whether the reflected sound arriving with the delay time can be detected.

[0402] Figure 13 This is a diagram illustrating an example of how a threshold is set. Figure 13 The horizontal axis in the figure corresponds to the arrival time of the reflected sound, specifically the time difference (T) between the direct sound and the reflected sound. Figure 13 The vertical axis corresponds to the volume of the reflected sound. Specifically, Figure 13 The vertical axis can correspond to either the volume of the reflected sound (volume ratio) which is set relatively to the volume of the direct tone, or the volume of the reflected sound which is determined absolutely without depending on the volume of the direct tone.

[0403] For example, in such Figure 9 When the listener is relatively far from the obstacle, the arrival time of the reflected sound is delayed, such as... Figure 13 As shown in C, the threshold is set low. As a result, in Figure 9 In such cases, reflected sounds are generated. On the other hand, in cases such as Figure 10 When the listener is relatively close to the obstacle, the arrival time of the reflected sound is... Figure 9 The condition occurs earlier, such as Figure 13 As shown in B, the threshold is set high. As a result, in Figure 10 In this case, no reflected sound is generated.

[0404] In addition, threshold data can also be stored in memory 1404 and retrieved from memory 1404 for use in selection processing.

[0405] Figure 14 This is a flowchart illustrating an example of the selection process. First, the selection unit 1302 selects the reflected tone detected by the analysis unit 1301 (S201). Next, the selection unit 1302 detects the volume ratio (L) of the direct tone and the reflected tone, as well as the time difference (T) between the direct tone and the reflected tone (S202 and S203).

[0406] The time difference (T) can be, for example, the time difference between the arrival time of the direct tone and the reflected tone at the listening position, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone, or the time difference between the end of the direct tone's sound and the arrival time of the reflected tone at the listening position. An example based on the time difference between the arrival time of the direct tone and the arrival time of the reflected tone is given here.

[0407] Specifically, the selection unit 1302 calculates the difference between the length of the direct sound path and the length of the reflected sound path based on the location information of the sound source object and the listener, as well as the location and shape information of the obstacle object. Furthermore, the selection unit 1302 detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position by dividing this length by the speed of sound.

[0408] The volume reaching the listener decreases proportionally (inversely proportional to distance) to the volume of the sound source. Therefore, the volume of the direct tone is obtained by dividing the volume of the sound source by the length of the direct tone's path. The volume of the reflected tone is obtained by dividing the volume of the sound source by the length of the reflected tone's path and then multiplying by the attenuation rate assigned to the virtual obstacle object. The selection unit 1302 detects the volume ratio by calculating the ratio of their volumes.

[0409] Furthermore, the selection unit 1302 uses threshold data to determine a threshold corresponding to the time difference (T) (S204). Next, the selection unit 1302 determines whether the detected volume ratio (L) is above the threshold (S205).

[0410] When the volume ratio (L) is above the threshold ("Yes" in S205), the selection unit 1302 selects the reflected sound as the reflected sound of the generation target (S206). When the volume ratio (L) is below the threshold ("No" in S205), the selection unit 1302 does not select the reflected sound as the reflected sound of the generation target (S207). That is, in this case, the selection unit 1302 determines the reflected sound as a reflected sound other than the generation target.

[0411] Then, the selection unit 1302 determines whether there is an unspecified reflected sound (S208). If there is an unspecified reflected sound ("Yes" in S208), the selection unit 1302 repeats the above process (S201 to S207). If there is no unspecified reflected sound ("No" in S208), the selection unit 1302 ends the process.

[0412] This selection process can be performed on all reflected sounds generated in the parsing process, or only on reflected sounds with high evaluation values.

[0413] (Details on the threshold storage method) The threshold data for this embodiment is stored in the memory 1404 of the sound signal processing apparatus 1001. The stored threshold data can be of any form and type. When multiple forms and types of thresholds are stored, it is possible to determine which form and type of threshold to use for the selection processing of reflected sounds during the selection process. The method for determining which threshold data to use for the selection process will be described later.

[0414] Furthermore, threshold data of multiple forms and types can be combined and stored. The combined threshold data can also be read from the spatial information management unit (1201, 1211) and set as the threshold used in the selection process. Alternatively, threshold data stored in memory 1404 can be stored in the spatial information management unit (1201, 1211).

[0415] Threshold data can also be stored, for example, as thresholds at various time differences, to depict... Figure 12C The threshold lines shown in [Example 1] and [Example 2] are examples of this.

[0416] In addition, threshold data can also be used as follows Figure 11 The diagram shows a table data storage system that establishes a correspondence between the threshold and the time difference (T). That is, the threshold data can also be stored as table data indexed by the time difference (T). Of course, Figure 11 The threshold shown is an example; the threshold is not limited to... Figure 11 Examples are provided. Alternatively, instead of storing the threshold itself, one can approximate the threshold with a function that takes the time difference (T) as a variable, and store the coefficients of that function. Furthermore, multiple approximations can be combined and stored.

[0417] For example, the time difference (T) can be set as timeDiff, and the threshold can be set as gainThresh, and the threshold data can be represented by the following formula.

[0418] [Mathematical Expression 1] The threshold is defined solely by the time range in which the priority effect occurs. Outside this time range (values ​​less than 1 ms or greater than 40 ms in Equation 1 above), a judgment based on gainThresh may not be performed; instead, the judgment may be made solely by a threshold representing the minimum volume reproduced in the virtual space, as described later. For example, outside this time range, reflected sounds may be selected as the target sound regardless of the priority effect. Alternatively, outside this time range, reflected sounds may be selected as the target sound if they are above the minimum volume, regardless of the priority effect.

[0419] Through experiments, the inventors discovered that, within the time frame where the preference effect occurs, it is preferable to use an upwardly convex function to approximate the threshold. Equation 1 above is an example of an approximation generated based on this experiment.

[0420] Figure 15 This is a diagram representing an auditory experiment involving preceding and following sounds related to the priority effect. Specifically, in Figure 15 The image shows a first virtual speaker, a second virtual speaker, a listener, and the listener's avatar.

[0421] In this listening experiment, the first virtual speaker outputs the direct sound as the preceding sound in virtual space. The direct sound is the sound heard in typical scenarios such as speeches or music. The second virtual speaker outputs the reflected sound as the following sound in virtual space. The reflected sound is the sound after the direct sound has been delayed and attenuated.

[0422] Furthermore, the first HRTF is the head-related transfer function when the sound output from the first virtual speaker reaches the listener's avatar. The second HRTF is the head-related transfer function when the sound output from the second virtual speaker reaches the listener's avatar. The first HRTF and the second HRTF respectively use the HRTF (Head-Related Transfer Function) disclosed in Non-Patent Document 2.

[0423] By sequentially applying various values ​​to the delay and attenuation used in generating the reflected sound, the volume ratio and time difference between the sound output from the first virtual speaker and the sound output from the second virtual speaker are set sequentially.

[0424] Figure 16 It is a graph representing the experimental parameters in the horizontal direction. Figure 17 This is a graph representing experimental parameters in the vertical plane. For example... Figure 16 and Figure 17 As shown, the experimental parameters include multiple versions. Furthermore, the volume ratio in the horizontal plane experimental parameters is set based on the time difference. Specifically, it is envisioned that if the time difference increases, the volume ratio threshold decreases, and if the time difference decreases, the volume ratio threshold increases. Therefore, to avoid unnecessary experiments, the scope of the experimental subjects was narrowed.

[0425] The listener, acting as the subject, listened to the voice of the avatar of the listener using headphones. The listener listened to the sound with the second virtual speaker on ON and with the second virtual speaker off, respectively, and answered whether they could distinguish the difference between ON and OFF. Here, "able to distinguish the difference" includes being able to recognize when the timbre (including impression level) of the direct tone changes due to the reflected tone, and when the direct tone and the reflected tone are heard as separate sounds.

[0426] Based on experimental results, the inventors discovered that the boundary between the difference between ON and OFF can be represented, for example, as shown in Equation 1 (gainThresh) above.

[0427] As shown in Equation 1 above, this boundary becomes an upward-convex curve (broken line) within the time range in which the priority effect occurs. This experimental result does not contradict the premise that the phenomenon is based on the priority effect. That is, it is hypothesized that no priority effect occurs when the time difference between the preceding and following sounds exceeds approximately 40 msec, which is consistent with the experimental result that the threshold decays more sharply as the time difference increases.

[0428] Furthermore, the domain used to represent this boundary can be simply the time range from which the priority effect occurs. Therefore, when the boundary is represented by a mathematical expression, the mathematical expression of the boundary can be represented using a smaller combination of mathematical expressions. Additionally, when the boundary is represented by a line (a collection of points connected at narrow intervals: table data), the line of the boundary can be represented with a smaller memory size.

[0429] This boundary is not significantly related to the placement of each virtual speaker. Therefore, Equation 1 does not include an azimuth factor. However, to more precisely control the selection of reflected sounds, an azimuth factor can be included. For example, an equation of the same type as Equation 1 can be set for each azimuth. Alternatively, an azimuth term can be added to Equation 1.

[0430] The memory 1404 may also store information related to the relational formulas representing the relationship between time difference (T) and threshold. That is, it may also store formulas that take time difference (T) as a variable. The threshold for each time difference (T) may also be approximated by a straight line or curve, and parameters representing the geometric shape of the straight line or curve may be stored. For example, if the geometric shape is a straight line, the starting point and slope of the straight line may also be stored.

[0431] Furthermore, the type and format of threshold data can be set and stored according to each property of the direct tone. Additionally, parameters used to adjust the threshold based on the property of the direct tone and for selection processing can be stored. The process of adjusting the threshold based on the property of the direct tone and for selection processing will be described later as a variation of the threshold setting method.

[0432] As an example of storing multiple threshold data combinations, it can also be like... Figure 12C As shown in [Example 3], the larger of the masking threshold and the echo detection limit threshold is stored for each time difference (T). Alternatively, as... Figure 12CAs shown in [Example 4], the value of the larger of the minimum volume and the echo detection limit threshold that are reproduced in the virtual space is stored for each time difference (T).

[0433] The combination of multiple types of threshold data is not limited to this. For example, information on the maximum value can also be stored in multiple threshold data for each time difference (T).

[0434] Furthermore, in the above, the information related to the threshold has time items as a one-dimensional index. The information related to the threshold can also have a two-dimensional or three-dimensional index that also includes variables related to the direction of arrival.

[0435] Figure 18 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold. For example, as... Figure 18 As shown, thresholds can also be pre-calculated based on the relationship between the direction of the direct tone (θ), the direction of the reflected tone (γ), the time difference (T), and the volume ratio (L).

[0436] The direction of the direct tone (θ) corresponds to the angle relative to the direction of arrival of the direct tone towards the listener. The direction of the reflected tone (γ) corresponds to the angle relative to the direction of arrival of the reflected tone towards the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of the direct tone and the arrival time of the reflected tone towards the listening position. The volume ratio (L) corresponds to the volume ratio of the volume of the direct tone at arrival to the volume of the reflected tone at arrival.

[0437] certainly, Figure 18 The threshold shown is an example; the threshold is not limited to... Figure 18 Examples. Furthermore, in Figure 18 The example primarily illustrates the threshold when the angle (θ) of the direction of arrival of the direct tone is 0 degrees. However, the threshold for cases where the direction of arrival of the direct tone (θ) is other than 0 degrees is also stored in memory 1404.

[0438] Furthermore, in the above, the threshold is stored as an arrangement where the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone are treated as independent variables or indices. However, the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone may also not be used as independent variables.

[0439] For example, the angle difference between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) can also be used. This angle difference corresponds to the angle formed by the direction of arrival of the direct tone and the direction of arrival of the reflected tone, and can also be expressed as the arrival angles of the direct tone and the reflected tone.

[0440] Figure 19This is a graph representing the relationship between angular difference, time difference, and threshold. For example, it can also be like... Figure 19 The example shown stores a pre-calculated threshold, using the angle difference (Φ) between the angle of arrival of the direct sound (θ) and the angle of arrival of the reflected sound (γ) as a variable. Of course, Figure 19 The threshold shown is an example; the threshold is not limited to... Figure 19 Examples.

[0441] exist Figure 19 In the example, the number of variables used in deriving the threshold can be reduced. Therefore, the number of thresholds stored in memory 1404 can be reduced. Consequently, the amount of data stored in memory 1404 can be reduced.

[0442] Furthermore, when using the angle difference (Φ) between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound, the threshold data can also be stored in a two-dimensional arrangement. Additionally, in the selection process, a three-dimensional arrangement can be used to calculate the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound.

[0443] The method of selecting reflected sounds using a threshold corresponding to the direction of arrival will be described later.

[0444] (First variation of the threshold setting method) exist Figure 12A , Figure 12B and Figure 12C In the example, multiple forms and types of thresholds can also be stored in the spatial information management department (1201, 1211). Furthermore, it can be determined which form and type of threshold among the multiple forms and types will be used for the selection processing of reflected sounds. Specifically, it can also be done as follows: Figure 12C As shown in Example 3, the highest threshold is used when the time difference (T) corresponds to the arrival time of the reflected sound.

[0445] Alternatively, as shown in Example 4, a masking threshold, an echo detection limit threshold, and a threshold representing the minimum volume reproduced in the virtual space can be stored. Furthermore, the highest threshold can be used for the time difference (T) corresponding to the arrival time of the reflected sound.

[0446] (Second variation of the threshold setting method) As another example of a method for setting a threshold, a method for setting a threshold based on the properties of a direct tone will be explained.

[0447] Figure 20 It means Figure 7 A block diagram of another configuration example of the rendering unit 1300 shown. Figure 20 The rendering unit 1300 and Figure 7 The rendering unit 1300 differs from the one in that it includes a threshold adjustment unit 1304. The description other than the threshold adjustment unit 1304 is the same as that in... Figure 7 The content described in the previous section is the same, so it is omitted.

[0448] The threshold adjustment unit 1304 selects a threshold to be used by the selection unit 1302 from the threshold data based on information representing the nature of the sound signal. Alternatively, the threshold adjustment unit 1304 may also adjust the threshold included in the threshold data based on information representing the nature of the sound signal.

[0449] Information indicating the nature of the sound signal can also be included in the input signal. Furthermore, the threshold adjustment unit 1304 can also obtain information indicating the nature of the sound signal from the input signal. Alternatively, the analysis unit 1301 can analyze the sound signal contained in the received input signal, derive the nature of the sound signal, and output information indicating the nature of the sound signal to the threshold adjustment unit 1304.

[0450] Information representing the nature of a sound signal can be obtained either before rendering begins or at any time during rendering.

[0451] Furthermore, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, or it may function as a threshold adjustment unit 1304 in another communication device. In this case, the parsing unit 1301 or the selection unit 1302 may also obtain information representing the nature of the sound signal, threshold data corresponding to the nature, or information for adjusting the threshold data according to the nature from other communication devices via the communication IF 1403.

[0452] Figure 21 This is a flowchart representing another example of the selection process. Figure 22 This is another flowchart illustrating the selection process. Figure 21 and Figure 22 In this context, a threshold is set based on the properties of the direct sound. Specifically, in... Figure 21 In this process, the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the properties of the sound signal. Figure 22 In the process, the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the properties of the sound signal.

[0453] The actions in each example are explained below. Additionally, regarding... Figure 14 The examples commonly use omitting explanations.

[0454] First, let me explain in Figure 21The following is an example of the processing. Here, threshold data is pre-stored in memory 1404 according to each property of the direct tone. Thus, multiple threshold data corresponding to multiple properties are pre-stored in memory 1404. Furthermore, the threshold adjustment unit 1304 determines the threshold data to be used in the selection processing of the reflected tone from the multiple threshold data.

[0455] For example, the threshold adjustment unit 1304 obtains the properties of the direct tone based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the properties of the direct tone that are associated with the input signal. Next, the threshold adjustment unit 1304 determines a threshold corresponding to the time difference (T) and the properties of the direct tone (S212).

[0456] In addition, such as Figure 22 As shown, the threshold adjustment unit 1304 can also adjust the threshold determined by the selection unit 1302 based on the properties of the direct tone (S221).

[0457] In any case, the input signal may include information representing the nature of the sound signal, information for adjusting the threshold according to the nature of the sound signal, or both. The threshold adjustment unit 1304 may also use one or both of these to adjust the threshold.

[0458] Furthermore, information indicating the nature of the sound signal, information used to adjust the threshold, or both, can be transmitted via an input signal different from the input signal containing the sound signal. In this case, information relating to an input signal different from the input signal can also be included in the input signal containing the sound signal, or the information relating to the input signal different from the input signal can be stored in memory 1404 along with information about the threshold.

[0459] exist Figure 21 and Figure 22 In the example, the threshold used in selecting the reflected tone is set based on the properties of the direct tone, i.e., the properties of the sound signal. This can be achieved as follows: Figure 21 Using pre-defined threshold data for each property, it is also possible to... Figure 22 In this way, the threshold can be adjusted based on the properties of the sound signal. Alternatively, the parameters of the threshold data can also be adjusted based on the properties of the sound signal.

[0460] Furthermore, the operation performed by the threshold adjustment unit 1304 can also be performed by the analysis unit 1301 or the selection unit 1302. For example, the analysis unit 1301 may obtain the properties of the sound signal. Alternatively, the selection unit 1302 may set the threshold based on the properties of the sound signal.

[0461] Next, the relationship between the properties of the sound signal and the threshold will be explained.

[0462] Two short sounds arriving at a listener's ear consecutively, if the time interval between them is sufficiently short, are perceived as a single sound. This phenomenon is called the priority effect. It is known that the priority effect occurs only for discontinuous, i.e., transient sounds (Non-Patent Document 1). Therefore, when the sound signal represents a stationary tone, the echo detection limit can be set lower compared to when the sound signal represents a non-stationary tone.

[0463] That is, based on the characteristics of this priority effect, for example, when the direct tone is a stable sound, the threshold is set to be relatively small. Alternatively, the higher the stability, the smaller the threshold can be set.

[0464] An example of processing when the sound signal is stationary will be explained. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the stationarity based on the amount of change in the frequency components of the sound signal over time. For example, if the amount of change is small, the stationarity is determined to be high. Conversely, if the amount of change is large, the stationarity is determined to be low. The result of the determination can be used to set a flag representing the level of stationarity, or a parameter representing stationarity can be set based on the amount of change.

[0465] Next, the threshold adjustment unit 1304 may adjust the threshold data or threshold based on information indicating the stability of the sound signal, such as a flag or parameter, and set the adjusted threshold data or threshold as the threshold data or threshold used in the selection unit 1302.

[0466] Alternatively, parameters for setting threshold data based on information representing the stability of the direct tone can be pre-stored in the memory 1404. In this case, the threshold adjustment unit 1304 can also determine the stability of the sound signal and set the threshold data used in selecting the reflected tone based on the information and parameters representing the stability.

[0467] Alternatively, multiple parameters of the threshold data may be pre-stored in the memory 1404 corresponding to multiple patterns of the direct tone's stability. In this case, the threshold adjustment unit 1304 may also determine the stability of the sound signal, select parameters of the threshold data based on the pattern of the direct tone's stability, and set the threshold data used in the selection of reflected tones based on the parameters of the threshold data.

[0468] In addition, the stability of a sound signal can be determined based on the amount of change in the frequency components of the sound signal each time a sound signal is input.

[0469] Alternatively, the stationarity of the sound signal can be determined based on information representing stationarity that has been pre-associated with the sound signal. That is, information representing the stationarity of the sound signal can be pre-associated with the sound signal and stored in the memory 1404. The parsing unit 1301 can also obtain the information representing stationarity associated with the sound signal each time an input sound signal is received. Furthermore, the threshold adjustment unit 1304 can adjust the threshold based on the information representing stationarity associated with the sound signal.

[0470] As another example of setting a threshold based on the properties of the sound signal, the application range of the echo detection limit can be set shorter when the sound signal represents a shorter sound (such as a click) compared to when the sound signal represents a longer sound. This processing is based on the characteristics of the priority effect.

[0471] It is known that, through the priority effect, two short sounds arriving consecutively at a listener's ear are perceived as a single sound if the time interval between them is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is approximately 5 ms for a click sound, and sometimes as high as 40 ms for complex sounds such as human voices or music (Non-Patent Document 1).

[0472] Based on this priority effect, for example, in the case of sounds with shorter direct tone durations, a shorter duration threshold is set. Furthermore, the shorter the direct tone duration, the shorter the duration threshold is set.

[0473] Setting a shorter time threshold means setting a threshold corresponding to the echo detection limit based on the priority effect characteristics within a range where the time difference (T) between the direct tone and the reflected tone is small. Outside this range, no threshold corresponding to the echo detection limit based on the priority effect characteristics is set. That is, outside this range, the threshold is small. Therefore, setting a shorter time threshold for shorter sounds corresponds to setting a smaller threshold for shorter sounds.

[0474] As another example of setting a threshold based on the nature of the direct sound, the threshold can be set lower when the direct sound is an intermittent sound (speech, etc.) compared to when the direct sound is a continuous sound (music, etc.).

[0475] For example, when the direct sound corresponds to speech, there are repeated vocal and non-vocal parts, and as a masking effect, only the aftermasking effect occurs in the non-vocal part. On the other hand, when the direct sound is a continuous sound like musical content, both the aftermasking effect and the simultaneous masking effect based on the sound produced at that time occur. Therefore, the comprehensive masking effect is higher in the case of music, etc., than in the case of speech, etc.

[0476] Based on the masking effect characteristics described above, the threshold can be set higher in the case of music, etc., compared to the case of speech, etc. Conversely, the threshold can be set lower in the case of speech, etc., compared to the case of music, etc. That is, the threshold can be set lower when there are many interruptions in the direct sound.

[0477] As mentioned above, information indicating the properties of a direct tone can also include information about its stability, discontinuity, and duration. Furthermore, information indicating the properties of a direct tone can be any combination of these characteristics. Additionally, information indicating the properties of a direct tone can be information about the temporal variation of any one of these characteristics, or information about the temporal variation of any combination thereof. In other words, information indicating the properties of a direct tone can also be information about the temporal variation of the direct tone.

[0478] For example, as shown in the explanation of stability determination, information representing the properties of a direct tone can also be time-series data of frequency characteristics. Here, frequency characteristics can also be expressed in conventional forms such as gain values ​​for each frequency band, Fourier series of the signal over the time axis, or LPC coefficients or cepstral coefficients used to calculate the frequency envelope.

[0479] Furthermore, information representing the properties of a direct tone can also be obtained by listing multiple groups in time sequence the duration of stable amplitude of the signal and the amplitude value of the signal during that period (the approximate shape of the amplitude envelope). Here, the amplitude value can also be expressed as a ratio relative to a reference volume.

[0480] Furthermore, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. For example, the information representing the properties of a direct tone can also be information representing the stability of the frequency characteristics of the direct tone. Specifically, the information representing the properties of a direct tone can also be information obtained by listing multiple groups of states with small frequency characteristic variations in a time series, along with the frequency characteristics of the signal during those periods (approximate shape of the spectrum). Here, the volume used as a reference for the aforementioned frequency characteristics can also be the aforementioned reference volume.

[0481] For example, information indicating the temporal variation of a direct tone is information representing the envelope of the direct tone. Information indicating the temporal variation of a direct tone can also be found in... Figure 12C The “minimum audible limit” described in [Example 4] is used when the threshold is a threshold. The signal compared with the minimum audible limit is the volume of the reflected sound.

[0482] The volume of the reflected sound is calculated geometrically based on the positional information of the sound source, the listener, and the reflecting object. Specifically, a reference volume of the reflected sound is obtained relative to a reference volume of the sound source. By using information about changes in the volume of the sound source as information representing the properties of the direct sound, and adjusting the reference volume of the reflected sound, the volume of the reflected sound at any given moment can be accurately determined. This is because changes in the volume of the sound source are reflected in changes in the volume of the reflected sound.

[0483] After adjusting the volume of the reflected sound, by comparing the volume of the reflected sound with the threshold, it is possible to more accurately and appropriately select the reflected sound that is audibly desired.

[0484] Of course, the same result can be obtained by adjusting the threshold based on the reciprocal of the change in the volume of the sound source, without adjusting the reference volume of the reflected sound, and then comparing the adjusted threshold with the reference volume of the reflected sound. In other words, the reference volume of the reflected sound can be adjusted using information about the change in the volume of the sound source, and the threshold can also be adjusted using the same information. The adjustment of the reference volume of the reflected sound corresponds to the adjustment of the threshold.

[0485] Depending on the surface composition of the object reflecting the sound, the reflectivity of the sound (and the attenuation rate of the reflected sound) varies for each frequency band. Therefore, as described later, a correlation can be established between the reflectivity (attenuation rate) of the sound and the object reflecting the sound, for each frequency band. Based on this reflectivity information and the information from the spectrogram, it is possible to more accurately determine whether to select the reflected sound. For example, the following processing can be performed.

[0486] Specifically, for example, information from the spectrogram shows that, within a certain time interval, high-frequency components are more dominant than low-frequency components. Additionally, for example, information from the reflectivity of sound shows that, in high-frequency components, the reflectivity is extremely low compared to low-frequency components.

[0487] In this case, even if the amplitude of the signal on the time axis of the sound source is large, the volume of the reflected sound is reduced by multiplying the frequency components represented by the information in the spectrogram with the attenuation rate of each frequency band represented by the information in the reflectivity, and the reflected sound may not be selected.

[0488] As mentioned above, information representing the properties of a direct tone can also be information representing the temporal variation of the direct tone. For example, information representing the properties of a direct tone can also be a value obtained by analyzing the direct tone over a predetermined time period.

[0489] Specifically, information representing the properties of a direct tone can also be obtained by calculating the average energy or average amplitude of the direct tone for each predetermined time length. Alternatively, information representing the properties of a direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each short time analysis length and then calculating a weighted average of the energy or average amplitude for each long time analysis length longer than the short time analysis length.

[0490] More specifically, for example, information representing the temporal variation of a direct tone could be obtained by calculating the energy or average amplitude of the direct tone for each pre-defined short time length (e.g., 5 ms, after which frames of that time length are represented as analysis frames). Alternatively, information representing the temporal variation of a direct tone could also be represented by a weighted average of the energy or average amplitude calculated over the past N-1 analysis frames.

[0491] Assuming the energy of the nth analysis frame is represented by E(n), the information I(n) representing the properties of the direct tone is obtained according to the following formula.

[0492] [Mathematical Expression 2] Here, the parameter a(i) represents the weighting coefficient. Typically, a(i) is set to a(i) ≥ 0 and the sum of a(i) is 1. However, the method of setting a(i) is not limited to this.

[0493] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the time variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.

[0494] Alternatively, the information I(n) representing the properties of direct sounds can be obtained using the following formula.

[0495] [Mathematical Expression 3] Here, the parameter b(i) represents the weighting coefficient. Typically, b(i) is set to b(i) ≥ 0 and the sum of b(i) is 1. However, the method of setting b(i) is not limited to this.

[0496] In this formula, information I(n) representing the properties of the direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with less computation.

[0497] Equations 2 and 3 above can be viewed as filters where E(n) is the input signal and I(n) is the output signal. In this case, Equation 2 is a filter for the moving average (MA) model, and Equation 3 is a filter for the autoregressive (AR) model, both exhibiting the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines both, can also be used.

[0498] Furthermore, the method for deriving information representing the temporal variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. As mentioned above, the information representing the temporal variation of the direct tone is a value obtained by analyzing the direct tone over a predetermined time length. The direct tone can also be analyzed from a perspective other than average energy.

[0499] Furthermore, as mentioned above, information indicating the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. Information related to the frequency characteristics of a direct tone can also be information calculated using the frequency characteristics of the direct tone. For example, information related to the frequency characteristics of a direct tone can also be obtained by averaging the low-frequency components of the direct tone over a predetermined analysis length to obtain the average energy of the low-frequency components.

[0500] Specifically, the low-frequency components of the direct tone are determined by applying a low-pass filter to the direct tone contained in the analysis frame length. Based on the energy or average amplitude of this low-frequency component, information representing the properties of the direct tone is derived in the same manner as in Equation 2 above.

[0501] Assume the energy of the low-frequency components in the nth analysis frame is determined by E L In the case of (n), the information I(n) representing the nature of the direct sound is obtained according to the following formula.

[0502] [Mathematical Expression 4] Here, the parameter c(i) represents the weighting coefficient. Typically, c(i) is set to c(i) ≥ 0 and the sum of c(i) is 1. However, the method of setting c(i) is not limited to this.

[0503] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the time variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.

[0504] In addition, similar to Equation 3, the information I(n) representing the properties of direct sounds can also be obtained according to the following formula.

[0505] [Mathematical Expression 5] Here, the parameter d(i) represents a weight coefficient. Usually, d(i) is set such that d(i) ≥ 0 and the sum of d(i) is 1. However, the method of setting d(i) is not limited to this.

[0506] In this formula, the information I(n) representing the properties of the direct sound is recursively obtained. Therefore, the average energy of a long time length can be calculated with a relatively small amount of computation.

[0507] The above formulas 4 and 5 can be regarded as filters with E(n) as the input signal and I(n) as the output signal. In this case, formula 4 is a filter of a moving average (MA) model, and formula 5 is a filter of an autoregressive (AR) model, both having the characteristics of a low-pass filter. In addition, a filter of an ARMA model formed by combining the two can also be used.

[0508] In the above, a filter with low-pass characteristics is used in the method for obtaining the low-frequency components of the direct sound, but the method for obtaining the low-frequency components of the direct sound is not limited to this. In addition, the method for deriving the information representing the temporal variation of the direct sound is not limited to the above calculation formulas or filters, and other well-known methods can also be used. For example, the spectrum of the direct sound can also be calculated by performing a frequency transformation on the direct sound. And the energy or average amplitude of the low-frequency components of this spectrum can also be calculated.

[0509] In addition, in the above, an MA model or an AR model is used in the derivation of the information representing the temporal variation of the direct sound. The coefficients of these models can be preset fixed values or variable values that are variable over time.

[0510] In addition, the relationship between the above analysis frame length and the occurrence interval of the above information update thread can also be as follows.

[0511] For example, when the time length of the analysis frame is TA (msec) and the occurrence interval of the information update thread is TU (msec), the value of N in the above (formula 2) and (formula 4) in the MA filter can also be around the value given by TU / TA. In addition, b(i) and d(i) (1 ≤ i < N) in the above (formula 3) and (formula 5) in the AR filter can also be values such that the time constant of this filter is around TU (msec).

[0512] The reason for the above setting is that it is desired for this filter to converge during the interval of information update.

[0513] On the other hand, in the above-described scenario where the value of the information representing the temporal variation of the direct tone changes excessively rapidly, I(n) can be pre-calculated. Furthermore, the pre-calculated I(n) can be applied to the selection processing of the reflected tone. For example, I(t+tau) can be used in the processing of the frame at time t. Here, tau is a value determined based on the convergence characteristics of the filter. In cases of slow convergence, the value of tau is larger compared to cases of fast convergence.

[0514] Additionally, auditory masking (frequency masking) information calculated based on the direct tone can be used as information representing the characteristics of the direct tone. The auditory masking information represents a threshold value for the amplitude in the frequency domain masked by the direct tone. Alternatively, the amplitude values ​​of reflected tones in the same frequency domain can be compared with the threshold, and reflected tones with amplitude values ​​less than the threshold can be excluded. The amplitude values ​​of reflected tones in the frequency domain can also be obtained by the analysis unit 1301 as information representing the characteristics of the reflected tones.

[0515] By setting the threshold used in selecting reflected sounds based on the properties of the direct sound, the reflected sounds that are audibly required can be appropriately selected, and auditory characteristics can be effectively reflected in the stereo sound reproduction system 1000. The processing of detecting the properties of the direct sound, determining the threshold based on the properties, and adjusting the threshold based on the properties can be performed either during the rendering process or before the rendering process begins.

[0516] For example, these processes can occur during virtual space creation (when the software is created), at the start of virtual space processing (when the software starts or rendering begins), or at timed intervals in information update threads that occur periodically during virtual space processing. Furthermore, virtual space creation can be timed to build the virtual space before the start of sound processing, or it can occur when virtual space information (spatial information) is acquired, or it can occur when the software acquires the information.

[0517] Here, in the information update thread, processing is performed to update spatial information managed by the Spatial Information Management Department (1201, 1211).

[0518] The information update thread is responsible for tasks such as updating the position and orientation of the listener's avatar within the virtual space based on the position and orientation of the VR goggles worn by the listener, or updating the position of moving objects within the virtual space. This processing is provided within a relatively low-frequency processing thread that starts at around tens of Hz.

[0519] In such low-frequency processing threads, the processing of updating information representing the properties of direct tones can still be performed. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. Therefore, the computational load of this processing can be relatively reduced. Furthermore, updating information at an unnecessarily fast frequency risks impulse noise. This risk can be avoided by updating information at a low frequency.

[0520] (The third variation of the threshold setting method) As another example of a method for setting the threshold, the threshold can also be set based on the computing resources (CPU power, memory resources, PC performance, or remaining battery power, etc.) used to process the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources and sets a higher threshold when the amount of computing resources is low. As a result, the volume of more reflected sounds is lower than the threshold, thus reducing the reflected sounds that are processed by both ears and reducing the amount of computing power.

[0521] Alternatively, in cases where signal processing is performed by battery-powered devices such as smartphones or VR headsets, it may be desirable to prioritize long processing times and conserve computing resources. In such cases, the threshold can be set high without needing to monitor the amount or remaining amount of computing resources.

[0522] (4th variation of the threshold setting method) As another example of a method for setting a threshold, the audio signal processing device 1001 or the audio prompting device 1002 may have a threshold setting unit (not shown), allowing the administrator or listener of the virtual space to set the threshold.

[0523] For example, the listener wearing the sound prompt device 1002 could choose between an "energy-saving mode" (fewer reflected sounds and less computation) and a "high-performance mode" (more reflected sounds and more computation). Alternatively, the administrator of the stereo sound reproduction system 1000 or the producer of the stereo sound content could choose the mode. Furthermore, instead of a mode, a threshold or threshold data could be directly selected.

[0524] (The first variation of the rendering department's actions) Figure 23 This is a flowchart illustrating a first variation of the operation of the sound signal processing device 1001. Figure 23 The diagram shows the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, volume compensation processing is added to the operation of the rendering unit 1300.

[0525] For example, the analysis unit 1301 acquires data (input signal) (S301). Next, the analysis unit 1301 analyzes the data (S302). Next, the selection unit 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, the synthesis unit 1303 performs volume compensation processing based on the unselected reflected sounds (S304). Next, the synthesis unit 1303 performs audio processing on both direct and reflected sounds (S305). Finally, the synthesis unit 1303 outputs the direct and reflected sounds as audio (S306).

[0526] In the processes described above (S301 to S306), the processes other than the volume compensation process (S304) are common to the other examples described above, so their descriptions are omitted.

[0527] Volume compensation processing is performed for reflected sounds that were not selected in the selection process. For example, by not selecting reflected sounds in the selection process, a lack of volume perception occurs. Volume compensation processing can suppress the unpleasantness that accompanies this lack of volume perception. Two methods are disclosed as examples of methods for compensating for volume perception. Either method can be used.

[0528] First, the method of compensating for the sense of volume by increasing the volume of the direct tone will be explained. The synthesis unit 1303 generates a direct tone by increasing the volume of the direct tone by an amount corresponding to the volume of the unselected reflected tone. As a result, the sense of volume lost due to the lack of generated reflected tone is compensated.

[0529] When increasing the volume, the synthesis unit 1303 can also increase the volume according to the frequency characteristics of the reflected sound, one frequency component at a time. To enable this process, a predetermined attenuation rate can be assigned to the volume of the reflecting object for each frequency band. Thus, the frequency characteristics of the reflected sound can be derived.

[0530] Next, a method for compensating for the sense of volume by synthesizing reflected sounds into direct sounds will be explained. In this method, the synthesis unit 1303 adds unselected reflected sounds to direct sounds to generate direct sounds, thereby compensating for the sense of volume caused by the lack of generated reflected sounds. The generated direct sounds reflect the volume (amplitude), frequency, and delay of the unselected reflected sounds.

[0531] In the case of increasing the volume of the direct tone, the computational workload of the compensation process is very small, but only the volume is compensated. In the case of synthesizing the reflected tone into the direct tone, the computational workload of the compensation process is larger compared to the method of increasing the volume of the direct tone, but the characteristics of the reflected tone are compensated more accurately.

[0532] In both cases, no reflected tones are generated, only direct tones, thus reducing the overall computational load. In particular, the computational load required for binaural processing, which includes convolutional HRTF processing, is reduced, resulting in a significant reduction in overall computational load. This is because the computational load required for binaural processing is far greater than that required for the aforementioned compensation processing.

[0533] In addition, if the reason for not selecting the reflected sound is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume will not be lost, the reflected sound can be removed without compensation.

[0534] (The second variation of the rendering department's actions) Figure 24 This is a flowchart illustrating a second variation of the operation of the sound signal processing device 1001. Figure 24 The text primarily describes the processing performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, the operation of the rendering unit 1300 includes left and right volume difference adjustment processing.

[0535] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the direction of sound arrival (S402). Next, the selection unit 1302 adjusts the volume difference of the sound perceived by the left and right ears (S403). Furthermore, the selection unit 1302 adjusts the time difference (delay) of the sound arrival perceived by the left and right ears (S404). Based on the adjusted sound information, the selection unit 1302 determines whether to select the reflected sound (S405).

[0536] In the above-described processes (S401 to S405), the processes other than left and right volume difference adjustment (S403) and delay adjustment (S404) are common to the other examples described above, so their descriptions are omitted.

[0537] Figure 25 This is a diagram illustrating the configuration of avatars, sound source objects, and obstacle objects. For example, in the case where the listener's facing direction is 0 degrees, such as... Figure 25 As shown, if the polarities (e.g., positive and negative) of the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) are different, the volume difference generated between the two ears is corrected.

[0538] Specifically, when the polarities of θ and γ are different, the ear that primarily (first) perceives the sound in the direct tone and the reflected tone is different. In this case, the selection unit 1302 adjusts the volume of the direct tone by matching the position of the ear that primarily perceives the reflected tone as a left-right volume difference adjustment (S403). For example, the selection unit 1302 attenuates the volume of the direct tone when it reaches the listener by multiplying the volume of the direct tone when it reaches the listener by (1.0 - 0.3sin(θ)) (0 ≤ θ ≤ 180).

[0539] The selection unit 1302 determines whether to select the reflected sound by calculating the volume ratio of the corrected direct sound volume to the reflected sound volume as described above, and comparing the calculated volume ratio with a threshold. This corrects the volume difference between the two ears, more accurately derives the volume of the direct sound that affects the reflected sound, and more accurately determines whether to select the reflected sound.

[0540] In addition to adjusting the left and right volume difference (S403), the selection unit 1302 can also perform delay adjustment (S404) to match the position of the ear that perceives the reflected sound, thus delaying the arrival time of the direct sound. Specifically, the selection unit 1302 can also delay the arrival time of the direct sound by adding (a(sinθ+θ) / c) ms (where a is the radius of the head and c is the speed of sound) to the arrival time of the direct sound.

[0541] (The third variation of the rendering department's actions) The method for setting a threshold corresponding to the direction of arrival is explained.

[0542] Figure 26 This is another flowchart illustrating the selection process. Regarding... Figure 14 Examples of this type of writing commonly involve omitting explanatory notes. Figure 26 In the example, the selection unit 1302 uses a threshold corresponding to the direction of arrival to select the reflected sound.

[0543] Specifically, the selection unit 1302 calculates the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) determined by using the incarnation's orientation as a reference, based on the direct sound arrival path (pd), the reflected sound arrival path (pr), and the incarnation's orientation information D calculated by the analysis unit 1301. That is, the selection unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The incarnation's orientation corresponds to the listener's orientation. The incarnation's orientation information D may also be included in the input signal.

[0544] The selection unit 1302 uses three indices, including the direct direction of arrival (θ) and the reflected direction of arrival (γ), as well as the time difference (T), according to... Figure 18 The three-dimensional arrangement shown determines the threshold used in the selection process (S232).

[0545] As an example, to illustrate in such Figure 25 The method shown describes the threshold setting method used in the selection process when an avatar, a sound source object, and an obstacle object are configured.

[0546] The position information of the avatar, sound source object, and obstacle object, as well as the orientation information D of the avatar, are obtained from the input signal. Using this position and orientation information D, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound (γ) are calculated when the orientation of the avatar is set to 0 degrees. Figure 25 In this case, the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees).

[0547] Next, refer to Figure 18 The threshold data shown is stored in a three-dimensional arrangement. The threshold is determined from the arrangement region corresponding to the values ​​of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. Even if there is no index corresponding to the calculated values ​​of (θ), (γ), and (T), the threshold corresponding to the nearest index can be determined.

[0548] Alternatively, the threshold can be determined by interpolation, extrapolation, or other processing based on one or more thresholds corresponding to one or more indices close to the calculated values ​​of (θ), (γ), and (T). For example, the threshold corresponding to (20 degrees, 265 degrees, T) can be determined based on four thresholds corresponding to the four indices (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).

[0549] The selection process based on the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound is explained.

[0550] For example, it can also be pre-made and set up as follows: Figure 19 The threshold data shown is obtained by arranging the angle difference (Φ) between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) and the time difference (T) as a two-dimensional index. In this case, the angle difference (Φ) and the time difference (T) are referenced in the selection process. Alternatively, the angle difference (Φ) between the angle of arrival of the direct tone (θ) and the angle of arrival of the reflected tone (γ) can be calculated in the selection process, and the calculated angle difference (Φ) can be used to determine the threshold.

[0551] Alternatively, a threshold data can be set to be used as an index for arranging the combination of the angle difference (Φ), the direction of arrival of the direct sound (θ), and the time difference (T), or the combination of the angle difference (Φ), the direction of arrival of the reflected sound (γ), and the time difference (T).

[0552] Alternatively, it can be set as follows: Figure 18 The threshold data shown is obtained by arranging the values ​​of (θ), (γ), and (T) as a three-dimensional index.

[0553] (The fourth variation of the rendering department's actions) The processing performed by the analysis unit 1301, selection unit 1302 and synthesis unit 1303 described above can also be performed as pipeline processing as described in Patent Document 3.

[0554] Figure 27 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit 1300.

[0555] Figure 27 The rendering unit 1300 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binocular processing unit 1316. These multiple components can also be... Figure 7 The rendering unit 1300 shown can be composed of multiple constituent elements, or it can be made of... Figure 5 It constitutes at least a portion of the multiple constituent elements of the sound signal processing device 1001 shown.

[0556] Pipeline processing refers to dividing the processing used to impart sound effects into multiple processes and executing these processes sequentially. Each process may perform signal processing on the audio signal or generate parameters used in the signal processing.

[0557] The rendering unit 1300 can also perform reverberation processing, initial reflection processing, distance attenuation processing, and binaural processing as a pipeline process. However, these are just examples; pipeline processing may include other processing methods, or it may exclude some of them. For example, pipeline processing may also include diffraction processing and occlusion processing. Furthermore, reverberation processing, for example, can be omitted if it is not needed.

[0558] Furthermore, each process can be represented as a stage. Additionally, the results of each process, and the generated sound signals such as reflected sounds, can be represented as rendering elements. The multiple stages in the pipeline processing and their order are not limited to... Figure 27 The example shown.

[0559] Here, the parameters used in the selection process (arrival path, arrival time, and volume ratio related to direct and reflected tones) can also be calculated in one of the multiple stages used to generate the render item. That is, the parameters used in the selection of reflected tones are calculated in a part of the pipeline processing used to generate the render item. Alternatively, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted, or they may be performed outside of the rendering unit 1300.

[0560] The reverberation processing, initial reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that may be included as stages in pipeline processing are described. Within each stage, metadata contained in the input signal can also be parsed to calculate the parameters used in generating the reflected tones.

[0561] In reverberation processing, the reverberation processing unit 1311 generates parameters used in the generation of a sound signal representing reverberant sound. Reverberant sound refers to the sound that arrives at the listener as reverberation after the direct tone. As an example, reverberant sound is the sound that arrives at the listener after a relatively late stage (e.g., from the arrival of the direct tone to about one hundred and several tens of ms) following the arrival of the initial reflected sound, as described later. It undergoes more reflections (e.g., dozens of times) than the initial reflected sound.

[0562] The reverberation processing unit 1311 refers to the sound signal and spatial information contained in the input signal and calculates the reverberation sound using a pre-prepared function that is used to generate the reverberation sound.

[0563] The reverberation processing unit 1311 can also apply known reverberation generation methods to the sound signal contained in the input signal to generate reverberation. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by spatial information in the application of known reverberation generation methods. Therefore, the reverberation processing unit 1311 can calculate the parameters used to generate the reverberation.

[0564] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating the initial reflection tone based on spatial information. The initial reflection tone is the reflection tone that reaches the listener after more than one reflection in a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct tone) after the direct tone reaches the listener from the sound source object.

[0565] The initial reflection processing unit 1312 calculates, for example, the path of the reflected sound from the sound source object to the listener via the reflection object, by referring to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflecting object such as the structure, and the reflectivity of the reflecting object can also be used in the path calculation.

[0566] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct tone. This path information may also be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected tone, or as a parameter for the selection unit 1314 to select the reflected tone.

[0567] In distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct tone and the reflected tone reaching the listener based on the path lengths of the direct tone and the reflected tone. The volume of the direct tone and the reflected tone reaching the listener is attenuated proportionally to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct tone by dividing the volume of the sound source by the path length of the direct tone, and can calculate the volume of the reflected tone by dividing the volume of the sound source by the path length of the reflected tone.

[0568] In the selection process, the selection unit 1314 selects the object to be reflected based on parameters calculated prior to the selection process. A selection method of this disclosure may also be used in the selection of the object to be reflected.

[0569] Selection processing can be performed on all reflected sounds, or, as described above, on only reflected sounds with high evaluation values. That is, reflected sounds with low evaluation values ​​are automatically disqualified without any selection processing. For example, reflected sounds with very low volume can also be considered as having low evaluation values ​​and are therefore disqualified.

[0570] Furthermore, for example, selection processing can be applied to all reflected sounds. Also, the evaluation value of the selected reflected sounds in the selection process can be determined, and reflected sounds with low evaluation values ​​can be re-selected as not selected.

[0571] Selection and evaluation processes can be executed independently or in combination. When combining selection and evaluation processes, either process can be executed first.

[0572] In the generation process, the generation unit 1315 generates direct tones and reflected tones. For example, the generation unit 1315 generates a direct tone based on the sound signal contained in the input signal, according to the arrival time and volume of the direct tone. Furthermore, regarding the reflected tone selected in the selection process, the generation unit 1315 generates a reflected tone based on the sound signal contained in the input signal, according to the arrival time and volume of the reflected tone.

[0573] In binaural processing, the binaural processing unit 1316 performs signal processing to make the sound signal of the direct tone perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing to make the reflected tone selected by the selection unit 1314 perceived as sound arriving at the listener from the reflecting object.

[0574] For example, the binaural processing unit 1316 performs HRIR DB processing based on the position and orientation of the listener in the sound space, so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.

[0575] Additionally, HRIR (Head-Related Impulse Responses) describes the response characteristics when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from its frequency domain representation to its time domain representation using a Fourier transform. This head-related transfer function represents the changes in sound produced by surrounding objects, including the auricle, head, and shoulders, as a transfer function. The HRIR DB is a database containing such information.

[0576] Furthermore, the position and orientation of the listener in the sound space can be, for example, the position and orientation of a virtual listener in a virtual sound space. Alternatively, the position and orientation of a virtual listener in a virtual sound space can change in accordance with the movement of the listener's head. Furthermore, the position and orientation of a virtual listener in a virtual sound space can also be determined based on information obtained from sensor 1405.

[0577] The program, spatial information, HRIR DB, threshold data or other parameters used in the above processing are obtained from the memory 1404 of the sound signal processing device 1001 or from outside the sound signal processing device 1001.

[0578] Furthermore, pipeline processing may also include other processing. Additionally, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in pipeline processing. For example, the rendering unit 1300 may also include a diffraction processing unit and an occlusion processing unit.

[0579] The diffraction processing unit performs processing to generate a sound signal representing a sound containing diffracted tones, which are caused by an obstacle object in a three-dimensional sound field (space) located between the listener and the sound source object. A diffracted tone is a sound that reaches the listener from the sound source object by bypassing an obstacle object when such an obstacle object exists between the sound source object and the listener.

[0580] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object, bypassing the obstacle object, to reach the listener, and generates the diffracted sound based on the path. In the path calculation, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object, can also be used.

[0581] When a sound source object exists on the opposite side of an obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on spatial information and information such as the material of the obstacle object.

[0582] (Example of a sound source object) In the above, the positional information assigned to the sound source object represents the position of the sound source object as a "point" in the virtual space. That is, in the above, the sound source is defined as a "point sound source".

[0583] On the other hand, a sound source in virtual space can also be defined as an object with length, size, and shape, that is, a non-point sound source that extends spatially. In this case, the distance between the listener and the sound source and the direction of sound arrival are uncertain. Therefore, reflected sounds caused by such sound sources do not need to be analyzed by the analysis unit 1301, or are limited to being selected by the selection unit 1302 regardless of the analysis result. As a result, sound quality degradation that may occur due to failure to select reflected sounds can be avoided.

[0584] Alternatively, a representative point, such as the object's center of gravity, can be determined, and the processing disclosed herein can be applied assuming that sound originates from that point. In this case, the threshold can also be adjusted based on the spatial extension information of the sound source.

[0585] (Examples of direct and reflected sounds) For example, a direct sound is a sound that is not reflected by a reflecting object, while a reflected sound is a sound that is reflected by a reflecting object. A direct sound can also be a sound that reaches the listener from the sound source without being reflected by a reflecting object, and a reflected sound can also be a sound that reaches the listener from the sound source after being reflected by a reflecting object.

[0586] Furthermore, direct tone and reflected tone are not limited to the sound that reaches the listener; they can also be the sound before it reaches the listener. For example, direct tone can also be the sound output from the sound source, or in other words, the sound of the sound source.

[0587] Figure 28 This is a diagram illustrating the transmission and diffraction of sound. For example... Figure 28 As shown, sometimes the direct sound does not reach the listener because of an obstacle object between the sound source and the listener. In this case, the sound emitted from the sound source, passing through the obstacle object, and reaching the listener can also be considered as the direct sound. Furthermore, the sound emitted from the sound source, diffracted through the obstacle object, and reaching the listener can also be considered as the reflected sound.

[0588] Furthermore, the two sounds compared in the selection process are not limited to the direct and reflected sounds of a sound emitted from a single sound source. For example, sound selection can also be performed by comparing two reflected sounds of a sound emitted from a single sound source. In this case, the direct sound in this disclosure can be replaced with the sound that arrives at the listener first, and the reflected sound in this disclosure can be replaced with the sound that arrives at the listener later.

[0589] (Example of bitstream construction) Bitstreams may contain, for example, audio signals and metadata. Audio signals are audio data representing sound, including information related to the frequency and intensity of the sound. Furthermore, metadata contains spatial information related to the space of the sound field, i.e., the sound space.

[0590] For example, spatial information is information relating to the space in which a listener is located when receiving sound based on a sound signal. Specifically, spatial information is information related to a predetermined location (location position) used to position a sound image at a specific location in sound space (e.g., a three-dimensional sound field), that is, information used to enable the listener to perceive sound arriving from a direction corresponding to that predetermined location. Spatial information may include, for example, information about the sound source and location information indicating the listener's position.

[0591] Sound source object information refers to the information about the sound source object that generates sound based on the sound signal. That is, sound source object information is information related to the object (sound source object) that reproduces the sound signal, and it is information related to a virtual sound source object configured in a virtual sound space. Here, the virtual sound space can also correspond to the real space where the object generating the sound is configured, and the sound source object in the virtual sound space can also correspond to the object generating sound in the real space.

[0592] Sound source object information can also represent the location of a sound source object configured in the sound space, the orientation of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living being, and whether the sound source object is a moving object. For example, a sound signal can be associated with one or more sound source objects represented by the sound source object information.

[0593] Bitstreams, for example, have a data structure consisting of metadata (control information) and sound signals.

[0594] The audio signal and metadata can be contained in a single bitstream or in multiple separate bitstreams. Furthermore, the audio signal and metadata can be contained in a single file or in multiple separate files.

[0595] Bitstreams can exist either per audio source or per playback time. When bitstreams exist per playback time, multiple bitstreams can be processed in parallel simultaneously.

[0596] Metadata can be assigned to each bitstream individually, or it can be assigned to multiple bitstreams together as information to control them. In this case, multiple bitstreams can also share the metadata. Alternatively, metadata can be assigned at each playback time.

[0597] In the presence of multiple bitstreams or multiple files, information indicating associated bitstreams or files may be included in more than one bitstream or more than one file. Alternatively, information indicating associated bitstreams or files may be included in each of the individual bitstreams or each of the individual files.

[0598] Here, associated bitstreams or associated files refer, for example, to bitstreams or files that may be used simultaneously during audio processing. Additionally, it may also include bitstreams or files that contain information indicating associated bitstreams or associated files.

[0599] Here, the information representing the associated bitstream or file can be, for example, an identifier representing the associated bitstream or file. Alternatively, the information representing the associated bitstream or file can be, for example, the filename, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier).

[0600] In this case, the acquisition unit can also determine and acquire the associated bitstream or associated file based on information representing the associated bitstream or associated file. Alternatively, information representing the associated bitstream or associated file can be included in the bitstream or file, and also in other bitstreams or other files.

[0601] Here, the file containing information representing the associated bitstream or associated file can also be a control file such as a declaration file for content distribution.

[0602] In addition, all or part of the metadata can be obtained from outside the audio signal bitstream. For example, metadata for controlling the audio and metadata for controlling the video can be obtained from outside the bitstream, or metadata for both can be obtained from outside the bitstream.

[0603] Furthermore, metadata for controlling the image may also be included in the bitstream acquired by the stereo sound reproduction system 1000. In this case, the stereo sound reproduction system 1000 may also output the metadata for controlling the image to a display device that displays the image or a stereo image reproduction device that reproduces the stereo image.

[0604] (Example of information contained in metadata) Metadata can also be information used in the description of a scene represented by a sound space. Here, a scene is a term that refers to the collection of all elements of a sound space, including three-dimensional images and sound events, modeled by a sound signal reproduction system using metadata.

[0605] That is, metadata includes not only information used to control audio processing, but also information used to control video processing. Metadata can contain only one of the information used to control audio processing or the information used to control video processing, or it can contain both.

[0606] The stereo sound reproduction system 1000 processes sound signals using metadata contained in the bitstream and interactive listener location information acquired through appending, to generate virtual sound effects. These sound effects can include initial reflection processing, obstacle removal, diffraction processing, masking, and reverberation processing, as well as other sound processing using metadata. For example, sound effects such as distance attenuation, localization, or Doppler effects can be added.

[0607] In addition, information can be attached to the metadata to toggle the on / off of all or some of the additional sound effects, or priority information for multiple processing of sound effects.

[0608] In addition, as an example, metadata includes information related to the sound space, including sound source objects and obstacle objects, and information related to the positioning location used to locate the sound image in a specified position within the sound space (i.e., to make the listener perceive the sound coming from a specified direction).

[0609] Here, an obstacle object is an object that may affect the listener's perception of sound by blocking or reflecting it before the sound emitted by the sound source reaches the listener. Besides stationary objects, obstacle objects can also include moving bodies such as animals or machines. Animals can also be people.

[0610] Furthermore, when multiple sound source objects exist in the sound space, for any given sound source object, the other sound source objects may become obstacle objects. That is, objects that do not emit sound, such as building materials or inanimate objects (i.e., non-sound-emitting objects), as well as sound source objects that emit sound, can all become obstacle objects.

[0611] The metadata contains all or part of the information representing the shape of the sound space, the shape and location of obstacle objects in the sound space, the shape and location of sound source objects in the sound space, and the location and orientation of the listener in the sound space.

[0612] The sound space can be either enclosed or open. Furthermore, the metadata can also include information about the reflectivity of obstacles within the sound space that can reflect sound. For example, the floor, walls, or ceiling that form the boundary of the sound space can also be considered obstacles.

[0613] Reflectivity is the energy ratio of reflected sound to incident sound, and it can be set for each frequency band of the sound. Of course, reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, parameters such as attenuation rate, diffraction tone, and initial reflection tone can be set uniformly, for example.

[0614] Metadata can also include information beyond reflectivity as parameters relating to obstacle or sound source objects. For example, metadata can also include information about the material of an object as parameters relating to both the sound source and the non-sound-producing object. Specifically, metadata can also include information such as diffusivity, transmissivity, and sound absorption.

[0615] Information related to a sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information representing the sound source region within the object. Reproduction conditions may, for example, specify whether the sound is a continuously flowing sound or an event-triggered sound. The sound source region within an object can be set based on the relative position of the listener and the object, or it can be set using the object as a reference.

[0616] For example, when the sound source area is set according to the relative position of the listener and the object, from the listener's perspective, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.

[0617] Furthermore, when using an object as a reference to define a sound source region, it is possible to fix which area of ​​the object emits which sound. For example, when the listener views the object from the front, the listener can perceive high frequencies from the right side of the object and low frequencies from the left side. And when the listener views the object from the back, the listener can perceive low frequencies from the right side of the object and high frequencies from the left side.

[0618] Spatial metadata can also include the time up to the initial reflection, reverberation time, and the ratio of direct to diffuse sound. When the ratio of direct to diffuse sound is zero, the listener can perceive only the direct sound.

[0619] (Determination and utilization of thresholds) For example, as described above, a threshold for the volume ratio of the direct tone to the reflected tone is determined based on the time difference between the direct tone and the reflected tone. The threshold can be determined by the selection unit 1302, the threshold adjustment unit 1304, or other components. For example, the threshold can also be determined by a threshold determination unit or a threshold determination device. The threshold determination unit or threshold determination device can be included in the selection unit 1302 or the threshold adjustment unit 1304, or it can be a component or device different from the selection unit 1302 and the threshold adjustment unit 1304.

[0620] Specifically, the threshold determination unit or threshold determination device may include an acquisition unit and a determination unit. The acquisition unit acquires the time difference between a preceding tone (such as a direct tone) and a subsequent tone (such as a reflected tone). The determination unit determines a threshold value for the volume ratio of the direct tone to the reflected tone based on the time difference between the direct tone and the reflected tone, and outputs the threshold value. The sound signal processing device 1001 may also operate as a threshold determination device. Furthermore, the multiple components included in the threshold determination unit or threshold determination device may be installed using a processor 1402 and a memory 1404, etc.

[0621] The threshold determining device can also determine the threshold using Equation 1 above. In this case, the time difference is the time difference between the arrival time of the previous tone and the arrival time of the subsequent tone generated corresponding to the previous tone. The threshold corresponds to the boundary between whether the subsequent tone is subject to the preferential effect of the previous tone. In the region having a first axis of time difference and a second axis of volume ratio, the determining unit determines the volume ratio corresponding to the time difference as the threshold on the boundary line determined as the boundary line.

[0622] Here, the domain of the boundary line relative to the first axis corresponds to the time range in which the priority effect occurs. With the first axis being the horizontal axis and the second axis the vertical axis, the boundary line is a curve that descends on the right shoulder and convexes upwards within the domain.

[0623] Figure 29 This is a graph representing the threshold. Specifically, the threshold determined by Equation 1 above is graphically represented. The horizontal axis corresponds to the time difference (ms) between the arrival time of the direct tone and the arrival time of its reflected tone, and the vertical axis corresponds to the volume ratio (dB) between the volume of the direct tone and the volume of its reflected tone. The boundary line shown in the graph represents the boundary between whether the subsequent tone is subject to the preferential effect of the previous tone, that is, the boundary (threshold) between whether the reflected tone is perceived.

[0624] For example, the threshold determination device will Figure 29 The information in the chart shown is stored in internal memory, and a threshold is determined based on the information stored in internal memory. The threshold can be used for the following purposes. The following processing for utilizing the threshold can be performed by the threshold determining device, or the following processing for utilizing the threshold can be performed by the sound signal processing device 1001.

[0625] Figure 30 This is a diagram illustrating the first application example of the threshold. Specifically, reflected sounds located below the boundary line are not perceived and therefore are not selected (i.e., invalidated). On the other hand, reflected sounds located above the boundary line are perceived and therefore selected (i.e., validated). This application example corresponds to the use of... Figure 14 Examples of descriptions, etc.

[0626] Figure 31 This is a diagram illustrating the second application of the threshold. In this example, the threshold is used in deciding which of the multiple reflected sounds generated in virtual space.

[0627] For example, given that the time difference between the i-th reflected tone and its direct tone is determined as D(i), and their volume ratio is determined as G(i), the distance between the point (D(i), G(i)) and the boundary line is derived. Furthermore, N reflected tones are selected (i.e., made valid) starting from those reflected tones that are longer than the distance upwards from the boundary line. Other reflected tones are not selected (i.e., invalidated).

[0628] Alternatively, N reflected notes can be selected starting from the one with the longest distance among the multiple reflected notes that are above the boundary line. Or, the distance between points (D(i), G(i)) and the boundary line can be assigned a positive sign, and the distance between points (D(i), G(i)) and the boundary line can be assigned a negative sign, and N reflected notes can be selected starting from the one with the longest distance.

[0629] Through the processing described above, it is possible to select reflective sounds in sequence, starting with those of higher perceived priority. Therefore, for example, even when the computing power of the computing unit implementing the virtual space is small, the number of reflective sounds to be selected is reduced, and the amount of computation is reduced, it is still possible to select reflective sounds that are perceptually effective.

[0630] Figure 32 This is a diagram illustrating the third application example of a threshold. In this example, in the assignment (selection) of reflected sounds for a performance in virtual space, a threshold is used to teach whether the reflected sounds of the assigned (selected) object have perceptual value.

[0631] For example, the time difference between the reflected tone and its direct tone assigned to an object can be defined as D(i), and their volume ratio can be defined as G(i). If the point (D(i), G(i)) is located below the boundary line, the reflected tone will not be perceived even if it is assigned, and thus it can be taught that the reflected tone has no value. Alternatively, if the time difference between the reflected tone and its direct tone is specified, the volume of the reflected tone that is effective for that time difference can also be taught.

[0632] The sound signal processing device 1001 may also, as described above, indicate to the administrator whether the reflected sound assigned to the object has perceptual value when a new reflected sound is assigned as an effect sound in the virtual space. Alternatively, the sound signal processing device 1001 may also indicate to the administrator whether the reflected sound generated according to the settings of objects, etc., in the virtual space has perceptual value.

[0633] Figure 33 This is a diagram illustrating the fourth use case of a threshold. In this example, in the assignment (selection) of reflected sounds for a performance in virtual space, a threshold is used to teach whether the reflected sounds assigned (selected) to the object are reflected sounds that compress the dynamic range.

[0634] For example, if more reflected sounds are assigned to an object, the combined amplitude of the multiple sounds generated in virtual space increases. Typically, the multiple sounds generated in virtual space are ultimately mixed into a left-right 2ch signal and listened to by headphones, etc. Therefore, excessive assignment of multiple reflected sounds can cause the combined amplitude of the mixed signal to overflow.

[0635] By using thresholds, it is possible to quantitatively teach, among multiple perceptually effective reflexes, reflexes that compress the dynamic range by imparting only a small number of reflexes, and reflexes that do not compress the dynamic range even when imparted with a large number of reflexes.

[0636] The sound signal processing device 1001 may also, as described above, indicate to the administrator whether the reflected sound assigned to the object compresses the dynamic range when a new reflected sound is assigned as an effect sound in the virtual space. Alternatively, the sound signal processing device 1001 may also indicate to the administrator whether the reflected sound generated according to the settings of the object, etc., in the virtual space compresses the dynamic range.

[0637] (Basic action example) Figure 34 This is a flowchart illustrating a first basic operation example of the sound signal processing apparatus 1001. For example, in the sound signal processing apparatus 1001, the processor 1402, which operates as a circuit, uses the memory 1404 to perform... Figure 34 The actions shown.

[0638] Specifically, processor 1402 acquires sound space information related to the sound space (S1401). Furthermore, based on the sound space information, processor 1402 acquires the volume ratio of the volume of the direct tone generated from the sound source in the sound space to the volume of the reflected tone generated corresponding to the direct tone in the sound space (S1402). Additionally, processor 1402 acquires information related to additional processing (S1403). Furthermore, based on the information related to additional processing and the volume ratio, processor 1402 controls whether to select the reflected tone (S1404).

[0639] Therefore, the sound signal processing device 1001 can appropriately control whether to select a reflected sound based on information related to additional processing and the volume ratio of the direct sound to the reflected sound. That is, it can appropriately control whether to select a reflected sound that corresponds to the direct sound generated from the sound source in the sound space. As a result, it can appropriately reduce the amount of computation and the computational load.

[0640] For example, the processor 1402 can also adjust the volume ratio based on information related to additional processing. Furthermore, the processor 1402 can also control whether to select reflected sounds based on the adjusted volume ratio. Thus, the sound signal processing device 1001 can appropriately control whether to select reflected sounds based on the volume ratio reflecting the effects of additional processing.

[0641] Additionally, for example, the processor 1402 can also calculate the time when the sound emitted from the sound source arrives as a reflected sound. Furthermore, the processor 1402 can control whether to select the reflected sound based on the calculated time, information related to additional processing, and the volume ratio. Thus, the sound signal processing device 1001 can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the arrival time of the reflected sound, information related to additional processing, and the volume ratio of the direct sound to the reflected sound.

[0642] Furthermore, for example, processor 1402 can also calculate the time difference between the arrival time of the direct tone and the arrival time of the reflected tone based on sound spatial information. Moreover, processor 1402 can also control whether to select the reflected tone based on the calculated time difference, information related to additional processing, and the volume ratio.

[0643] Therefore, the sound signal processing device 1001 can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, information related to additional processing, and the volume ratio of the direct sound to the reflected sound. Thus, the sound signal processing device 1001 can more appropriately select the reflected sound that has a greater impact on the listener's perception based on a priority effect.

[0644] Furthermore, for example, processor 1402 can also calculate the time difference between the end time of the direct tone and the arrival time of the reflected tone based on sound spatial information. Moreover, processor 1402 can also control whether to select the reflected tone based on the calculated time difference, information related to additional processing, and the volume ratio.

[0645] Therefore, the sound signal processing device 1001 can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the time difference between the end time of the direct tone and the arrival time of the reflected sound, information related to additional processing, and the volume ratio of the direct tone to the reflected sound. Thus, the sound signal processing device 1001 can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the back-shielding effect.

[0646] Alternatively, for example, the additional processing could be the processing of increasing or decreasing the amplitude of the reflected sound. Furthermore, the information related to the additional processing could be information indicating the amount of amplitude increase or decrease in the processing of increasing or decreasing the amplitude of the reflected sound. The processor 1402 could also adjust the volume ratio based on the amount of amplitude increase or decrease. Furthermore, the processor 1402 could also control whether to select the reflected sound based on the adjusted volume ratio.

[0647] Therefore, the sound signal processing device 1001 can appropriately adjust the volume ratio based on the amplitude increased or decreased through additional processing. Furthermore, the sound signal processing device 1001 can appropriately control whether to select reflected sound based on the volume ratio adjusted through additional processing.

[0648] Alternatively, for example, the additional processing could be the processing of increasing or decreasing the delay accompanying the reflected sound. Furthermore, the information related to the additional processing could be information indicating the amount of delay increase or decrease in the processing accompanying the delayed reflection. The processor 1402 could also adjust the time difference based on the amount of delay increase or decrease. Furthermore, the processor 1402 could also control whether to select the reflected sound based on the adjusted time difference and the volume ratio.

[0649] Therefore, the sound signal processing device 1001 can appropriately adjust the time difference based on the delay increased or decreased through additional processing. Furthermore, the sound signal processing device 1001 can appropriately control whether to select reflected sound based on the time difference adjusted through additional processing.

[0650] Alternatively, for example, the additional processing can be processing that is not applied to the direct tone. Therefore, the sound signal processing apparatus 1001 can apply additional processing to the reflected tone without applying additional processing to the direct tone. That is, the sound signal processing apparatus 1001 can apply additional processing to the reflected tone that is not applied to the direct tone. Thus, the sound signal processing apparatus 1001 can apply additional processing suitable for the reflected tone to the reflected tone.

[0651] Alternatively, for example, additional processing could be filtering that emphasizes the diffusion of sound audibly. Thus, the sound signal processing device 1001 can appropriately control whether to select reflected sound based on information related to filtering that emphasizes the diffusion of sound audibly, and the volume ratio of direct sound to reflected sound.

[0652] Furthermore, for example, if the information related to additional processing indicates that no additional processing will be performed, the volume ratio may not be adjusted. Therefore, the sound signal processing device 1001 can appropriately control whether to adjust the volume ratio based on whether additional processing is performed.

[0653] Furthermore, for example, if the information related to additional processing indicates that no additional processing is performed, the time difference may not need to be adjusted. Therefore, the sound signal processing apparatus 1001 can appropriately control whether to adjust the time difference based on whether additional processing is performed.

[0654] Figure 35 This is a flowchart illustrating a second basic operation example of the sound signal processing apparatus 1001. For example, in the sound signal processing apparatus 1001, the processor 1402, which operates as a circuit, uses the memory 1404 to perform... Figure 35 The actions shown.

[0655] Specifically, the processor 1402 acquires sound space information related to the sound space (S1501). Furthermore, the processor 1402 determines whether a sound source in the sound space is a point sound source based on the sound space information (S1502).

[0656] Here, if the sound source is determined to be a point sound source ("Yes" in S1502), the processor 1402 obtains the volume ratio of the volume of the direct sound generated from the sound source to the volume of the reflected sound generated in the sound space corresponding to the direct sound, based on the sound space information (S1503). Furthermore, the processor 1402 calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, based on the sound space information (S1504). Finally, the processor 1402 controls whether to select the reflected sound based on the volume ratio and the time difference (S1505).

[0657] On the other hand, if it is determined that the sound source is not a point sound source ("No" in S1502), the processor 1402 selects the reflected sound regardless of the volume ratio and time difference (S1506).

[0658] Therefore, even if the distance between the listener and the sound source is uncertain, making it difficult to properly calculate the volume ratio and time difference between the direct sound and the reflected sound, the sound signal processing device 1001 can select the reflected sound regardless of the volume ratio and time difference. Consequently, the sound signal processing device 1001 can suppress sound quality degradation that may occur when the sound source is not a point sound source.

[0659] Figure 36 This is a flowchart illustrating a third basic operation example of the sound signal processing apparatus 1001. For example, in the sound signal processing apparatus 1001, the processor 1402, which operates as a circuit, uses the memory 1404 to perform... Figure 36 The actions shown.

[0660] Specifically, the processor 1402 obtains the time difference between the arrival time of the previous tone and the arrival time of the subsequent tone generated corresponding to the previous tone (S1601). Furthermore, based on the time difference, the processor 1402 determines and outputs a threshold, which is a threshold for the volume ratio of the subsequent tone to the volume of the previous tone, and is a threshold corresponding to the boundary of whether the subsequent tone is subject to the priority effect brought by the previous tone (S1602).

[0661] In determining the threshold, within the region of the first axis with a time difference and the second axis with a volume ratio, on the boundary line of the line that is determined as the boundary indicating whether subsequent sounds are subject to a priority effect, the volume ratio corresponding to the obtained time difference is determined as the threshold. Here, the domain defined relative to the boundary line of the first axis is the time range in which the priority effect occurs.

[0662] Therefore, the sound signal processing device 1001 can appropriately determine the threshold corresponding to the boundary of the priority effect based on the relationship between the time difference between the previous tone and the subsequent tone, and the volume ratio between the previous tone and the subsequent tone. In addition, the sound signal processing device 1001 can efficiently determine the threshold for the time range in which the priority effect occurs.

[0663] For example, when the first axis is the horizontal axis and the second axis is the vertical axis, the boundary line can also be a curve that slopes down from the right shoulder and convexes upward within the defined domain. Therefore, the audio signal processing device 1001 can appropriately determine the threshold based on the boundary characteristic that a larger time difference results in a smaller volume ratio, and that the decrease in volume ratio is proportional to the increase in time difference. Consequently, the audio signal processing device 1001 can provide an appropriate threshold for determining whether a subsequent sound is subject to a priority effect.

[0664] Furthermore, for example, the boundary line can also be defined by a mathematical formula with time difference as a variable. And the threshold can also be calculated based on the mathematical formula. Therefore, the audio signal processing device 1001 can efficiently determine the threshold based on the mathematical formula. Consequently, the audio signal processing device 1001 can efficiently provide the threshold.

[0665] Here, in the sound signal processing device 1001, the processor 1402, which operates as a circuit, uses the memory 1404 for processing. Figure 36 The action shown can also be performed in other devices such as threshold determination devices, where the circuit uses memory. Figure 36 The actions shown. Furthermore, the determined threshold can be provided to the sound signal processing device 1001 or other devices.

[0666] In the examples described above, memory 1404 can store programs, bitstreams, sound spatial information, or threshold data (boundary data), etc. Furthermore, processor 1402 can refer to information stored in memory 1404 during operation, and can also retrieve information from memory 1404. Additionally, processor 1402 can also store information obtained through operation into memory 1404.

[0667] (Replenish) Furthermore, the solutions disclosed herein are not limited to specific implementation methods and can be implemented with various modifications.

[0668] For example, in an implementation, a process that is performed by a specific component may be performed by other components instead of that specific component. Furthermore, the order of multiple processes may be changed, or multiple processes may be performed in parallel.

[0669] Furthermore, the ordinal numbers 1, 2, etc., used in the description can be appropriately replaced, removed, or newly assigned. These ordinal numbers do not necessarily correspond to a meaningful order and can also be used for element identification.

[0670] Furthermore, for example, in comparisons of thresholds, "above the threshold" and "greater than the threshold" can be interchanged. Similarly, "below the threshold" and "less than the threshold" can be interchanged. Additionally, for example, "time" and "moment" can be interchanged.

[0671] Furthermore, in the process of selecting one or more target sounds from multiple sounds, if no sound meets the conditions, then none of the sounds may be selected as target sounds. That is, the process of selecting one or more target sounds from multiple sounds may also include the case of not selecting any target sounds.

[0672] Furthermore, such a representation, for example, of at least one of the first, second, and third elements, can correspond to the first, second, third elements, or any combination thereof.

[0673] Furthermore, the embodiments have described, for example, cases in which the solutions based on this disclosure are implemented as audio processing apparatus, encoding apparatus, or decoding apparatus. However, the solutions based on this disclosure are not limited to these, and can also be implemented as software for performing audio processing methods, encoding methods, or decoding methods.

[0674] For example, the program used to execute the aforementioned audio processing, encoding, or decoding methods can be pre-stored in ROM. Furthermore, the CPU can operate according to this program.

[0675] Alternatively, the program used to perform the aforementioned audio processing, encoding, or decoding methods can be stored in a computer-readable recording medium. Furthermore, the computer can also record the program stored in the recording medium into its RAM and operate according to that program.

[0676] Furthermore, the aforementioned components can typically be implemented as integrated circuits, i.e., LSIs, which have input and output terminals. They can be formed as a single chip or as a chip incorporating all or some of the components of the implementation. Depending on the level of integration, LSIs can also be classified as ICs, system LSIs, super LSIs, or very large-scale LSIs.

[0677] Furthermore, it is not limited to LSIs; dedicated circuits or general-purpose processors can also be used. Additionally, FPGAs that can be programmed after LSI manufacturing, or reconfigurable processors capable of reconfiguring the connection or configuration of circuit units within the LSI, can also be used. Moreover, if advancements in semiconductor technology or other derived technologies lead to integrated circuit technologies that replace LSIs, then these technologies can certainly be used for the integration of constituent elements. This could include applications in biotechnology, among others.

[0678] Furthermore, an FPGA or CPU can download all or part of the software used to implement the audio processing, encoding, or decoding methods described in this disclosure via wireless or wired communication. Additionally, all or part of the software for updates can be downloaded via wireless or wired communication. Moreover, the digital signal processing described in this disclosure can be executed by storing the downloaded software in a memory using an FPGA or CPU and operating based on the stored software.

[0679] At this time, devices equipped with FPGAs or CPUs can also be connected to the signal processing device wirelessly or via wired connection, or connected to the signal processing server via a network. Furthermore, the device and the signal processing device or signal processing server can perform the audio processing methods, encoding methods, or decoding methods described in this disclosure.

[0680] For example, the audio processing apparatus, encoding apparatus, or decoding apparatus of this disclosure may also include an FPGA or a CPU. Furthermore, the audio processing apparatus, encoding apparatus, or decoding apparatus may also include an interface for obtaining software from an external source to operate the FPGA or CPU, and a memory for storing the obtained software. Moreover, the FPGA or CPU may also execute the signal processing described in this disclosure based on the stored software operations.

[0681] Alternatively, the server may provide software related to the audio processing, encoding, or decoding processes disclosed herein. Furthermore, a terminal or device may operate as the audio processing, encoding, or decoding apparatus described in this disclosure by installing the software. Alternatively, the terminal or device may install the software by connecting to the server via a network.

[0682] Alternatively, a device other than the terminal or device may obtain the software installation data via a network connection to a server, and install the software on the terminal or device by providing the software installation data to the terminal or device through this other device. Another example of software could be VR software or AR software used to enable the terminal or device to execute the audio processing method described in the implementation method.

[0683] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0684] The above description describes apparatuses and the like with respect to one or more embodiments based on the implementation methods, but the solutions available in this disclosure are not limited to the implementation methods. Solutions that can be obtained by applying various modifications to the implementation methods that can be conceived by those skilled in the art, as well as solutions constructed by combining the constituent elements of different modifications, are also included within the scope of one or more solutions.

[0685] (Postscript) Based on the description of the above embodiments, the following technology is disclosed.

[0686] (Technology 1) An audio processing apparatus comprising a circuit and a memory; the circuit using the memory to acquire sound space information relating to a sound space; based on the sound space information, acquiring characteristics relating to a first sound generated from a sound source in the sound space; and based on the characteristics relating to the first sound, controlling whether to select a second sound generated in the sound space corresponding to the first sound.

[0687] (Technology 2) The sound processing apparatus as described in Technology 1, wherein the first sound is a direct sound and the second sound is a reflected sound.

[0688] (Technology 3) The sound processing apparatus as described in Technology 2, wherein the characteristic related to the first sound is the volume ratio of the volume of the direct sound to the volume of the reflected sound; the circuit calculates the volume ratio based on the sound spatial information; and controls whether to select the reflected sound based on the volume ratio.

[0689] (Technology 4) The sound processing apparatus as described in Technology 3, when the reflected sound is selected, the circuit generates sounds that reach the listener's ears respectively by applying binaural processing to the reflected sound and the direct sound.

[0690] (Technology 5) The audio processing apparatus as described in Technology 3 or 4, wherein the circuit calculates the time difference between the end time of the direct tone and the arrival time of the reflected tone based on the sound spatial information; and controls whether to select the reflected tone based on the time difference and the volume ratio.

[0691] (Technology 6) The audio processing apparatus of Technology 5, wherein the circuit selects the reflected sound when the volume ratio is above a threshold; the first threshold used as the threshold when the time difference is a first value is greater than the second threshold used as the threshold when the time difference is a second value greater than the first value.

[0692] (Technology 7) The audio processing apparatus as described in Technology 3 or 4, wherein the circuit calculates the time difference between the arrival time of the direct tone and the arrival time of the reflected tone based on the sound spatial information; and controls whether to select the reflected tone based on the time difference and the volume ratio.

[0693] (Technology 8) The audio processing apparatus of Technology 7, wherein the circuit selects the reflected sound when the volume ratio is above a threshold; the first threshold used as the threshold when the time difference is a first value is greater than the second threshold used as the threshold when the time difference is a second value greater than the first value.

[0694] (Technology 9) The audio processing apparatus of Technology 8, wherein the circuit adjusts the threshold based on the direction of arrival of the direct sound and the direction of arrival of the reflected sound.

[0695] (Technology 10) The audio processing apparatus of any one of Technologies 2 to 9, wherein, without selecting the reflected tone, the circuit corrects the volume of the direct tone based on the volume of the reflected tone.

[0696] (Technology 11) The audio processing apparatus of any one of Technologies 2 to 9, wherein the circuit synthesizes the reflected sound into the direct sound without selecting the reflected sound.

[0697] (Technology 12) The sound processing apparatus according to any one of Technologies 3 to 9, wherein the volume ratio is the volume ratio of the direct sound at a first moment to the volume of the reflected sound at a second moment different from the first moment.

[0698] (Technology 13) The audio processing apparatus as described in Technology 1 or 2, wherein the circuit sets a threshold based on characteristics related to the first sound, and controls whether to select the second sound based on the threshold.

[0699] (Technology 14) The sound processing apparatus of any one of Technologies 1, 2 and 13, wherein the characteristic related to the first sound is one or a combination of two or more of the volume of the sound source, the visuality of the sound source and the localization of the sound source.

[0700] (Technology 15) The sound processing apparatus of any one of Technologies 1, 2 and 13, wherein the characteristic related to the first sound is the frequency characteristic of the first sound.

[0701] (Technology 16) The sound processing apparatus of any one of Technologies 1, 2 and 13, wherein the characteristic related to the first sound is a characteristic representing the discontinuity of the amplitude of the first sound.

[0702] (Technology 17) The sound processing apparatus of any one of Technologies 1, 2, 13 and 16, wherein the characteristic related to the first sound is a characteristic representing the duration of the audible portion of the first sound or the duration of the silent portion of the first sound.

[0703] (Technology 18) The sound processing apparatus of any one of Technologies 1, 2, 13, 16 and 17 has the characteristic of representing the duration of the audible portion of the first sound and the duration of the silent portion of the first sound in a time sequence.

[0704] (Technology 19) The sound processing apparatus of any one of Technologies 1, 2, 13 and 15, wherein the characteristic related to the first sound is a characteristic representing the variation of the frequency characteristics of the first sound.

[0705] (Technology 20) The sound processing apparatus of any one of Technologies 1, 2, 13, 15 and 19, wherein the characteristic related to the first sound is a characteristic representing the smoothness of the frequency characteristics of the first sound.

[0706] (Technology 21) The audio processing apparatus of any one of Technologies 1, 2 and 13 to 20 obtains characteristics related to the first sound from the bit stream.

[0707] (Technology 22) The audio processing apparatus of any one of Technologies 1, 2 and 13 to 21, wherein the circuit calculates characteristics related to the second sound; and controls whether to select the second sound based on the characteristics related to the first sound and the characteristics related to the second sound.

[0708] (Technology 23) The audio processing apparatus of Technology 22, wherein the circuit obtains a threshold representing the volume corresponding to the boundary of whether a sound can be heard; and controls whether to select the second sound based on characteristics related to the first sound, characteristics related to the second sound, and the threshold.

[0709] (Technology 24) The sound processing apparatus as described in Technology 23, wherein the characteristic related to the second sound is the volume of the second sound.

[0710] (Technology 25) The sound processing apparatus of Technology 1 or 2, wherein the sound space information includes information about the location of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit controls whether to select each of the plurality of second sounds based on characteristics related to the first sound, thereby selecting one or more processing target sounds from the plurality of second sounds to be applied to binaural processing.

[0711] (Technology 26) In the audio processing apparatus of Technology 25, the timing for obtaining the characteristics related to the first sound is at least one of the following: when the sound space is created, when the processing of the sound space begins, and when an information update thread occurs during the processing of the sound space.

[0712] (Technology 27) The sound processing apparatus as described in Technology 26 periodically acquires characteristics related to the first sound after the processing of the sound space begins.

[0713] (Technology 28) The audio processing apparatus as described in Technology 1 or 2, wherein the characteristic associated with the first sound is the volume of the first sound; the circuit calculates an evaluation value of the second sound based on the volume of the first sound; and controls whether to select the second sound based on the evaluation value.

[0714] (Technology 29) The audio processing apparatus as described in Technology 28, wherein the volume of the first sound has a change.

[0715] (Technology 30) The audio processing apparatus of Technology 28 or 29, wherein the circuit calculates the evaluation value such that the greater the volume of the first sound, the easier it is to select the second sound.

[0716] (Technology 31) The sound processing apparatus as described in Technology 1 or 2, wherein the sound space information is scene information including information about the sound source in the sound space and information about the location of the listener in the sound space; the second sound is each of a plurality of second sounds generated in the sound space corresponding to the first sound; the circuit acquires the signal of the first sound; calculates the plurality of second sounds based on the scene information and the signal of the first sound; acquires characteristics related to the first sound from the information about the sound source; controls whether to select each of the plurality of second sounds as a sound not to be processed by binaural processing based on the characteristics related to the first sound, thereby selecting one or more second sounds from the plurality of second sounds that are not to be processed by binaural processing.

[0717] (Technology 32) The audio processing apparatus of Technology 31, wherein the scene information is updated based on the input information; and characteristics related to the first sound are obtained based on the update of the scene information.

[0718] (Technology 33) The audio processing apparatus as described in Technology 31 or 32 obtains the scene information and characteristics related to the first sound from the metadata contained in the bitstream.

[0719] (Technology 34) An audio processing method, comprising: a step of obtaining sound space information relating to a sound space; a step of obtaining, based on the sound space information, characteristics relating to a first sound generated from a sound source in the sound space; and a step of controlling, based on the characteristics relating to the first sound, whether to select a second sound generated in the sound space corresponding to the first sound.

[0720] (Technology 35) A program for causing a computer to perform the audio processing method described in Technology 34.

[0721] (Technology 36) The audio processing apparatus according to any one of Technologies 3 to 9 and 12 further comprises: an additional unit for applying additional processing to the second sound; an additional increment / decrease acquisition unit for acquiring an amplitude increment / decrease in the additional unit; and a volume ratio adjustment unit for adjusting the volume according to the increment / decrease acquired by the additional increment / decrease acquisition unit.

[0722] (Technology 37) The audio processing apparatus according to any one of Technologies 5 to 9 comprises: an additional unit for applying additional processing to the second sound; an additional delay amount acquisition unit for acquiring a delay increase or decrease amount in the additional unit; and a time difference adjustment unit for adjusting the time difference based on the increase or decrease amount acquired by the additional delay amount acquisition unit.

[0723] (Technology 38) The audio processing apparatus according to Technology 36 or 37, wherein the additional processing is the processing not applied to the first sound.

[0724] (Technology 39) In the audio processing apparatus according to Technology 36 or 37, the additional unit is a filter unit that emphasizes the diffusion of sound in an auditory sense.

[0725] (Technology 40) An audio processing apparatus comprising a circuit and a memory, wherein the circuit uses the memory to acquire sound space information relating to a sound space, acquires, based on the sound space information, a volume ratio of the volume of a direct sound generated from a sound source in the sound space to the volume of a reflected sound generated in the sound space corresponding to the direct sound, acquires information relating to additional processing, and controls whether to select the reflected sound based on the information relating to additional processing and the volume ratio.

[0726] (Technology 41) According to the audio processing apparatus of Technology 40, the circuit adjusts the volume ratio based on the information related to the additional processing, and controls whether to select the reflected sound based on the adjusted volume ratio.

[0727] (Technology 42) According to the audio processing apparatus of Technology 40 or 41, the circuit calculates the time when the sound emitted from the sound source arrives as the reflected sound, and controls whether to select the reflected sound based on the time, the information related to the additional processing, and the volume ratio.

[0728] (Technology 43) The sound processing apparatus according to any one of Technologies 40 to 42, wherein the circuit calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound based on the sound spatial information, and controls whether to select the reflected sound based on the time difference, the information related to additional processing, and the volume ratio.

[0729] (Technology 44) The sound processing apparatus according to any one of Technologies 40 to 42, wherein the circuit calculates the time difference between the end time of the direct tone and the arrival time of the reflected tone based on the sound spatial information, and controls whether to select the reflected tone based on the time difference, the information related to additional processing, and the volume ratio.

[0730] (Technology 45) The audio processing apparatus according to any one of Technologies 40 to 44, wherein the additional processing is a processing that accompanies the increase or decrease of the amplitude of the reflected sound, the information related to the additional processing is information indicating the amount of increase or decrease in amplitude during the processing that accompanies the increase or decrease in amplitude of the reflected sound, the circuit adjusts the volume ratio according to the amount of increase or decrease in amplitude, and controls whether to select the reflected sound based on the adjusted volume ratio.

[0731] (Technology 46) In the audio processing apparatus according to Technology 43 or 44, the additional processing is a process of increasing or decreasing the delay of the reflected sound, the information related to the additional processing is information indicating the amount of delay increase or decrease in the process of increasing or decreasing the delay of the reflected sound, the circuit adjusts the time difference according to the amount of delay increase or decrease, and controls whether to select the reflected sound based on the adjusted time difference and the volume ratio.

[0732] (Technology 47) The audio processing apparatus according to any one of Technologies 40 to 46, wherein the additional processing is a processing not applied to the direct sound.

[0733] (Technology 48) The audio processing apparatus according to any one of Technologies 40 to 47, wherein the additional processing is a filtering process that emphasizes the diffusion of sound in terms of auditory perception.

[0734] (Technology 49) The audio processing apparatus according to Technology 41 or 45 does not adjust the volume ratio when the information related to the additional processing indicates that the additional processing is not to be performed.

[0735] (Technology 50) The audio processing apparatus according to Technology 46 does not adjust the time difference when the information related to the additional processing indicates that the additional processing is not to be performed.

[0736] (Technology 51) An audio processing apparatus: comprising a circuit and a memory, wherein the circuit uses the memory to acquire sound space information relating to a sound space, and based on the sound space information, determines whether a sound source in the sound space is a point sound source, and if the sound source is determined to be the point sound source, (i) based on the sound space information, acquires a volume ratio of the volume of a direct sound generated from the sound source to the volume of a reflected sound generated in the sound space corresponding to the direct sound, (ii) based on the sound space information, calculates a time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and (iii) based on the volume ratio and the time difference, controls whether to select the reflected sound, and if the sound source is determined not to be the point sound source, selects the reflected sound regardless of the volume ratio and the time difference.

[0737] (Technology 52) A threshold determining apparatus comprising a circuit and a memory, the circuit using the memory to obtain a time difference between the arrival time of a previous tone and the arrival time of a subsequent tone generated corresponding to the previous tone, determining a threshold based on the time difference and outputting the threshold, the threshold being a threshold of the volume ratio of the subsequent tone to the volume of the previous tone, and a threshold corresponding to a boundary of whether the subsequent tone is subject to a priority effect caused by the previous tone, wherein in determining the threshold, in a region having a first axis having the time difference and a second axis having the volume ratio, on a boundary line determined to represent the boundary, the volume ratio corresponding to the time difference is determined as the threshold, and the domain of the boundary line relative to the first axis is the time range in which the priority effect occurs.

[0738] (Technology 53) According to the threshold determination device of Technology 52, when the first axis is the horizontal axis and the second axis is the vertical axis, the boundary line is a curve that descends on the right shoulder and convexes upward in the defined domain.

[0739] (Technology 54) According to the threshold determination apparatus of Technology 52 or 53, the boundary line is defined by a mathematical formula having the time difference as a variable, and the threshold is calculated based on the mathematical formula.

[0740] (Technology 55) An audio processing method, comprising: a step of acquiring sound space information relating to a sound space; a step of acquiring, based on the sound space information, a volume ratio of a direct sound generated from a sound source in the sound space to a reflected sound generated in the sound space corresponding to the direct sound; a step of acquiring information relating to additional processing; and a step of controlling whether to select the reflected sound based on the information relating to additional processing and the volume ratio.

[0741] (Technology 56) A program for causing a computer to perform the audio processing method described in Technology 55.

[0742] Industrial applicability This disclosure includes, for example, methods applicable to audio processing apparatus, encoding apparatus, decoding apparatus, or terminals or devices having any of these apparatuses.

[0743] Label Explanation 1000 Stereo Sound Reproduction System 1001 Sound signal processing device (audio processing device) 1002 Sound Prompt Device 1100, 1120, 1500 encoding devices Input data for 1101 and 1113 1102 Encoder 1103 Encoded Data 1104, 1114, 1404, 1503 memory 1110, 1130 Decoding Devices 1111 sound signal 1112, 1200, 1210 decoders 1121 Sending Department 1122 Send signal 1131 Receiving Department 1132 Received signal Spatial Information Management Department, 1201 & 1211 1202 Audio Data Decoder Rendering Departments 1203, 1213, and 1300 1301 Analysis Department 1302, 1314 Selection Department 1303 Synthesis Department 1304 Threshold Adjustment Section 1311 Reverb Processing Department 1312 Initial Reflection Processing Unit 1313 Distance Attenuation Processing Unit 1315 Production Department 1316 Binocular Processing Unit 1401 Speaker 1402 and 1501 processors 1403, 1502 Communication IF 1405 sensor

Claims

1. An audio processing device, wherein, Equipped with circuitry and memory, The circuit uses the memory. Obtain sound space information related to sound space. Based on the sound space information, the volume ratio of the direct tone generated from the sound source in the sound space to the volume of the reflected tone generated in the sound space corresponding to the direct tone is obtained. Obtain information related to additional processing. Based on the information related to the additional processing and the volume ratio, control is made to select the reflected sound.

2. The audio processing device according to claim 1, wherein, The circuit adjusts the volume ratio based on the information related to the additional processing. Based on the adjusted volume ratio, control whether to select the reflected sound.

3. The audio processing apparatus according to claim 1, wherein, The circuit calculates the time when the sound emitted from the sound source arrives as the reflected sound, and controls whether to select the reflected sound based on the time, the information related to the additional processing, and the volume ratio.

4. The audio processing apparatus according to claim 1, wherein, The circuit, Based on the aforementioned sound spatial information, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone is calculated. Based on the time difference, the information related to additional processing, and the volume ratio, control is made to select the reflected sound.

5. The audio processing apparatus according to claim 1, wherein, The circuit, Based on the aforementioned sound spatial information, the time difference between the end time of the direct tone and the arrival time of the reflected tone is calculated. Based on the time difference, the information related to additional processing, and the volume ratio, control is made to select the reflected sound.

6. The audio processing apparatus according to claim 1, wherein, The additional processing is the processing that accompanies the increase or decrease of the amplitude of the reflected sound. The information related to the additional processing refers to information about the amount of amplitude increase or decrease in the processing accompanying the increase or decrease in the amplitude of the reflected sound. The circuit, Adjust the volume ratio based on the increase or decrease in amplitude. Based on the adjusted volume ratio, control whether to select the reflected sound.

7. The sound processing apparatus according to claim 4 or 5, wherein, The additional processing is the processing that accompanies the increase or decrease of the delay of the reflected sound. The information related to the additional processing refers to information about the amount of delay increase or decrease in the processing accompanying the increase or decrease of the delay of the reflected sound. The circuit, Adjust the time difference based on the increase or decrease in delay. Based on the adjusted time difference and volume ratio, control whether to select the reflected sound.

8. The audio processing apparatus according to claim 1, wherein, The additional processing is the processing not applied to the direct tone.

9. The audio processing apparatus according to claim 1, wherein, The additional processing is a filtering process that emphasizes the diffusion of sound in terms of auditory perception.

10. The sound processing apparatus according to claim 2 or 6, wherein, If the information related to additional processing indicates that the additional processing is not performed, the volume ratio will not be adjusted.

11. The audio processing apparatus according to claim 7, wherein, If the information related to the additional processing indicates that the additional processing will not be performed, the time difference will not be adjusted.

12. An audio processing device, wherein, Equipped with circuitry and memory, The circuit uses the memory. Obtain sound space information related to sound space. Based on the sound space information, determine whether the sound source in the sound space is a point sound source. If the sound source is determined to be the point sound source... (i) Based on the sound space information, obtain the volume ratio of the volume of the direct tone generated from the sound source to the volume of the reflected tone generated in the sound space corresponding to the direct tone. (ii) Based on the sound spatial information, calculate the time difference between the arrival time of the direct tone and the arrival time of the reflected tone. (iii) Based on the volume ratio and the time difference, control whether to select the reflected sound. If it is determined that the sound source is not the point sound source, the reflected sound is selected regardless of the volume ratio and the time difference.

13. A threshold determination device, wherein, Equipped with circuitry and memory, The circuit uses the memory. Obtain the time difference between the arrival time of the preceding sound and the arrival time of the subsequent sound generated corresponding to the preceding sound. Based on the time difference, a threshold is determined and output. This threshold is the volume ratio of the subsequent sound to the previous sound, corresponding to the boundary of whether the subsequent sound is subject to the priority effect of the previous sound. In determining the threshold, within the region having the first axis of the time difference and the second axis of the volume ratio, on the boundary line determined to represent the boundary, the volume ratio corresponding to the time difference is determined as the threshold. The domain of the boundary line relative to the first axis is the time range in which the priority effect occurs.

14. The threshold determination device according to claim 13, wherein, When the first axis is the horizontal axis and the second axis is the vertical axis, the boundary line is a curve that descends on the right shoulder and convexes upward within the defined domain.

15. The threshold determination device according to claim 13 or 14, wherein, The boundary line is defined by a mathematical formula that uses the time difference as a variable. The threshold is calculated based on the mathematical formula.

16. A sound processing method, wherein, include: Steps for obtaining sound space information related to sound space; Based on the sound space information, the step of obtaining the volume ratio of the volume of the direct sound generated from the sound source in the sound space to the volume of the reflected sound generated in the sound space corresponding to the direct sound; The steps for obtaining information related to additional processing; and The step of controlling whether to select the reflected sound based on the information related to the additional processing and the volume ratio.

17. A program for causing a computer to perform the audio processing method of claim 16.

Citation Information

Patent Citations

  • Signal processor

    JP2019022049A

  • Apparatus and method for rendering a sound scene using pipeline stages

    WO2021180938A1

  • Sound processing apparatus, decoder, encoder, bitstream and corresponding methods

    WO2023083780A2