Acoustic processing device and acoustic processing method

By using sound space information to control the selection of reflected sound in the audio processing device, the problem of reflective sound processing of multiple sound sources in the virtual space is solved, and the calculation amount and load are reduced, which improves the efficiency and effect of sound processing.

CN120077428APending Publication Date: 2025-05-30PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380071403.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-05
Filing Date
2023-10-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process the reflected sound generated by multiple sound sources in the virtual space, resulting in large computing volume and heavy operation load, and it is difficult to reduce the number of sound lines without damaging sound positioning and space mastery.

Method used

By using circuits and memory in the audio processing device, sound space information is obtained, and based on this information, whether to select the reflected sound generated in the sound space, and the calculation amount and calculation load are appropriately reduced.

Benefits of technology

It realizes that the calculation amount and calculation load are reduced without damaging sound positioning and space mastery is improved, and the efficiency and effect of audio processing are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077428A_ABST
    Figure CN120077428A_ABST
Patent Text Reader

Abstract

The sound processing device (1001) is provided with a circuit (1402) and a memory (1404). A circuit (1402) acquires sound space information relating to a sound space using the memory (1404); acquiring, on the basis of the sound space information, a characteristic relating to a first sound generated from the sound source in the sound space; on the basis of the characteristics related to the first sound, whether to select a second sound generated in the sound space corresponding to the first sound is controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio processing device and the like. Background Art

[0002] In recent years, products and services utilizing ER (Extended Reality) (which may also be expressed as XR) including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality) have been spreading. Along with this, the importance of audio processing technology that provides immersive audio by imparting an acoustic effect corresponding to the environment of the space to the sound emitted from a virtual sound source to a listener in a virtual space or a real space has increased.

[0003] In addition, a listener may also be expressed as a user. Furthermore, technologies related to the audio processing device and audio processing method of the present disclosure are shown in Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1.

[0004] Prior Art Documents

[0005] Patent Documents

[0006] Patent Document 1: Japanese Patent No. 6288100

[0007] Patent Document 2: Japanese Unexamined Patent Application Publication No. 2019-22049

[0008] Patent Document 3: International Publication No. 2021 / 180938

[0009] Non-Patent Documents

[0010] Non-Patent Document 1: "Outline of Auditory Psychology" by B.C.J. Moore, Seishin Shobo, April 20, 1994, Chapter 6: Spatial Perception, p. 225 Summary of the Invention

[0011] Problems to be Solved by the Invention

[0012] For example, in Patent Document 1, a technology for performing signal processing on an object audio signal and presenting it to a listener is disclosed. With the spread of ER technology and the diversification of services using ER technology, for example, audio processing corresponding to differences in audio quality required by each service, signal processing capabilities of terminals used, and sound quality that can be provided by sound presentation devices is required. In addition, further improvement of audio processing technology is required to provide these.

[0013] Here, the improvement of the audio processing technology refers to the change of the existing audio processing. For example, the improvement of the audio processing technology provides a process for imparting a new audio effect, a reduction in the processing amount of the audio processing, an improvement in the quality of the sound obtained through the audio processing, a reduction in the data amount of the information used in the implementation of the audio processing, or an ease of acquisition or generation of the information used in the implementation of the audio processing. Or, the improvement of the audio processing technology can also provide a combination of any two or more of these.

[0014] In particular, these improvements are required in devices or services where the listener can move freely in a virtual space. However, the above effects obtained by the improvement of the audio processing technology are only examples. One or more technical solutions grasped based on the present disclosure can also be technical solutions conceived from different viewpoints from the above, technical solutions for achieving different purposes from the above, or technical solutions capable of obtaining different effects from the above.

[0015] Means for solving the problem

[0016] The audio device related to one technical solution grasped based on the present disclosure includes a circuit and a memory; the circuit uses the memory to obtain sound space information related to a sound space; based on the sound space information, obtains characteristics related to a first sound generated from a sound source in the sound space; based on the characteristics related to the first sound, controls whether to select a second sound generated corresponding to the first sound in the sound space.

[0017] In addition, these inclusive or specific technical solutions can also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of these.

[0018] Effect of the invention

[0019] One technical solution of the present disclosure can, for example, provide a process for imparting a new audio effect, a reduction in the processing amount of the audio processing, an improvement in the sound quality of the sound obtained through the audio processing, a reduction in the data amount of the information used in the implementation of the audio processing, or an ease of acquisition or generation of the information used in the implementation of the audio processing. Or, one technical solution of the present disclosure can provide any combination of these. As a result, one technical solution of the present disclosure provides audio processing suitable for the usage environment of the listener and can contribute to the improvement of the listener's audio experience.

[0020] In particular, the above effects can be obtained in devices or services that allow listeners to move freely within a virtual space. However, the above effects are merely examples of the effects of various technical solutions grasped based on the present disclosure. Each of one or more technical solutions grasped based on the present disclosure may also be a technical solution conceived from a different perspective from the above, a technical solution achieving a different purpose from the above, or a technical solution capable of obtaining a different effect from the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. is an example showing direct sound and reflected sound generated in a sound space.

[0022] Figure 2 FIG. is an example showing a stereophonic reproduction system according to an embodiment.

[0023] Figure 3A FIG. is a block diagram showing a configuration example of an encoding device according to an embodiment.

[0024] Figure 3B FIG. is a block diagram showing a configuration example of a decoding device according to an embodiment.

[0025] Figure 3C FIG. is a block diagram showing another configuration example of an encoding device according to an embodiment.

[0026] Figure 3D FIG. is a block diagram showing another configuration example of a decoding device according to an embodiment.

[0027] Figure 4A FIG. is a block diagram showing a configuration example of a decoder according to an embodiment.

[0028] Figure 4B FIG. is a block diagram showing another configuration example of a decoder according to an embodiment.

[0029] Figure 5 FIG. is an example showing a physical configuration of a sound signal processing device according to an embodiment.

[0030] Figure 6 FIG. is an example showing a physical configuration of an encoding device according to an embodiment.

[0031] Figure 7 FIG. is a block diagram showing a configuration example of a rendering unit according to an embodiment.

[0032] Figure 8 FIG. is a flowchart showing an operation example of a sound signal processing device according to an embodiment.

[0033] Figure 9 FIG. shows a positional relationship where the listener and the obstacle object are relatively far apart.

[0034] Figure 10It is a diagram showing the positional relationship where the listener is relatively close to the obstacle object.

[0035] Figure 11 It is a diagram showing the relationship between the time difference between the direct sound and the reflected sound and the threshold value.

[0036] Figure 12A It is a diagram showing a part of an example of the setting method of the threshold data.

[0037] Figure 12B It is a diagram showing a part of an example of the setting method of the threshold data.

[0038] Figure 12C It is a diagram showing a part of an example of the setting method of the threshold data.

[0039] Figure 13 It is a diagram showing an example of the setting method of the threshold value.

[0040] Figure 14 It is a flowchart showing an example of the selection process.

[0041] Figure 15 It is a diagram showing the relationship between the direction of the direct sound, the direction of the reflected sound, the time difference, and the threshold value.

[0042] Figure 16 It is a diagram showing the relationship between the angle difference, the time difference, and the threshold value.

[0043] Figure 17 It is a block diagram showing another configuration example of the rendering unit.

[0044] Figure 18 It is a flowchart showing another example of the selection process.

[0045] Figure 19 It is a flowchart showing yet another example of the selection process.

[0046] Figure 20 It is a flowchart showing the first modification example of the operation of the sound signal processing device according to the embodiment.

[0047] Figure 21 It is a flowchart showing the second modification example of the operation of the sound signal processing device according to the embodiment.

[0048] Figure 22 It is a diagram showing an example of the configuration of the avatar, the sound source object, and the obstacle object.

[0049] Figure 23 It is a flowchart showing yet another example of the selection process.

[0050] Figure 24 It is a block diagram showing a configuration example for the rendering unit to perform pipeline processing.

[0051] Figure 25 This is a diagram showing the transmission and diffraction of sound. Detailed implementation

[0052] (Understanding as the basis of the present disclosure)

[0053] Figure 1 This is a diagram showing an example of direct sound and reflected sound generated in the sound space. In audio processing for expressing the characteristics of a virtual space with sound, in order to express the spaciousness of the space, the material of the wall surface, etc., and in order to correctly grasp the position of the sound source (sound image localization), it is effective to reproduce not only the direct sound but also the reflected sound.

[0054] For example, in the case of listening to sound in a rectangular room such as Figure 1 Six primary reflected sounds corresponding to the six wall surfaces are generated for one sound source. The reproduction of these reflected sounds becomes a clue for appropriate understanding related to the space and the sound image. Furthermore, for each reflected sound, secondary reflected sounds are generated on the surfaces other than the reflecting surface where the reflected sound is generated. These reflected sounds also become perceptually effective clues.

[0055] However, even when considering only up to secondary reflections, for one sound source, one direct sound and 36 (6 + 6×5) reflected sounds are generated, so 37 sound rays are generated, and a considerable amount of computational power is required to process these sound rays.

[0056] In addition, in recent application products envisioned for the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, there will inevitably be multiple sound sources, so an even more substantial amount of computational power is required.

[0057] In addition, listeners who listen to sound in a virtual space use head-mounted headphones or VR goggles. In order to provide stereophonic sound to such listeners, binaural processing is performed on each sound ray to reproduce the direction of arrival and the sense of distance of the sound by giving the sound pressure ratio and phase difference between the two ears. Therefore, if all the generated reflected sounds are to be reproduced, the computational amount is very large.

[0058] On the other hand, as the battery of the VR goggles worn by listeners experiencing in a virtual space, a small rechargeable battery is sometimes used for its convenience. In order to extend its battery life, the computational load required for the above-mentioned processing is preferably small. For this purpose, it is desired to reduce the number of sound rays generated on the order of several hundred within a range that does not impair the localization of sound and the understanding of the space.

[0059] In addition, in a reproduced sound system, sometimes degrees of freedom such as 6DoF (6 Degrees of Freedom) are allowed for the position and orientation of the listener. In this case, the positional relationship among the listener, the sound source, and the object that reflects sound cannot be determined unless it is during reproduction (rendering). Therefore, the reflected sound also cannot be determined unless it is during reproduction. Thus, it is difficult to pre-determine the reflected sound to be processed.

[0060] Therefore, appropriately selecting one or more reflected sounds to be processed or not to be processed among the multiple reflected sounds generated in the sound space during reproduction is beneficial for appropriately reducing the amount of calculation and the computational load.

[0061] Therefore, an object of the present disclosure is to provide an audio processing device or the like that can appropriately control whether to select the sound generated in the sound space.

[0062] In addition, controlling whether to select a sound corresponds to determining whether to select the sound. Moreover, selecting a sound can be either selecting the sound as the sound to be processed or selecting the sound as the sound not to be processed.

[0063] (Summary of the disclosure)

[0064] The audio processing device according to the first aspect of the present disclosure includes a circuit and a memory; the circuit uses the memory to obtain sound space information related to the sound space; based on the sound space information, obtains the characteristics related to the first sound generated from the sound source in the sound space; based on the characteristics related to the first sound, controls whether to select the second sound generated corresponding to the first sound in the sound space.

[0065] The device according to the above aspect can appropriately control whether to select the second sound generated corresponding to the first sound in the sound space based on the characteristics related to the first sound generated in the sound space. That is, it can appropriately control whether to select the sound generated in the sound space. Thus, the amount of calculation and the computational load can be appropriately reduced.

[0066] The audio processing device according to the second aspect of the present disclosure may be such that, in the audio processing device according to the first aspect, the first sound is the direct sound and the second sound is the reflected sound.

[0067] The device according to the above aspect can appropriately control whether to select the reflected sound based on the characteristics related to the direct sound.

[0068] The audio processing device according to the third aspect of the present disclosure may be such that, in the audio processing device according to the second aspect, the characteristic related to the first sound is the volume ratio of the direct sound to the reflected sound; the circuit calculates the volume ratio based on the sound space information and controls whether to select the reflected sound based on the volume ratio.

[0069] The device of the above technical solution can appropriately select a reflected sound that has a greater impact on the listener's perception based on the volume ratio of the direct sound to the reflected sound.

[0070] The audio processing device related to the fourth technical solution learned based on the present disclosure may also be that in the audio processing device of the third technical solution, when a reflected sound is selected, the circuit generates sounds that reach the two ears of the listener respectively by applying binaural processing to the reflected sound and the direct sound.

[0071] The device of the above technical solution can appropriately select a reflected sound that has a greater impact on the listener's perception and apply binaural processing to the selected reflected sound.

[0072] The audio processing device related to the fifth technical solution learned based on the present disclosure may also be that in the audio processing device of the third or fourth technical solution, the circuit calculates the time difference between the end time of the direct sound and the arrival time of the reflected sound based on the sound spatial information, and controls whether to select the reflected sound based on the time difference and the volume ratio.

[0073] The device of the above technical solution can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the time difference between the end time of the direct sound and the arrival time of the reflected sound, and the volume ratio of the volume of the direct sound to the volume of the reflected sound. Therefore, the device of the above technical solution can more appropriately select a reflected sound that has a greater impact on the listener's perception based on the post-masking effect.

[0074] The audio processing device related to the sixth technical solution learned based on the present disclosure may also be that in the audio processing device of the fifth technical solution, the circuit selects the reflected sound when the volume ratio is equal to or greater than the threshold value; the first threshold value used as the threshold value when the time difference is the first value is greater than the second threshold value used as the threshold value when the time difference is the second value greater than the first value.

[0075] The device of the above technical solution can increase the possibility of selecting a reflected sound with a large time difference between the end time of the direct sound and the arrival time of the reflected sound. Therefore, the device of the above technical solution can more appropriately select a reflected sound that has a greater impact on the listener's perception.

[0076] The audio processing device related to the seventh technical solution learned based on the present disclosure may also be that in the audio processing device of the third or fourth technical solution, the circuit calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound based on the sound spatial information; and controls whether to select the reflected sound based on the time difference and the volume ratio.

[0077] The device of the above technical solution can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and the volume ratio of the volume of the direct sound to the volume of the reflected sound. Therefore, the device of the above technical solution can more appropriately select the reflected sound that has a greater impact on the listener's perception based on the precedence effect.

[0078] The audio processing device related to the eighth technical solution known from the present disclosure may also be that in the audio processing device of the seventh technical solution, the circuit selects the reflected sound when the volume ratio is equal to or greater than the threshold value, and the first threshold value used as the threshold value when the time difference is the first value is greater than the second threshold value used as the threshold value when the time difference is the second value greater than the first value.

[0079] The device of the above technical solution can increase the possibility of selecting the reflected sound with a large time difference between the arrival time of the direct sound and the arrival time of the reflected sound. Therefore, the device of the above technical solution can appropriately select the reflected sound that has a greater impact on the listener's perception.

[0080] The audio processing device related to the ninth technical solution known from the present disclosure may also be that in the audio processing device of the sixth or eighth technical solution, the circuit adjusts the threshold value based on the arrival direction of the direct sound and the arrival direction of the reflected sound.

[0081] The device of the above technical solution can appropriately select the reflected sound that has a greater impact on the listener's perception based on the arrival direction of the direct sound and the arrival direction of the reflected sound.

[0082] The audio processing device related to the tenth technical solution known from the present disclosure may also be that in any one of the audio processing devices of the first to ninth technical solutions, when the second sound is not selected, the circuit corrects the volume of the first sound based on the volume of the second sound.

[0083] The device of the above technical solution can appropriately reduce the sense of incongruity caused by the lack of the volume of the second sound due to the non - selection of the second sound with a relatively small amount of computation.

[0084] The audio processing device related to the eleventh technical solution known from the present disclosure may also be that in any one of the audio processing devices of the first to ninth technical solutions, when the second sound is not selected, the circuit synthesizes the second sound into the first sound.

[0085] The device of the above technical solution can more correctly reflect the characteristics of the second sound into the first sound. Therefore, the device of the above technical solution can reduce the sense of incongruity caused by the lack of the volume of the second sound due to the non - selection of the second sound.

[0086] The audio processing device related to the 12th technical solution according to the present disclosure may also be an audio processing device according to any one of the 3rd to 9th technical solutions, where the volume ratio is the volume ratio of the direct sound at the first moment to the volume of the reflected sound at the second moment different from the first moment.

[0087] When the moments at which the direct sound is perceived and the reflected sound is perceived are different in the device of the above technical solution, it is possible to appropriately select the reflected sound that has a greater influence on the perception of the listener based on the volume ratio of the direct sound and the reflected sound at different moments.

[0088] The audio processing device related to the 13th technical solution according to the present disclosure may also be an audio processing device according to the 1st or 2nd technical solution, where the circuit sets a threshold based on the characteristics related to the first sound and controls whether to select the second sound based on the threshold.

[0089] The device of the above technical solution can appropriately control whether to select the second sound based on the threshold set according to the characteristics related to the first sound.

[0090] The audio processing device related to the 14th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, and 13th technical solutions, where the characteristics related to the first sound are one or a combination of two or more of the volume of the sound source, the visuality of the sound source, and the localizability of the sound source.

[0091] The device of the above technical solution can appropriately control whether to select the second sound based on the volume of the sound source, the visuality of the sound source, or the localizability of the sound source.

[0092] The audio processing device related to the 15th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, and 13th technical solutions, where the characteristics related to the first sound are the frequency characteristics of the first sound.

[0093] The device of the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the frequency characteristics of the first sound.

[0094] The audio processing device related to the 16th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, and 13th technical solutions, where the characteristics related to the first sound are the characteristics indicating the intermittency of the amplitude of the first sound.

[0095] The device of the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the characteristics indicating the intermittency of the amplitude of the first sound.

[0096] The audio processing device related to the 17th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, 13th, and 16th technical solutions, wherein the characteristic related to the 1st sound is a characteristic representing the duration of the sounding part of the 1st sound or the duration of the silent part of the 1st sound.

[0097] The device according to the above technical solution can appropriately control whether to select the 2nd sound generated corresponding to the 1st sound based on the characteristic representing the duration of the sounding part of the 1st sound or the duration of the silent part of the 1st sound.

[0098] The audio processing device related to the 18th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, 13th, 16th, and 17th technical solutions, wherein the characteristic related to the 1st sound is a characteristic representing the duration of the sounding part of the 1st sound and the duration of the silent part of the 1st sound in a time series.

[0099] The device according to the above technical solution can appropriately control whether to select the 2nd sound generated corresponding to the 1st sound based on the characteristic representing the duration of the sounding part of the 1st sound and the duration of the silent part of the 1st sound in a time series.

[0100] The audio processing device related to the 19th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, 13th, and 15th technical solutions, wherein the characteristic related to the 1st sound is a characteristic representing the variation of the frequency characteristic of the 1st sound.

[0101] The device according to the above technical solution can appropriately control whether to select the 2nd sound generated corresponding to the 1st sound based on the characteristic representing the variation of the frequency characteristic of the 1st sound.

[0102] The audio processing device related to the 20th technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, 13th, 15th, and 19th technical solutions, wherein the characteristic related to the 1st sound is a characteristic representing the smoothness of the frequency characteristic of the 1st sound.

[0103] The device according to the above technical solution can appropriately control whether to select the 2nd sound generated corresponding to the 1st sound based on the characteristic representing the smoothness of the frequency characteristic of the 1st sound.

[0104] The audio processing device related to the 21st technical solution according to the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, and 13th to 20th technical solutions, and obtains the characteristic related to the 1st sound from the bitstream.

[0105] The device of the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the information obtained from the bitstream.

[0106] The audio processing device related to the 22nd technical solution understood based on the present disclosure may also be an audio processing device of any one of the 1st, 2nd, and 13th to 21st technical solutions, where the circuit calculates the characteristics related to the second sound; based on the characteristics related to the first sound and the characteristics related to the second sound, it controls whether to select the second sound.

[0107] The device of the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the characteristics related to the first sound and the characteristics related to the second sound.

[0108] The audio processing device related to the 23rd technical solution understood based on the present disclosure may also be an audio processing device of the 22nd technical solution, where the circuit obtains a threshold value representing the volume corresponding to the boundary of whether a sound can be heard; based on the characteristics related to the first sound, the characteristics related to the second sound, and the threshold value, it controls whether to select the second sound.

[0109] The device of the above technical solution can appropriately control whether to select the second sound based on, in addition to the characteristics related to the first sound and the characteristics related to the second sound, also a threshold value related to whether a sound can be heard.

[0110] The audio processing device related to the 24th technical solution understood based on the present disclosure may also be an audio processing device of the 22nd or 23rd technical solution, where the characteristics related to the second sound are the volume of the second sound.

[0111] The device of the above technical solution can appropriately control whether to select the second sound based on the volume of the second sound.

[0112] The audio processing device related to the 25th technical solution understood based on the present disclosure may also be an audio processing device of any one of the 1st to 24th technical solutions, where the sound space information includes information on the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space; the circuit controls whether to select each of the plurality of second sounds based on the characteristics related to the first sound, and selects one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds.

[0113] The device of the above technical solution can appropriately control whether to select each of the plurality of second sounds generated corresponding to the first sound in the sound space based on the characteristics related to the first sound generated in the sound space. And, the device of the above technical solution can appropriately select one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds.

[0114] The sound processing device related to the 26th technical solution according to the present disclosure may also be that, in the sound processing device of any one of the 1st to 25th technical solutions, the timing for obtaining the characteristics related to the first sound is at least one of the time of creating the sound space, the start of processing the sound space, and the generation of the information update thread during the processing of the sound space.

[0115] The device of the above technical solution can appropriately select one or more sound objects to which binaural processing is applied based on the information obtained at an adaptive timing.

[0116] The sound processing device related to the 27th technical solution according to the present disclosure may also be that, in the sound processing device of any one of the 1st to 26th technical solutions, the characteristics related to the first sound are obtained periodically after the start of processing the sound space.

[0117] The device of the above technical solution can appropriately select one or more sound objects to which binaural processing is applied based on the information obtained periodically.

[0118] The sound processing device related to the 28th technical solution according to the present disclosure may also be that, in the sound processing device of any one of the 1st, 2nd, and 25th to 27th technical solutions, the characteristics related to the first sound are the volume of the first sound; the circuit calculates an evaluation value of the second sound based on the volume of the first sound, and controls whether to select the second sound based on the evaluation value.

[0119] The device of the above technical solution can appropriately control whether to select the second sound based on the evaluation value calculated for the second sound according to the volume of the first sound.

[0120] The sound processing device related to the 29th technical solution according to the present disclosure may also be that, in the sound processing device of the 28th technical solution, the volume of the first sound has a transition.

[0121] The device of the above technical solution can appropriately control whether to select the second sound based on the evaluation value calculated according to the volume having a transition.

[0122] The sound processing device related to the 30th technical solution according to the present disclosure may also be that, in the sound processing device of the 28th or 29th technical solution, the circuit calculates an evaluation value such that the larger the volume of the first sound, the more likely the second sound is to be selected.

[0123] The device of the above technical solution can appropriately control whether to select the second sound based on the evaluation value set such that the larger the volume of the first sound, the more likely the second sound is to be selected.

[0124] The audio processing device related to the 31st technical solution based on the present disclosure may also be an audio processing device according to any one of the 1st to 30th technical solutions, where the sound space information is scene information including information on sound sources in the sound space and information on the positions of listeners in the sound space; the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space; the circuit acquires a signal of the first sound; calculates a plurality of second sounds based on the scene information and the signal of the first sound; acquires characteristics related to the first sound from the information on the sound source; and controls whether to select each of the plurality of second sounds as a sound not to be subjected to binaural processing based on the characteristics related to the first sound, thereby selecting one or more second sounds not to be subjected to binaural processing from the plurality of second sounds.

[0125] The device according to the above technical solution can appropriately select one or more second sounds not to be subjected to binaural processing from a plurality of second sounds generated corresponding to the first sound in the sound space based on the characteristics related to the first sound.

[0126] The audio processing device related to the 32nd technical solution based on the present disclosure may also be an audio processing device according to the 31st technical solution, where the scene information is updated based on input information; and corresponding to the update of the scene information, characteristics related to the first sound are acquired.

[0127] The device according to the above technical solution can appropriately select one or more second sounds not to be subjected to binaural processing based on the information acquired corresponding to the update of the scene information.

[0128] The audio processing device related to the 33rd technical solution based on the present disclosure may also be an audio processing device according to the 31st or 32nd technical solution, where the scene information and the characteristics related to the first sound are acquired from the metadata included in the bitstream.

[0129] The device according to the above technical solution can appropriately select one or more second sounds not to be subjected to binaural processing based on the information acquired from the metadata included in the bitstream.

[0130] The audio processing device related to the 34th technical solution based on the present disclosure may also be an audio processing device according to any one of the 1st, 2nd, 13th, 16th to 18th, 25th to 27th, and 31st to 33rd technical solutions, where the characteristics related to the first sound are characteristics represented by a time series of a plurality of groups each composed of a duration with the amplitude value of the first sound as the representative amplitude value and a group of representative amplitude values in the duration.

[0131] The device according to the above technical solution can appropriately control whether to select the second sound generated corresponding to the first sound based on the information on the duration and the time series of the representative amplitude values.

[0132] The audio processing device according to the 35th technical solution mastered based on the present disclosure may also be that in the audio processing device according to the 34th technical solution, the representative amplitude value is the ratio of the volume of the first sound to a preset reference volume.

[0133] The device according to the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the representative amplitude value corresponding to the ratio to the reference volume.

[0134] The audio processing device according to the 36th technical solution mastered based on the present disclosure may also be that in the audio processing device according to any one of the 1st, 2nd, 13th, 15th, 19th, and 20th technical solutions, the characteristic related to the first sound is a characteristic representing the duration during which the variation amount of the frequency characteristic is lower than a preset threshold value.

[0135] The device according to the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the duration during which the variation amount of the frequency characteristic is lower than a preset threshold value.

[0136] The audio processing device according to the 37th technical solution mastered based on the present disclosure may also be that in the audio processing device according to any one of the 1st, 2nd, 13th, 15th, 19th, 20th, and 36th technical solutions, the characteristic related to the first sound is a characteristic representing a plurality of groups composed of the duration during which the variation amount of the frequency characteristic is lower than a preset threshold value and the frequency characteristic in the duration, represented in time series.

[0137] The device according to the above technical solution can appropriately control whether to select the second sound corresponding to the first sound based on the duration during which the variation amount of the frequency characteristic is lower than a preset threshold value and the time series of the frequency characteristic.

[0138] The audio processing device according to the 38th technical solution mastered based on the present disclosure may also be that in the audio processing device according to any one of the 1st, 2nd, 13th to 24th, and 34th to 37th technical solutions, the circuit obtains a threshold value representing the volume corresponding to the boundary of whether a sound can be heard, calculates the volume of the second sound based on the characteristic related to the first sound, and selects the second sound when the volume of the second sound is greater than the threshold value.

[0139] The device according to the above technical solution can appropriately select the second sound when the volume of the second sound is greater than the threshold value corresponding to whether a sound can be heard.

[0140] The audio processing apparatus according to the 39th technical solution mastered based on the present disclosure may also be an audio processing apparatus according to any one of the 1st, 2nd, 13th to 20th, and 31st to 38th technical solutions, wherein the sound space information is scene information including information on sound sources in the sound space and information on the positions of listeners in the sound space, the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space, the circuit obtains a signal of the first sound, calculates the plurality of second sounds based on the scene information and the signal of the first sound, obtains characteristics related to the first sound from the information on the sound source, and controls whether to select each of the plurality of second sounds as a sound to which binaural processing is applied based on the characteristics related to the first sound, thereby selecting one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds. The scene information is updated based on input information, the characteristics related to the first sound are obtained according to the update of the scene information, and the update of the scene information is performed at a frequency lower than the frequency at which binaural processing is applied to one or more processing target sounds.

[0141] The apparatus according to the above technical solution can appropriately select one or more second sounds to which binaural processing is applied based on information obtained according to the update of scene information updated at a relatively low frequency.

[0142] The audio processing method according to the 40th technical solution mastered based on the present disclosure includes: a step of obtaining sound space information related to a sound space; a step of obtaining characteristics related to a first sound generated from a sound source in the sound space based on the sound space information; and a step of controlling whether to select a second sound generated corresponding to the first sound in the sound space based on the characteristics related to the first sound.

[0143] The method according to the above technical solution can achieve the same effect as the audio processing apparatus described in the 1st technical solution.

[0144] The program according to the 41st technical solution mastered based on the present disclosure is a program for causing a computer to execute the audio processing method according to the 40th technical solution.

[0145] The program according to the above technical solution can cause a computer to achieve the same effect as the audio processing method according to the 40th technical solution.

[0146] In addition, these inclusive or specific technical solutions can be implemented by a system, apparatus, method, integrated circuit, computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, apparatus, method, integrated circuit, computer program, or recording medium.

[0147] Hereinafter, an audio processing device, an encoding device, a decoding device, and a stereophonic sound reproduction system according to the present disclosure will be described in detail with reference to the accompanying drawings. The stereophonic sound reproduction system can also be expressed as a sound signal reproduction system.

[0148] In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, and the order of the steps shown in the following embodiments are examples, and do not limit the gist of the technical solutions grasped based on the present disclosure. In addition, regarding the constituent elements in the following embodiments that are not included in the basic technical solutions described in the present disclosure or are not described in the independent claims representing the most general concept, they are described as arbitrary constituent elements.

[0149] (Embodiment)

[0150] (Example of Stereophonic Sound Reproduction System)

[0151] Figure 2 FIG. is an example showing a stereophonic sound reproduction system. Specifically, Figure 2 FIG. shows a stereophonic sound reproduction system 1000 as an example of a system to which audio processing or decoding processing according to the present disclosure can be applied. Stereophonic sound is also expressed as Immersive Audio. The stereophonic sound reproduction system 1000 includes a sound signal processing device 1001 and a sound presentation device 1002.

[0152] The sound signal processing device 1001, which is also expressed as an audio processing device, performs audio processing on the sound signal emitted from a virtual sound source and generates a sound signal after audio processing to be presented to a listener. The sound signal is not limited to speech and can be any audible sound. The audio processing is, for example, signal processing performed on the sound signal in order to reproduce one or more effects that the sound undergoes during the period from being generated by the sound source to reaching the listener.

[0153] The sound signal processing device 1001 performs audio processing based on spatial information that describes the causes of the above effects. The spatial information is, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, and parameters related to the propagation of sound. The sound signal processing device 1001 is, for example, a PC (Personal Computer), a smart phone, a tablet computer, or a game console.

[0154] The signal after audio processing is presented to the listener from the sound presentation device 1002. The sound presentation device 1002 is connected to the sound signal processing device 1001 via wireless or wired communication. The audio processed sound signal generated by the sound signal processing device 1001 is transmitted to the sound presentation device 1002 via wireless or wired communication.

[0155] In the case where the sound presentation device 1002 is composed of multiple devices such as a device for the right ear and a device for the left ear, etc., sounds are presented synchronously by the multiple devices through communication between the multiple devices or communication of each of the multiple devices with the sound signal processing device 1001. The sound presentation device 1002 is, for example, a headphone worn on the listener's head, earplugs, a head-mounted display, or a surround speaker composed of multiple fixed speakers.

[0156] In addition, the stereophonic reproduction system 1000 can also be used in combination with an image presentation device or a stereoscopic image presentation device that visually provides an ER experience including AR / VR. For example, the space processed by the spatial information is a virtual space, and the positions of the sound source, the listener, and the object in this space are virtual positions of a virtual sound source, a virtual listener, and a virtual object in the virtual space. This space can also be represented as a sound space. In addition, the spatial information can also be represented as sound spatial information.

[0157] In addition, Figure 2 The system configuration example shows that the sound signal processing device 1001 and the sound presentation device 1002 are different devices, but the stereophonic reproduction system 1000 that can apply the audio processing method or decoding method of the present disclosure is not limited to Figure 2 this configuration. For example, the sound signal processing device 1001 may be included in the sound presentation device 1002, and the sound presentation device 1002 performs both audio processing and sound presentation.

[0158] In addition, the sound signal processing device 1001 and the sound presentation device 1002 may share the audio processing described in the present disclosure. In addition, a part or all of the audio processing described in the present disclosure may be implemented by a server connected to the sound signal processing device 1001 or the sound presentation device 1002 via a network.

[0159] In addition, the sound signal processing device 1001 may also perform audio processing by decoding a bitstream generated by encoding at least a part of the sound signal and the data of the spatial information for audio processing. Therefore, the sound signal processing device 1001 can also be represented as a decoding device.

[0160] (Examples of encoding devices)

[0161] Figure 3AIt is a block diagram showing a configuration example of an encoding device. Specifically, Figure 3A It shows the configuration of an encoding device 1100 as an example of the encoding device of the present disclosure.

[0162] Input data 1101 is encoding target data including spatial information and / or audio signals input to an encoder 1102. Details of the spatial information will be described later.

[0163] The encoder 1102 encodes the input data 1101 to generate encoded data 1103. The encoded data 1103 is, for example, a bitstream generated through an encoding process.

[0164] A memory 1104 stores the encoded data 1103. The memory 1104 can be, for example, a hard disk or an SSD (Solid-State Drive), or other memories.

[0165] In addition, in the above description, a bitstream generated through an encoding process is listed as an example of the encoded data 1103 stored in the memory 1104, but the encoded data 1103 can also be data other than a bitstream. For example, the encoding device 1100 can also store, in the memory 1104, transformed data generated by transforming the bitstream into a specified data format. The transformed data can also be, for example, a file or a multiplexed stream corresponding to one or more bitstreams.

[0166] Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format). In addition, the encoded data 1103 can also be in the form of multiple packets generated by splitting the above-mentioned bitstream or file.

[0167] For example, the bitstream generated by the encoder 1102 can also be transformed into data different from the bitstream. In this case, the encoding device 1100 includes a transformation unit (not shown), and the transformation process can be performed either by the transformation unit or by a CPU (Central Processing Unit), which is an example of a processor described later.

[0168] (Example of a decoding device)

[0169] Figure 3B It is a block diagram showing a configuration example of a decoding device. Specifically, Figure 3B It shows the configuration of a decoding device 1110 as an example of the decoding device of the present disclosure.

[0170] The memory 1114 stores, for example, the same data as the encoded data 1103 generated by the encoding device 1100. The stored data is read out from the memory 1114 and input into the decoder 1112 as input data 1113. The input data 1113 is, for example, a bitstream to be decoded. The memory 1114 can be, for example, a hard disk or an SSD, or other memories.

[0171] Alternatively, the decoding device 1110 may not input the data read out from the memory 1114 as the input data 1113 into the decoder 1112 as it is, but transform the read-out data and input the transformed data into the decoder 1112 as the input data 1113. The data before transformation can be, for example, multiplexed data containing one or more bitstreams. Here, the multiplexed data can also be, for example, a file having a file format such as ISOBMFF.

[0172] In addition, the data before transformation can also be a plurality of packets generated by splitting the above-mentioned bitstream or file. Data different from the bitstream can also be read out from the memory 1114 and transformed into a bitstream. In this case, the decoding device 1110 may include a transformation unit (not shown), and the transformation process may be performed by the transformation unit, or may be performed by a CPU as an example of the processor described later.

[0173] The decoder 1112 decodes the input data 1113 and generates a sound signal 1111 representing the sound to be presented to the listener.

[0174] (Another example of the encoding device)

[0175] Figure 3C is a block diagram showing another configuration example of the encoding device. Specifically, Figure 3C shows the configuration of an encoding device 1120 as another example of the encoding device of the present disclosure. In Figure 3C for the components identical to the components of Figure 3A the same reference numerals as those of Figure 3A are given, and the description of these components is omitted.

[0176] The encoding device 1100 stores the encoded data 1103 in the memory 1104. On the other hand, the encoding device 1120 is different from the encoding device 1100 in that it includes a transmission unit 1121 that sends the encoded data 1103 to the outside.

[0177] The transmission unit 1121 transmits a transmission signal 1122 generated based on the encoded data 1103 or data obtained by transforming the encoded data 1103 into another data form to other devices or servers. The data used in the generation of the transmission signal 1122 is, for example, the bitstream, multiplexed data, file, or packet described in the encoding device 1100.

[0178] (Another example of a decoding device)

[0179] Figure 3D It is a block diagram showing another configuration example of a decoding device. Specifically, Figure 3D It shows the configuration of a decoding device 1130 as another example of the decoding device of the present disclosure. In Figure 3D for components that are the same as the components of Figure 3B the same reference numerals as those of Figure 3B are given, and the description of these components is omitted.

[0180] The decoding device 1110 reads the input data 1113 from the memory 1114. On the other hand, the decoding device 1130 is different from the decoding device 1110 in that it has a receiving unit 1131 that receives the input data 1113 from the outside.

[0181] The receiving unit 1131 receives the received signal 1132 to obtain received data and outputs the input data 1113 input to the decoder 1112. The received data may be the same as the input data 1113 input to the decoder 1112 or data in a data form different from the input data 1113.

[0182] When the data form of the received data is different from the data form of the input data 1113, the receiving unit 1131 may transform the received data into the input data 1113. Alternatively, the received data may be transformed into the input data 1113 by a transformation unit (not shown) or the CPU of the decoding device 1130. The received data is, for example, the bitstream, multiplexed data, file, or packet described in the encoding device 1120.

[0183] (Examples of decoders)

[0184] Figure 4A It is a block diagram showing a configuration example of a decoder. Specifically, Figure 4A It shows the configuration of a decoder 1200 as an example of the decoder 1112 in Figure 3B or Figure 3D The input data 1113 is an encoded bitstream and includes encoded audio data as an encoded sound signal and metadata used in audio processing.

[0185]

[0186] The spatial information management unit 1201 acquires the metadata included in the input data 1113 and analyzes the metadata. The metadata includes information describing elements acting on sound configured in the sound space. The spatial information management unit 1201 manages the spatial information used in audio processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1203.

[0187] In addition, in the present disclosure, the information used in audio processing is expressed as spatial information, but other expressions may also be used. For example, the information used in audio processing may be expressed as sound space information or may be expressed as scene information. In addition, when the information used in audio processing changes over time, the spatial information input to the rendering unit 1203 may also be information expressed as a spatial state, a sound space state, or a scene state.

[0188] In addition, the spatial information can be managed for each sound space or for each scene. For example, when a plurality of different rooms are respectively represented as virtual spaces, the plurality of rooms can be managed as a plurality of different scenes. In addition, even in the same space, the spatial information can be managed as different scenes according to the represented situation.

[0189] Therefore, it is also possible to manage a plurality of spatial information for a plurality of sound spaces or a plurality of scenes. In the management of the plurality of spatial information, an identifier for respectively identifying the plurality of spatial information may be assigned to the spatial information.

[0190] The data of the spatial information may also be included in the bitstream as an example of the input data 1113. Alternatively, it may be that the bitstream includes an identifier of the spatial information, and the data of the spatial information is obtained from an information source outside the bitstream. Specifically, when the bitstream only includes the identifier of the spatial information, in rendering, the identifier of the spatial information can be used to obtain the data of the spatial information stored in the memory in the device or an external server as the input data 1113.

[0191] In addition, the information managed by the spatial information management unit 1201 is not limited to the information included in the bitstream. For example, in the input data 1113, as data not included in the bitstream, data representing the characteristics and structure of the space obtained from software or a server providing VR or AR may also be included.

[0192] In addition, the input data 1113 may also include data representing the characteristics and positions of the listener or the object. In addition, the input data 1113 may also include information obtained by sensors provided in the terminal including the decoding devices (1110, 1130) regarding the position of the listener, and may also include information representing the position of the terminal inferred based on the information obtained by the sensors.

[0193] That is, the spatial information management unit 1201 can also communicate with an external system or server to obtain spatial information and the listener's position. The spatial information management unit 1201 can obtain clock synchronization information from an external system and perform clock synchronization processing with the rendering unit 1203.

[0194] In addition, the space in the above description can be either a virtual space, i.e., a VR space, or a real space or a virtual space corresponding to the real space, i.e., an AR space or an MR space. In addition, the virtual space can also be represented as a sound field or a sound space. In addition, the information representing the position in the above description can be information such as coordinate values representing positions in the space, information representing relative positions with respect to a specified reference position, or information representing the movement or acceleration of positions in the space.

[0195] The sound data decoder 1202 decodes the encoded sound data included in the input data 1113 to obtain a sound signal.

[0196] The encoded sound data obtained by the stereophonic reproduction system 1000 is, for example, a bitstream encoded in a specified format such as MPEG-H 3D Audio (ISO / IEC 23008-3). In addition, MPEG-H 3D Audio is just an example of an encoding method that can be used when generating the encoded sound data included in the bitstream. The encoded sound data can also be a bitstream encoded in other encoding methods.

[0197] For example, the encoding method can also be non-reversible encoding and decoding such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Or, the encoding method can also be reversible encoding and decoding such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).

[0198] Or, any encoding method other than the above can be used. For example, PCM (pulse code modulation) data can be a type of encoded sound data. In this case, for example, when the quantization bit number of the PCM data is N, the decoding process can also be a process of converting an N-bit binary number into a number format (e.g., floating-point format) that the rendering unit 1203 can process.

[0199] The rendering unit 1203 acquires a sound signal and spatial information, performs acoustic processing on the sound signal using the spatial information, and outputs the sound signal after acoustic processing (sound signal 1111).

[0200] Before starting rendering, the spatial information management unit 1201 reads the metadata of the input signal, detects rendering items such as objects and sounds specified by the spatial information, and sends them to the rendering unit 1203. After the start of rendering, the spatial information management unit 1201 grasps the changes over time of the spatial information and the position of the listener, updates and manages the spatial information. Then, the updated spatial information is sent to the rendering unit 1203.

[0201] Based on the sound signal included in the input data 1113 and the spatial information received from the spatial information management unit 1201, the rendering unit 1203 generates and outputs a sound signal with acoustic processing added.

[0202] The update process of the spatial information and the output process of the sound signal with acoustic processing added can also be executed by the same thread. The spatial information management unit 1201 and the rendering unit 1203 can allocate processing to their respective independent threads. When the spatial information management unit 1201 and the rendering unit 1203 perform the update process of the spatial information and the output process of the sound signal with acoustic processing added in different threads, they can respectively set the start frequency of the threads or execute the processing in parallel.

[0203] When the spatial information management unit 1201 and the rendering unit 1203 perform processing in different independent threads, computing resources can be preferentially allocated to the rendering unit 1203. Thereby, it is possible to safely execute sound output processing that does not allow minute delays, such as generating a popping noise at a delay of 1 sample (0.02 msec).

[0204] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared with the output process of the sound signal, the update of the spatial information is a low-frequency process (for example, a process such as updating the facial orientation of the listener), so it does not have to be instantaneous like the output process of the sound signal. Therefore, even if the allocation of computing resources is restricted, it will not have a great impact on the audio quality.

[0205] The update of the spatial information can be executed regularly at every pre-set time or period, or can be executed when pre-set conditions are met. In addition, the update of the spatial information can be manually executed by the listener or the manager of the sound space, or can be executed using the change of an external system as a trigger.

[0206] For example, the controller can also be operated by the listener to update the spatial information when the standing position of the listener's own avatar is instantaneously distorted, or when the time instantaneously advances or returns. Alternatively, the spatial information can also be updated when a performance such as a sudden change of the venue is implemented by the administrator of the virtual space. In these cases, the thread for updating the spatial information managed by the spatial information management unit 1201 can be started not only regularly but also as a one-shot interrupt process.

[0207] Figure 4B It is a block diagram showing another configuration example of the decoder. Specifically, Figure 4B represents Figure 3B or Figure 3D the configuration of the decoder 1210 which is another example of the decoder 1112 in

[0208] Figure 4B It is different in that the input data 1113 does not contain encoded audio data but an unencoded audio signal as compared with Figure 4A . The input data 1113 includes a bitstream containing metadata and an audio signal.

[0209] Since the spatial information management unit 1211 is the same as the spatial information management unit 1201 of Figure 4A , the description thereof is omitted.

[0210] Since the rendering unit 1213 is the same as the rendering unit 1203 of Figure 4A , the description thereof is omitted.

[0211] In addition, the decoders 1112, 1200, and 1210 can also be represented as an audio processing unit that performs audio processing. Further, the decoding devices 1110 and 1130 can also be the audio signal processing device 1001 and can be represented as an audio processing device.

[0212] (Physical configuration of the audio signal processing device)

[0213] Figure 5 It is a diagram showing an example of the physical configuration of the audio signal processing device 1001. In addition, Figure 5 the audio signal processing device 1001 of Figure 3B can also be the decoding device 1110 of Figure 3D or the decoding device 1130 of Figure 3B or Figure 3D The multiple constituent elements shown in Figure 5 can also be installed by the multiple constituent elements shown in

[0214] Figure 5The voice signal processing device 1001 includes a processor 1402, a memory 1404, a communication IF (Interface) 1403, a sensor 1405, and a speaker 1401.

[0215] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The audio processing or decoding processing of the present disclosure can also be implemented by executing a program stored in the memory 1404 by the CPU, DSP, or GPU. In addition, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 can also be a dedicated circuit that performs signal processing on the voice signal including the audio processing of the present disclosure.

[0216] The memory 1404 is constituted by, for example, a RAM (Random Access Memory) or a ROM (ReadOnly Memory). The memory 1404 can also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1404 can also be an internal memory built in the CPU or GPU. In addition, in the memory 1404, spatial information managed by the spatial information management unit (1201, 1211) can also be stored. In addition, threshold data described later can also be stored.

[0217] The communication IF 1403 is, for example, a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The voice signal processing device 1001 communicates with other communication devices via the communication IF 1403, for example, and acquires a bitstream to be decoded. The acquired bitstream is stored in the memory 1404, for example.

[0218] The communication IF 1403 is constituted by, for example, a signal processing circuit corresponding to the communication method and an antenna. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), and can also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark), etc.

[0219] In addition, the communication method is not limited to the wireless communication method as described above. The communication method may also be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).

[0220] The sensor 1405 performs sensing for estimating the position and orientation of the listener. Specifically, the sensor 1405 estimates the position and / or orientation of the listener based on one or more detection results among the position, orientation, movement, speed, angular velocity, and acceleration of a part or the whole of the body, and generates position / orientation information indicating the position and / or orientation of the listener.

[0221] Alternatively, a device external to the sound signal processing device 1001 may include the sensor 1405. A part of the body may be the listener's head or the like. The position / orientation information may be information indicating the position and / or orientation of the listener in the real space, or may be information indicating the displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a specified time point. In addition, the position / orientation information may be information indicating the relative position and / or orientation with respect to the stereophonic reproduction system 1000 or an external device including the sensor 1405.

[0222] The sensor 1405 is, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Light Detection And Ranging). The sensor 1405 may also capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. In addition, a device that estimates the position using wireless in an arbitrary frequency band such as millimeter waves may be used as the sensor 1405.

[0223] In addition, the sound signal processing device 1001 may also acquire the position information from an external device including the sensor 1405 via the communication IF 1403. In this case, the sound signal processing device 1001 may not include the sensor 1405. Here, the external device is, for example, the sound prompting device 1002 or the stereoscopic image reproduction device worn on the listener's head as described in Figure 2 At this time, the sensor 1405 is constituted by combining various sensors such as a gyro sensor and an acceleration sensor, for example.

[0224] For example, as the speed of movement of the listener's head, the sensor 1405 can detect the angular velocity of rotation about at least one of three axes orthogonal to each other in the sound space, or can also detect the acceleration of displacement in which at least one of the above three axes is the displacement direction.

[0225] For example, as the amount of movement of the listener's head, the sensor 1405 can detect the amount of rotation about at least one of three axes orthogonal to each other in the sound space, or can also detect the amount of displacement in which at least one of the above three axes is the displacement direction. Specifically, the sensor 1405 detects the 6DoF position (x, y, z) and angles (yaw, pitch, roll) as the position of the listener. The sensor 1405 is constituted by combining various sensors for motion detection such as a gyro sensor and an acceleration sensor.

[0226] In addition, the sensor 1405 can also be implemented by a camera or a GPS (Global Positioning System) receiver for detecting the position of the listener, etc. It is also possible to use the position information obtained by performing self-position estimation using LiDAR or the like as the sensor 1405. For example, when the stereophonic reproduction system 1000 is implemented by a smart phone, the sensor 1405 is built in the smart phone.

[0227] In addition, the sensor 1405 may also include a temperature sensor such as a thermocouple for detecting the temperature of the sound signal processing device 1001. In addition, the sensor 1405 may also include a battery included in the sound signal processing device 1001, or a sensor for detecting the remaining amount of the battery connected to the sound signal processing device 1001, etc.

[0228] The speaker 1401 has, for example, a driving mechanism such as a diaphragm, a magnet, or a voice coil, and an amplifier, and presents the sound after sound processing to the listener as sound. The speaker 1401 operates the driving mechanism according to the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), and the driving mechanism vibrates the diaphragm. In this way, the diaphragm that vibrates corresponding to the sound signal generates sound waves, and the sound waves propagate in the air and are transmitted to the listener's ears, and the listener perceives the sound.

[0229] In addition, an example in which the sound signal processing device 1001 includes the speaker 1401 and presents the sound signal after sound processing via the speaker 1401 is listed here, but the sound signal presentation mechanism is not limited to the above configuration.

[0230] For example, it is also possible to output the sound signal after sound processing to an external sound prompting device 1002 connected via a communication module. The communication via the communication module can be either wired or wireless. In addition, as another example, the sound signal processing device 1001 has a terminal for outputting an analog signal of sound, and a cable such as an earplug is connected to the terminal to prompt the sound signal from the earplug or the like.

[0231] In the above case, the sound prompting device 1002 can also be a headset, earplug, head-mounted display, neck speaker, or wearable speaker that is worn on a part of the listener's head or body. Alternatively, the sound prompting device 1002 can also be a surround speaker composed of a plurality of fixed speakers. And the sound prompting device 1002 can also reproduce the sound signal.

[0232] (Physical configuration of the encoding device)

[0233] Figure 6 is a diagram showing an example of the physical configuration of the encoding device. Figure 6 The encoding device 1500 can also be Figure 3A the encoding device 1100 or Figure 3C the encoding device 1120, and can also Figure 3A or Figure 3C The multiple components shown by Figure 6 are installed by the multiple components shown by

[0234] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503, and a communication IF 1502.

[0235] The processor 1501 is, for example, a CPU, DSP, or GPU. The encoding process of the present disclosure can also be implemented by executing a program stored in the memory 1503 by the CPU, DSP, or GPU. In addition, the processor 1501 is, for example, a circuit for information processing. The processor 1501 can also be a dedicated circuit for signal processing of the sound signal including the encoding process of the present disclosure.

[0236] The memory 1503 is composed of, for example, RAM or ROM. The memory 1503 can also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1503 can also be an internal memory embedded in the CPU or GPU.

[0237] The communication IF 1502 is, for example, a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with other communication devices via the communication IF 1502, for example, and transmits the encoded bitstream.

[0238] The communication IF 1502 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), and may also be LTE, NR, Wi-Fi (registered trademark), etc. In addition, the communication method is not limited to wireless communication methods. The communication method may also be a wired communication method such as Ethernet (registered trademark), USB, or HDMI (registered trademark).

[0239] (Configuration of the rendering unit)

[0240] Figure 7 It is a block diagram showing a configuration example of the rendering unit. Specifically, Figure 7 represents Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213.

[0241] The rendering unit 1300 is composed of an analysis unit 1301, a selection unit 1302, and a synthesis unit 1303, and performs audio processing on the sound data included in the input signal and outputs it.

[0242] The input signal is composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bitstream composed of sound data and metadata (control information), and in this case, spatial information may also be included in the metadata.

[0243] The spatial information is information related to the sound space (three-dimensional sound field) formed by the stereophonic reproduction system 1000, and is composed of information related to the objects included in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources and non-sounding objects that do not emit sound. The sound source object may also be simply expressed as a sound source.

[0244] The non-sounding object functions as an obstacle object that reflects the sound emitted by the sound source object, but there are also cases where the sound source object functions as an obstacle object that reflects the sound emitted by other sound source objects. The obstacle object may also be expressed as a reflection object.

[0245] As information commonly given to the sound source object and the non-sounding object, there are position information, shape information, and the attenuation rate of the volume when the object reflects sound, etc.

[0246] The position information is represented by the coordinate values of three axes, such as the X-axis, Y-axis, and Z-axis, in the Euclidean space, but it does not necessarily have to be three-dimensional information. For example, the position information may also be two-dimensional information represented by the coordinate values of the two axes, the X-axis and the Y-axis. The position information of the object is determined by the representative position of the shape represented by a grid or a voxel.

[0247] The shape information may also include information related to the material of the surface.

[0248] The attenuation rate can be represented by a real number between 0 and 1, or by a negative decibel value. Since the volume does not increase through reflection in real space, the attenuation rate is set to a negative decibel value. However, for example, in order to present a sense of horror in an unrealistic space, an attenuation rate of 1 or more, that is, a positive decibel value, can be deliberately set.

[0249] In addition, the attenuation rate can also be set to different values for each of the multiple frequency bands that make up the frequency bands, or can be set independently for each frequency band. In addition, when setting the attenuation rate for each type of material of the object surface, the corresponding attenuation rate value can also be used based on the information related to the material of the surface.

[0250] In addition, the space information may also include information indicating whether the object belongs to a living being and information indicating whether the object is a moving body, etc. When the object is a moving body, the position represented by the position information can also move with time. In this case, the information on the changed position or the amount of change is transmitted to the rendering unit 1300.

[0251] The information related to the sound source object includes, in addition to the information commonly given to the sound source object and the non-sounding object, sound data, and information required to radiate the sound data into the sound space. The sound data is data representing information related to the frequency and intensity of the sound, etc., and is data representing the sound perceived by the listener.

[0252] Typically, the sound data is a PCM signal, but it can also be data compressed using an encoding method such as MP3. In this case, at least decoding is required before the signal reaches the synthesis unit 1303, so the rendering unit 1300 may also include a decoding unit (not shown). Alternatively, the signal can also be decoded by the sound data decoder 1202.

[0253] For one sound source object, one sound data can be set, or multiple sound data can be set. In addition, identification information for identifying each sound data can be given to the sound data, and the information related to the sound source object can also include the identification information of the sound data.

[0254] The information required to radiate the sound data into the sound space may include, for example, information on the reference volume used as a reference in the reproduction of the sound data, information indicating the nature (also called characteristics) of the sound data, information related to the position of the sound source object, and information related to the orientation of the sound source object (that is, information related to the directivity of the sound emitted by the sound source object), etc.

[0255] The information on the reference volume is, for example, the effective value of the amplitude value of the sound data at the sound source position when radiating the sound data into the sound space, and can also be represented in floating point as a decibel (dB) value.

[0256] For example, when the reference volume is 0 dB, it can also mean radiating the sound into the sound space from the position indicated by the information related to the position of the sound source object at the original volume without increasing or decreasing the volume of the signal level represented by the sound data. Additionally, when the reference volume is -6 dB, it can also mean setting the volume of the signal level represented by the sound data to approximately half and radiating the sound into the sound space from the position indicated by the information related to the position of the sound source object.

[0257] The information on the reference volume can be assigned to each sound data, or can be uniformly assigned to multiple sound data.

[0258] The information indicating the nature of the sound data can be, for example, information related to the volume of the sound source, and is information representing the change in the volume of the sound source over time series.

[0259] For example, when the sound space is a virtual meeting room and the sound source is a speaker, the volume changes intermittently in a short period of time. That is, the audible part and the inaudible part alternate. Additionally, when the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain length of time. Additionally, when the sound space is a battlefield and the sound source is an explosive, the volume of the explosion sound only becomes large for an instant and then remains silent or in a small state.

[0260] In this way, the information on the volume of the sound source not only includes the information on the loudness of the sound, but can also include the information on the change in the loudness of the sound. Such information can also be used as the information indicating the nature of the sound data.

[0261] The information on the change can also be represented by data representing the frequency characteristics in time series. The information on the change can also be represented by data representing the duration length of the audible interval. The information on the change can also be represented by time series data representing the duration length of the audible interval and the duration length of the inaudible interval. The information on the change can also be represented by multiple sets of data that list in time series the duration during which the amplitude of the sound signal can be regarded as stable (can be regarded as approximately constant) and the amplitude value of the signal during that period, etc.

[0262] The information on the change can also be represented by data on the duration during which the frequency characteristics of the sound signal can be regarded as stable. The information on the change can also be represented by listing in time series multiple sets of data on the duration during which the frequency characteristics of the sound signal can be regarded as stable and the frequency characteristics during that period. The information on the change can also be represented, for example, in the form of data representing the approximate shape of the spectrogram.

[0263] In addition, the volume used as a reference for the above frequency characteristics may also be the above reference volume. Information on the reference volume and information indicating the nature of the sound data can be used for the calculation process of the volume of the direct sound or reflected sound perceived by the listener, and can also be used for the selection process of whether to make the listener perceive it. Other examples and utilization methods of information indicating the nature of the sound data will be described later.

[0264] Information related to the orientation of the sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted, and the orientation information of the sound source object can be represented by the azimuth angle (yaw) and the elevation angle (pitch). The orientation information of the sound source object can also change over time, and when it changes, it is transmitted to the rendering unit 1300.

[0265] Information related to the listener is information related to the position and orientation of the listener in the sound space. The information related to the position (position information) is represented by the positions of the XYZ axes in the Euclidean space, but it does not necessarily have to be three-dimensional information and can also be two-dimensional information. The information related to the orientation of the listener (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted, and the orientation information of the listener can be represented by the azimuth angle (yaw) and the elevation angle (pitch).

[0266] The position information and orientation information of the listener can also change over time, and when they change, they are transmitted to the rendering unit 1300.

[0267] The sensor information includes information such as the rotation amount or displacement amount detected by the sensor 1405 worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering unit 1300, and the rendering unit 1300 updates the information on the position and orientation of the listener based on the sensor information. The sensor information can also include, for example, the position information obtained by the portable terminal performing self-position estimation using GPS, a camera, LiDAR, etc.

[0268] In addition, instead of the sensor 1405, information obtained from the outside via the communication module can be detected as sensor information. Information indicating the temperature of the sound signal processing device 1001 and information indicating the remaining battery level can also be obtained from the sensor 1405. In addition, the computing resources (CPU capabilities, memory resources, PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 can be obtained in real time.

[0269] The analysis unit 1301 analyzes the sound signal included in the input signal and the spatial information received from the spatial information management units (1201, 1211), and detects the information required for generating the direct sound and the reflected sound, and the information required for selecting whether to generate the reflected sound.

[0270] The information required for generating the direct sound and the reflected sound is, for example, values related to the path to the listening position, the time required to reach, and the volume at the time of arrival, respectively, for the direct sound and the reflected sound.

[0271] The information required for selecting the output reflected sound is information indicating the relationship between the direct sound and the reflected sound, such as a value related to the time difference between the direct sound and the reflected sound, and a value related to the volume ratio between the direct sound and the reflected sound at the listening position.

[0272] In addition, when the volume is expressed in units of decibels on a logarithmic axis (when the volume is represented in the decibel region), the volume ratio of the two signals is of course represented by the difference in decibel values. Specifically, the volume ratio of the two signals can be the difference when the amplitude values of the respective signals are represented in the decibel region. This value can also be calculated based on energy values or power values, etc. In addition, this difference can be referred to as the difference in gain or simply as the gain difference in the decibel region.

[0273] That is, since the volume ratio in the present disclosure is substantially the amplitude ratio of the signals, it can also be expressed as Soundvolume ratio, Volume ratio, Amplitude ratio, Soundlevel ratio, Sound intensity ratio, or Gain ratio, etc. In addition, when the unit of volume is decibel, the volume ratio in the present disclosure can of course also be referred to as the volume difference.

[0274] In the present disclosure, the "volume ratio" typically refers to the gain difference when the volumes of two sounds are expressed in units of decibels. In the example of the embodiment, the threshold data is also typically defined by the gain difference represented in the decibel region. However, the volume ratio is not limited to the gain difference in the decibel region. When using a volume ratio represented outside the decibel region, the threshold data defined in the decibel region can be converted to the unit of the calculated volume ratio for use. Alternatively, the threshold data defined in each unit can be stored in the memory in advance.

[0275] That is, even if a ratio such as an energy value or a power value is used instead of the volume ratio, it is obvious that the algorithm in the present disclosure can be applied to solve the problems in the present disclosure.

[0276] The time difference between the direct sound and the reflected sound is, for example, the time difference between the arrival time (arrival moment) of the direct sound and the arrival time (arrival moment) of the reflected sound. The time difference between the direct sound and the reflected sound can also be the time difference between the moments when the direct sound and the reflected sound respectively reach the listening position, the difference in the times required for the direct sound and the reflected sound to respectively reach the listening position, or the time difference between the moment when the generation of the direct sound ends and the moment when the reflected sound reaches the listening position. The calculation methods for these values will be described later.

[0277] The selection unit 1302 uses the information calculated by the analysis unit 1301 and the threshold data to select whether to generate a reflected sound. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound to be generated. In still other words, the selection unit 1302 selects which of the multiple reflected sounds to generate.

[0278] The threshold data is, for example, represented as the boundary (threshold) at which the reflected sound is perceived or not perceived in a curve graph having the value of the time difference between the direct sound and the reflected sound on the horizontal axis and the volume ratio between the direct sound and the reflected sound on the vertical axis. The threshold data can also be expressed by an approximate formula having the value of the time difference between the direct sound and the reflected sound as a variable, or can be expressed by an array having the value of the time difference between the direct sound and the reflected sound as an index and having corresponding thresholds.

[0279] For example, when the volume ratio of the volume of the direct sound at the arrival time to the volume of the reflected sound at the arrival time among the values of the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is a value larger than the threshold set with reference to the threshold data, the selection unit 1302 selects to generate a reflected sound.

[0280] The time difference between the arrival time of the direct sound and the arrival time of the reflected sound is, in other words, the difference in the times required for the direct sound and the reflected sound to respectively reach the listening position. In addition, the time difference between the time point when the generation of the direct sound ends and the time point when the reflected sound reaches the listening position can also be used as the time difference between the direct sound and the reflected sound. In this case, threshold data different from the threshold data set using the time difference between the arrival time of the direct sound and the arrival time of the reflected sound as a reference can also be used, or common threshold data can be used.

[0281] Regarding the threshold data, it can be obtained from the memory 1404 of the sound signal processing device 1001, or can be obtained from an external storage device via the communication module. The storage method of the threshold data and the setting method of the threshold will be described later.

[0282] The synthesis unit 1303 synthesizes the sound signal of the direct sound and the sound signal of the reflected sound selected and generated by the selection unit 1302.

[0283] Specifically, based on the information of the direct sound arrival time and the volume at the time of direct sound arrival calculated by the analysis unit 1301, the synthesis unit 1303 processes the input sound signal to generate a direct sound. In addition, based on the information of the reflected sound arrival time and the volume at the time of reflected sound arrival of the reflected sound selected by the selection unit 1302, the synthesis unit 1303 processes the input sound signal to generate a reflected sound. Then, the synthesis unit 1303 synthesizes and outputs the generated direct sound and reflected sound.

[0284] (Operation of the rendering unit)

[0285] Figure 8 is a flowchart showing an operation example of the sound signal processing apparatus 1001. In Figure 8 the processing mainly executed by the rendering unit 1300 of the sound signal processing apparatus 1001 is shown.

[0286] In the analysis process of the input signal ( Figure 8 S101), the analysis unit 1301 analyzes the input signal input to the sound signal processing apparatus 1001 and detects the direct sound and the reflected sound that can be generated in the sound space. The reflected sound detected here is a reflected sound candidate that is selected by the selection unit 1302 as the reflected sound to be finally generated by the synthesis unit 1303. In addition, the analysis unit 1301 analyzes the input signal and calculates the information required for generating the direct sound and the reflected sound and the information required for selecting the reflected sound to be generated.

[0287] First, the characteristics of the direct sound and the reflected sound are calculated. Specifically, the arrival time and the volume at the time of arrival when the direct sound and the reflected sound reach the listener are calculated. In the case where there are multiple objects in the sound space as reflection objects, the characteristics of the reflected sound are calculated for each of the multiple objects.

[0288] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is a path connecting the position information S(xs, ys, zs) of the sound source object and the position information A(xa, ya, za) of the listener. The direct sound arrival time (td) is a value obtained by dividing the length of the path connecting the position information S(xs, ys, zs) and the position information A(xa, ya, za) by the speed of sound (about 340 m / s).

[0289] For example, the length (X) of the path is obtained by ((xs - xa)^2 + (ys - ya)^2 + (zs - za)^2)^0.5. The volume attenuates inversely with the distance. Therefore, when the volume at the position information S(xs, ys, zs) of the sound source object is N and the unit distance is U, the volume at the time of direct sound arrival (ld) is obtained by ld = N * U / X.

[0290] The volume N at the sound source position may also be the reference volume described above.

[0291] The reflection arrival time (tr) is calculated based on the reflection arrival path (pr). The reflection arrival path (pr) is the path connecting the position of the sound image of the reflected sound and the position information A(xa, ya, za).

[0292] In addition, for example, the "mirror method" or "ray tracing method" can be used to derive the position of the sound image of the reflected sound, and any other method for deriving the sound image position can also be used. The mirror method is a method of simulating the sound image by assuming that the reflected wave on the wall surface in the room exists as a mirror image at a position symmetric to the wall surface with respect to the sound source, and assuming that sound waves are radiated from the position of this mirror image. The ray tracing method is a method of simulating the image (sound image) observed at a certain point by tracing waves that propagate in a straight line such as light rays or sound rays.

[0293] Figure 9 It is a diagram showing the positional relationship where the listener and the obstacle object are relatively far apart. Figure 10 It is a diagram showing the positional relationship where the listener and the obstacle object are relatively close. That is, Figure 9 and Figure 10 respectively show examples where the sound image of the reflected sound is formed at positions symmetric to the sound source position across the wall. By obtaining the position of the sound image of the reflected sound on the xyz axis based on such a relationship, the arrival time of the reflected sound can be obtained in the same way as the method for calculating the arrival time of the direct sound.

[0294] The reflection arrival time (tr) is a value obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound and the position information A(xa, ya, za) by the speed of sound (about 340 m / s). The volume attenuates in inverse proportion to the distance. Therefore, when the volume at the sound source position is N, the unit distance is U, and the attenuation rate of the volume during reflection is G, the volume (lr) when the reflected sound arrives is obtained by lr = N * G * U / Y.

[0295] As described above, the attenuation rate G can be represented by a real number between 0 and 1, or by a negative decibel value. In this case, the volume of the entire signal attenuates by an amount corresponding to G. In addition, the attenuation rate can also be set for each of the multiple frequency bands that make up the signal. In this case, the analysis unit 1301 applies the specified attenuation rate to each frequency component of the signal. In addition, in order to reduce the amount of calculation, the analysis unit 1301 can also use the representative value or average value of the multiple attenuation rates of the multiple frequency bands as the overall attenuation rate, and attenuate the volume of the entire signal accordingly.

[0296] Next, the analysis unit 1301 calculates the volume ratio (L), which is the ratio of the volume (ld) at the arrival of the direct sound to the volume (lr) at the arrival of the reflected sound, which is required for the selection of the target reflected sound to be generated, and the time difference (T) between the direct sound and the reflected sound.

[0297] The volume ratio (L), which is the ratio of the volume (ld) at the arrival of the direct sound to the above-mentioned lr, is obtained, for example, by L = (N * G * U / Y) / (N * U / X) = G * X / Y. Since the obtained value is the ratio of volumes, the values of N and U can be any preset values.

[0298] The time difference (T) between the direct sound and the reflected sound can also be, for example, the time difference between the times required for the direct sound and the reflected sound to reach the listening position respectively. For example, the difference (T) between the times required for the direct sound and the reflected sound to reach the listening position respectively is obtained by T = tr - td.

[0299] In addition, the time difference (T) can also be the difference in the arrival times of the direct sound and the reflected sound at the listening position. In addition, the time difference (T) can also be the time difference between the end time of the direct sound speech and the arrival time of the reflected sound at the listening position. That is, the time difference (T) can also be the time difference between the end time of the direct sound and the start time of the reflected sound at the listening position.

[0300] Next, in the selection process of the reflected sound ( Figure 8 S102), the selection unit 1302 selects whether to generate the reflected sound calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound to be generated. In the case of multiple reflected sounds, the selection unit 1302 selects whether to generate each reflected sound. The result of the selection unit 1302's selection of whether to generate each reflected sound can be to select one or more target reflected sounds from the multiple reflected sounds, or not to select any target reflected sound.

[0301] In addition, the selection unit 1302 is not limited to the generation process, and can also select the reflected sound as the application object of other processes. For example, the selection unit 1302 can also select the reflected sound as the application object of binaural processing. In addition, the selection unit 1302 basically only selects one or more reflected sounds of the processing object. However, the selection unit 1302 can also only select one or more reflected sounds that are not the processing object. And it is also possible to apply processing to one or more reflected sounds that are not selected.

[0302] For example, the selection of the reflected sound is based on the volume ratio (L) and the time difference (T) calculated by the analysis unit 1301. By performing the selection process based on the time difference (T) between the direct sound and the reflected sound, compared with the case of performing the selection process only based on the volume difference between the direct sound and the reflected sound, it is possible to more appropriately select the reflected sound that has a greater impact on the listener's perception.

[0303] Specifically, the selection of whether to generate a reflected sound is performed, for example, by comparing the volume ratio of the direct sound to the reflected sound corresponding to the time difference between the direct sound and the reflected sound with a preset threshold value. The threshold value is set with reference to threshold data. The threshold data is an index representing the boundary at which the reflected sound of the direct sound is perceived by the listener, and is defined by the ratio of the volume (Id) at the arrival time of the direct sound to the volume (lr) at the arrival time of the reflected sound.

[0304] In addition, the threshold value corresponds to a value expressed by a numerical value set corresponding to the time difference (T), etc. The threshold data corresponds to the relationship between the time difference (T) and the threshold value, and corresponds to table data or a relational expression used to determine or calculate the threshold value at the time difference (T). The form and type of the threshold data are not limited to table data or a relational expression.

[0305] Figure 11 is a graph showing the relationship between the time difference between the direct sound and the reflected sound and the threshold value. For example, it is also possible to refer to, as Figure 11 shown, the threshold data of the volume ratio preset for each value of the time difference between the direct sound and the reflected sound. Or, it is also possible to refer to the threshold data obtained by interpolation or extrapolation, etc., from the Figure 11 shown threshold data.

[0306] And, the threshold value of the volume ratio at the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. And, the selection unit 1302 determines whether to select the reflected sound as the generation target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than the threshold value.

[0307] By performing the selection process using the threshold data of the volume ratio preset for each value of the time difference between the direct sound and the reflected sound, it is possible to implement a selection process that takes into account post-masking or the precedence effect. Details regarding the type, form, storage method, and setting method, etc., of the threshold data will be described later.

[0308] Next, in the generation process of the direct sound and the reflected sound ( Figure 8 S103), the synthesis unit 1303 generates the sound signal of the direct sound and the sound signal of the reflected sound selected as the generation target reflected sound by the selection unit 1302 and synthesizes them.

[0309] The sound signal of the direct sound is generated by applying the arrival time (td) and the arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object included in the input information. Specifically, a process of delaying the sound data by the amount of the arrival time (td) and multiplying it by the arrival volume (ld) is performed. The process of delaying the sound data is a process of moving the position of the sound data back and forth on the time axis. For example, a process of delaying the sound data without deteriorating the sound quality as disclosed in Patent Document 2 may also be applied.

[0310] The sound signal of the reflected sound is generated in the same way as the direct sound by applying the arrival time (tr) and the arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object.

[0311] However, the arrival volume (lr) in the generation of the reflected sound is different from the arrival volume of the direct sound, and is a value to which the attenuation rate G of the volume in the reflection is applied. G may be an attenuation rate applied to the entire frequency band together. Or, in order to reflect the bias of the frequency components generated by the reflection, the reflection rate may be specified for each specified frequency band. In this case, the process of applying the arrival volume (lr) may also be implemented as a process of multiplying by the attenuation rate for each frequency band, that is, a process of a frequency equalizer.

[0312] In the above example, the path lengths when the direct sound and the reflected sound candidates reach the listener are calculated respectively. Furthermore, the arrival time and the arrival volume are calculated based on each path length. And, based on their time difference and volume ratio, a selection process of the reflected sound candidate is performed.

[0313] In addition, as another example, the selection process may also be performed based on the path lengths when the direct sound and the reflected sound reach the listener respectively, omitting the calculation of the arrival time and the arrival volume of the direct sound and the reflected sound and the calculation of the time difference and the volume ratio. In this case, a threshold corresponding to the path length difference may also be preset for the path length ratio. And, the selection process may also be performed according to whether the calculated path length ratio is above the threshold corresponding to the calculated path length difference. Thus, it is possible to perform the selection process based on the path length difference corresponding to the time difference while reducing the amount of calculation.

[0314] In addition, in addition to the path length difference, a value of a parameter representing the sound propagation speed or a value of a parameter that affects the parameter of the sound propagation speed may also be used.

[0315] (Details of the selection process)

[0316] The details of the selection process for whether to generate the reflected sound will be described.

[0317] The selection of the reflected sound is performed by comparing a threshold value of the volume ratio, which is the ratio of the volume of the direct sound when it arrives to the volume of the reflected sound when the time difference (T) between the direct sound and the reflected sound is set, with the volume ratio (L) calculated by the analysis unit 1301. For example, the threshold value of the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced among the threshold values ​​of the volume ratio pre-set for each value of the time difference between the direct sound and the reflected sound. And, depending on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value, it is determined whether the reflected sound is selected as the generated target reflected sound.

[0318] The time difference (T) may be, for example, the difference between the time when the direct sound and the reflected sound arrive at the listening position, the time difference between the time required for the direct sound and the reflected sound to arrive at the listening position, and the time difference between the time when the direct sound ends and the time when the reflected sound arrives at the listening position. Here, the end time of the direct sound may also be obtained by, for example, adding the duration of the direct sound to the arrival time of the direct sound.

[0319] Threshold data can also be determined by the auditory nerve function or the cognitive function of the brain, more specifically, by the priority effect described later, the temporal masking phenomenon described later, or a combination thereof, based on the listener's perception to detect the minimum time difference between two sounds. The specific value can be derived from the research results of known temporal masking effects, priority effects, or echo detection limits, or can be obtained through a listening experiment based on the application to the virtual space.

[0320] Figure 12A , Figure 12B and Figure 12C FIG. 2 is a diagram showing an example of a method for setting threshold data. Figure 12A , Figure 12B and Figure 12C As shown, the threshold data is represented by the boundary (threshold) where the reflected sound is perceived or not perceived in a graph having the time difference between the direct sound and the reflected sound on the horizontal axis and the volume ratio of the direct sound to the reflected sound on the vertical axis.

[0321] The threshold data may also be expressed by an approximate expression having the time difference between the direct sound and the reflected sound as a variable. Figure 11 The index of the time difference between the direct sound and the reflected sound and the arrangement of the threshold values ​​corresponding to the index are stored in the area of ​​the memory 1404.

[0322] In addition, Figure 12CWhen the height (minimum audible limit) of the line parallel to the horizontal axis in Example 4 is used as the threshold, instead of comparing the volume ratio (L) of the reflected sound to the direct sound with the threshold, the volume of the reflected sound itself is compared with the threshold. This is because this threshold represents the boundary volume at which it can be perceived by the listener and is used to determine that sounds with a volume ratio smaller than this threshold are non-reproduced sounds. That is, the threshold corresponding to the minimum audible limit is not a threshold for the ratio of the volume of the reflected sound to the volume of the direct sound.

[0323] When the minimum audible limit is used as the threshold, the threshold is constant regardless of the time difference (T), so the time difference (T) does not have to be calculated either.

[0324] In addition, when multiple reflected sounds are generated through the analysis process ( Figure 8 S101), the selection process can be performed on all the reflected sounds, or the selection process can be performed only on the reflected sounds with high evaluation values based on the evaluation values derived for each reflected sound through a preset evaluation method. Here, the evaluation value of the reflected sound corresponds to the perceptual importance of the reflected sound. In addition, a high evaluation value corresponds to a large evaluation value, and these expressions can be used interchangeably.

[0325] The selection unit 1302 can also calculate the evaluation value of the reflected sound, for example, through a preset evaluation method corresponding to the volume of the sound source, the visuality of the sound source, the localizability of the sound source, the visuality of the reflection object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.

[0326] Specifically, the higher the volume of the sound source, the higher the evaluation value. In addition, in order to make the visual localization consistent with the acoustic localization, the evaluation value can also be high when the sound source object or the reflection object (obstacle object) can be visually observed by the listener or when the localizability of the sound source object is high.

[0327] In addition, the opening of the arrival angles of the direct sound and the reflected sound and the difference in the arrival times of the direct sound and the reflected sound have a great impact on the perception of space. Therefore, the evaluation value can also be high when the opening of the arrival angles of the direct sound and the reflected sound is large, or when the difference in the arrival times of the direct sound and the reflected sound is large.

[0328] The information on the volume of the sound source can also represent the reference volume set for each content, the temporal transition of the volume, or both.

[0329] For example, when the virtual space is a virtual meeting room and the direct sound is the voice of conversation, the volume changes intermittently in a short period of time. That is, the audible part and the inaudible part alternate. In addition, when the virtual space is a concert hall and the direct sound is the performance of a piece of music, the volume is maintained for a certain length of time. In addition, when the virtual space is a battlefield and the direct sound is an explosion sound, the volume only increases for an instant and then continues to be in a silent or low state.

[0330] In this way, the information on the volume of the sound source not only includes the information on the reference volume corresponding to the volume setting when the sound is radiated into the virtual space, but can also include the information on the change in the size of the sound.

[0331] The information on the change can also be represented by data indicating the frequency characteristics in a time series. The information on the change can also be represented by data indicating the duration length of the audible section. The information on the change can also be represented by time series data indicating the duration length of the audible section and the duration length of the inaudible section. The information on the change can also be represented by multiple sets of data that list in a time series the duration during which the amplitude of the sound signal can be regarded as stationary (can be regarded as approximately constant) and the amplitude value of the signal during that period, etc.

[0332] The information on the change can also be represented by data on the duration during which the frequency characteristics of the sound signal can be regarded as stationary. The information on the change can also be represented by multiple sets of data that list in a time series the duration during which the frequency characteristics of the sound signal can be regarded as stationary and the frequency characteristics during that period, etc.

[0333] In addition, countermeasures for using the temporal change in the frequency characteristics of a signal for audio processing in a virtual space have been widely carried out (Patent Document 1, etc.). In view of such prior art, the above-mentioned group can of course also be a group of the duration length with constant frequency characteristics and its frequency characteristics.

[0334] The geometric relationship can also be the relationship among the positions of the sound source, the listener, and the reflection object in the virtual space. Through their relationship, the path lengths through which the direct sound and the reflected sound arrive can be geometrically calculated. Therefore, if the relationship that the volume is inversely proportional to the distance is utilized, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.

[0335] In the calculation of the reference volume of the reflected sound, the reflection coefficient of the reflection object can also be used. In addition, as the reflection coefficient, typical values that are commonly used can also be used. On the other hand, in the case of special conditions such as the reflection object being covered with a sound-absorbing material, etc., as the reflection coefficient of the reflection object, a specially given reflection coefficient can also be used.

[0336] The reflected sound can also be evaluated based on the volume of the reflected sound. The volume of the reflected sound can also be calculated based on the geometric relationship between the direct sound and the reflected sound as described above, and the index given to the reflection object. The reflected sound can also be evaluated by comparing the volume with a preset threshold.

[0337] Furthermore, information indicating the temporal transition of the volume of the sound source may be reflected in the evaluation. For example, when the information indicating the temporal transition of the volume of the sound source indicates the duration of the sound interval, the evaluation value of the reflected sound may be maintained as it is when the moment is within the sound interval. On the other hand, when the moment is outside the sound interval, even if the reference volume of the reflected sound exceeds the threshold, the evaluation value of the reflected sound may be reduced or set to zero.

[0338] Alternatively, the information indicating the temporal change in the volume of the sound source may be data that lists in time series a duration during which the amplitude of the sound signal is considered to be substantially constant and a plurality of sets of amplitude values ​​of the signal during the duration. In this case, the reflected sound may be evaluated by changing the reference volume of the reflected sound in conjunction with the change in the amplitude value in the data.

[0339] In addition, as information indicating the volume of direct sound, both the reference volume information and the temporally transitioned volume information may be used. For example, after calculating the evaluation value based on the reference volume information, the evaluation value may be corrected using the transitioned volume information.

[0340] In the evaluation of reflected sound, all of the above methods may be executed, or only a part of them may be executed. For example, reflected sound may be evaluated by a plurality of evaluation methods, or by a single evaluation method.

[0341] When the reflected sound is evaluated by a plurality of evaluation methods, whether or not to select the reflected sound may be determined based on evaluation values ​​comprehensively determined by the plurality of evaluation methods, or may be determined based on evaluation values ​​of each of the plurality of evaluation methods.

[0342] The sound signal processing device 1001 may also select the sound when all the evaluation results based on the multiple evaluation methods indicate that the sound is selected, when determining whether to select the reflected sound based on each of the multiple evaluation methods. Alternatively, the sound signal processing device 1001 may also select the sound when any one of the multiple evaluation results based on the multiple evaluation methods indicates that the sound is selected.

[0343] In addition, for example, priorities may also be set for the first to third evaluation methods. Moreover, when the sound signal processing device 1001 determines not to select a sound by the first evaluation method, it may finally determine not to select the sound without relying on the determination results of the second and third evaluation methods. In addition, when the sound processing device determines not to select a sound by one of the second and third evaluation methods but determines to select the sound by the other, it may finally determine to select the sound.

[0344] In addition, the selection process and the evaluation process may be executed independently, or only one of them may be executed. In addition, the evaluation process may be performed only on the reflected sound determined to be selected in the selection process, and whether to select the reflected sound may be determined again in the evaluation process. Or, the evaluation process may be performed only on the reflected sound determined not to be selected in the selection process, and whether to select the reflected sound may be determined again in the evaluation process.

[0345] The above selection process can be interpreted as a process of selecting a reflected sound based on the properties of the direct sound. For example, in the process of selecting a reflected sound based on the properties of the direct sound, a threshold used in the selection of the reflected sound is set or adjusted according to the properties of the direct sound. Or, an evaluation value used in the selection of the reflected sound is calculated based on one or more of the volume of the sound source, the visibility of the sound source, the localizability of the sound source, the visibility of the reflection object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.

[0346] In addition, in the process of selecting a reflected sound based on the properties of the direct sound, it is not limited to the process of setting or adjusting a threshold according to the properties of the direct sound and the process of calculating an evaluation value used in the selection of the reflected sound to be processed, and other processes may also be performed. In addition, when performing the process of setting or adjusting a threshold according to the properties of the direct sound or the process of calculating an evaluation value used in the selection of the reflected sound to be processed, a part of the process may be changed, or a new process may be added.

[0347] In addition, setting a threshold may also include adjusting the threshold and changing the threshold, etc.

[0348] (Method for setting threshold)

[0349] The threshold data used in the selection process may also be set, for example, with reference to known values of echo detection limits based on the precedence effect or masking thresholds based on the post-masking effect.

[0350] The precedence effect refers to the phenomenon where, when hearing sounds from two locations, the sound source is identified as being at the location where the sound was heard first in time. If two short sounds fuse and are heard as one sound, the perceived location of the overall sound (the localization location) is generally mainly determined by the location of the initial sound. The echo detection limit is a phenomenon that occurs due to the precedence effect and is the minimum time difference at which a listener's perception can detect a deviation between two sounds.

[0351] In Figure 12C Example 2, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the reflected sound that can be detected relative to the direct sound, specifically, the threshold for whether the reflected sound arriving with the delay time can be detected.

[0352] Figure 13 It is a diagram showing an example of the method for setting the threshold. Figure 13 The horizontal axis in Figure 13 corresponds to the arrival time of the reflected sound, specifically, the time difference (T) between the direct sound and the reflected sound. Figure 13 The vertical axis in

[0353] corresponds to the volume of the reflected sound. Specifically, Figure 9 the vertical axis in Figure 13 can either correspond to the volume of the reflected sound set relative to the volume of the direct sound (volume ratio) or correspond to the volume of the reflected sound determined absolutely regardless of the volume of the direct sound. Figure 9 For example, in the case where the listener is relatively far from the obstacle object as shown in Figure 10 the arrival time of the reflected sound is later, and as shown in Figure 9 C of Figure 13 the threshold is set low. As a result, in Figure 10 this case, a reflected sound is generated. On the other hand, in the case where the listener is relatively close to the obstacle object as shown in

[0354] the arrival time of the reflected sound is earlier than in

[0355] Figure 14 the case of

[0356] The time difference (T) can be, for example, any one of the time differences between the times when the direct sound and the reflected sound respectively reach the listening position, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound, and the time difference between the time point when the generation of the direct sound ends and the time point when the reflected sound reaches the listening position. Here, an example based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound will be described.

[0357] Specifically, the selection unit 1302 calculates the difference between the length of the path of the direct sound and the length of the path of the reflected sound based on the position information of the sound source object and the listener and the position information and shape information of the obstacle object. Further, the selection unit 1302 detects the time difference (T) between the time when the direct sound reaches the listener position and the time when the reflected sound reaches the listener position by dividing the difference in lengths by the speed of sound.

[0358] The volume reaching the listener decays in proportion to the distance from the sound source (inversely proportional to the distance) relative to the volume of the sound source. Thus, the volume of the direct sound is obtained by dividing the volume of the sound source by the length of the path of the direct sound. The volume of the reflected sound is obtained by dividing the volume of the sound source by the length of the path of the reflected sound and then multiplying by the attenuation rate given to the virtual obstacle object. The selection unit 1302 detects the volume ratio by calculating the ratio of their volumes.

[0359] In addition, the selection unit 1302 determines a threshold value (S204) corresponding to the time difference (T) using threshold data. Next, the selection unit 1302 determines whether the detected volume ratio (L) is equal to or greater than the threshold value (S205).

[0360] When the volume ratio (L) is equal to or greater than the threshold value (Yes in S205), the selection unit 1302 selects the reflected sound as the reflected sound to be generated (S206). When the volume ratio (L) is less than the threshold value (No in S205), the selection unit 1302 does not select the reflected sound as the reflected sound to be generated (S207). That is, in this case, the selection unit 1302 determines the reflected sound as a reflected sound outside the generation target.

[0361] Then, the selection unit 1302 determines whether there is an unassigned reflected sound (S208). If there is an unassigned reflected sound (Yes in S208), the selection unit 1302 repeats the above processing (S201 to S207). If there is no unassigned reflected sound (No in S208), the selection unit 1302 ends the processing.

[0362] This selection process can be performed for all the reflected sounds generated in the analysis process, or can be performed only for the above-described reflected sounds with high evaluation values.

[0363] (Details of the method for storing the threshold value)

[0364] The threshold data related to this embodiment is stored in the memory 1404 of the sound signal processing device 1001. The form and type of the stored threshold data can be any form and any type. When storing multiple forms and multiple types of thresholds, in the selection process, it is possible to determine which form and which type of threshold to use for the selection process of the reflected sound. The method for determining which threshold data to use for the selection process will be described later.

[0365] In addition, it is also possible to store a combination of multiple forms and multiple types of threshold data. It is also possible to read out the combined threshold data from the spatial information management unit (1201, 1211) and set the threshold used in the selection process. Additionally, it is also possible to store the threshold data stored in the memory 1404 in the spatial information management unit (1201, 1211).

[0366] The threshold data can also be stored, for example, as thresholds at each time difference to depict Figure 12C the lines of the thresholds shown in [Example 1] and [Example 2].

[0367] In addition, the threshold data can also be stored as Figure 11 table data in which the threshold is associated with the time difference (T). That is, the threshold data can also be stored as table data having the time difference (T) as an index. Of course, Figure 11 the threshold shown is an example, and the threshold is not limited to Figure 11 the example. In addition, instead of storing the threshold itself, it is also possible to approximate the threshold with a function having the time difference (T) as a variable and store the coefficients of the function. Additionally, it is also possible to store a combination of multiple approximation formulas.

[0368] In the memory 1404, it is also possible to store information related to the relational expression representing the relationship between the time difference (T) and the threshold. That is, it is also possible to store an expression having the time difference (T) as a variable. The thresholds at each time difference (T) can also be approximated by a straight line or a curve, and the parameters representing the geometric shape of the straight line or the curve can be stored. For example, when the geometric shape is a straight line, it is also possible to store the starting point and the slope used to represent the straight line.

[0369] In addition, it is also possible to set the type and form of the threshold data according to each property of the direct sound and store it. In addition, it is also possible to store parameters used to adjust the threshold according to the property of the direct sound and used for the selection process. The process of adjusting the threshold according to the property of the direct sound and using it for the selection process will be described later as a modified example of the threshold setting method.

[0370] As an example of storing a combination of multiple types of threshold data, it is also possible to store it as Figure 12CAs shown in [Exemplification 3], the value of the larger one between the masking threshold and the echo detection limit threshold is stored for each time difference (T). It is also possible to store, as shown in Figure 12C the [Exemplification 4], the value of the larger one between the minimum volume reproduced in the virtual space and the echo detection limit threshold for each time difference (T).

[0371] The combination of multiple types of threshold data is not limited to this. For example, it is also possible to store the information of the maximum value for each time difference (T) among multiple threshold data.

[0372] In addition, in the above, the information related to the threshold has an item of time as a one-dimensional index. The information related to the threshold can also have a two-dimensional or three-dimensional index that further includes variables related to the arrival direction.

[0373] Figure 15 is a diagram showing the relationship between the direction of the direct sound, the direction of the reflected sound, the time difference, and the threshold. For example, as shown in Figure 15 , it is also possible to store the threshold pre-calculated according to the relationship of the direction of the direct sound (θ), the direction of the reflected sound (γ), the time difference (T), and the volume ratio (L).

[0374] The direction of the direct sound (θ) corresponds to the angle of the arrival direction of the direct sound relative to the listener. The direction of the reflected sound (γ) corresponds to the angle of the arrival direction of the reflected sound relative to the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of the direct sound at the listening position and the arrival time of the reflected sound. The volume ratio (L) corresponds to the volume ratio of the volume of the direct sound at arrival to the volume of the reflected sound at arrival.

[0375] Of course, Figure 15 the threshold shown is an example, and the threshold is not limited to the Figure 15 example. In addition, in Figure 15 , the threshold in the case where the angle (θ) of the arrival direction of the direct sound is 0 degrees is mainly exemplified. However, the threshold in the case where the arrival direction (θ) of the direct sound is other than 0 degrees is also stored in the memory 1404.

[0376] In addition, in the above, the threshold is stored as an arrangement having the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound as independent variables or indices. However, the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound may not be used as independent variables.

[0377] For example, the angular difference (θ) between the arrival direction of the direct sound and the angular difference (γ) between the arrival direction of the reflected sound can also be used. This angular difference corresponds to the angle formed by the arrival direction of the direct sound and the arrival direction of the reflected sound, and can also be expressed as the arrival angles of the direct sound and the reflected sound.

[0378] Figure 16 is a diagram showing the relationship between the angular difference, the time difference, and the threshold value. For example, it is also possible to store, as shown in Figure 16 the example shown, the threshold value pre-calculated using the angular difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound as a variable. Of course, Figure 16 the threshold value shown is an example, and the threshold value is not limited to Figure 16 the example.

[0379] In Figure 16 the example, the number of variables used in the derivation of the threshold value can be reduced. Therefore, the number of threshold values stored in the memory 1404 can be reduced. Thus, the amount of data stored in the memory 1404 can be reduced.

[0380] In addition, when using the angular difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound, the threshold data can also be stored in a two-dimensional arrangement. In addition, in the selection process, a three-dimensional arrangement can also be used to calculate the difference between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound.

[0381] The method of selecting the reflected sound using the threshold value corresponding to the arrival direction will be described later.

[0382] (First Variation Example of Threshold Setting Method)

[0383] In Figure 12A , Figure 12B and Figure 12C the example, multiple forms and multiple types of threshold values can also be stored in the spatial information management unit (1201, 1211). And it is also possible to determine which form and which type of threshold value among the multiple forms and multiple types of threshold values are used for the selection process of the reflected sound. Specifically, as shown in Figure 12C Example 3, the highest threshold value is adopted at the time difference (T) corresponding to the arrival time of the reflected sound.

[0384] In addition, as shown in Example 4, the masking threshold value, the threshold value of the echo detection limit, and the threshold value indicating the minimum volume reproduced in the virtual space can also be stored. And it is also possible to adopt the highest threshold value at the time difference (T) corresponding to the arrival time of the reflected sound.

[0385] (Second Variation Example of Threshold Setting Method)

[0386] As another example of the threshold setting method, a method of setting a threshold according to the nature of the direct sound will be described.

[0387] Figure 17 represents Figure 7 A block diagram showing another configuration example of the rendering unit 1300 shown in Figure 17 The rendering unit 1300 of Figure 7 is different from the rendering unit 1300 of Figure 7 in that it includes a threshold adjustment unit 1304. The description of the parts other than the threshold adjustment unit 1304 is the same as that described in

[0388] Based on the information indicating the nature of the sound signal, the threshold adjustment unit 1304 selects the threshold to be used by the selection unit 1302 from the threshold data. Alternatively, the threshold adjustment unit 1304 may also adjust the threshold included in the threshold data based on the information indicating the nature of the sound signal.

[0389] The information indicating the nature of the sound signal may also be included in the input signal. Also, the threshold adjustment unit 1304 may obtain the information indicating the nature of the sound signal from the input signal. Alternatively, the analysis unit 1301 may analyze the sound signal included in the received input signal, derive the nature of the sound signal, and output the information indicating the nature of the sound signal to the threshold adjustment unit 1304.

[0390] The information indicating the nature of the sound signal may be obtained either before starting the rendering process or at any time during the rendering.

[0391] In addition, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, and other communication devices may have the function of the threshold adjustment unit 1304. In this case, the analysis unit 1301 or the selection unit 1302 may also obtain the information indicating the nature of the sound signal, the threshold data corresponding to the nature, or the information used to adjust the threshold data according to the nature from other communication devices via the communication IF 1403.

[0392] Figure 18 A flowchart showing another example of the selection process. Figure 19 A flowchart showing still another example of the selection process. In Figure 18 and Figure 19 the threshold is set according to the nature of the direct sound. Specifically, in Figure 18 the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the nature of the sound signal. In Figure 19 the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the nature of the sound signal.

[0393] The operations of each example will be described below. In addition, descriptions of processes common to the examples of Figure 14 will be omitted.

[0394] First, an example of the process shown in Figure 18 will be described. Here, threshold data is pre-stored in the memory 1404 for each property of the direct sound. Thus, a plurality of threshold data corresponding to a plurality of properties are pre-stored in the memory 1404. Further, the threshold adjustment unit 1304 determines the threshold data to be used in the selection process of the reflected sound from among the plurality of threshold data.

[0395] For example, the threshold adjustment unit 1304 obtains the property of the direct sound based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the property of the direct sound associated with the input signal. Next, the threshold adjustment unit 1304 determines the threshold corresponding to the time difference (T) and the property of the direct sound (S212).

[0396] In addition, as shown in Figure 19 , the threshold adjustment unit 1304 may also adjust the threshold determined by the selection unit 1302 based on the property of the direct sound (S221).

[0397] In any case, information indicating the property of the sound signal, information for adjusting the threshold according to the property of the sound signal, or both of them may be included in the input signal. The threshold adjustment unit 1304 may also use one or both of them to adjust the threshold.

[0398] In addition, information indicating the property of the sound signal, information for adjusting the threshold, or both of them may also be transmitted through an input signal different from the input signal including the sound signal. In this case, information associated with an input signal different from the input signal may also be included in the input signal including the sound signal, and information associated with an input signal different from the input signal may also be stored in the memory 1404 together with the information related to the threshold.

[0399] In Figure 18 and Figure 19 , the threshold used in the selection of the reflected sound is set according to the property of the direct sound, that is, the property of the sound signal. Either the threshold data set in advance for each property as in Figure 18 may be used, or the threshold may be adjusted according to the property of the sound signal as in Figure 19 . In addition, the parameters of the threshold data may also be adjusted according to the property of the sound signal.

[0400] In addition, the operation performed by the threshold adjustment unit 1304 can also be performed by the analysis unit 1301 or the selection unit 1302. For example, it can also be that the analysis unit 1301 obtains the properties of the sound signal. In addition, it can also be that the selection unit 1302 sets a threshold according to the properties of the sound signal.

[0401] Next, the relationship between the properties of the sound signal and the threshold will be described.

[0402] If the time interval between two short sounds continuously reaching the listener's ear is sufficiently short, they are heard as one sound. This phenomenon is called the precedence effect. It is known that the precedence effect occurs only for discontinuous, i.e., transient, sounds (Non-Patent Document 1). Therefore, when the sound signal represents a steady sound, the echo detection limit can be set lower compared to the case where the sound signal represents a non-steady sound.

[0403] That is, according to the characteristics of such a precedence effect, for example, when the direct sound is a steady sound, the threshold is set smaller. In addition, the higher the steadiness, the smaller the threshold can be set.

[0404] An example of the processing when the properties of the sound signal are steady will be described. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the steadiness based on the amount of change in the frequency components of the sound signal over time. For example, when the amount of change is small, it is determined that the steadiness is high. On the contrary, when the amount of change is large, it is determined that the steadiness is low. As a result of the determination, a flag indicating the level of steadiness can be set, or a parameter indicating the steadiness can be set according to the amount of change.

[0405] Next, the threshold adjustment unit 1304 can also adjust the threshold data or the threshold based on the information indicating the steadiness of the sound signal, such as a flag or a parameter indicating the steadiness of the sound signal, and set the adjusted threshold data or threshold as the threshold data or threshold used in the selection unit 1302.

[0406] Alternatively, parameters for setting threshold data according to the information indicating the steadiness of the direct sound can be stored in the memory 1404 in advance. In this case, the threshold adjustment unit 1304 can also determine the steadiness of the sound signal and set the threshold data used in the selection of the reflected sound based on the information indicating the steadiness and the parameters.

[0407] Alternatively, multiple parameters of the threshold data can be stored in the memory 1404 in correspondence with multiple patterns of the steadiness of the direct sound. In this case, the threshold adjustment unit 1304 can also determine the steadiness of the sound signal, select the parameters of the threshold data based on the pattern of the steadiness of the direct sound, and set the threshold data used in the selection of the reflected sound based on the parameters of the threshold data.

[0408] In addition, regarding the smoothness of the sound signal, it can also be determined based on the amount of change in the frequency components of the sound signal every time the sound signal is input.

[0409] Alternatively, regarding the smoothness of the sound signal, it can also be determined based on the information indicating smoothness that has been previously associated with the sound signal. That is, the information indicating the smoothness of the sound signal can be associated with the sound signal and stored in the memory 1404 in advance. The analysis unit 1301 can also obtain the information indicating smoothness associated with the sound signal every time the sound signal is input. And the threshold adjustment unit 1304 can also adjust the threshold based on the information indicating smoothness associated with the sound signal.

[0410] As another example of setting the threshold according to the nature of the sound signal, when the sound signal represents a short sound (such as a click sound), the application range of the echo detection limit can be set shorter compared to the case where the sound signal represents a long sound. This processing is based on the characteristics of the precedence effect.

[0411] It is known that through the precedence effect, two short sounds that continuously reach the listener's ear are heard as one sound if their time interval is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is about 5 ms for a click sound, and sometimes 40 ms for a complex sound such as a human voice or music (Non-Patent Document 1).

[0412] According to the characteristics of such a precedence effect, for example, in the case of a sound with a short duration of the direct sound, a threshold with a shorter time length is set. In addition, the shorter the duration of the direct sound, the shorter the time length of the threshold set.

[0413] Setting a threshold with a shorter time length means setting a threshold corresponding to the echo detection limit based on the characteristics of the precedence effect in a range where the time difference (T) between the direct sound and the reflected sound is small. Outside this range, a threshold corresponding to the echo detection limit based on the characteristics of the precedence effect is not set. That is, outside this range, the threshold is small. Therefore, setting a threshold with a shorter time length for a short sound can correspond to setting a smaller threshold for a short sound.

[0414] As another example of setting the threshold according to the nature of the direct sound, when the direct sound is an intermittent sound (such as speech), the threshold can be set lower compared to the case where the direct sound is a continuous sound (such as music).

[0415] For example, in the case where the direct sound corresponds to speech, there are repeatedly voiced and unvoiced parts. As a masking effect, only post-masking occurs in the unvoiced part. On the other hand, in the case where the direct sound is a continuous sound such as music content, two masking effects occur: post-masking and simultaneous masking based on the sound generated at this time. Therefore, the comprehensive masking effect is higher in the case of music or the like than in the case of speech or the like.

[0416] According to the characteristics of the masking effect as described above, the threshold can also be set higher in the case of music or the like compared to the case of speech or the like. Conversely, the threshold can also be set lower in the case of speech or the like compared to the case of music or the like. That is, the threshold can also be set smaller in the case where there are many intermittent parts in the direct sound.

[0417] As described above, the information indicating the nature of the direct sound can also be information indicating the smoothness, intermittency, and duration of the direct sound, etc. In addition, the information indicating the nature of the direct sound can also be any combination of them. In addition, the information indicating the nature of the direct sound can be information indicating the time variation of any one of them, or information indicating the time variation of any combination of them. That is, the information indicating the nature of the direct sound can also be information indicating the time variation of the direct sound.

[0418] For example, as shown in the description of the determination of smoothness, the information indicating the nature of the direct sound can also be time-series data of frequency characteristics. Here, the frequency characteristics can also be expressed in conventional forms such as the gain value for each frequency band, the Fourier series for the time-axis signal, or the LPC coefficients or cepstrum coefficients used to obtain the frequency envelope.

[0419] Furthermore, the information indicating the nature of the direct sound can also be information that lists, in time series, multiple groups of the duration of the signal with stable amplitude and the amplitude value of the signal during that period as information indicating the intermittency of the direct sound (approximate shape of the amplitude envelope). Here, the amplitude value can also be expressed as a ratio to the reference volume.

[0420] And the information indicating the nature of the direct sound can also be information related to the frequency characteristics of the direct sound. For example, the information indicating the nature of the direct sound can also be information indicating the smoothness of the frequency characteristics of the direct sound. Specifically, the information indicating the nature of the direct sound can be information that lists, in time series, the duration of the state with small changes in frequency characteristics and multiple groups of the frequency characteristics of the signal during that period (approximate shape of the spectrogram). Here, the volume used as the reference for the above frequency characteristics can also be the above reference volume.

[0421] For example, the information indicating the time variation of the direct sound is information indicating the envelope of the direct sound. The information indicating the time variation of the direct sound can also be inFigure 12C It is used when the "minimum audible limit" described in [Example 4] is the threshold. The signal compared with the minimum audible limit is the volume of the reflected sound.

[0422] The volume of the reflected sound is obtained by geometric calculation based on the information of the sound source, the listener, and the position of the reflecting object. Specifically, the reference volume of the reflected sound relative to the reference volume of the sound source is obtained. By using the information on the change in the magnitude of the sound of the sound source as the information representing the nature of the direct sound to increase or decrease the reference volume of the reflected sound, the volume of the reflected sound at any given moment can be accurately obtained. The reason is that the change in the volume of the sound source is reflected in the change in the volume of the reflected sound.

[0423] After adjusting the volume of the reflected sound, by comparing the volume of the reflected sound with the threshold, it is possible to more accurately and appropriately select the reflected sound required aurally.

[0424] Of course, without adjusting the reference volume of the reflected sound, but based on the reciprocal of the information on the change in the magnitude of the sound of the sound source to adjust the threshold, and comparing the adjusted threshold with the reference volume of the reflected sound, the same result can also be obtained. That is, the information on the change in the magnitude of the sound of the sound source can be used to adjust the reference volume of the reflected sound, or the information on the change in the magnitude of the sound of the sound source can be used to adjust the threshold. The adjustment of the reference volume of the reflected sound and the adjustment of the threshold correspond to each other.

[0425] Depending on the composition of the surface of the object that reflects the sound, the sound reflectivity (the attenuation rate of the sound accompanying reflection) is different for each frequency band. Therefore, as described later, it is also possible to associate the sound reflectivity (attenuation rate) for each frequency band with the object that reflects the sound. Based on such information on the reflectivity and the information on the spectrogram, it is possible to more accurately determine whether to select the reflected sound. For example, the following processing is performed.

[0426] Specifically, for example, according to the information on the spectrogram, it is indicated that the frequency components in the high-frequency band are more dominant than the frequency components in the low-frequency band in a certain time interval. Additionally, for example, according to the information on the sound reflectivity, it is indicated that among the frequency components in the high-frequency band, the reflectivity is extremely small compared to the frequency components in the low-frequency band.

[0427] In this case, even if the amplitude of the signal of the sound source is large on the time axis, the volume of the reflected sound obtained by multiplying the frequency components represented by the information on the spectrogram by the attenuation rate of each frequency band represented by the information on the reflectivity becomes small, and the reflected sound may not be selected.

[0428] As described above, the information representing the nature of the direct sound can also be the information representing the temporal change of the direct sound. For example, the information representing the nature of the direct sound can also represent the value obtained by analyzing the direct sound with a preset time length.

[0429] Specifically, the information indicating the property of the direct sound may also be information obtained by calculating the average energy or average amplitude of the direct sound for each preset time length. Additionally, the information indicating the property of the direct sound may also be information obtained by calculating the weighted average of the energy or average amplitude of the direct sound for each short-time analysis length and calculating the energy or average amplitude for each long-time analysis length longer than the short-time analysis length.

[0430] More specifically, for example, the information indicating the temporal variation of the direct sound may also be information obtained by calculating the energy or average amplitude of the direct sound for each preset short-time length (e.g., 5 ms, and hereinafter the frames of this time length will be expressed as analysis frames). Additionally, the information indicating the temporal variation of the direct sound may also be information represented by the weighted average of the energy or average amplitude calculated in the past N - 1 analysis frames.

[0431] Assuming that the energy of the n-th analysis frame is represented by E(n), the information I(n) indicating the property of the direct sound is obtained according to the following formula.

[0432] [Mathematical formula 1]

[0433]

[0434] Here, the parameter a(i) represents the weight coefficient. Usually, a(i) is set such that a(i) ≥ 0 and the sum of a(i) is 1. However, the setting method of a(i) is not limited to this.

[0435] Furthermore, every time the direct sound is taken in for 5 ms, the information I(n) indicating the property of the direct sound is calculated. That is, the temporal variation of the information I(n) indicating the property of the direct sound can be calculated with low latency. Therefore, this method is suitable for applications that require real-time performance.

[0436] Additionally, the information I(n) indicating the property of the direct sound can also be obtained according to the following formula.

[0437] [Mathematical formula 2]

[0438]

[0439] Here, the parameter b(i) represents the weight coefficient. Usually, b(i) is set such that b(i) ≥ 0 and the sum of b(i) is 1. However, the setting method of b(i) is not limited to this.

[0440] In this formula, the information I(n) indicating the property of the direct sound is obtained recursively. Therefore, the average energy of a long time length can be calculated with a small amount of computation.

[0441] The above-mentioned Formula 1 and Formula 2 can be regarded as filters with E(n) as the input signal and I(n) as the output signal. In this case, Formula 1 is a filter of a moving average (MA) model, and Formula 2 is a filter of an autoregressive (AR) model, both having the characteristics of a low-pass filter. In addition, a filter of an ARMA model formed by combining the two can also be used.

[0442] In addition, the method for deriving the information representing the temporal variation of the direct sound is not limited to the above-mentioned calculation formulas or filters, and other well-known methods can also be used. As described above, the information representing the temporal variation of the direct sound represents a value obtained by analyzing the direct sound with a preset time length. The direct sound can also be analyzed from viewpoints other than the average energy.

[0443] In addition, as described above, the information representing the property of the direct sound can also be information related to the frequency characteristics of the direct sound. The information related to the frequency characteristics of the direct sound can also be information calculated using the frequency characteristics of the direct sound. For example, the information related to the frequency characteristics of the direct sound can be information obtained by averaging the low-frequency components of the direct sound with a preset analysis length as the average energy of the low-frequency components.

[0444] Specifically, by applying a filter with a low-pass characteristic to the direct sound included in the analysis frame length, the low-frequency components of the direct sound are obtained. Based on the energy or average amplitude of the low-frequency components, information representing the property of the direct sound is derived in the same manner as Formula 1 above.

[0445] Assuming that the energy of the low-frequency components of the nth analysis frame is represented by E L (n), the information I(n) representing the property of the direct sound is obtained according to the following formula.

[0446] [Mathematical Formula 3]

[0447]

[0448] Here, the parameter c(i) represents a weighting coefficient. Usually, c(i) is set such that c(i) ≥ 0 and the sum of c(i) is 1. However, the setting method of c(i) is not limited to this.

[0449] In addition, every time the direct sound is taken in for 5 ms, the information I(n) representing the property of the direct sound is calculated. That is, the temporal variation of the information I(n) representing the property of the direct sound can be calculated with low latency. Therefore, this method is suitable for applications that require real-time performance.

[0450] In addition, similar to Formula 2, the information I(n) representing the property of the direct sound can also be obtained according to the following formula.

[0451] [Formula 4]

[0452]

[0453] Here, the parameter d(i) represents a weighting coefficient. Usually, d(i) is set such that d(i) ≥ 0 and the sum of d(i) is 1. However, the method of setting d(i) is not limited to this.

[0454] In this formula, the information I(n) representing the properties of the direct sound is recursively obtained. Therefore, the average energy of a long time length can be calculated with a small amount of computation.

[0455] The above formulas 3 and 4 can be regarded as filters with E(n) as the input signal and I(n) as the output signal. In this case, formula 3 is a filter of a moving average (MA) model, and formula 4 is a filter of an autoregressive (AR) model, both having the characteristics of a low-pass filter. In addition, a filter of an ARMA model formed by combining the two can also be used.

[0456] In the above, a filter with low-pass characteristics is used in the method of obtaining the low-frequency component of the direct sound, but the method of obtaining the low-frequency component of the direct sound is not limited to this. In addition, the method of deriving the information representing the temporal variation of the direct sound is not limited to the above calculation formulas or filters, and other well-known methods can also be used. For example, the spectrum of the direct sound can be calculated by performing a frequency transformation on the direct sound. And the energy or average amplitude of the low-frequency component of the spectrum can be calculated.

[0457] In addition, in the above, an MA model or an AR model is used in the derivation of the information representing the temporal variation of the direct sound. The coefficients of these models can be preset fixed values or time-variable values, that is, variable values.

[0458] In addition, the relationship between the above analysis frame length and the generation interval of the above information update thread can also be the following relationship.

[0459] For example, when the time length of the analysis frame is TA (msec) and the generation interval of the information update thread is TU (msec), the value of N in the above (formula 1) and (formula 3) in the MA filter can also be around the value given by TU / TA. In addition, b(i) and d(i) (1 ≤ i < N) in the above (formula 2) and (formula 4) in the AR filter can also be values such that the time constant of this filter becomes around TU (msec).

[0460] The reason for the above setting is that it is expected that this filter converges during the information update interval.

[0461] On the other hand, in the above settings, when the value in the information indicating the time variation of the direct sound changes too abruptly, I(n) can also be calculated in advance. Moreover, the pre-calculated I(n) can also be applied to the selection process of the reflected sound. For example, in the processing of the frame at time t, I(t + tau) can also be used. Here, tau is a value determined according to the convergence characteristics of the filter. In the case of slow convergence, the value of tau is larger than that in the case of fast convergence.

[0462] In addition, as information indicating the characteristics of the direct sound, information on auditory masking (frequency masking) calculated based on the direct sound can also be used. The information on auditory masking represents the threshold of the amplitude value in the frequency domain masked by the direct sound. It is also possible to perform a process of comparing the amplitude value of the reflected sound in the same frequency domain with the threshold and not selecting the reflected sound whose amplitude value is smaller than the threshold. The amplitude value of the reflected sound in the frequency domain can also be obtained by the analysis unit 1301 as information indicating the characteristics of the reflected sound.

[0463] By setting the threshold used in the selection of the reflected sound according to the properties of the direct sound in this way, it is possible to appropriately select the reflected sound required in terms of sound perception and effectively reflect the auditory characteristics into the stereophonic reproduction system 1000. The process of detecting the properties of the direct sound, the process of determining the threshold according to the properties, and the process of adjusting the threshold according to the properties can be performed either during the rendering process or before the start of the rendering process.

[0464] For example, these processes can be performed at the time of virtual space production (during software production), at the start of virtual space processing (when the software is started or when rendering starts), or at the timing of an information update thread that occurs regularly during virtual space processing. In addition, the time of virtual space production can be either the timing of constructing the virtual space before the start of audio processing, the timing of obtaining the information of the virtual space (spatial information), or the timing of obtaining the software.

[0465] Here, in the information update thread, a process for updating the spatial information managed by the spatial information management unit (1201, 1211) is executed.

[0466] The role played by the information update thread is, for example, a process of updating the position and orientation of the listener's avatar arranged in the virtual space based on the position and orientation of the VR goggles worn by the listener, or an update of the position of an object moving in the virtual space. Such a process is provided in a processing thread that starts at a relatively low frequency of about several tens of Hz.

[0467] In such a processing thread with a low occurrence frequency, it is also possible to perform processing for updating information representing the nature of the direct sound. The reason is that the frequency of change in the nature of the direct sound is lower than the occurrence frequency of the audio processing frames used for audio output. Thus, the computational load of this processing can be relatively reduced. Additionally, if information is updated at an unnecessarily high frequency, there is a risk of generating impulse noise. By updating information at a low frequency, such a risk can also be avoided.

[0468] (Third Variant Example of Threshold Setting Method)

[0469] As another example of the threshold setting method, the threshold can also be set according to the computing resources (such as CPU capabilities, memory resources, PC performance, or remaining battery level) for processing the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources, and sets the threshold high when the amount of computing resources is small. Thus, the volume of more reflected sounds is smaller than the threshold, so the reflected sounds to be binaurally processed can be reduced, and the computational amount can be reduced.

[0470] Or, when the signal processing is performed by a device driven by a battery such as a smartphone or VR goggles, it is desirable to prioritize making the processing last for a long time and save computing resources. In such a case, the threshold can also be set high without detecting the amount or remaining amount of computing resources.

[0471] (Fourth Variant Example of Threshold Setting Method)

[0472] As another example of the threshold setting method, the threshold can also be set by a virtual space administrator or listener through a threshold setting unit (not shown) provided in the sound signal processing device 1001 or the sound notification device 1002.

[0473] For example, it can be such that the listener wearing the sound notification device 1002 can select either an "energy-saving mode" with less target reflected sound and less computational amount or a "high-performance mode" with more target reflected sound and more computational amount. Or, it can be such that the administrator of the stereophonic reproduction system 1000 or the producer of stereophonic content can select the mode. In addition, instead of a mode, the threshold or threshold data can also be directly selected.

[0474] (First Variant Example of the Operation of the Rendering Unit)

[0475] Figure 20 It is a flowchart showing the first variant example of the operation of the sound signal processing device 1001. In Figure 20 it shows the processing mainly executed by the rendering unit 1300 of the sound signal processing device 1001. In this variant example, volume compensation processing is added to the operation of the rendering unit 1300.

[0476] For example, the analysis unit 1301 acquires data (input signal) (S301). Next, the analysis unit 1301 analyzes the data (S302). Next, the selection unit 1302 determines whether to select the reflected sound based on the analysis result (S303). Next, the synthesis unit 1303 performs a volume compensation process based on the unselected reflected sound (S304). Next, the synthesis unit 1303 performs acoustic processing on the direct sound and the reflected sound (S305). And, the synthesis unit 1303 outputs the direct sound and the reflected sound as audio (S306).

[0477] In the above processing (S301 to S306), the processing other than the volume compensation process (S304) is the same as that in the other examples above, so their descriptions are omitted.

[0478] The volume compensation process is executed corresponding to the unselected reflected sound in the selection process. For example, if the reflected sound is not selected in the selection process, a lack of volume feeling occurs. Through the volume compensation process, the sense of discomfort accompanying such a lack of volume feeling can be suppressed. As examples of methods for compensating the volume feeling, the following two methods are disclosed. Either method can be used.

[0479] First, a method of compensating the volume feeling by increasing the volume of the direct sound will be described. The synthesis unit 1303 increases the volume of the direct sound by an amount corresponding to the volume of the unselected reflected sound to generate the direct sound. Thereby, the volume feeling lost due to the non-generation of the reflected sound is compensated.

[0480] When the synthesis unit 1303 increases the volume, it may increase the volume for each frequency component according to the frequency characteristics of the reflected sound. In order to perform such processing, an attenuation rate of attenuating the volume of the reflection target may be given for each prescribed frequency band. Thereby, the frequency characteristics of the reflected sound can be derived.

[0481] Next, a method of compensating the volume feeling by synthesizing the reflected sound into the direct sound will be described. In this method, the synthesis unit 1303 adds the unselected reflected sound to the direct sound to generate the direct sound, thereby compensating for the volume feeling caused by the non-generation of the reflected sound. The volume (amplitude), frequency, delay, etc. of the unselected reflected sound are reflected in the generated direct sound.

[0482] In the case of the method of increasing the volume of the direct sound, the computational amount of the compensation process is very slight, but only the volume is compensated. In the case of the method of synthesizing the reflected sound into the direct sound, the computational amount of the compensation process is larger than that of the method of increasing the volume of the direct sound, but the characteristics of the reflected sound are compensated more correctly.

[0483] In any case, no reflected sound is generated and only direct sound is generated, so the overall computational load is reduced. In particular, the computational load required for binaural processing including convolution HRTF processing is reduced, so the overall computational load is significantly reduced. The reason is that the computational load required for binaural processing is much larger than that required for the above compensation processing.

[0484] In addition, when the reason for not selecting the reflected sound is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume is not lost, the compensation processing may not be performed and only the reflected sound may be removed.

[0485] (Second modified example of the operation of the rendering unit)

[0486] Figure 21 is a flowchart showing a second modified example of the operation of the sound signal processing apparatus 1001. In Figure 21 mainly shows the processing performed by the rendering unit 1300 of the sound signal processing apparatus 1001. In this modified example, a left-right volume difference adjustment process is added to the operation of the rendering unit 1300.

[0487] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the arrival direction of the sound (S402). Next, the selection unit 1302 adjusts the volume difference of the sound perceived by the left and right ears (S403). In addition, the selection unit 1302 adjusts the arrival time difference (delay) of the sound perceived by the left and right ears (S404). The selection unit 1302 determines whether to select the reflected sound based on the adjusted sound information (S405).

[0488] In the above processing (S401 to S405), the processing other than the left-right volume difference adjustment (S403) and the delay adjustment (S404) is the same as that in the above other examples, so their descriptions are omitted.

[0489] Figure 22 is a diagram showing a configuration example of an avatar, a sound source object, and an obstacle object. For example, when the front direction of the listener is 0 degrees, as Figure 22 shown, if the polarities (e.g., positive and negative) of the arrival direction (θ) of the direct sound and the arrival direction (γ) of the reflected sound are different, the volume difference generated between the two ears is corrected.

[0490] Specifically, when the polarities of θ and γ are different, the ears that mainly (first) perceive the sound in the direct sound and the reflected sound are different. In this case, in the selection unit 1302, as the left-right volume difference adjustment (S403), the volume of the direct sound is adjusted to match the position of the ear that mainly perceives the reflected sound. For example, the selection unit 1302 attenuates the volume of the direct sound when it reaches the listener by multiplying the volume of the direct sound when it reaches the listener by (1.0 - 0.3sin(θ)) (0 ≤ θ ≤ 180).

[0491] The selection unit 1302 determines whether to select the reflected sound by calculating the volume ratio of the volume of the direct sound corrected as described above to the volume of the reflected sound and comparing the calculated volume ratio with a threshold value. Thereby, the volume difference generated between the two ears is corrected, the volume of the direct sound affecting the reflected sound is derived more accurately, and the determination of whether to select the reflected sound is performed more accurately.

[0492] In addition, in addition to the left-right volume difference adjustment (S403), the selection unit 1302 can also, as the delay adjustment (S404), delay the arrival time of the direct sound to match the position of the ear that perceives the reflected sound. Specifically, the selection unit 1302 can also delay the arrival time of the direct sound by adding (a(sinθ + θ) / c) ms (a is the radius of the head, c is the speed of sound) to the arrival time of the direct sound.

[0493] (The third modification example of the operation of the rendering unit)

[0494] A method for setting a threshold corresponding to the arrival direction will be described.

[0495] Figure 23 is a flowchart showing still another example of the selection process. Regarding the processing common to the example of Figure 14 will be omitted. In the example of Figure 23 the selection unit 1302 uses a threshold corresponding to the arrival direction to select the reflected sound.

[0496] Specifically, the selection unit 1302 calculates the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) determined using the orientation of the avatar as a reference based on the direct sound arrival path (pd), the reflected sound arrival path (pr), and the orientation information D of the avatar calculated by the analysis unit 1301. That is, the selection unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. The orientation information D of the avatar may also be included in the input signal.

[0497] The selection unit 1302 uses three indexes including the time difference (T) in addition to the direct sound arrival direction (θ) and the reflected sound arrival direction (γ), according to Figure 15For the three-dimensional arrangement shown, determine the threshold value used in the selection process (S232).

[0498] As an example, the method for setting the threshold value used in the selection process in the case where an avatar, a sound source object, and an obstacle object are arranged as shown in Figure 22 will be described.

[0499] Obtain the position information of the avatar, the sound source object, and the obstacle object and the orientation information D of the avatar from the input signal. Using this position information and the orientation information D, calculate the direction (θ) of the direct sound and the direction (γ) of the sound image of the reflected sound when the orientation of the avatar is set to 0 degrees. In the case of Figure 22 , the direction (θ) of the direct sound is about 20 degrees, and the direction (γ) of the sound image of the reflected sound is about 265 degrees (-95 degrees).

[0500] Next, referring to the threshold value data stored in a three-dimensional arrangement as shown in Figure 15 , determine the threshold value from the arrangement area corresponding to the values of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. In the case where there is no index corresponding to the calculated values of (θ), (γ), and (T), it is also possible to determine the threshold value corresponding to the closest index.

[0501] As another method, it is also possible to determine the threshold value by performing processing such as interpolation, interpolation, or extrapolation based on one or more threshold values corresponding to one or more indices close to the calculated values of (θ), (γ), and (T). For example, it is also possible to determine the threshold value corresponding to (20 degrees, 265 degrees, T) based on the four threshold values corresponding to the four indices of (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).

[0502] The selection process based on the difference between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound will be described.

[0503] For example, it is also possible to pre-create and set threshold value data having the angle difference (Φ) between the arrival direction (θ) of the direct sound and the arrival direction (γ) of the reflected sound and the time difference (T) arranged as two-dimensional indices as shown in Figure 16 . In this case, refer to the angle difference (Φ) and the time difference (T) in the selection process. Alternatively, it is also possible to calculate the angle difference (Φ) between the angle (θ) of the arrival direction of the direct sound and the angle (γ) of the arrival direction of the reflected sound in the selection process and use the calculated angle difference (Φ) for determining the threshold value.

[0504] Alternatively, it is also possible to set threshold data that has a combination of the angular difference (Φ), the arrival direction (θ) of the direct sound, and the time difference (T), or a combination of the angular difference (Φ), the arrival direction (γ) of the reflected sound, and the time difference (T) as an index arrangement.

[0505] Alternatively, it is also possible to set Figure 15 threshold data that has the values of (θ), (γ), and (T) as a three-dimensional index arrangement as shown.

[0506] (Fourth Modification Example of the Operation of the Rendering Unit)

[0507] The processing performed by the above-described analysis unit 1301, selection unit 1302, and synthesis unit 1303 can also be performed as, for example, a pipeline process described in Patent Document 3.

[0508] Figure 24 is a block diagram showing a configuration example for the rendering unit 1300 to perform a pipeline process.

[0509] Figure 24 The rendering unit 1300 of includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components can also be composed of Figure 7 multiple components of the rendering unit 1300 shown in, or can be composed of at least a part of multiple components of the sound signal processing device 1001 shown in Figure 5 .

[0510] The pipeline process refers to dividing the process for imparting sound effects into multiple processes and sequentially executing the multiple processes one by one. In each of the multiple processes, for example, signal processing of a sound signal or generation of parameters used in the signal processing is performed.

[0511] The rendering unit 1300 can also perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, etc. as a pipeline process. However, these processes are just examples, and the pipeline process can also include processes other than these, or can exclude some of the processes. For example, the pipeline process can also include diffraction processing and occlusion processing. In addition, for example, reverberation processing can be omitted when not needed.

[0512] In addition, each process can also be represented as a stage. In addition, the results of each process, sound signals such as the generated reflected sound, etc. can also be represented as rendering items. The multiple stages and their order in the pipeline process are not limited to Figure 24 the example shown.

[0513] Here, the parameters used in the selection process (the arrival path, arrival time, and volume ratio related to the direct sound and the reflected sound) can also be calculated in one of the multiple stages used to generate the rendering item. That is, the parameters used in the selection of the reflected sound are calculated in a part of the pipeline process for generating the rendering item. In addition, it is not necessary for the rendering unit 1300 to perform all stages. For example, some stages can be omitted, or they can be performed by entities other than the rendering unit 1300.

[0514] The reverberation processing, early reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that can be included as stages in the pipeline process will be described. In each stage, it is also possible to analyze the metadata included in the input signal and calculate the parameters used in the generation of the reflected sound.

[0515] In the reverberation processing, the reverberation processing unit 1311 generates a sound signal representing the reverberation sound or the parameters used in the generation of the sound signal. The reverberation sound is the sound that reaches the listener as reverberation after the direct sound. As an example, the reverberation sound is the sound that reaches the listener after the early reflection sound described later, at a relatively late stage (for example, around 100 to 100+ ms from the arrival time of the direct sound), after being reflected more times (for example, dozens of times) than the early reflection sound.

[0516] The reverberation processing unit 1311 calculates the reverberation sound by referring to the sound signal and the spatial information included in the input signal and using a prescribed function prepared in advance as a function for generating the reverberation sound.

[0517] The reverberation processing unit 1311 can also apply a known reverberation generation method to the sound signal included in the input signal to generate the reverberation sound. An example of a known reverberation generation method is the Schroeder method, but the known reverberation generation method is not limited to the Schroeder method. In addition, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by the spatial information in the application of the known reverberation generation method. Thereby, the reverberation processing unit 1311 can calculate the parameters used to generate the reverberation sound.

[0518] In the early reflection processing, the early reflection processing unit 1312 calculates the parameters used to generate the early reflection sound based on the spatial information. The early reflection sound is the reflected sound that reaches the listener after being reflected one or more times at a relatively early stage (for example, around several tens of ms from the arrival time of the direct sound) after the direct sound reaches the listener from the sound source object.

[0519] The early reflection processing unit 1312, for example, refers to the sound signal and the metadata and calculates the path of the reflected sound that reaches the listener after being reflected by the reflection object from the sound source object. For example, in the calculation of the path, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflection object such as a structure, and the reflectivity of the reflection object can also be used.

[0520] In addition, the early reflection processing unit 1312 may also calculate the path of the direct sound. The information on this path may also be used as a parameter for the early reflection processing unit 1312 to generate early reflection sounds, and may also be used as a parameter for the selection unit 1314 to select reflection sounds.

[0521] In the distance attenuation processing, the distance attenuation processing unit 1313 calculates the volumes of the direct sound and the reflection sound reaching the listener based on the lengths of the paths of the direct sound and the reflection sound. The volumes of the direct sound and the reflection sound reaching the listener are attenuated in proportion to the distance of the path to the listener (inversely proportional to the distance) with respect to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct sound by dividing the volume of the sound source by the length of the path of the direct sound, and can calculate the volume of the reflection sound by dividing the volume of the sound source by the length of the path of the reflection sound.

[0522] In the selection processing, the selection unit 1314 selects the target reflection sound to be generated based on the parameters calculated before the selection processing. In the selection of the target reflection sound to be generated, a certain selection method of the present disclosure may also be used.

[0523] The selection processing may be performed on all the reflection sounds, or may be performed only on the reflection sounds with high evaluation values based on the evaluation processing as described above. That is, for the reflection sounds with low evaluation values, it is determined not to be selected without performing the selection processing. For example, for the reflection sounds with very small volumes, the evaluation values of the reflection sounds may be regarded as low and it is determined not to be selected.

[0524] In addition, for example, the selection processing may be performed on all the reflection sounds. Also, the evaluation value of the reflection sound selected in the selection processing may be determined, and the reflection sound with a low evaluation value determined may be re-determined not to be selected.

[0525] The selection processing and the evaluation processing may be performed independently of each other, or may be performed in combination. In the case where the selection processing and the evaluation processing are performed in combination, either one of the two processes may be performed first.

[0526] In the generation processing, the generation unit 1315 generates the direct sound and the reflection sound. For example, the generation unit 1315 generates the direct sound based on the arrival time and the volume at arrival of the direct sound according to the sound signal included in the input signal. In addition, for the reflection sound selected in the selection processing, the generation unit 1315 generates the reflection sound based on the arrival time and the volume at arrival of the reflection sound according to the sound signal included in the input signal.

[0527] In binaural processing, the binaural processing unit 1316 performs signal processing so that the sound signal of the direct sound is perceived as a sound arriving at the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing so that the reflected sound selected by the selection unit 1314 is perceived as a sound arriving at the listener from the reflection object.

[0528] For example, based on the position and orientation of the listener in the sound space, the binaural processing unit 1316 performs processing of applying the HRIR DB so that the sound arrives at the listener from the position of the sound source object or the position of the obstacle object.

[0529] In addition, HRIR (Head-Related Impulse Responses) is the response characteristic when a pulse is generated. Specifically, HRIR is a response characteristic obtained by transforming the head-related transfer function from the representation in the frequency domain to the representation in the time domain by Fourier transform. The head-related transfer function represents the change in sound generated by peripheral objects including the auricle, human head, and shoulders as a transfer function. The HRIR DB is a database containing such information.

[0530] Furthermore, the position and orientation of the listener in the sound space are, for example, the position and orientation of the virtual listener in the virtual sound space. It is also possible that the position and orientation of the virtual listener in the virtual sound space change in accordance with the movement of the listener's head. In addition, the position and orientation of the virtual listener in the virtual sound space can also be determined based on the information obtained from the sensor 1405.

[0531] The program, spatial information, HRIR DB, threshold data, or other parameters used in the above processing are obtained from the memory 1404 included in the sound signal processing device 1001 or from outside the sound signal processing device 1001.

[0532] In addition, the pipeline processing may also include other processing. And the rendering unit 1300 may also include an unillustrated processing unit for performing other processing included in the pipeline processing. For example, the rendering unit 1300 may include a diffraction processing unit and an occlusion processing unit.

[0533] The diffraction processing unit performs processing of generating a sound signal representing a sound including a diffracted sound caused by an obstacle object between the listener and the sound source object in a three-dimensional sound field (space). The diffracted sound is a sound that bypasses the obstacle object and arrives at the listener from the sound source object when there is an obstacle object between the sound source object and the listener.

[0534] The diffraction processing unit calculates, for example with reference to the sound signal and the metadata, the path of the diffracted sound that reaches the listener from the sound source object bypassing the obstacle object, and generates the diffracted sound based on this path. In the calculation of the path, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object, etc. can also be used.

[0535] When there is a sound source object on the opposite side of the obstacle object, the occlusion processing unit generates a sound signal of the sound that leaks out from the sound source object through the obstacle object based on the spatial information and information such as the material of the obstacle object.

[0536] (Examples of sound source objects)

[0537] In the above, the position information given to the sound source object represents the "point" in the virtual space as the position of the sound source object. That is, in the above, the sound source is defined as a "point sound source".

[0538] On the other hand, the sound source in the virtual space can also be defined as an object having a length, size, shape, etc., that is, a spatially extended sound source defined as a non-point sound source. In this case, the distance between the listener and the sound source and the arrival direction of the sound are uncertain. Thus, regarding the reflected sound caused by such a sound source, it is not parsed by the parsing unit 1301, or regardless of the parsing result, it is limited to being selected by the selection unit 1302. Thereby, it is possible to avoid deterioration of the sound quality that may occur due to non-selection of the reflected sound.

[0539] Alternatively, a representative point such as the center of gravity of the object can be determined, and it is assumed that sound is generated from this representative point and the processing of the present disclosure is applied. In this case, the threshold can also be adjusted according to the spatial extension information of the sound source.

[0540] (Examples of direct sound and reflected sound)

[0541] For example, the direct sound is the sound that is not reflected by the reflection object, and the reflected sound is the sound that is reflected by the reflection object. The direct sound can also be the sound that reaches the listener without being reflected by the reflection object from the sound source, and the reflected sound can also be the sound that reaches the listener after being reflected by the reflection object from the sound source.

[0542] In addition, the direct sound and the reflected sound are not limited to the sound that reaches the listener respectively, and can also be the sound before reaching the listener. For example, the direct sound can also be the sound output from the sound source, in other words, the sound of the sound source.

[0543] Figure 25 It is a diagram showing the transmission and diffraction of sound. As Figure 25As shown, sometimes, due to the presence of an obstacle object between the sound source object and the listener, the direct sound does not reach the listener. In this case, the sound that is emitted from the sound source object, penetrates the obstacle object, and reaches the listener can also be regarded as the direct sound. Also, the sound that is emitted from the sound source object, diffracts through the obstacle object, and reaches the listener can also be regarded as the reflected sound.

[0544] In addition, the two sounds compared in the selection process are not limited to the direct sound and the reflected sound based on the sound emitted from one sound source. For example, it is also possible to compare two reflected sounds based on the sound emitted from one sound source to perform sound selection. In this case, the direct sound in the present disclosure can also be replaced with the sound that reaches the listener first, and the reflected sound in the present disclosure can also be replaced with the sound that reaches the listener later.

[0545] (Example of the structure of the bitstream)

[0546] In the bitstream, for example, it includes a sound signal and metadata. The sound signal is sound data that represents sound, and represents information related to the frequency and intensity of the sound, etc. In addition, the metadata includes spatial information representing the space of the sound field, that is, the sound space.

[0547] For example, the spatial information is information related to the space where the listener who listens to the sound based on the sound signal is located. Specifically, the spatial information is information related to a specified position (positioning position) used to position the sound image in the sound space (for example, a three-dimensional sound field), that is, used to make the listener perceive the sound coming from the direction corresponding to the specified position. In the spatial information, for example, it includes sound source object information and position information representing the position of the listener.

[0548] The sound source object information is information about the sound source object that generates the sound based on the sound signal. That is, the sound source object information is information related to the object (sound source object) that reproduces the sound signal, and is information related to the virtual sound source object configured in the virtual sound space. Here, the virtual sound space can also correspond to the real space where the object that generates the sound is configured, and the sound source object in the virtual sound space can also correspond to the object that generates the sound in the real space.

[0549] The sound source object information can also represent the position of the sound source object configured in the sound space, the orientation of the sound source object, the directivity of the sound emitted from the sound source object, whether the sound source object belongs to a living being, and whether the sound source object is a moving body, etc. For example, the sound signal is associated with one or more sound source objects represented by the sound source object information.

[0550] The bitstream, for example, has a data structure composed of metadata (control information) and a sound signal.

[0551] The audio signal and metadata can be included in a single bitstream, or can be included separately in multiple bitstreams. In addition, the audio signal and metadata can be included in a single file, or can be included separately in multiple files.

[0552] The bitstream can also exist for each sound source, or can exist for each reproduction time. In the case where the bitstream exists for each reproduction time, multiple bitstreams can also be processed in parallel simultaneously.

[0553] Metadata can also be assigned to each bitstream, or can be assigned to multiple bitstreams together as information for controlling multiple bitstreams. In this case, multiple bitstreams can also share the metadata. In addition, metadata can also be assigned for each reproduction time.

[0554] In the case where there are multiple bitstreams or multiple files, information indicating associated bitstreams or associated files can also be included in one or more bitstreams or one or more files. Alternatively, information indicating associated bitstreams or associated files can also be included in each of all the bitstreams or each of all the files.

[0555] Here, associated bitstreams or associated files refer to, for example, bitstreams or files that may be used simultaneously during audio processing. In addition, a bitstream or a file that collectively describes information indicating associated bitstreams or associated files can also be included.

[0556] Here, the information indicating associated bitstreams or associated files can, for example, also be an identifier indicating associated bitstreams or associated files. In addition, the information indicating associated bitstreams or associated files can, for example, also be a file name, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier) of associated bitstreams or associated files, etc.

[0557] In this case, the acquisition unit can also determine and acquire associated bitstreams or associated files based on the information indicating associated bitstreams or associated files. In addition, information indicating associated bitstreams or associated files can be included in a bitstream or a file, and information indicating associated bitstreams or associated files can also be included in other bitstreams or other files.

[0558] Here, a file that includes information indicating associated bitstreams or associated files can, for example, also be a control file such as a manifest file for content distribution.

[0559] In addition, all or part of the metadata can also be obtained from outside the bitstream of the audio signal. For example, it is also possible to obtain either the metadata for controlling the audio or the metadata for controlling the video from outside the bitstream, or it is also possible to obtain both metadata from outside the bitstream.

[0560] In addition, metadata for controlling the video can also be included in the bitstream obtained by the stereophonic reproduction system 1000. In this case, the stereophonic reproduction system 1000 can also output the metadata for controlling the video to a display device that displays an image or a stereoscopic video reproduction device that reproduces a stereoscopic image.

[0561] (Examples of information included in the metadata)

[0562] The metadata can also be information used in the description of a scene represented by sound. Here, a scene is a term that represents the entire set of elements of a three-dimensional image and audio events in a sound space modeled by a sound signal reproduction system using metadata.

[0563] That is, the metadata includes not only information for controlling audio processing, but can also include information for controlling video processing. In the metadata, it is possible to include only one of the information for controlling audio processing and the information for controlling video processing, or it is possible to include both.

[0564] The stereophonic reproduction system 1000 performs audio processing on the audio signal by using the metadata included in the bitstream and by additionally obtaining the position information of the interactive listener, etc., to generate virtual audio effects. In the audio effects, it is possible to perform early reflection processing, obstacle processing, diffraction processing, blocking processing, and reverberation processing, or it is also possible to perform other audio processing using the metadata. For example, it is also possible to add audio effects such as distance attenuation effects, localization, or Doppler effects.

[0565] In addition, it is also possible to attach information for switching the on / off of all or part of the additional audio effects to the metadata, or priority information for multiple processes of the audio effects.

[0566] In addition, as an example, the metadata includes information related to a sound space including a sound source object and an obstacle object, and information related to a localization position for localizing a sound image at a specified position in the sound space (that is, making a listener perceive sound coming from a specified direction).

[0567] Here, an obstacle object is an object that may affect the sound perceived by a listener, such as blocking or reflecting sound during the period from when the sound emitted by the sound source object reaches the listener. In addition to stationary objects, the obstacle object can also include moving objects such as animals or machinery. The animal can also be a person, etc.

[0568] In addition, in the case where there are multiple sound source objects in the sound space, for any sound source object, other sound source objects may become obstacle objects. That is, both non-sounding objects that do not emit sound such as building materials or inanimate objects, and sounding sound source objects may become obstacle objects.

[0569] The metadata includes information representing all or part of the shape of the sound space, the shape and position of the obstacle objects in the sound space, the shape and position of the sound source objects in the sound space, and the position and orientation of the listener in the sound space.

[0570] The sound space can be either a closed space or an open space. In addition, the metadata may also include information representing the reflectivity of the obstacle objects that can reflect sound in the sound space. For example, the floor, walls, or ceiling that form the boundary of the sound space can also become obstacle objects.

[0571] The reflectivity is the energy ratio of the reflected sound to the incident sound, and can also be set for each frequency band of the sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, in the case where the sound space is an open space, for example, parameters such as the uniformly set attenuation rate, diffracted sound, and early reflected sound can also be used.

[0572] The metadata may also include information other than the reflectivity as parameters related to the obstacle objects or sound source objects. For example, the metadata may include information related to the material of the object as a parameter related to both the sound source object and the non-sounding object. Specifically, the metadata may include information such as diffusivity, transmittance, and sound absorption rate.

[0573] The information related to the sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and types of sound sources in one object, and information representing the sound source area in the object, etc. In the reproduction conditions, for example, it can also be set whether the sound is continuously flowing out or event-triggered. The sound source area in the object can be set either according to the relative relationship between the position of the listener and the position of the object, or by using the object as a reference.

[0574] For example, in the case where the sound source area is set according to the relative relationship between the position of the listener and the position of the object, when viewed from the listener, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.

[0575] In addition, when an object is used as a reference to set a sound source area, the object can be used as a reference to fix which sound is emitted from which area of the object. For example, when a listener views the object from the front, the listener can perceive a high pitch from the right side of the object and a low pitch from the left side of the object. Also, when a listener views the object from the back, the listener can perceive a low pitch from the right side of the object and a high pitch from the left side of the object.

[0576] Spatial metadata may also include the time up to the early reflection sound, the reverberation time, and the ratio of the direct sound to the diffused sound. When the ratio of the direct sound to the diffused sound is zero, the listener can perceive only the direct sound.

[0577] (Supplementary)

[0578] In addition, the forms grasped based on the present disclosure are not limited to the embodiments, and various modifications can be made and implemented.

[0579] For example, the processing performed by a specific component in the embodiment can be performed by other components instead of the specific component. In addition, the order of multiple processes can be changed, and multiple processes can be executed in parallel.

[0580] In addition, the ordinal numbers such as the first and the second used in the description can be appropriately replaced, removed, or newly assigned. These ordinal numbers do not necessarily correspond to a meaningful order and can also be used for identifying elements.

[0581] In addition, for example, in the comparison with a threshold value, "above the threshold value" and "greater than the threshold value" can be replaced with each other. Similarly, "below the threshold value" and "less than the threshold value" can be replaced with each other. In addition, for example, "time" and "moment" can be replaced with each other.

[0582] In addition, in the process of selecting one or more processing target sounds from multiple sounds, if there is no sound that satisfies the condition, none of the sounds may be selected as the processing target sound. That is, in the process of selecting one or more processing target sounds from multiple sounds, the case where no processing target sound is selected may also be included.

[0583] In addition, expressions such as "at least one of the first element, the second element, and the third element" can correspond to the first element, the second element, the third element, or any combination thereof.

[0584] In addition, for example, in the embodiment, the form grasped based on the present disclosure is described as being implemented as an audio processing device, an encoding device, or a decoding device. However, the form grasped based on the present disclosure is not limited to these and can also be implemented as software for executing an audio processing method, an encoding method, or a decoding method.

[0585] For example, a program for executing the above-described audio processing method, encoding method, or decoding method may also be pre-stored in a ROM. Further, it may also be that the CPU operates according to the program.

[0586] In addition, a program for executing the above-described audio processing method, encoding method, or decoding method may also be stored in a computer-readable recording medium. Further, the computer may also record the program stored in the recording medium into the RAM of the computer and operate according to the program.

[0587] Moreover, each of the above-described components may typically also be implemented as an integrated circuit (LSI) having an input terminal and an output terminal. They may be formed as a single chip either individually or in a manner that includes all or a part of the components of the embodiment. Depending on the degree of integration, the LSI may also be expressed as an IC, a system LSI, a super LSI, or an ultra-large scale LSI.

[0588] In addition, not limited to LSI, a dedicated circuit or a general-purpose processor may also be used. Further, an FPGA that can be programmed after the LSI is manufactured, or a reconfigurable processor that can perform connection or setting of circuit units inside the LSI may also be used. Furthermore, if an integrated circuit technology alternative to the LSI appears due to the progress of semiconductor technology or other derived technologies, then of course, such technology may also be used for the integration of components. It may be the application of biotechnology or the like.

[0589] In addition, it may also be that all or a part of the software for implementing the audio processing method, encoding method, or decoding method described in the present disclosure is downloaded by an FPGA or a CPU or the like through wireless communication or wired communication. Further, all or a part of the software for updating may also be downloaded through wireless communication or wired communication. Moreover, it may also be that the digital signal processing described in the present disclosure is executed by storing the downloaded software in a memory by an FPGA or a CPU or the like and operating based on the stored software.

[0590] At this time, a device having an FPGA or a CPU or the like may also be connected to the signal processing device wirelessly or by wire, or may be connected to the signal processing server via a network. Moreover, the device and the signal processing device or the signal processing server may also perform the audio processing method, encoding method, or decoding method described in the present disclosure.

[0591] For example, the audio processing device, encoding device, or decoding device of the present disclosure may also include an FPGA, a CPU, or the like. Furthermore, the audio processing device, encoding device, or decoding device may also include an interface for obtaining software for operating the FPGA, CPU, or the like from the outside, and a memory for storing the obtained software. Moreover, the FPGA, CPU, or the like may execute the signal processing described in the present disclosure by operating based on the stored software.

[0592] It may also be that the server provides software related to the audio processing, encoding processing, or decoding processing of the present disclosure. And it may also be that the terminal or device operates as the audio processing device, encoding device, or decoding device described in the present disclosure by installing the software. Additionally, it may also be that the terminal or device is connected to the server via a network to install the software.

[0593] In addition, it may also be that a device different from the terminal or device is connected to the server via a network to obtain data for installing the software, and the software is installed in the terminal or device by the other device providing the data for installing the software to the terminal or device. Additionally, examples of the software may be VR software or AR software for causing the terminal or device to execute the audio processing method described in the embodiments.

[0594] Furthermore, in the above embodiments, each component may be constituted by dedicated hardware or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[0595] As described above, the apparatus and the like related to one or more aspects have been described based on the embodiments, but the aspects grasped based on the present disclosure are not limited to the embodiments. As long as it does not deviate from the gist of the present disclosure, aspects obtained by applying various modifications conceivable by those skilled in the art to the embodiments, and aspects constructed by combining constituent elements in different modification examples are also included within the scope of one or more aspects.

[0596] (Supplementary Note)

[0597] Through the description of the above embodiments, the following technology is disclosed.

[0598] (Technology 1) An audio processing device includes a circuit and a memory; the circuit uses the memory to obtain sound space information related to a sound space; based on the sound space information, obtains characteristics related to a first sound generated from a sound source in the sound space; and based on the characteristics related to the first sound, controls whether to select a second sound corresponding to the first sound generated in the sound space.

[0599] (Technique 2) For the audio processing device described in Technique 1, the first sound is the direct sound; the second sound is the reflected sound.

[0600] (Technique 3) For the audio processing device described in Technique 2, the characteristic related to the first sound is the volume ratio of the volume of the direct sound to the volume of the reflected sound; the circuit calculates the volume ratio based on the sound space information; and controls whether to select the reflected sound based on the volume ratio.

[0601] (Technique 4) For the audio processing device described in Technique 3, when the reflected sound is selected, the circuit generates sounds that reach the two ears of the listener respectively by applying binaural processing to the reflected sound and the direct sound.

[0602] (Technique 5) For the audio processing device described in Technique 3 or 4, the circuit calculates the time difference between the end time of the direct sound and the arrival time of the reflected sound based on the sound space information; and controls whether to select the reflected sound based on the time difference and the volume ratio.

[0603] (Technique 6) For the audio processing device described in Technique 5, the circuit selects the reflected sound when the volume ratio is above the threshold; the first threshold used when the time difference is the first value is greater than the second threshold used when the time difference is the second value greater than the first value.

[0604] (Technique 7) For the audio processing device described in Technique 3 or 4, the circuit calculates the time difference between the arrival time of the direct sound and the arrival time of the reflected sound based on the sound space information; and controls whether to select the reflected sound based on the time difference and the volume ratio.

[0605] (Technique 8) For the audio processing device described in Technique 7, the circuit selects the reflected sound when the volume ratio is above the threshold; the first threshold used when the time difference is the first value is greater than the second threshold used when the time difference is the second value greater than the first value.

[0606] (Technique 9) For the audio processing device described in Technique 6 or 8, the circuit adjusts the threshold based on the arrival direction of the direct sound and the arrival direction of the reflected sound.

[0607] (Technique 10) For the audio processing device described in any one of Techniques 1 to 9, when the second sound is not selected, the circuit corrects the volume of the first sound based on the volume of the second sound.

[0608] (Technique 11) For the audio processing device described in any one of Techniques 1 to 9, when the second sound is not selected, the circuit synthesizes the second sound into the first sound.

[0609] (Technique 12) For the audio processing device described in any one of Techniques 3 to 9, the volume ratio is the volume ratio of the direct sound at the first moment to the volume of the reflected sound at the second moment different from the first moment.

[0610] (Technique 13) For the audio processing device described in Technique 1 or 2, the circuit sets a threshold based on the characteristics related to the first sound, and controls whether to select the second sound based on the threshold.

[0611] (Technique 14) For the audio processing device described in any one of Techniques 1, 2, and 13, the characteristics related to the first sound are one or a combination of two or more of the volume of the sound source, the visuality of the sound source, and the localizability of the sound source.

[0612] (Technique 15) For the audio processing device described in any one of Techniques 1, 2, and 13, the characteristics related to the first sound are the frequency characteristics of the first sound.

[0613] (Technique 16) For the audio processing device described in any one of Techniques 1, 2, and 13, the characteristics related to the first sound are the characteristics indicating the intermittency of the amplitude of the first sound.

[0614] (Technique 17) For the audio processing device described in any one of Techniques 1, 2, 13, and 16, the characteristics related to the first sound are the characteristics indicating the duration of the sounding part of the first sound or the duration of the silent part of the first sound.

[0615] (Technique 18) For the audio processing device described in any one of Techniques 1, 2, 13, 16, and 17, the characteristics related to the first sound are the characteristics representing the duration of the sounding part of the first sound and the duration of the silent part of the first sound in time series.

[0616] (Technique 19) For the audio processing device described in any one of Techniques 1, 2, 13, and 15, the characteristics related to the first sound are the characteristics indicating the variation of the frequency characteristics of the first sound.

[0617] (Technique 20) For the audio processing device described in any one of Techniques 1, 2, 13, 15, and 19, the characteristics related to the first sound are the characteristics indicating the smoothness of the frequency characteristics of the first sound.

[0618] (Technique 21) The audio processing device according to any one of Techniques 1, 2, and 13 to 20 obtains characteristics related to the first sound from the bitstream.

[0619] (Technique 22) The audio processing device according to any one of Techniques 1, 2, and 13 to 21, wherein the circuit calculates characteristics related to the second sound; and controls whether to select the second sound based on the characteristics related to the first sound and the characteristics related to the second sound.

[0620] (Technique 23) The audio processing device according to Technique 22, wherein the circuit obtains a threshold value representing the volume corresponding to the boundary of whether a sound can be heard; and controls whether to select the second sound based on the characteristics related to the first sound, the characteristics related to the second sound, and the threshold value.

[0621] (Technique 24) The audio processing device according to Technique 22 or 23, wherein the characteristics related to the second sound are the volume of the second sound.

[0622] (Technique 25) The audio processing device according to any one of Techniques 1 to 24, wherein the sound space information includes information on the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space; and the circuit controls whether to select each of the plurality of second sounds based on the characteristics related to the first sound, thereby selecting one or more processing target sounds to which binaural processing is applied from the first sound and the plurality of second sounds.

[0623] (Technique 26) The audio processing device according to any one of Techniques 1 to 25, wherein the timing of obtaining the characteristics related to the first sound is at least one of when the sound space is created, when the processing of the sound space starts, and when an information update thread occurs during the processing of the sound space.

[0624] (Technique 27) The audio processing device according to any one of Techniques 1 to 26, wherein the characteristics related to the first sound are obtained periodically after the start of the processing of the sound space.

[0625] (Technique 28) The audio processing device according to any one of Techniques 1, 2, and 25 to 27, wherein the characteristics related to the first sound are the volume of the first sound; the circuit calculates an evaluation value of the second sound based on the volume of the first sound; and controls whether to select the second sound based on the evaluation value.

[0626] (Technique 29) The audio processing device according to Technique 28, wherein the volume of the first sound has a transition.

[0627] (Technique 30) The audio processing device as described in Technique 28 or 29, wherein the circuit calculates the evaluation value such that the louder the volume of the first sound, the more likely the second sound is to be selected.

[0628] (Technique 31) The audio processing device as described in any one of Techniques 1 to 30, wherein the sound space information is scene information including information on the sound source in the sound space and information on the position of the listener in the sound space; the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space; the circuit obtains a signal of the first sound; calculates the plurality of second sounds based on the scene information and the signal of the first sound; obtains characteristics related to the first sound from the information on the sound source; and controls whether each of the plurality of second sounds is selected as a sound not to be subjected to binaural processing based on the characteristics related to the first sound, thereby selecting one or more second sounds not to be subjected to the binaural processing from the plurality of second sounds.

[0629] (Technique 32) The audio processing device as described in Technique 31, wherein the scene information is updated based on input information; and characteristics related to the first sound are obtained according to the update of the scene information.

[0630] (Technique 33) The audio processing device as described in Technique 31 or 32, wherein the scene information and the characteristics related to the first sound are obtained from metadata included in a bitstream.

[0631] (Technique 34) The audio processing device as described in any one of Techniques 1, 2, 13, 16 to 18, 25 to 27, and 31 to 33, wherein the characteristics related to the first sound are characteristics of a plurality of groups represented in time series and each composed of a duration with the amplitude value of the first sound as a representative amplitude value and a group of the representative amplitude values during the duration.

[0632] (Technique 35) The audio processing device as described in Technique 34, wherein the representative amplitude value is a ratio of the volume of the first sound to a preset reference volume.

[0633] (Technique 36) The audio processing device as described in any one of Techniques 1, 2, 13, 15, 19, and 20, wherein the characteristics related to the first sound are characteristics representing the duration during which the variation amount of the frequency characteristic is lower than a preset threshold value.

[0634] (Technique 37) The audio processing device as described in any one of Techniques 1, 2, 13, 15, 20, and 36, wherein the characteristics related to the first sound are characteristics of a plurality of groups represented in time series and each composed of a duration during which the variation amount of the frequency characteristic is lower than a preset threshold value and a group of the frequency characteristics during the duration.

[0635] (Technique 38) The audio processing apparatus according to any one of Techniques 1, 2, 13 to 24, and 34 to 37, wherein the circuit obtains a threshold value of a volume corresponding to a boundary between being able to hear a sound or not, calculates the volume of the second sound based on a characteristic related to the first sound, and selects the second sound when the volume of the second sound is greater than the threshold value.

[0636] (Technique 39) The audio processing apparatus according to any one of Techniques 1, 2, 13 to 20, and 31 to 38, wherein the sound space information is scene information including information on the sound source in the sound space and information on the position of the listener in the sound space, the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space, the circuit obtains a signal of the first sound, calculates the plurality of second sounds based on the scene information and the signal of the first sound, obtains a characteristic related to the first sound from the information on the sound source, controls whether to select each of the plurality of second sounds as a sound to which binaural processing is applied based on the characteristic related to the first sound, thereby selecting one or more processing target sounds to which the binaural processing is applied from the first sound and the plurality of second sounds, the scene information is updated based on input information, the characteristic related to the first sound is obtained according to the update of the scene information, and the update of the scene information is performed at a frequency lower than the frequency at which the binaural processing is applied to the one or more processing target sounds.

[0637] (Technique 40) An audio processing method, comprising: a step of obtaining sound space information related to a sound space; a step of obtaining a characteristic related to a first sound generated from a sound source in the sound space based on the sound space information; and a step of controlling whether to select a second sound generated corresponding to the first sound in the sound space based on the characteristic related to the first sound.

[0638] (Technique 41) A program for causing a computer to execute the audio processing method described in Technique 40.

[0639] Industrial Applicability

[0640] The present disclosure includes, for example, forms that can be applied to an audio processing apparatus, an encoding apparatus, a decoding apparatus, or a terminal or device having some of these apparatuses.

[0641] Reference Signs

[0642] 1000 Stereo Reproduction System

[0643] 1001 Sound Signal Processing Apparatus (Audio Processing Apparatus)

[0644] 1002 Sound prompting device

[0645] 1100, 1120, 1500 Encoding devices

[0646] 1101, 1113 Input data

[0647] 1102 Encoder

[0648] 1103 Encoded data

[0649] 1104, 1114, 1404, 1503 Memories

[0650] 1110, 1130 Decoding devices

[0651] 1111 Sound signal

[0652] 1112, 1200, 1210 Decoders

[0653] 1121 Transmitting unit

[0654] 1122 Transmitted signal

[0655] 1131 Receiving unit

[0656] 1132 Received signal

[0657] 1201, 1211 Spatial information management units

[0658] 1202 Sound data decoder

[0659] 1203, 1213, 1300 Rendering units

[0660] 1301 Analysis unit

[0661] 1302, 1314 Selection units

[0662] 1303 Synthesis unit

[0663] 1304 Threshold adjustment unit

[0664] 1311 Reverberation processing unit

[0665] 1312 Early reflection processing unit

[0666] 1313 Distance attenuation processing unit

[0667] 1315 Generation unit

[0668] 1316 Binaural processing unit

[0669] 1401 Speaker

[0670] 1402 and 1501 Processors

[0671] 1403 and 1502 Communication IF

[0672] 1405 Sensors

Claims

1. An audio processing device, wherein, it includes a circuit and a memory, the circuit uses the memory, obtains sound space information related to a sound space, based on the sound space information, obtains characteristics related to a first sound generated from a sound source in the sound space, based on the characteristics related to the first sound, controls whether to select a second sound generated corresponding to the first sound in the sound space.

2. The audio processing device according to claim 1, wherein, the characteristics related to the first sound are characteristics of a plurality of groups represented in time series and each composed of a duration with the amplitude value of the first sound as a representative amplitude value and a group of the representative amplitude values during the duration.

3. The audio processing device according to claim 2, wherein, the representative amplitude value is a ratio of the volume of the first sound to a preset reference volume.

4. The audio processing device according to claim 1, wherein, the characteristics related to the first sound are characteristics representing the duration of a state where the variation amount of the frequency characteristics is lower than a preset threshold.

5. The audio processing device according to claim 1, wherein, the characteristics related to the first sound are characteristics of a plurality of groups represented in time series and each composed of a duration of a state where the variation amount of the frequency characteristics is lower than a preset threshold and a group of the frequency characteristics during the duration.

6. The audio processing device according to any one of claims 1 to 5, wherein, the circuit, obtains a threshold representing the volume corresponding to the boundary of whether a sound can be heard, calculates the volume of the second sound based on the characteristics related to the first sound, selects the second sound when the volume of the second sound is greater than the threshold.

7. The audio processing device according to claim 1, wherein, the sound space information is scene information including information on the sound source in the sound space and information on the position of a listener in the sound space, the second sound is each of a plurality of second sounds generated corresponding to the first sound in the sound space, the circuit, obtains a signal of the first sound, calculates the plurality of second sounds based on the scene information and the signal of the first sound, obtains characteristics related to the first sound from the information on the sound source, based on the characteristics related to the first sound, controls whether to respectively select the plurality of second sounds as sounds to which binaural processing is applied, thereby selecting one or more processing target sounds to which the binaural processing is applied from the first sound and the plurality of second sounds, the scene information is updated based on input information, the characteristics related to the first sound are obtained according to the update of the scene information, the update of the scene information is performed at a frequency lower than the frequency of applying the binaural processing to the one or more processing target sounds.

8. An audio processing method, wherein, it includes: a step of obtaining sound space information related to a sound space; a step of obtaining characteristics related to a first sound generated from a sound source in the sound space based on the sound space information; and A step of controlling whether to select a second sound generated corresponding to the first sound in the sound space based on characteristics related to the first sound.

9. A program for causing a computer to execute the audio processing method according to claim 8.

Citation Information

Patent Citations

  • Signal processor

    JP2019022049A

  • Apparatus and method for rendering a sound scene using pipeline stages

    WO2021180938A1