Acoustic processing device and acoustic processing method
By calculating the evaluation value of the reflected sound in the audio processing device and selecting the processing object, the problem of excessive load on the reflected sound processing operation in the virtual space is solved, and the efficiency of the audio processing and battery life are improved.
Patent Information
- Application Number
- CN202380071402.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-05
- Filing Date
- 2023-10-06
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively process a large number of reflected sounds in the virtual space, resulting in excessive computational load, affecting the efficiency of audio processing and battery life.
By using circuits and memory in the audio processing device, sound space information is obtained, the evaluation value of the reflected sound is calculated, and the reflected sound of the processing object is selected based on the evaluation value to reduce the computational load.
It realizes appropriate processing of reflected sound in the virtual space, reducing the computational load of audio processing, and improving sound quality and battery life.
Smart Images

Figure CN119999236A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio processing device and the like. Background Art
[0002] In recent years, products and services using ER (Extended Reality) (also known as XR), including VR (Virtual Reality), AR (Augmented Reality) and MR (Mixed Reality) have become popular. As a result, the importance of sound processing technology that provides immersive audio (Immersive Audio) to listeners by giving the sound emitted by virtual sound sources in virtual space or real space an acoustic effect corresponding to the environment of the space has increased.
[0003] In addition, the listener may also be expressed as a listener or a user. Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1 disclose technologies related to the sound processing device and the sound processing method of the present disclosure.
[0004] Prior art literature
[0005] Patent Literature
[0006] Patent Document 1: Japanese Patent No. 6288100
[0007] Patent Document 2: Japanese Patent Application Publication No. 2019-22049
[0008] Patent Document 3: International Publication No. 2021 / 180938
[0009] Non-patent literature
[0010] Non-patent literature 1: Appl. Sci. 2019, 9, 2854; doi: 10.3390 / app9142854 "Psychoacoustic Models for Perceptual Audio Coding-A Tutorial Review" Summary of the invention
[0011] Problems to be solved by the invention
[0012] For example, Patent Document 1 discloses a technology for performing signal processing on a target audio signal and presenting it to a listener. With the popularization of ER technology and the diversification of services using ER technology, there is a demand for audio processing corresponding to differences in, for example, the audio quality required by each service, the signal processing capability of the terminal used, and the sound quality that can be provided by the audio presentation device. In addition, in order to provide these, further improvements in the audio processing technology are required.
[0013] Here, the improvement of the sound processing technology refers to the change of the existing sound processing. For example, the improvement of the sound processing technology provides a process for giving a new sound effect, a reduction in the amount of sound processing, an improvement in the quality of the sound obtained by the sound processing, a reduction in the amount of data of information used in the implementation of the sound processing, or an facilitation of obtaining or generating information used in the implementation of the sound processing, etc. Alternatively, the improvement of the sound processing technology may provide a combination of any two or more of these.
[0014] In particular, these improvements are required in devices or services that allow listeners to move freely in a virtual space. However, the above-mentioned effects obtained by improving the sound processing technology are just examples. One or more technical solutions grasped based on the present disclosure may also be technical solutions conceived based on different viewpoints from the above, technical solutions that achieve different purposes from the above, or technical solutions that can obtain different effects from the above.
[0015] Means used to solve problems
[0016] An audio device related to a technical solution grasped based on the present disclosure includes a circuit and a memory; the circuit uses the memory to obtain sound space information, the sound space information including information of a sound source in the sound space, information of an object in the sound space, and information of a listener's position in the sound space; and uses the sound space information to calculate an evaluation value of a reflected sound generated corresponding to the sound generated from the sound source.
[0017] Furthermore, these inclusive or specific technical solutions may also be implemented by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or by any combination of these.
[0018] Effects of the Invention
[0019] A technical solution of the present disclosure can provide, for example, processing for imparting new sound effects, reduction in the amount of processing for sound processing, improvement in the sound quality of sound obtained through sound processing, reduction in the amount of data of information used in the implementation of sound processing, or facilitation of obtaining or generating information used in the implementation of sound processing, etc. Alternatively, a technical solution of the present disclosure can provide any combination of these. As a result, a technical solution of the present disclosure provides sound processing suitable for the use environment of the listener, and can contribute to improving the sound experience of the listener.
[0020] In particular, the above-mentioned effect can be obtained in a device or service that allows the listener to move freely in a virtual space. However, the above-mentioned effect is only an example of the effect of various technical solutions grasped by the present disclosure. One or more technical solutions grasped by the present disclosure may also be technical solutions conceived based on different viewpoints from the above, technical solutions that achieve different purposes from the above, or technical solutions that can obtain different effects from the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a diagram showing an example of direct sound and reflected sound generated in an acoustic space.
[0022] Figure 2 It is a diagram showing an example of a stereophonic sound reproduction system according to an embodiment.
[0023] Figure 3A This is a block diagram showing a configuration example of an encoding device according to an embodiment.
[0024] Figure 3B This is a block diagram showing a configuration example of a decoding device according to an embodiment.
[0025] Figure 3C This is a block diagram showing another configuration example of the encoding device according to the embodiment.
[0026] Figure 3D This is a block diagram showing another configuration example of the decoding device according to the embodiment.
[0027] Figure 4A This is a block diagram showing an example configuration of a decoder according to an embodiment.
[0028] Figure 4B This is a block diagram showing another configuration example of a decoder according to an embodiment.
[0029] Figure 5 It is a diagram showing an example of the physical structure of the sound signal processing device according to the embodiment.
[0030] Figure 6 This is a diagram showing an example of the physical structure of the encoding device according to the embodiment.
[0031] Figure 7 This is a block diagram showing a configuration example of a rendering unit according to an embodiment.
[0032] Figure 8 This is a flowchart showing an operation example of the sound signal processing device according to the embodiment.
[0033] Fig. 9 This diagram shows the positional relationship between the listener and the obstacle object, which is relatively far away.
[0034] Fig.10 This diagram shows the positional relationship between the listener and the obstacle object, which is relatively close.
[0035] Fig.11 This is a flowchart showing an example of selection processing in the embodiment.
[0036] Fig.12 This is a flowchart showing an example of evaluation processing in the embodiment.
[0037] Fig.13 It is a diagram showing examples of arrival angles of direct sound and reflected sound.
[0038] Fig.14 This is a diagram showing an example of a method of setting threshold data based on the temporal masking phenomenon.
[0039] Fig.15 It is a diagram showing an example of threshold data.
[0040] Fig.16 This is a diagram showing the relationship between the time difference between direct sound and reflected sound and the threshold value.
[0041] Fig.17 This is a block diagram showing an example of a configuration for a rendering unit to perform pipeline processing. DETAILED DESCRIPTION
[0042] (Understanding that is the basis of the present disclosure)
[0043] Figure 1 This is a diagram showing an example of direct sound and reflected sound generated in a sound space. In the audio processing that expresses the characteristics of a virtual space with sound, in order to express the width of the space and the material of the wall, etc., and to accurately grasp the position of the sound source (localization of the sound image), it is effective to reproduce not only the direct sound but also the reflected sound.
[0044] For example, in Figure 1When listening to sound in a rectangular room, six primary reflections corresponding to the six walls are generated for one sound source. The reproduction of these reflections becomes a clue to the proper understanding of space and sound image. Furthermore, for each reflection, a secondary reflection is generated on a surface other than the reflection surface that generated the reflection. These reflections also become effective clues in perception.
[0045] However, even if only the second reflection is considered, one direct sound and 36 (6+6×5) reflected sounds are generated for one sound source, so 37 sound rays are generated. A considerable amount of calculation is required to process these sound rays.
[0046] In addition, in recent years' application products based on the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, there will inevitably be multiple sound sources, so even larger amounts of computing power will be required.
[0047] In addition, listeners who listen to sounds in a virtual space use headphones or VR goggles. In order to provide stereo sound to such listeners, binaural processing is performed on each sound line to reproduce the direction and sense of distance of the sound by giving a sound pressure ratio and phase difference between the two ears. Therefore, if all the reflected sounds generated are to be reproduced, the amount of calculation is very large.
[0048] On the other hand, as batteries for VR goggles worn by listeners experiencing in a virtual space, small storage batteries are sometimes used for their convenience. In order to extend the battery life, the computational load required for the above-mentioned processing is preferably small. For this reason, it is desirable to reduce the number of sound rays generated on a scale of several hundred within a range that does not impair the localization of the sound and the grasp of the space.
[0049] In addition, in a system that reproduces sound, sometimes 6DoF (6 Degrees of Freedom) or the like is allowed for the position and orientation of the listener. In this case, the positional relationship between the listener, the sound source, and the object that reflects the sound cannot be determined unless it is reproduced (rendered). Therefore, the reflected sound cannot be determined unless it is reproduced. Therefore, it is difficult to predetermine the reflected sound of the processing object.
[0050] That is, the number of sound rays for expressing the characteristics of the virtual space by sound and the transition of the volume of the sound of each sound ray are calculated during rendering. Therefore, it is not easy to reduce the amount of calculation during rendering.
[0051] As a method of reducing the number of sound rays in a space, for example, Patent Document 1 discloses a method of detecting the importance of an audio object and not reproducing sounds caused by an audio object with a low importance.
[0052] However, when the degrees of freedom such as 6DoF are allowed with respect to the position and orientation of the listener, the listener understands the positional relationship between the listener, the sound source, and the object that reflects the sound based on the direct sound generated from the sound source and the reflected sound generated by the direct sound being reflected by the object. Therefore, it is sometimes difficult to accurately understand the positioning and space of the sound by reducing the direct sound and reflected sound caused by a specific sound source.
[0053] Therefore, an object of the present disclosure is to provide a sound processing device or the like that can grasp the localization and space of sound and can suppress the calculation load for processing reflected sound.
[0054] (Public Summary)
[0055] The sound processing device based on the first technical solution mastered by the present disclosure includes a circuit and a memory. The circuit uses the memory to obtain sound space information, which includes information about a sound source in the sound space, information about an object in the sound space, and information about a listener's position in the sound space; and uses the sound space information to calculate an evaluation value of a reflected sound generated corresponding to a sound generated from the sound source.
[0056] The device of the above technical solution can use the sound space information to appropriately calculate the evaluation value of the reflected sound that depends on the information of the sound source, the information of the object, and the information of the position of the listener. Therefore, based on the evaluation value of the reflected sound, the reflected sound of the processing object can be appropriately selected. Therefore, the positioning and space of the sound can be grasped, and the calculation load for the reflected sound can be suppressed.
[0057] A sound processing device according to a second technical aspect grasped by the present disclosure may be the sound processing device according to the first technical aspect, wherein the circuit controls whether to select the reflected sound based on the evaluation value.
[0058] The device of the above-mentioned technical solution can appropriately select the reflected sound to be processed based on the evaluation value of the reflected sound.
[0059] The sound processing device according to a third technical solution grasped based on the present disclosure may be the sound processing device according to the second technical solution, wherein, when the reflected sound is not selected, the circuit does not perform binaural processing on the reflected sound.
[0060] The device of the above technical solution can suppress the calculation load for reflected sound by omitting binaural processing.
[0061] A sound processing device according to a fourth technical solution grasped by the present disclosure may be the sound processing device according to any one of the first to third technical solutions, wherein the circuit calculates the volume of the reflected sound and calculates the evaluation value when the volume exceeds a predetermined threshold value.
[0062] The device of the above-mentioned technical solution can omit the calculation of the evaluation value of the reflected sound when the volume of the reflected sound is equal to or less than a predetermined threshold value. Therefore, the device of the above-mentioned technical solution can suppress the calculation load for the reflected sound.
[0063] The sound processing device based on the fifth technical solution mastered by the present disclosure may also be, in the sound processing device of the second technical solution, when a reflected sound is selected based on an evaluation value, the circuit calculates the total computational load of one or more selected reflected sounds including the reflected sound, and when the total computational load exceeds a predetermined upper limit, the selection of the reflected sound is terminated.
[0064] The device of the above technical solution can suppress the total calculation load from exceeding a predetermined upper limit. Therefore, the device of the above technical solution can suppress the calculation load for reflected sound.
[0065] The sound processing device according to the sixth technical solution grasped based on the present disclosure may be the sound processing device according to the fifth technical solution, wherein the total calculation load is defined by the number of one or more selected reflected sounds or the processing amount of one or more selected reflected sounds.
[0066] The device of the above technical solution can suppress the number of one or more selected reflected sounds or the processing amount of one or more selected reflected sounds from exceeding a predetermined upper limit. Therefore, the device of the above technical solution can suppress the calculation load for the reflected sounds.
[0067] The sound processing device based on the 7th technical solution mastered by the present disclosure may also be, in the sound processing device of any one of the 1st to 6th technical solutions, the circuit calculates the volume of each of the multiple reflected sounds generated as reflected sounds in the sound space, and calculates the evaluation value of each of the more than one reflected sounds among the multiple reflected sounds having a volume above a predetermined threshold.
[0068] The device of the above technical solution can omit calculation of the evaluation value of each of the plurality of reflected sounds when the volume of the reflected sound is lower than a predetermined threshold value. Therefore, the device of the above technical solution can reduce the calculation load for the reflected sound.
[0069] The sound processing device of the 8th technical solution mastered based on the present disclosure may also be that, in the sound processing device of the 7th technical solution, the circuit calculates the total computational load of one or more reflected sounds, and when the total computational load exceeds a predetermined upper limit, calculates the evaluation value of each of the more than one reflected sounds.
[0070] The device of the above-mentioned technical solution can omit the calculation of the evaluation value of the reflected sound when the total calculation load is equal to or less than a predetermined upper limit. Therefore, the device of the above-mentioned technical solution can suppress the calculation load for the reflected sound.
[0071] The sound processing device based on the 9th technical solution mastered by the present disclosure may also be, in the sound processing device of any one of the 1st to 8th technical solutions, the circuit calculates the evaluation value of each of the multiple reflected sounds generated as reflected sounds in the sound space, and adds the computational load of the reflected sound to the total computational load for each of the multiple reflected sounds in order of the evaluation values from high to low, and each time the computational load of the reflected sound is added to the total computational load, the total computational load is compared with a predetermined upper limit, and when the total computational load obtained by adding the computational loads of the reflected sounds does not exceed the predetermined upper limit, the reflected sound is selected, and when the total computational load obtained by adding the computational loads of the reflected sounds exceeds the predetermined upper limit, one or more remaining reflected sounds after the reflected sound among the multiple reflected sounds are not selected.
[0072] The device of the above technical solution can exclude the remaining reflected sounds from selection when the total calculation load after adding the calculation loads in sequence exceeds a predetermined upper limit. Therefore, the device of the above technical solution can limit the reflected sounds to be processed and suppress the calculation load.
[0073] The sound processing device of the 10th technical solution mastered based on the present disclosure may also be, in the sound processing device of any one of the 1st to 9th technical solutions, the evaluation value is the total value of at least one of the index values related to the volume, the visual index value, the index value related to the object, and the index value representing the relationship between the direct sound corresponding to the reflected sound and the reflected sound.
[0074] The device of the above technical solution can calculate the total value of at least one of the index value related to the volume, the index value of the visuality, the index value related to the object, and the index value indicating the relationship between the direct sound and the reflected sound as the evaluation value. Therefore, based on the index value related to the volume, the visuality, the index value related to the object, or the index value indicating the relationship between the direct sound and the reflected sound, the reflected sound of the processing object can be appropriately selected.
[0075] The sound processing device of the eleventh technical solution grasped based on the present disclosure may be the sound processing device of the tenth technical solution, wherein the circuit increases the index value related to the volume as the volume of the sound generated from the sound source increases.
[0076] The device of the above technical solution can calculate a higher evaluation value as the volume of the sound generated from the sound source increases. Therefore, based on the higher evaluation value as the volume of the sound generated from the sound source increases, the reflected sound to be processed can be appropriately selected.
[0077] The sound processing device of the 12th technical solution grasped based on the present disclosure may also be, in the sound processing device of the 10th or 11th technical solution, the circuit makes the visual index value larger when the sound source enters the listener's field of view than when the sound source does not enter the listener's field of view.
[0078] The device of the above technical solution can calculate a higher evaluation value when the sound source enters the field of view than when the sound source does not enter the field of view. Therefore, based on the higher evaluation value when the sound source enters the field of view than when the sound source does not enter the field of view, the reflected sound of the processing object can be appropriately selected.
[0079] The sound processing device of a thirteenth technical solution grasped based on the present disclosure may be the sound processing device of any one of the tenth to twelfth technical solutions, wherein the circuit increases the visibility index value as the moving speed of the sound source is slower.
[0080] The device of the above technical solution can calculate a higher evaluation value as the moving speed of the sound source is slower. Therefore, based on the higher evaluation value as the moving speed of the sound source is slower, the reflected sound to be processed can be appropriately selected.
[0081] The sound processing device of the 14th technical solution grasped by the present disclosure may be the sound processing device of any one of the 10th to 13th technical solutions, wherein the index value related to the object is assigned to each object in the sound space and included in the sound space information.
[0082] The device of the above-mentioned technical solution can calculate the evaluation value based on the index value assigned to each object. Therefore, based on the index value assigned to each object, the reflected sound of the processing object can be appropriately selected.
[0083] The sound processing device based on the 15th technical solution mastered by the present disclosure can also be a sound processing device in any one of the 10th to 14th technical solutions, wherein the larger the angle between the direction of arrival of the direct sound and the direction of arrival of the reflected sound, the larger the index value representing the relationship between the direct sound and the reflected sound is made by the circuit.
[0084] The device of the above technical solution can calculate the evaluation value that becomes higher as the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound increases. Therefore, the reflected sound to be processed can be appropriately selected based on the evaluation value that becomes higher as the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound increases.
[0085] The sound processing device based on the 16th technical solution mastered by the present disclosure can also be, in the sound processing device of any one of the 10th to 15th technical solutions, the greater the difference between the distance from the sound source to the listener of the direct sound and the distance from the sound source to the listener of the reflected sound after reflection, the larger the circuit makes the index value representing the relationship between the direct sound and the reflected sound larger.
[0086] The device of the above technical solution can calculate a higher evaluation value as the difference between the distance of the direct sound and the distance of the reflected sound increases. Therefore, the reflected sound to be processed can be appropriately selected based on the higher evaluation value as the difference between the distance of the direct sound and the distance of the reflected sound increases.
[0087] The sound processing device based on the 17th technical solution mastered by the present disclosure can also be, in the sound processing device of any one of the 10th to 16th technical solutions, the more the amplitude value of the reflected sound exceeds the temporal masking threshold, the larger the circuit makes the index value representing the relationship between the direct sound and the reflected sound, and the temporal masking threshold is the threshold of the temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is below the threshold.
[0088] The device of the above technical solution can calculate a higher evaluation value as the amplitude of the reflected sound exceeds the temporal masking threshold. Therefore, the reflected sound to be processed can be appropriately selected based on the higher evaluation value as the amplitude of the reflected sound exceeds the temporal masking threshold.
[0089] The sound processing device based on the 18th technical solution mastered by the present disclosure may also be, in the sound processing device of any one of the 10th to 17th technical solutions, the circuit repeatedly implements the following processing: among the multiple reflected sounds generated as reflected sounds in the sound space, reducing the index value related to the object related to the selected reflected sound, calculating the evaluation value for the reflected sound that has not yet been selected, and selecting the reflected sound in descending order of the evaluation value; when the total computational load of one or more selected reflected sounds among the multiple reflected sounds exceeds a predetermined upper limit, the repeated processing is terminated.
[0090] The device of the above technical solution can end the processing of the newly selected reflected sound when the total calculation load of the one or more selected reflected sounds exceeds a predetermined upper limit. Therefore, the device of the above technical solution can limit the reflected sounds to be processed and suppress the calculation load.
[0091] The sound processing device based on the 19th technical solution mastered by the present disclosure has a circuit and a memory. The circuit uses the memory to obtain information on the volume of the sound output from the sound source, uses the volume information to correct the evaluation value of the reflected sound corresponding to the sound, and controls whether to select the reflected sound based on the corrected evaluation value.
[0092] The device of the above-mentioned technical solution can appropriately correct the evaluation value of the reflected sound corresponding to the sound by using the information of the volume of the sound, and can appropriately control the selection of the reflected sound.
[0093] The sound processing device according to the 20th technical solution grasped based on the present disclosure may be the sound processing device according to the 19th technical solution, wherein the volume has a transition.
[0094] The device of the above technical solution can appropriately correct the evaluation value of the reflected sound corresponding to the sound by using the information of the converted volume, and can appropriately control the selection of the reflected sound.
[0095] The sound processing method of the 21st technical solution mastered based on the present disclosure includes: a step of obtaining sound space information, the sound space information includes information about the sound source in the sound space, information about the object in the sound space, and information about the position of the listener in the sound space; and using the sound space information, calculating the evaluation value of the reflected sound generated corresponding to the sound generated from the sound source.
[0096] The method of the above technical solution can achieve the same effect as the sound processing device described in the first technical solution.
[0097] The program according to the 22nd technical solution grasped based on the present disclosure is a program for causing a computer to execute the sound processing method according to the 21st technical solution.
[0098] The program of the above technical solution can achieve the same effect as the sound processing method of the 21st technical solution using a computer.
[0099] In addition, these inclusive or specific technical solutions can be implemented by systems, devices, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, or recording media.
[0100] Hereinafter, the sound processing device, the encoding device, the decoding device, and the stereo sound reproduction system of the present disclosure will be described in detail with reference to the accompanying drawings. The stereo sound reproduction system can also be expressed as a sound signal reproduction system.
[0101] In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, configuration positions and connection forms of constituent elements, steps and the order of steps, etc. shown in the following embodiments are examples and do not limit the main purpose of the technical solution grasped by this disclosure. In addition, regarding the constituent elements of the following embodiments, for example, constituent elements not included in the basic technical solution described in this disclosure or constituent elements not recorded in the independent claims representing the highest concept, they are set as arbitrary constituent elements for description.
[0102] (Implementation Method)
[0103] (Example of Stereo Sound Reproduction System)
[0104] Figure 2 is a diagram showing an example of a stereophonic sound reproduction system. Specifically, Figure 2 A stereo sound reproduction system 1000 is shown as an example of a system to which the sound processing or decoding processing of the present disclosure can be applied. Stereo sound is also expressed as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound presentation device 1002 .
[0105] The sound signal processing device 1001 is also represented as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio signal after the audio processing for the listener. The sound signal is not limited to speech, as long as it is audible. For example, the audio processing is a signal processing performed on the sound signal in order to reproduce one or more effects to which the sound is subjected during the period from the sound source to the listener.
[0106] The sound signal processing device 1001 performs sound processing based on the spatial information describing the causes of the above-mentioned effects. The spatial information is, for example, information indicating the positions of the sound source, the listener, and surrounding objects, information indicating the shape of the space, and parameters related to the propagation of the sound. The sound signal processing device 1001 is, for example, a PC (Personal Computer), a smart phone, a tablet computer, or a game console.
[0107] The acoustically processed signal is presented to the listener from the audio presentation device 1002. The audio presentation device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The acoustically processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio presentation device 1002 via wireless or wired communication.
[0108] When the sound prompting device 1002 is composed of a plurality of devices such as a device for the right ear and a device for the left ear, the plurality of devices synchronously prompt the sound through communication between the plurality of devices or communication between the plurality of devices and the sound signal processing device 1001. The sound prompting device 1002 is, for example, a headset, an earplug, a head-mounted display, or a surround speaker composed of a plurality of fixed speakers worn on the head of the listener.
[0109] In addition, the stereo sound reproduction system 1000 can also be used in combination with an image prompting device or a stereoscopic image prompting device that visually provides an ER experience including AR / VR. For example, the space handled by the spatial information is a virtual space, and the positions of the sound source, listener, and object in the space are virtual positions of the virtual sound source, virtual listener, and virtual object in the virtual space. The space can also be expressed as a sound space. In addition, the spatial information can also be expressed as sound space information.
[0110] also, Figure 2 The system configuration example in which the sound signal processing device 1001 and the sound prompting device 1002 are different devices is shown. However, the stereo sound reproduction system 1000 to which the sound processing method or decoding method disclosed in the present invention can be applied is not limited to Figure 2 For example, the audio signal processing device 1001 may be included in the audio prompting device 1002, and the audio prompting device 1002 may perform both audio processing and audio prompting.
[0111] Furthermore, the audio signal processing device 1001 and the audio prompting device 1002 may share and implement the audio processing described in this disclosure. Furthermore, a server connected to the audio signal processing device 1001 or the audio prompting device 1002 via a network may implement part or all of the audio processing described in this disclosure.
[0112] Furthermore, the audio signal processing device 1001 may perform audio processing by decoding a bit stream generated by encoding at least a part of the data of the audio signal and the spatial information used for the audio processing. Therefore, the audio signal processing device 1001 may be expressed as a decoding device.
[0113] (Example of encoding device)
[0114] Figure 3A is a block diagram showing an example of the configuration of an encoding device. Specifically, Figure 3A The configuration of an encoding device 1100 is shown as an example of an encoding device of the present disclosure.
[0115] Input data 1101 is encoding target data including spatial information and / or a sound signal input to encoder 1102. The details of spatial information will be described later.
[0116] The encoder 1102 encodes the input data 1101 to generate encoded data 1103. The encoded data 1103 is, for example, a bit stream generated by encoding processing.
[0117] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or may be other memory.
[0118] In the above description, a bit stream generated by encoding is listed as an example of the encoded data 1103 stored in the memory 1104, but the encoded data 1103 may be data other than a bit stream. For example, the encoding device 1100 may store converted data generated by converting a bit stream into a predetermined data format in the memory 1104. The converted data may be, for example, a file or a multiplexed stream corresponding to one or more bit streams.
[0119] Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format) etc. In addition, the encoded data 1103 may be in the form of a plurality of packets generated by dividing the above-mentioned bit stream or file.
[0120] For example, the bit stream generated by the encoder 1102 may be transformed into data different from the bit stream. In this case, the encoding device 1100 includes a transform unit (not shown), and the transform process may be performed by the transform unit or by a CPU (Central Processing Unit) as an example of a processor described later.
[0121] (Example of decoding device)
[0122] Figure 3B is a block diagram showing an example of the structure of a decoding device. Specifically, Figure 3B The configuration of a decoding device 1110 is shown as an example of a decoding device of the present disclosure.
[0123] The memory 1114 stores, for example, the same data as the coded data 1103 generated by the coding device 1100. The stored data is read out from the memory 1114 and input to the decoder 1112 as input data 1113. The input data 1113 is, for example, a bit stream to be decoded. The memory 1114 may be, for example, a hard disk or an SSD, or other memory.
[0124] In addition, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 directly to the decoder 1112, but may transform the data read and input the transformed data to the decoder 1112 as input data 1113. The data before the transformation may be, for example, multiplexed data including one or more bit streams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF.
[0125] In addition, the data before conversion may be a plurality of packets generated by dividing the bit stream or file described above. Data different from the bit stream may be read from the memory 1114 and converted into the bit stream. In this case, the decoding device 1110 may also include a conversion unit (not shown) to perform the conversion process, or a CPU as an example of a processor described later may perform the conversion process.
[0126] The decoder 1112 decodes the input data 1113 and generates a sound signal 1111 representing a sound to be presented to the listener.
[0127] (Another example of an encoding device)
[0128] Figure 3C is a block diagram showing another example of the configuration of the encoding device. Specifically, Figure 3C 1 shows the structure of an encoding device 1120 which is another example of the encoding device of the present disclosure. Figure 3C In Figure 3A The same constituent elements as the constituent elements of Figure 3A The same reference numerals as those in the figure are used to represent the components, and descriptions of these components are omitted.
[0129] The coding device 1100 stores the coded data 1103 in the memory 1104. On the other hand, the coding device 1120 is different from the coding device 1100 in that it includes a transmission unit 1121 that transmits the coded data 1103 to the outside.
[0130] The transmission unit 1121 transmits a transmission signal 1122 generated based on the coded data 1103 or data converted from the coded data 1103 to another data format to another device or server. The data used to generate the transmission signal 1122 is, for example, the bit stream, multiplexed data, file, or packet described in the coding device 1100.
[0131] (Another example of a decoding device)
[0132] Figure 3D is a block diagram showing another example of the structure of a decoding device. Specifically, Figure 3D 1 shows the structure of a decoding device 1130 which is another example of the decoding device of the present disclosure. Figure 3D In Figure 3B The same constituent elements as the constituent elements of Figure 3B The same reference numerals as those in the figure are used to represent the components, and descriptions of these components are omitted.
[0133] The decoding device 1110 reads the input data 1113 from the memory 1114. On the other hand, the decoding device 1130 is different from the decoding device 1110 in that it includes a receiving unit 1131 that receives the input data 1113 from the outside.
[0134] The receiving unit 1131 receives the reception signal 1132 to obtain reception data, and outputs input data 1113 to be input to the decoder 1112. The reception data may be the same as the input data 1113 to be input to the decoder 1112, or may be data of a data format different from that of the input data 1113.
[0135] When the data format of the received data is different from the data format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the received data may be converted into the input data 1113 by a conversion unit (not shown) or a CPU of the decoding device 1130. The received data is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 1120.
[0136] (Decoder example)
[0137] Figure 4A is a block diagram showing an example of the configuration of a decoder. Specifically, Figure 4A Indicate as Figure 3B or Figure 3D The configuration of a decoder 1200 is an example of a decoder 1112 in FIG.
[0138] The input data 1113 is a coded bit stream, and includes coded audio data, which is a coded audio signal, and metadata used in audio processing.
[0139] The spatial information management unit 1201 obtains metadata included in the input data 1113 and analyzes the metadata. The metadata includes information describing elements that act on the sound and are arranged in the sound space. The spatial information management unit 1201 manages the spatial information used in the sound processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1203.
[0140] In addition, in the present disclosure, the information used in the sound processing is expressed as spatial information, but other expressions may be used. For example, the information used in the sound processing may be expressed as sound spatial information or scene information. In addition, when the information used in the sound processing changes over time, the spatial information input to the rendering unit 1203 may be information expressed as a spatial state, a sound spatial state, or a scene state.
[0141] In addition, spatial information can be managed for each sound space or each scene. For example, when multiple different rooms are represented as virtual spaces, the multiple rooms can be managed as multiple different scenes. In addition, even in the same space, spatial information can be managed as different scenes according to the conditions represented.
[0142] Therefore, a plurality of spatial information may be managed for a plurality of sound spaces or a plurality of scenes. In the management of a plurality of spatial information, identifiers for identifying the plurality of spatial information may be given to the spatial information.
[0143] The spatial information data may also be included in a bitstream as an example of input data 1113. Alternatively, the bitstream may include an identifier of the spatial information, and the spatial information data may be obtained from an information source other than the bitstream. Specifically, when the bitstream includes only an identifier of the spatial information, the spatial information data stored in a memory in the device or an external server may be obtained as input data 1113 during rendering using the identifier of the spatial information.
[0144] The information managed by the space information management unit 1201 is not limited to the information included in the bitstream. For example, the input data 1113 may include data representing the characteristics and structure of the space obtained from software or a server providing VR or AR as data not included in the bitstream.
[0145] In addition, the input data 1113 may also include data indicating characteristics and positions of the listener or object, etc. In addition, the input data 1113 may also include information about the position of the listener obtained by a sensor included in the terminal including the decoding device (1110, 1130), and may also include information indicating the position of the terminal estimated based on the information obtained by the sensor.
[0146] That is, the spatial information management unit 1201 may communicate with an external system or server to obtain spatial information and listener positions. The spatial information management unit 1201 may obtain clock synchronization information from an external system and perform processing synchronized with the clock of the rendering unit 1203 .
[0147] In addition, the space in the above description can be a virtually formed space, i.e., a VR space, or a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR space. In addition, the virtual space can also be represented as a sound field or a sound space. In addition, the information indicating the position in the above description can be information such as coordinate values indicating the position in the space, information indicating the relative position relative to a predetermined reference position, or information indicating the movement or acceleration of the position in the space.
[0148] The audio data decoder 1202 decodes the encoded audio data included in the input data 1113 to obtain an audio signal.
[0149] The coded audio data obtained by the stereo sound reproduction system 1000 is, for example, a bit stream coded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3) or the like. MPEG-H 3D Audio is merely an example of a coding method that can be used when generating coded audio data included in a bit stream. The coded audio data may also be a bit stream coded in another coding method.
[0150] For example, the encoding method may be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method may be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0151] Alternatively, any encoding method other than the above may be used. For example, PCM (pulse code modulation) data may be a type of encoded sound data. In this case, for example, when the number of quantization bits of the PCM data is N, the decoding process may be a process of converting an N-bit binary number into a number format (e.g., floating point format) that can be processed by the rendering unit 1203.
[0152] The rendering unit 1203 acquires the sound signal and the spatial information, performs acoustic processing on the sound signal using the spatial information, and outputs the sound signal after the acoustic processing (sound signal 1111 ).
[0153] Before starting rendering, the spatial information management unit 1201 reads metadata of the input signal, detects rendering items such as objects and sounds specified by the spatial information, and sends them to the rendering unit 1203. After starting rendering, the spatial information management unit 1201 grasps the changes in the spatial information and the position of the listener over time, updates and manages the spatial information, and sends the updated spatial information to the rendering unit 1203.
[0154] The rendering unit 1203 generates and outputs a sound signal to which acoustic processing has been added based on the sound signal included in the input data 1113 and the space information received from the space information management unit 1201 .
[0155] The updating process of the spatial information and the output process of the sound signal with the added sound processing may be executed by the same thread. The spatial information management unit 1201 and the rendering unit 1203 may allocate the processing to their own independent threads. When the spatial information management unit 1201 and the rendering unit 1203 execute the updating process of the spatial information and the output process of the sound signal with the added sound processing in different threads, the startup frequency of the threads may be set separately, or the processing may be executed in parallel.
[0156] When the spatial information management unit 1201 and the rendering unit 1203 execute processing in different independent threads, computing resources can be preferentially allocated to the rendering unit 1203. This makes it possible to safely perform sound output processing that does not allow for even a slight delay, such as producing a noise such as a popping sound with a delay of 1 sample (0.02 msec).
[0157] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared with the output processing of the sound signal, the update of the spatial information is a low-frequency process (for example, the update of the face orientation of the listener), so it does not need to be performed instantly like the output processing of the sound signal. Therefore, even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality.
[0158] The updating of spatial information may be performed regularly at a preset time or period, or when a preset condition is satisfied. In addition, the updating of spatial information may be performed manually by the listener or the manager of the sound space, or may be triggered by a change in an external system.
[0159] For example, the listener may operate the controller to update the space information when the standing position of the listener's avatar is instantly twisted, or when the listener moves forward or backward. Alternatively, the space information may be updated when the manager of the virtual space performs a performance in which the environment of the venue is suddenly changed. In these cases, the thread for updating the space information managed by the space information management unit 1201 may be started as a one-shot interrupt process in addition to being started regularly.
[0160] For example, the updating process of the spatial information managed by the spatial information management unit 1201 is executed by the information updating thread.
[0161] The role of the information update thread is, for example, to update the position and orientation of the listener's avatar arranged in the virtual space based on the position and orientation of the VR goggles worn by the listener, or to update the position of objects moving in the virtual space, etc. Such processing is provided in the processing thread that is started at a relatively low frequency of about tens of Hz.
[0162] In such a processing thread with a low occurrence frequency, the processing of updating the information representing the properties of the direct sound can also be performed. The reason is that the frequency of changes in the properties of the direct sound is lower than the frequency of generation of the audio processing frame used for audio output. As a result, the computational load of the processing can be relatively reduced. In addition, if the information is updated at an unnecessarily fast frequency, there is a risk of generating impulse noise. By updating the information at a low frequency, such a risk can also be avoided.
[0163] Figure 4B is a block diagram showing another example of the configuration of a decoder. Specifically, Figure 4B Indicate as Figure 3B or Figure 3D The structure of decoder 1210 is another example of decoder 1112 in FIG.
[0164] Figure 4B The input data 1113 does not include coded audio data but includes an uncoded audio signal. Figure 4A The input data 1113 includes a bit stream including metadata and an audio signal.
[0165] The spatial information management unit 1211 is related to Figure 4A Since the spatial information management unit 1201 is the same as that of the spatial information management unit 1201, the description thereof will be omitted.
[0166] The rendering unit 1213 is related to Figure 4A The rendering unit 1203 is the same as that of FIG. 1 , so the description is omitted.
[0167] In addition, the decoders 1112, 1200, and 1210 may be expressed as an audio processing unit that performs audio processing. In addition, the decoding devices 1110 and 1130 may be the audio signal processing device 1001, or may be expressed as an audio processing device.
[0168] (Physical Structure of Sound Signal Processing Device)
[0169] Figure 5 1 is a diagram showing an example of the physical structure of the sound signal processing device 1001. Figure 5 The sound signal processing device 1001 may also be Figure 3B The decoding device 1110 or Figure 3D Decoding device 1130. Figure 3B or Figure 3D The multiple components shown can also be Figure 5 The multiple components shown are installed. In addition, part of the components described here can also be equipped in the sound prompting device 1002.
[0170] Figure 5 The sound signal processing device 1001 includes a processor 1402 , a memory 1404 , a communication IF (Interface) 1403 , a sensor 1405 , and a speaker 1401 .
[0171] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The audio processing or decoding processing of the present disclosure may be implemented by executing a program stored in the memory 1404 by the CPU, DSP or GPU. In addition, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 may also be a dedicated circuit that performs signal processing of sound signals including the audio processing of the present disclosure.
[0172] The memory 1404 is composed of, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory). The memory 1404 may also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1404 may also be an internal memory built into a CPU or a GPU. In addition, the memory 1404 may also store spatial information managed by the spatial information management unit (1201, 1211). In addition, threshold data described later may also be stored.
[0173] The communication IF 1403 is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices via the communication IF 1403 to obtain a bit stream to be decoded. The obtained bit stream is stored in the memory 1404, for example.
[0174] The communication IF 1403 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and Wi-Fi (registered trademark), and may also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark).
[0175] The communication method is not limited to the wireless communication method described above, but may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0176] The sensor 1405 performs sensing for estimating the position and orientation of the listener. Specifically, the sensor 1405 estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, speed, angular velocity, and acceleration of a part or the whole of the body, and generates position / orientation information indicating the position and / or orientation of the listener.
[0177] Alternatively, a device external to the sound signal processing device 1001 may include the sensor 1405. The body part may be the head of the listener, etc. The position / orientation information may be information indicating the position and / or orientation of the listener in real space, or information indicating the displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time point. Alternatively, the position / orientation information may be information indicating the relative position and / or orientation to the stereophonic reproduction system 1000 or the external device including the sensor 1405.
[0178] The sensor 1405 is, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Light Detection And Ranging). The sensor 1405 may also capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. In addition, a device that performs position estimation using wireless in any frequency band such as millimeter waves may also be used as the sensor 1405.
[0179] In addition, the sound signal processing device 1001 may obtain the position information from an external device having the sensor 1405 via the communication IF 1403. In this case, the sound signal processing device 1001 may not include the sensor 1405. Here, the external device is, for example, Figure 2 The sensor 1405 may be a combination of various sensors such as a gyro sensor and an acceleration sensor, or the like.
[0180] For example, as the speed of movement of the listener's head, the sensor 1405 can detect the angular velocity of rotation about at least one of the three axes orthogonal to each other in the sound space, or can detect the acceleration of displacement in the direction of displacement about at least one of the three axes.
[0181] For example, as the amount of movement of the listener's head, the sensor 1405 can detect the amount of rotation with at least one of the three axes orthogonal to each other in the sound space as the rotation axis, or can detect the amount of displacement with at least one of the three axes as the displacement direction. Specifically, the sensor 1405 detects the position (x, y, z) and angle (yaw, pitch, roll) of 6DoF as the position of the listener. The sensor 1405 is composed of a combination of various sensors for motion detection, such as a gyro sensor and an acceleration sensor.
[0182] The sensor 1405 may be realized by a camera or a GPS (Global Positioning System) receiver for detecting the position of the listener. Position information obtained by performing self-position estimation using LiDAR or the like as the sensor 1405 may also be used. For example, when the stereo sound reproduction system 1000 is realized by a smartphone, the sensor 1405 is built into the smartphone.
[0183] The sensor 1405 may include a temperature sensor such as a thermocouple for detecting the temperature of the sound signal processing device 1001. The sensor 1405 may include a battery included in the sound signal processing device 1001 or a sensor for detecting the remaining amount of a battery connected to the sound signal processing device 1001.
[0184] The speaker 1401 has a driving mechanism such as a vibration plate, a magnet or a voice coil, and an amplifier, and presents the sound signal after the sound processing to the listener as a sound. The speaker 1401 operates the driving mechanism according to the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), and the driving mechanism vibrates the vibration plate. In this way, the vibration plate vibrates according to the sound signal to generate sound waves, which propagate in the air and are transmitted to the listener's ears, and the listener perceives the sound.
[0185] In addition, although the example in which the sound signal processing device 1001 includes the speaker 1401 and presents the sound signal after the audio processing via the speaker 1401 is shown here, the presentation mechanism of the sound signal is not limited to the above-mentioned configuration.
[0186] For example, the sound signal after the sound processing can also be output to the external sound prompt device 1002 connected through the communication module. The communication performed through the communication module can be either wired or wireless. In addition, as another example, the sound signal processing device 1001 has a terminal for outputting an analog signal of sound, and a cable such as an earplug is connected to the terminal to prompt the sound signal from the earplug.
[0187] In the above case, the sound prompt device 1002 may also be a headset, earplug, head-mounted display, neck speaker, or wearable speaker, etc., which is worn on the head or a part of the body of the listener. Alternatively, the sound prompt device 1002 may also be a surround speaker composed of a plurality of fixed speakers, etc. Furthermore, the sound prompt device 1002 may also reproduce a sound signal.
[0188] (Physical Structure of Encoding Device)
[0189] Figure 6 This is a diagram showing an example of the physical structure of an encoding device. Figure 6 The encoding device 1500 may also be Figure 3A The encoding device 1100 or Figure 3C The encoding device 1120 can also Figure 3A or Figure 3C The multiple components shown are Figure 6 The multiple components shown are installed.
[0190] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503 and a communication IF 1502.
[0191] The processor 1501 is, for example, a CPU, a DSP, or a GPU. The encoding process of the present disclosure may also be implemented by executing a program stored in the memory 1503 by the CPU, DSP, or GPU. In addition, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 may also be a dedicated circuit that performs signal processing of a sound signal including the encoding process of the present disclosure.
[0192] The memory 1503 is composed of, for example, a RAM or a ROM. The memory 1503 may also include a magnetic recording medium represented by a hard disk or a semiconductor memory represented by an SSD. In addition, the memory 1503 may also be an internal memory embedded in a CPU or a GPU.
[0193] The communication IF 1502 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WIGI (registered trademark). The encoding device 1500 communicates with another communication device via the communication IF 1502, for example, and transmits an encoded bit stream.
[0194] Communication IF1502 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and Wi-Fi (registered trademark), and may also be LTE, NR, or Wi-Fi (registered trademark). In addition, the communication method is not limited to the wireless communication method. The communication method may also be a wired communication method such as Ethernet (registered trademark), USB, or HDMI (registered trademark).
[0195] (Composition of the rendering unit)
[0196] Figure 7 is a block diagram showing an example of the configuration of a rendering unit. Specifically, Figure 7 Representation and Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213 .
[0197] The rendering unit 1300 is composed of an analyzing unit 1301 , a selecting unit 1302 , and a synthesizing unit 1303 , and performs acoustic processing on audio data included in an input signal and outputs the resultant signal.
[0198] The input signal is composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bit stream composed of sound data and metadata (control information), and in this case, the metadata may also include spatial information.
[0199] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000, and is composed of information related to objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound emitting objects that do not emit sound. The sound source object can also be simply expressed as a sound source.
[0200] A non-sound-generating object acts as an obstacle object that reflects the sound emitted by a sound source object, but a sound source object may also act as an obstacle object that reflects the sound emitted by another sound source object. An obstacle object may also be represented as a reflecting object.
[0201] The information given to both the sound source object and the non-sound emitting object includes position information, shape information, and a volume attenuation rate when the object reflects sound.
[0202] The position information is represented by the coordinate values of three axes, such as the X-axis, the Y-axis, and the Z-axis, in the Euclidean space, but it is not necessarily three-dimensional information. For example, the position information may be two-dimensional information represented by the coordinate values of two axes, the X-axis and the Y-axis. The position information of the object is determined by the representative position of the shape represented by the grid or voxel.
[0203] The shape information may also include information about the material of the surface.
[0204] The attenuation rate can be expressed by a real number greater than 0 and less than 1, or by a negative decibel value. Since the volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value, but for example, in order to present a sense of horror in an unreal space, an attenuation rate greater than 1, that is, a positive decibel value, can be deliberately set.
[0205] In addition, the attenuation rate may be set to a different value for each frequency band constituting the plurality of frequency bands, or may be set to a value independently for each frequency band. In addition, when the attenuation rate is set for each type of material on the surface of the object, the corresponding attenuation rate value may be used based on information related to the material on the surface.
[0206] In addition, the spatial information may also include information indicating whether the object is a living being and information indicating whether the object is a moving body. If the object is a moving body, the position indicated by the position information may also move over time. In this case, the information of the changed position or the amount of change is transmitted to the rendering unit 1300.
[0207] The information related to the sound source object includes not only the information given to the sound source object and the non-sound-generating object, but also the sound data and the information required to radiate the sound data into the sound space. The sound data is data representing information related to the frequency and strength of the sound, and is data representing the sound perceived by the listener.
[0208] The audio data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal needs to be decoded at least before it reaches the synthesizer 1303, so the rendering unit 1300 may also include a decoding unit not shown. Alternatively, the signal may also be decoded by the audio data decoder 1202.
[0209] For one sound source object, one sound data or a plurality of sound data may be set. In addition, identification information for identifying each sound data may be given to the sound data, and the information related to the sound source object may also include the identification information of the sound data.
[0210] The information required to radiate sound data into the sound space may include, for example, information on a reference volume used as a reference in the reproduction of the sound data, information related to the position of the sound source object, and information related to the orientation of the sound source object (i.e., information related to the directionality of the sound emitted by the sound source object).
[0211] The reference volume information is, for example, an effective value of the amplitude value of the sound data at the sound source position when the sound data is radiated into the sound space, and may be expressed as a decibel (db) value in a floating point.
[0212] For example, when the reference volume is 0db, it may be indicated that the sound is radiated to the sound space from the position indicated by the information related to the position of the sound source object without increasing or decreasing the volume of the signal level indicated by the sound data. Also, when the reference volume is -6db, it may be indicated that the sound is radiated to the sound space from the position indicated by the information related to the position of the sound source object with the volume of the signal level indicated by the sound data being set to about half.
[0213] The information of the reference volume may be given to each sound data or may be given collectively to a plurality of sound data.
[0214] Information necessary for radiating audio data into the audio space may include, for example, information indicating a time-series change in the volume of a sound source as volume information.
[0215] For example, when the sound space is a virtual conference room and the sound source is a speaker, the volume changes intermittently in a short period of time. That is, the sound part and the silent part are produced alternately. In addition, when the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain length of time. In addition, when the sound space is a battlefield and the sound source is an explosive, the volume of the explosion sound increases only for a moment, and then continues to be silent or small.
[0216] In this way, the information of the volume of the sound source includes not only the information of the size of the sound but also the information of the transition of the size of the sound. Such information can also be used as information indicating the nature of the sound data.
[0217] The information of the transition may also be represented by data representing the frequency characteristics in a time series. The information of the transition may also be represented by data representing the duration of the voiced interval. The information of the transition may also be represented by data representing the duration of the voiced interval and the duration of the silent interval in a time series. The information of the transition may also be represented by data that lists in a time series a duration during which the amplitude of the sound signal can be considered stable (can be considered approximately constant) and a plurality of sets of amplitude values of the signal during the period.
[0218] The information of the transition may also be expressed by data of a duration during which the frequency characteristic of the sound signal can be regarded as stable. The information of the transition may also be expressed by data of a duration during which the frequency characteristic of the sound signal can be regarded as stable and a plurality of sets of the frequency characteristic during the duration listed in a time series. The information of the transition may also be expressed in the form of data representing the general shape of a spectrogram, for example.
[0219] In addition, the volume used as a reference for the frequency characteristics may also be the reference volume. The reference volume information and the information indicating the nature of the sound data may be used for calculation processing of the volume of the direct sound or the reflected sound perceived by the listener, and may also be used for selection processing of whether to make the listener perceive it. Other examples of the information indicating the volume and the method of using it will be described later.
[0220] The information related to the orientation of the sound source object (orientation information) is typically expressed by yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted and the orientation information of the sound source object may be expressed by azimuth (yaw) and elevation (pitch). The orientation information of the sound source object may also change over time and, if changed, is transmitted to the rendering unit 1300.
[0221] Information related to the listener is information related to the position and orientation of the listener in the sound space. Information related to the position (position information) is represented by the position of the XYZ axis of the Euclidean space, but it is not necessarily three-dimensional information and can also be two-dimensional information. Information related to the orientation of the listener (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted and the orientation information of the listener can be represented by azimuth (yaw) and elevation (pitch).
[0222] The position information and orientation information of the listener may also change over time, and if changed, the information is transmitted to the rendering unit 1300 .
[0223] The sensor information is information including the amount of rotation or displacement detected by the sensor 1405 worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering unit 1300, and the rendering unit 1300 updates the information on the position and orientation of the listener based on the sensor information. The sensor information may also include, for example, location information obtained by the portable terminal through self-position estimation using GPS, camera, LiDAR, etc.
[0224] In addition, instead of the sensor 1405, information obtained from the outside via the communication module may be detected as sensor information. Information indicating the temperature of the sound signal processing device 1001 and information indicating the remaining amount of the battery may be obtained from the sensor 1405. In addition, the computing resources (CPU capacity, memory resources, PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 may be obtained in real time.
[0225] The analyzing unit 1301 analyzes the sound signal included in the input signal and the space information received from the space information management unit (1201, 1211), and detects information required for generating direct sound and reflected sound, and information required for selecting whether to generate reflected sound.
[0226] The information required for the generation of direct sound and reflected sound is, for example, information on the characteristics of direct sound and reflected sound that may be generated in the sound space. The reflected sound detected here is a reflected sound candidate selected by the selection unit 1302 as the reflected sound to be finally generated by the synthesis unit 1303. The characteristics of direct sound and reflected sound refer to, for example, the arrival time (arrival moment) and the volume when the direct sound and reflected sound reach the listener respectively. In the case where multiple objects exist in the sound space as reflection objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0227] The information required for selecting the reflected sound to be output may be, for example, information indicating the evaluation value of the reflected sound and the upper limit of the computing resources, or information used to calculate the information indicating the evaluation value of the reflected sound and the upper limit of the computing resources. That is, the analysis unit 1301 may obtain the evaluation value of the reflected sound from an external device, a storage unit, or an input signal. Alternatively, the analysis unit 1301 or the selection unit 1302 may calculate the information indicating the evaluation value of the reflected sound and the upper limit of the computing resources using the information obtained by the analysis unit 1301 from an external device, a storage unit, or an input signal.
[0228] The selection unit 1302 determines whether to select the reflected sound based on the evaluation value of the reflected sound. That is, the selection unit 1302 selects the reflected sound with a high evaluation value more preferentially than the reflected sound with a low evaluation value. The evaluation value of the reflected sound is the value of the reflected sound, and corresponds to, for example, the perceived importance of the reflected sound. The higher the perceived importance of the reflected sound, the higher the evaluation value. The perceived importance of the reflected sound refers to the degree of necessity of the reflected sound used in order for the listener to correctly grasp the positioning of the sound source object in the sound space and the width of the space.
[0229] By preferentially selecting and processing reflected sounds with high evaluation values, i.e., reflected sounds with high perceptual importance, the listener can understand the localization of the sound image, such as the direction of arrival of the sound and the sense of distance to the sound source object, and understand the size and texture of the space.
[0230] Furthermore, by deciding not to select reflected sounds before starting the reflected sound generation process, it is possible to decide not to perform the processing after the processing of imparting acoustic effects to the reflected sounds. Therefore, compared with the case where it is decided whether to perform binaural processing after imparting acoustic effects to all detected reflected sounds, or the case where binaural processing is performed on all detected reflected sounds, the computational load can be reduced.
[0231] That is, by determining the reflected sounds not to be selected based on the perceptual importance of the reflected sounds, it is possible to reduce the calculation load used in generating the reflected sounds while preventing the listener's understanding of the sound localization and space from being impaired.
[0232] The selection unit 1302 evaluates the perceived importance of the reflected sound, and calculates the evaluation value, for example, based on the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflection object (obstacle object), information about the material of the reflection object, and the geometric relationship between the direct sound and the reflected sound. Other indicators may also be used to evaluate the perceived importance of the reflected sound. The evaluation value of the reflected sound may be calculated based on any one of a plurality of indicators related to the perceived importance of the reflected sound, or the evaluation value of the reflected sound may be calculated comprehensively using a plurality of indicators.
[0233] In addition, the selection unit 1302 may obtain the evaluation value of the reflected sound from an external device or a storage unit, or may obtain the evaluation value from the input signal.
[0234] Specifically, the louder the sound source volume, the higher the evaluation value. In addition, in order to make the visual positioning consistent with the sound positioning, the evaluation value may be high when the sound source object or the reflection object (obstacle object) can be identified by the listener, or when the sound source object has high localizability.
[0235] In addition, the opening of the arrival angle of direct sound and reflected sound, and the difference in arrival time of direct sound and reflected sound have a great influence on the grasp of space. Therefore, the evaluation value may be high when the opening of the arrival angle of direct sound and reflected sound is large, or when the difference in arrival time of direct sound and reflected sound is large.
[0236] Alternatively, the evaluation value of the reflected sound may be calculated using information on the difference in arrival time between the direct sound and the reflected sound. In this case, for example, a masking threshold in the well-known phenomenon of temporal masking (post-masking) may be used.
[0237] The specific reflected sound evaluation method and selection method of the selection unit 1302 will be described later.
[0238] The synthesizing unit 1303 synthesizes the audio signal of the direct sound and the audio signal of the reflected sound selected by the selecting unit 1302 to be generated.
[0239] Specifically, the synthesizing unit 1303 processes the input sound signal to generate the direct sound based on the information of the arrival time of the direct sound and the volume of the direct sound when it arrives calculated by the analyzing unit 1301. In addition, the synthesizing unit 1303 processes the input sound signal to generate the reflected sound based on the information of the arrival time of the reflected sound and the volume of the reflected sound when it arrives selected by the selecting unit 1302. Then, the synthesizing unit 1303 synthesizes the generated direct sound and reflected sound and outputs them.
[0240] (Rendering unit actions)
[0241] Figure 8 1 is a flowchart showing an example of the operation of the audio signal processing device 1001. Figure 8 2 shows the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001.
[0242] (Detection of direct sound and reflected sound)
[0243] In the analysis of the input signal ( Figure 8In S101 of FIG. 1 , the analyzing unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct sound and reflected sound that can be generated in the sound space. The reflected sound detected here is a reflected sound candidate selected by the selecting unit 1302 as the reflected sound to be finally generated by the synthesizing unit 1303. In addition, the analyzing unit 1301 analyzes the input signal and calculates the information required for generating the direct sound and the reflected sound and the information required for selecting the generated target reflected sound.
[0244] First, the characteristics of the direct sound and the reflected sound are calculated. Specifically, the arrival time and volume of the direct sound and the reflected sound when they reach the listener are calculated. When there are multiple objects in the sound space as reflection objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0245] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is the path connecting the position information S (xs, ys, zs) of the sound source object and the position information A (xa, ya, za) of the listener. The direct sound arrival time (td) is the value obtained by dividing the length of the path connecting the position information S (xs, ys, zs) and the position information A (xa, ya, za) by the speed of sound (approximately 340 m / sec).
[0246] For example, the length of the path (X) is calculated by (xs-xa)^2+(ys-ya)^2+(zs-za)^2)^0.5. The volume attenuates inversely proportional to the distance. Therefore, when the volume in the location information S(xs, ys, zs) of the sound source object is N and the unit distance is U, the volume (ld) when the direct sound arrives is calculated by ld=N*U / X.
[0247] The volume N at the sound source position may be the reference volume described above.
[0248] The reflected sound arrival time (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is a path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za).
[0249] In addition, the position of the sound image of the reflected sound can be derived using, for example, the "mirror method" or the "ray tracing method", or any other method for deriving the sound image position. The mirror method is a method of assuming that the reflected wave on the wall surface in the room has a mirror image at a position symmetrical to the wall surface and the sound source, and assuming that the sound wave is radiated from the position of the mirror image to simulate the sound image. The ray tracing method is a method of simulating an image (sound image) observed at a certain point by tracing a wave propagating in a straight line such as a light ray or a sound line.
[0250] Fig. 9This diagram shows the positional relationship between the listener and the obstacle object, which is relatively far away. Fig.10 is a diagram showing the positional relationship between the listener and the obstacle object. Fig. 9 and Fig.10 Each shows an example of forming a sound image of reflected sound at a position symmetrical to the sound source position across a wall. By finding the position of the sound image of the reflected sound on the xyz axis based on such a relationship, the arrival time of the reflected sound can be found in the same way as the method of calculating the arrival time of the direct sound.
[0251] The arrival time of the reflected sound (tr) is the value obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za) by the speed of sound (about 340 m / sec). The volume decays inversely with the distance. Therefore, when the volume at the sound source position is N, the unit distance is U, and the attenuation rate of the volume in the reflection is G, the volume (lr) when the reflected sound arrives is calculated by lr = N*G*U / Y.
[0252] As described above, the attenuation rate G can be expressed by a real number greater than 0 and less than 1, or by a negative decibel value. In this case, the volume of the entire signal is attenuated by an amount corresponding to G. In addition, the attenuation rate can also be set for each frequency band constituting a plurality of frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate to each frequency component of the signal. In addition, in order to reduce the amount of calculation, the analysis unit 1301 can also use a representative value or an average value of a plurality of attenuation rates of a plurality of frequency bands as the overall attenuation rate, so that the volume of the entire signal is attenuated accordingly.
[0253] (Selective processing of reflected sound)
[0254] Next, in the selection process of reflected sound ( Figure 8 In S102 of FIG. 1 , the selection unit 1302 selects whether to generate the reflected sound calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether to select the reflected sound as the generated object reflected sound. In the case where there are multiple reflected sounds, the selection unit 1302 selects whether to generate each reflected sound. As a result of the selection unit 1302 selecting whether to generate each reflected sound, more than one generated object reflected sound may be selected from the multiple reflected sounds, or none of the generated object reflected sounds may be selected.
[0255] In addition, the selection unit 1302 is not limited to the generation process, and may also select the reflected sound of the object to which other processes are applied. For example, the selection unit 1302 may also select the reflected sound of the object to which binaural processing is applied. In addition, the selection unit 1302 basically selects only one or more reflected sounds of the processing object. However, the selection unit 1302 may also select only one or more reflected sounds that are not the processing object. Furthermore, the processing may also be applied to one or more reflected sounds that are not selected.
[0256] For example, the selection of the reflected sound is performed based on the permissible computational load and the perceived importance of the reflected sound. Fig.11 The flowchart of will explain the process of selecting and processing the reflected sound.
[0257] Fig.11 2 is a flowchart showing an example of the selection process of the reflected sound. In this example, the selection process is performed based on the calculation load and the perceptual importance of the reflected sound, but the selection process may be performed based on only one of them.
[0258] (Acquisition of information indicating the upper limit of computation load)
[0259] First, the selection unit 1302 obtains information indicating the upper limit of the calculation load in the audio signal processing device 1001 (S201). The information indicating the upper limit of the calculation load may be determined in advance by the listener or may be obtained from the input signal.
[0260] Here, the information indicating the upper limit of the computational load may indicate the number of (one or more) reflected sounds as the upper limit, or may indicate the processing amount of (one or more) reflected sounds. When the information indicating the upper limit of the computational load indicates that the number of reflected sounds is the upper limit, the predicted value of the computational load of the reflected sound candidate described later also uses the predicted value of the number of reflected sounds, and thus the processing amount of the selection unit 1302 can be reduced compared to calculating the predicted value of the processing amount of the reflected sounds.
[0261] When the information indicating the upper limit of the computational load indicates that the processing amount of the reflected sound is the upper limit, the predicted value of the computational load of the reflected sound candidate described later also uses the predicted value of the processing amount of the reflected sound, so that a more accurate computational load can be predicted. In addition, the processing amount of (one or more) reflected sounds is, for example, the processing amount required for generating (one or more) reflected sounds, which is the total computational load required for processing to generate (one or more) reflected sounds.
[0262] The reflected sound processing is, for example, processing for generating reflected sound, and is a processing included in pipeline processing, which includes, for example, reverberation processing, initial reflection processing, distance attenuation processing, binaural processing, diffraction processing, and occlusion processing.
[0263] However, these processes are examples, and pipeline processing may include processes other than these or may not include some of the processes. For example, the rendering unit 1300 may perform diffraction processing and occlusion processing as pipeline processing. In addition, for example, if reverberation processing is not required, reverberation processing may be omitted.
[0264] The information indicating the upper limit of the computational load may also be determined based on the computational resources (CPU capability, memory resources, PC performance, or battery level, etc.) of the sound signal processing device 1001 or the sound prompting device 1002. For example, generally speaking, the processing power of the CPU increases in the order of head mounted display, VR / AR goggles, smartphone, notebook PC, desktop PC, and supercomputer, so the upper limit of the computational load may also be set to increase in the same order.
[0265] In addition, the selection unit 1302 may also obtain information indicating the temperature of the device or information indicating the remaining battery level from the sensor 1405 provided in the sound signal processing device 1001 or the sound prompting device 1002. In addition, the selection unit 1302 may also obtain the computing resources (CPU capacity, memory resources, PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 in real time.
[0266] In the above case, the selection unit 1302 may acquire the information indicating the upper limit of the calculation load in real time, or may acquire the information regularly at each timing when the space information management unit (1201, 1211) updates the space information.
[0267] Furthermore, the information indicating the upper limit of the calculation load may be set according to the battery life of the audio signal processing device 1001 or the audio prompting device 1002 .
[0268] Alternatively, the upper limit of the computation load may be set for each mode, such as a "power saving mode" in which the amount of computation is small and the device can be used for a long time, or a "high performance mode" in which the amount of computation is large but more reflected sounds can be heard. In this case, the listener, the administrator who manages the stereo reproduction system 1000, or the producer of the stereo content may specify the desired battery life or the desired mode. Alternatively, the upper limit of the computation load may be directly input without selecting a mode.
[0269] In addition, information indicating the upper limit of the computational load may be set for each content reproduced by the stereo sound reproduction system 1000. For example, for content that is more important for immersion, the upper limit of the computational load may be set high, and more reflected sounds may be selected. For content that is important for real-time performance, the upper limit of the computational load may be set low in order to avoid delays associated with an increase in the amount of processing. This suppresses the selection of more reflected sounds.
[0270] The input signal including the content may include information indicating the upper limit of the computation load. In addition, the selection unit 1302 may determine the upper limit of the computation load based on information indicating the type of content or the type of mode included in the input signal. Alternatively, without being limited to the information indicating the type of content or the type of mode, the selection unit 1302 may determine the upper limit of the computation load based on other flags or parameters included in the input signal.
[0271] (Extraction of reflected sound with a volume above a threshold)
[0272] Next, the selection unit 1302 extracts as selection candidates (S202) the (one or more) reflected sounds whose arrival volume is greater than the threshold value from among the (one or more) reflected sounds detected by the analysis unit 1301. That is, the selection unit 1302 decides not to perform subsequent processing on the (one or more) reflected sounds whose arrival volume is less than the threshold value.
[0273] In addition, when the volume of the direct sound is less than the threshold value, the selection unit 1302 may not extract the reflected sound caused by the direct sound. The volume of the reflected sound is smaller than the volume of the direct sound. Therefore, when the volume of the direct sound is less than the threshold value, the volume of the reflected sound caused by the direct sound is also less than the threshold value.
[0274] Therefore, the selection unit 1302 may extract the reflected sound having an arrival volume greater than the threshold value from the reflected sound caused by the direct sound having an arrival volume greater than the threshold value.
[0275] That is, the selection unit 1302 may first compare the volume of the direct sound at the time of arrival with the threshold. Thus, when the volume of the direct sound at the time of arrival is less than the threshold, it can be determined not to extract the multiple reflected sounds caused by the direct sound. Therefore, compared with the case where the volume of the reflected sound at the time of arrival is calculated for each of the multiple reflected sounds caused by the direct sound and whether to extract the reflected sound is determined, the amount of calculation can be reduced.
[0276] Here, the threshold value compared with the volume of the direct sound or the reflected sound at the time of arrival may also be the minimum volume reproduced in the sound space. In other words, the threshold value may also be the minimum audible limit of the volume indicating the boundary of whether the sound can be perceived by the listener. Moreover, for example, a sound with a volume lower than the threshold value may be regarded as a sound that cannot be perceived by the listener and not reproduced in the virtual space.
[0277] In addition, the threshold value can be determined in advance by the listener or obtained from the input signal. The threshold value of the volume at the time of arrival can also be determined based on the computing resources (CPU power, memory resources, PC performance or battery remaining, etc.) of the sound signal processing device 1001 or the sound prompt device 1002. For example, generally in the order of head-mounted display, VR / AR goggles, smartphones, notebook PCs, desktop PCs, and supercomputers, the processing power of the CPU increases, so the threshold value of the volume at the time of arrival can also be set to increase in the same order.
[0278] In addition, the selection unit 1302 may also obtain information indicating the temperature of the device or information indicating the remaining battery level from the sensor 1405 provided in the sound signal processing device 1001 or the sound prompting device 1002. In addition, the selection unit 1302 may also obtain the computing resources (CPU capacity, memory resources, PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 in real time.
[0279] In the above case, the selection unit 1302 may acquire the threshold value of the volume at the time of arrival in real time, or may acquire it periodically at each timing when the space information management unit (1201, 1211) updates the space information.
[0280] In addition, the threshold of the volume at the time of arrival may also be set according to the battery life of the sound signal processing device 1001 or the sound prompting device 1002 .
[0281] Alternatively, the threshold value of the volume at the time of arrival may be set for each mode, such as a "power saving mode" in which the amount of calculation is small and the device can be used for a long time, or a "high performance mode" in which the amount of calculation is large but more reflected sounds can be heard. In this case, the listener, the administrator who manages the stereo reproduction system 1000, or the producer of the stereo content may specify the desired battery duration or the desired mode. In addition, the threshold value of the volume at the time of arrival may be directly input without selecting a mode.
[0282] In addition, a threshold value of the volume at the time of arrival may be set for each content reproduced by the stereo sound reproduction system 1000. For example, for content that is more immersive, the threshold value of the volume at the time of arrival may be set higher to select more reflected sounds. For content that is important in terms of real-time performance, the threshold value of the volume at the time of arrival may also be set lower to avoid delays associated with an increase in the amount of processing. This suppresses the selection of more reflected sounds.
[0283] The input signal including the content may include a threshold value of the volume when it arrives. In addition, the selection unit 1302 may determine the threshold value of the volume when it arrives based on information indicating the type of content or the type of mode included in the input signal. Alternatively, without being limited to the information indicating the type of content or the type of mode, the selection unit 1302 may determine the threshold value of the volume when it arrives based on other flags or parameters included in the input signal.
[0284] (Calculation of the predicted value of the total computing load)
[0285] Next, the selection unit 1302 calculates the predicted value of the total computational load of all reflected sounds extracted as selection candidates whose arrival volume is above the threshold (S203). Here, the predicted value of the computational load may be the number of (one or more) reflected sounds or the predicted value of the processing amount of (one or more) reflected sounds.
[0286] As for whether to use the predicted value of the number of reflected sounds or the predicted value of the amount of reflected sound processed as the predicted value of the computational load, it can be determined by indicating which of the number of reflected sounds or the amount of reflected sound processed as the upper limit based on the information indicating the upper limit of the computational load.
[0287] When the predicted value of the computational load is the number of reflected sounds, the amount of processing performed by the selection unit 1302 can be reduced compared to when the predicted value of the computational load is the predicted value of the amount of processing of the reflected sounds. When the predicted value of the computational load is the predicted value of the amount of processing of the reflected sounds, the computational load can be predicted more accurately by calculating the total amount of computation required to generate (one or more) reflected sounds.
[0288] As described above, the process of reflected sound is, for example, a process for generating reflected sound, and is a process included in the pipeline process.
[0289] Whether each process included in the pipeline processing is necessary or not differs depending on the nature of the reflected sound. Therefore, the predicted value of the amount of computation for the pipeline processing (ie, the amount of processing for one reflected sound) may differ for each reflected sound.
[0290] In addition, in order to reduce the processing load of the calculation amount for predicting pipeline processing, it is also possible to assume that the same processing is performed on each reflected sound and calculate the predicted value of the processing amount of all reflected sounds. In other words, the predicted value of the processing amount of all reflected sounds can be calculated by applying the same predicted value to the predicted value of the processing amount of each reflected sound.
[0291] In calculating the predicted value of the total calculation load, the predicted value of the total calculation load of a plurality of reflected sounds may be calculated, or the predicted value of the total calculation load of one reflected sound may be calculated.
[0292] In addition, the reflected sounds used to calculate the predicted value of the total computing load may be all the reflected sounds extracted as selection candidates, or may be only a part of the reflected sounds extracted as selection candidates. In the case where only a part of the reflected sounds extracted as selection candidates is used for calculating the predicted value of the total computing load, the predicted value of the number or processing amount of the part of the reflected sounds may be used as the predicted value of the total computing load.
[0293] (Comparison between the predicted value of the total computing load and the upper limit of the computing load)
[0294] Next, the selection unit 1302 compares the calculated predicted value of the total computing load with the upper limit of the computing load to determine whether the predicted value of the total computing load exceeds the upper limit of the computing load (S204). When the predicted value of the total computing load exceeds the upper limit of the computing load ("Yes" in S204), the selection unit 1302 performs selection processing based on the evaluation value (S205 to S211). When the predicted value of the total computing load does not exceed the upper limit of the computing load ("No" in S204), the selection unit 1302 selects all the reflected sounds extracted as selection candidates and ends the processing.
[0295] (Selection of reflected sound based on evaluation value)
[0296] In the selection process based on the evaluation value, the selection unit 1302 calculates the evaluation value of the reflected sound based on the perceived importance for each reflected sound of the selection candidate, and controls whether to select the reflected sound based on the evaluation value. For example, the selection unit 1302 selects the reflected sound in order from the reflected sound with the highest evaluation value. The specific method for calculating the evaluation value of the reflected sound will be described later. Here, an example of the selection process of selecting the reflected sound based on the evaluation value is described.
[0297] The selection unit 1302 performs, for example, a loop process of sequentially adding the computational loads of the selected reflected sounds, and terminates the selection process (S205 to S211) when the cumulative value of the computational loads of the selected one or more reflected sounds exceeds the upper limit of the computational load ("Yes" in S209), the selection unit 1302 determines the remaining reflected sounds that have not been determined as non-selected reflected sounds, and terminates the selection process.
[0298] Specifically, the selection unit 1302 first sets the count of the total computational load to zero (S205). In addition, the selection unit 1302 calculates the evaluation value of each extracted reflected sound (S206). Then, the selection unit 1302 determines to select the reflected sound with a high evaluation value (S207). In addition, the selection unit 1302 adds the computational load of the reflected sound determined to be selected to the total computational load (S208).
[0299] Then, when the total computation load exceeds the upper limit of the computation load ("Yes" in S209), the selection unit 1302 determines the remaining reflected sounds that have not been determined as reflected sounds that are not selected, and ends the selection process. In this case, the selection unit 1302 may also re-determine the reflected sounds that were last determined to be selected as not selected. In this way, the total computation load can be suppressed to below the upper limit of the computation load.
[0300] The processes after the selection process are applied to the reflected sounds determined as non-selected reflected sounds. That is, it is determined that the remaining reflected sounds are not generated.
[0301] If there are undecided or unselected reflected sounds among the reflected sounds extracted as selection candidates ("Yes" in S210), the selection unit 1302 repeats the process (S207 to S209). If there are no undecided reflected sounds ("No" in S210), the selection process ends.
[0302] In addition, for the selected reflected sound, a process may be performed to reduce the importance of the sound source object and the reflecting object that generate the reflected sound by a predetermined amount ( S211 ).
[0303] Thus, when the reflected sound caused by the sound source object and the reflection object is selected, other undetermined reflected sounds caused by the same sound source object or the same reflection object are unlikely to be selected in the next selection process. In other words, when no reflected sound caused by the sound source object and the reflection object is selected, any reflected sound caused by the sound source object or the reflection object is likely to be selected in the next selection process.
[0304] As a result, only the reflected sound caused by a specific sound source object or reflection object is suppressed from being selected, and only the presence of the specific sound source object or reflection object is suppressed from increasing while the presence of other sound source objects or reflection objects disappears in the sound space.
[0305] That is, when a certain reflected sound is selected, the value of the "sound source" of the reflected sound can also be reduced. As a result, in the next round, reflected sounds related to other sound sources are more likely to be selected. In addition, when a certain reflected sound is selected, the value of the "wall" (reflection object) that produces the reflected sound can also be reduced. As a result, in the next round, reflected sounds produced by other walls are more likely to be selected.
[0306] For example, if there are three sound sources (direct sounds) in a cubic room, theoretically 18 (3 x 6 faces) reflected sounds are generated. However, due to the problem of computational load, it is difficult to generate all the reflected sounds that are theoretically generated in the sound space. In the case where it is difficult to select all 18 reflected sounds, the reflected sounds are selected in a way that evenly reflects the influence of the three "sound sources" and six "walls" without deviation. In this way, the presence of the three "sound sources" and six "walls" can be maintained, and the amount of computation used to generate reflected sounds can be reduced.
[0307] In the above example, for example, the sound source objects are represented as X, Y, and Z, the walls are represented as R1 to R6, and the reflected sounds are represented as x1 to x6, y1 to y6, and z1 to z6. Furthermore, although there is only the amount of computation required to generate six reflected sounds, when the reflected sounds of x1 to x6 are selected, the presence of the sound source objects Y and Z in the sound space becomes lacking.
[0308] In addition, when the six reflected sounds x1, y1, z1, x2, y2, and z2 are selected, the actual presence of the walls R3 to R6 in the sound space is no longer reproduced. On the other hand, for example, when the volume of the sound source Y is almost zero, it is not so important to express the actual presence of the sound source Y. Therefore, the evaluation value of the reflected sounds y1 to y6 can also be low. In this way, when the computing resources are limited, the reflected sounds can be selected not randomly, but evenly selected based on the importance from the acoustic, auditory, and visual perspectives.
[0309] In addition, the method for determining the reflected sound to be selected is not limited to the method of determining in order from the reflected sound with the highest evaluation value. For example, the reflected sound with the evaluation value above the threshold value may be selected, and the reflected sound with the evaluation value below the threshold value may not be selected. In addition, the reflected sound of the layer with the high evaluation value may be selected at a predetermined ratio. Alternatively, the reflected sound of the layer with the low evaluation value may not be selected at a predetermined ratio. In these cases, the loop processing of sequentially adding the computational load of the reflected sound may not be performed.
[0310] (Evaluation Processing)
[0311] Fig.12 This is a flowchart showing an example of evaluation processing. Fig.12 The flowchart shown explains a specific method for determining the evaluation value.
[0312] The selection unit 1302 may calculate the evaluation value of the reflected sound by a pre-set evaluation method corresponding to the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.
[0313] Specifically, the selection unit 1302 acquires a plurality of reflected sounds extracted as selection candidates, and calculates an evaluation value of the reflected sound based on the perceptual importance of the reflected sound for each of the plurality of reflected sounds.
[0314] For example, it is also possible to assign evaluation points to the reflected sound respectively with respect to the multiple indicators described below, and assign evaluation values to the reflected sound based on the evaluation points. Of course, the multiple indicators used for evaluation are not limited to the multiple indicators described below. In addition, any one of the multiple indicators can be used, any two or more of the multiple indicators can be used, or all of the multiple indicators can be used. In addition, the order of evaluation related to the multiple indicators can also be determined based on the priority order of the predetermined indicators.
[0315] (Calculation of evaluation points)
[0316] Specifically, as an evaluation index of reflected sound, an index related to a sound source object may also be used. In addition, when a reflected sound caused by a sound source object is selected as described above, the value of the sound source object may also be reduced. Thus, reflected sounds caused by a specific sound source object are not emphasized, and reflected sounds caused by more sound source objects can be reproduced uniformly. Thus, clues for the listener to correctly perceive the positioning of each sound source can be ensured.
[0317] For example, when selecting 30 reflected sounds from 300 reflected sounds caused by 10 sound sources, selecting 30 reflected sounds caused by a specific sound source makes it difficult to grasp the location of other sound sources. In addition, assigning 3 reflected sounds to each of the 10 sound sources is not necessarily optimal. Therefore, it is also possible to assign an evaluation score to the reflected sound caused by the sound source object based on the importance of the sound source object or the importance of the direct sound emitted by the sound source object.
[0318] For example, as described later, the importance of the sound source object, that is, the importance of the direct sound, may be evaluated based on the audibility of the direct sound or the recognizability of the sound source object. This evaluation may also be used as an evaluation of the reflected sound caused by the direct sound generated from the sound source object. That is, based on the audibility of the direct sound or the recognizability of the sound source object, the reflected sound may be given an evaluation score of an index related to the sound source object.
[0319] In addition, the evaluation of the sound source object and the direct sound is not only used for the selection of the reflected sound, but can also be used for the selection of the direct sound.
[0320] The selection unit 1302 may also evaluate the sound source object by using the audibility of the direct sound, i.e., the audibility, and use the evaluation as an evaluation index for the reflected sound (S301). For example, the sound source object (direct sound) and the reflected sound may be given an evaluation score A obtained by evaluating the audibility using information related to the size of the direct sound.
[0321] Specifically, a sound source object with a loud volume may be assigned a higher evaluation score A than a sound source object with a small volume. Similarly, a reflected sound caused by a sound source object with a loud volume may be assigned a higher evaluation score A than a reflected sound caused by a sound source object with a small volume.
[0322] In addition, the size of a sound is usually determined by the volume or amplitude value, so of course the amplitude value can be used instead of the volume. That is, the information related to the size of the sound can be the volume (decibel value) or the amplitude value. Usually, the volume or amplitude value of the sound changes all the time, so the information related to the size of the sound used for evaluation can be the reference volume assigned to the sound source object, and of course it can also be information indicating the size of the sound that changes over time.
[0323] In addition, as information indicating the magnitude of the direct sound, both the information of the reference volume and the information of the volume that changes over time may be used. For example, after calculating the evaluation score of the sound source object based on the information of the reference volume, the evaluation score may be corrected using the information indicating the magnitude of the sound that changes, thereby calculating the evaluation score of the direct sound. Of course, after first calculating the evaluation score of the direct sound using the information indicating the magnitude of the sound that changes, the evaluation score of the direct sound may be corrected using the reference volume assigned to the sound source object.
[0324] Alternatively, the evaluation score of the sound source object (direct sound) may be calculated using only one of the information on the reference volume and the information on the volume that changes over time.
[0325] For example, when the virtual space is a virtual conference room and the direct sound is the sound of a conversation, the volume changes intermittently in a short period of time. That is, the sound part and the silent part are produced alternately. In addition, when the virtual space is a concert hall and the direct sound is the performance of a piece of music, the volume is maintained for a certain length of time. In addition, when the virtual space is a battlefield and the direct sound is an explosion sound, the volume only increases for a moment and then continues to be silent or low.
[0326] Thus, the information of the volume of the sound source includes not only the information of the volume of the sound, but also the information of the change of the volume of the sound. For example, the information may be information of a plurality of groups of the time lengths during which the volume is substantially constant and the volume values within the time interval, listed in time series.
[0327] In addition, measures for using the temporal change of the frequency characteristics of a signal for the sound processing of a virtual space have been widely used (Patent Document 1, etc.) In view of such prior art, the above-mentioned set may of course be a set of a time length with a constant frequency characteristic and its frequency characteristic.
[0328] The selection unit 1302 may evaluate the sound source object using the discriminability and use the evaluation as an evaluation index of the reflected sound ( S302 ).
[0329] Specifically, the selection unit 1302 may detect a sound source object that can be recognized by the listener from the video provided by the video providing device in synchronization with the sound provided by the sound presentation device 1002 .
[0330] That is, a sound source object included in an image provided by the image providing device in synchronization with the sound provided by the sound prompting device 1002 may be detected as a recognizable sound source object. Whether or not it is recognizable may be determined based on the update processing of the spatial information managed by the spatial information management unit (1201, 1211), that is, based on the processing in the information update thread.
[0331] Then, the selection unit 1302 may assign a higher evaluation score V to the sound source object detected as a recognizable object than to the sound source object that is not recognizable to the listener.
[0332] Similarly, the selection unit 1302 may assign a higher evaluation score V to direct sound and reflected sound caused by a sound source object that can be recognized by the listener than to direct sound and reflected sound caused by a sound source object that cannot be recognized by the listener.
[0333] Furthermore, the method of detecting an object recognizable to the listener is not limited to the method based on the video provided synchronously with the sound as described above. For example, the recognizable object may be determined based on the relationship between the position of the listener and the position of the object in the sound space.
[0334] That is, based on the spatial information managed by the spatial information management unit (1201, 1211), when there is no obstacle object that serves as an occluder between the position of the listener and the position of the sound source object, it can be determined that the sound source object can be recognized by the listener. More specifically, when there is no obstacle object on the propagation path of the direct sound and reflected sound that may be generated in the sound space calculated by the analysis unit 1301, it can be determined that the sound source object or the reflected object can be recognized by the listener.
[0335] Alternatively, it may be determined that the sound source object located within a predetermined distance range relative to the position of the listener is recognizable by the listener.
[0336] Furthermore, the selection unit 1302 may evaluate that the importance of the reflected sound caused by the sound source object determined to be recognizable by the listener is high, and assign a high evaluation score V to such reflected sound.
[0337] By using the index of the recognizability of the sound source object, it is possible to appropriately select the reflected sound that makes the visual positioning in the image consistent with the auditory positioning (acoustic positioning) in the sound. When the visual positioning of the sound source object that the listener can recognize is inconsistent with the acoustic positioning based on the direct sound, reflected sound and their relationship provided by the sound prompt device 1002, the sense of positioning becomes unnatural, the listener feels a sense of disharmony, and the sense of immersion decreases.
[0338] On the other hand, for an unidentifiable sound source object, even if the sound localization is slightly deviated from the original localization, since there is little sense of discomfort, the evaluation score V of the reflected sound caused by the unidentifiable sound source object may be low.
[0339] In addition, the sound prompting device 1002 and the image providing device can be the same device like VR goggles and a head-mounted display, or can be different devices like headphones and a smartphone.
[0340] The selection unit 1302 may evaluate the sound source object based on the localization, and use the evaluation as an evaluation index of the reflected sound ( S303 ).
[0341] Specifically, the selection unit 1302 may detect the moving speed of the sound source object that can be recognized by the listener from the image provided by the image providing device in synchronization with the sound provided by the sound prompting device 1002. In addition, the selection unit 1302 may assign a higher evaluation score S to the sound source object with a slow moving speed than to the sound source object with a fast moving speed. Similarly, the selection unit 1302 may assign a higher evaluation score S to the direct sound and the reflected sound caused by the sound source object with a slow moving speed than to the direct sound and the reflected sound caused by the sound source object with a fast moving speed.
[0342] In addition, when the sound source object is stopped, the selection unit 1302 may also assign the highest evaluation score S to the direct sound and reflected sound caused by the sound source object in the localization index of the sound source object. For example, the selection unit 1302 may assign a higher evaluation score to the reflected sound caused by the sound emitted by the stopped sound source object than to the reflected sound caused by the sound emitted by the moving sound source object.
[0343] Furthermore, for example, the selection unit 1302 may assign higher evaluation points to the direct sound and the reflected sound caused by the sound source object as the moving speed of the sound source object is slower.
[0344] In localizing a sound source moving at high speed, visual localization and the direction of arrival of direct sound are dominant. Therefore, the selection unit 1302 may assign a low evaluation score to the reflected sound so as not to select the reflected sound from the sound source moving at high speed.
[0345] By using the localization index of the sound source object, it is possible to appropriately select the reflected sound that makes the visual localization and the acoustic localization consistent. This can prevent the localization feeling from becoming unnatural due to the inconsistency between the visual localization and the acoustic localization.
[0346] The selection unit 1302 may use the importance of the reflection object as an evaluation index of the reflected sound. That is, the selection unit 1302 may evaluate the importance of the reflection object (S304).
[0347] For example, the spatial information may include information related to the reflecting object. Furthermore, the selection unit 1302 may evaluate the importance of the reflecting object based on the information related to the reflecting object. Furthermore, the selection unit 1302 may also assign an evaluation score to the reflected sound caused by the reflecting object based on the importance of the reflecting object.
[0348] For example, the selection unit 1302 may determine the importance of the reflection object based on information included in the input signal or metadata included in the bitstream. In addition, the selection unit 1302 may determine the importance of the reflection object based on other flags or parameters included in the input signal.
[0349] For example, the importance of the reflection object may also be determined based on the visibility of the reflection object (obstacle object) or information related to the material of the reflection object, etc. For example, the importance may be determined to be high according to the visibility of the reflection object (obstacle object), that is, for the sound source object that can be identified by the listener.
[0350] Specifically, a reflective object (obstacle object) that can be identified by the listener may be detected from the image provided by the image providing device in synchronization with the sound provided by the sound prompting device 1002. Furthermore, the selection unit 1302 may increase the importance of the reflective object detected as an identifiable object compared to a reflective object that cannot be identified by the listener.
[0351] Similarly, the selection unit 1302 may also assign a higher evaluation score V to the reflected sound caused by the reflecting object that can be recognized by the listener than to the reflected sound caused by the reflecting object that cannot be recognized by the listener. That is, the selection unit 1302 may evaluate that the importance of the reflected sound caused by the reflecting object that enters the field of view of the listener is high, and assign a high evaluation score V to such reflected sound.
[0352] In addition, as a method of detecting a reflection object that can be recognized by a listener, a method similar to the above-mentioned method of detecting a recognizable sound source object can be used.
[0353] Here, a method of using information about the material of a reflecting object as an indicator for evaluating the perceived importance of reflected sound is described. For example, multiple parameters such as reflection coefficient (reflectivity), diffusion rate, transmittance, and sound absorption rate can be obtained from metadata as information about the material of the reflecting object. In addition, the perceived importance of reflected sound can also be evaluated based on the ratio of each parameter.
[0354] Specifically, for example, when the ratio of reflectivity or diffusivity among the multiple parameters related to the materials that can be set for the reflective surface of the reflective object is high, the volume of the reflected sound that is reflected from the reflective surface and reaches the listener becomes louder than when the ratio of transmittance or sound absorption is high. In this case, the possibility that the perceived importance becomes high is high. Therefore, when the ratio of reflectivity or diffusivity among the multiple parameters related to the materials set for the reflective surface of the reflective object is high, the evaluation value for the reflected sound reflected by the reflective object may also be high.
[0355] In addition, the information related to the material of the reflection object is not limited to the reflection coefficient (reflectivity), diffusion rate, transmittance, and sound absorption rate, etc., and may also be information that can determine the importance of the material. For example, a group of multiple parameters such as the reflection coefficient (reflectivity), diffusion rate, transmittance, and sound absorption rate may be obtained from metadata as information for determining the material. In addition, the importance may be predefined according to the identifier of each material. Furthermore, the evaluation value of the reflected sound may also be calculated based on the importance associated with the identifier of the material.
[0356] Furthermore, instead of considering all of the plurality of parameters, the importance of the material of the reflection object may be determined based on only some of the parameters.
[0357] In addition, the information for specifying (identifying) a material is not limited to information for uniquely identifying a material (material identification information), and may be, for example, information for classifying a material (material classification information). The information for classifying a material may be, for example, information classified by a classification method pre-set by a content creator.
[0358] By using the index of the importance of the reflection object, it is possible to appropriately select the reflected sound caused by the reflection object with high importance based on the importance of the reflection object.
[0359] In addition, when the reflected sound caused by the reflective object is selected, the importance of the reflective object can also be updated to be low. As a result, the reflected sound caused by a specific reflective object (such as a specific wall or a specific ceiling) will not be emphasized, and the reflected sound caused by more reflective objects can be reproduced evenly. Therefore, it is possible to ensure clues for the listener to correctly perceive the width of the sound space.
[0360] For example, when selecting 30 reflected sounds from 300 reflected sounds generated in the sound space, selecting 30 reflected sounds reflected by a specific wall surface makes it difficult to grasp the overall space. In addition, allocating 5 reflected sounds to each of the 6 walls is not necessarily optimal. Therefore, as described above, it is also possible to control whether to select reflected sounds caused by each reflection object (e.g., wall) based on the value (importance) of the reflection object.
[0361] The selection unit 1302 may also use the relationship between the direct sound and the reflected sound (e.g., geometric relationship) as an evaluation index of the reflected sound. Specifically, the selection unit 1302 may also evaluate the geometric relationship between the direct sound and the reflected sound by the arrival angles of the direct sound and the reflected sound, and use the evaluation as an evaluation index of the reflected sound (S305). Here, the arrival angles of the direct sound and the reflected sound correspond to the angle formed by the arrival direction of the direct sound and the arrival direction of the reflected sound, and correspond to the angle difference between the angle of the arrival direction of the direct sound relative to the reference direction and the angle of the arrival direction of the reflected sound relative to the reference direction.
[0362] The angle between the direction in which the direct sound arrives and the direction in which the reflected sound arrives may be detected, and the larger the angle is, the higher the evaluation score is given to the reflected sound.
[0363] For example, the analysis unit 1301 calculates the direct sound arrival path (pd) and the reflected sound arrival direction path (pr). The analysis unit 1301 or the selection unit 1302 calculates the arrival direction of the direct sound and the arrival direction of the reflected sound based on the direct sound arrival path (pd), the reflected sound arrival direction path (pr), and the direction information (D) of the avatar (listener) included in the input signal. The arrival direction of the direct sound and the arrival direction of the reflected sound are expressed based on the direction of the listener.
[0364] Furthermore, the selection unit 1302 calculates the evaluation score of the reflected sound based on the angle formed by the arrival direction of the direct sound and the arrival direction of the reflected sound.
[0365] Fig.13 is a diagram showing examples of the arrival angles of direct sound and reflected sound. Fig.13 The configuration shown includes an avatar, a sound source object, and an obstacle object. The position information of the avatar, the sound source object, and the obstacle object and the direction information (D) of the avatar are obtained from the input signal. And, based on this information, the direction of the avatar is regarded as 0 degrees, and the direction of the direct sound (θ) and the direction of the sound image of the reflected sound (γ) are calculated.
[0366] exist Fig.13 In the case of , the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees). In this case, the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound is about 115 degrees.
[0367] When the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound is large, a high evaluation score is given to the reflected sound. Thus, for example, the evaluation score of the reflected sound of the sound source visible in front of the listener and heard from behind the listener becomes high. As a result, the reflected sound that helps the listener to foresee the existence of a large object behind can be preferentially selected, and a sense of occlusion and urgency can be expressed.
[0368] The selection unit 1302 may also evaluate the relationship between the direct sound and the reflected sound by the time difference between the direct sound and the reflected sound, and use the evaluation as an evaluation index of the reflected sound (S306). For example, the selection unit 1302 may also give a higher evaluation score to the reflected sound with a larger difference in arrival time between the direct sound and the reflected sound than to the reflected sound with a smaller difference in arrival time. For example, the echo returned when shouting "Yoho" on the top of a mountain has a decisive influence on the grasp of the space. Therefore, such a reflected sound may also be given a high evaluation score.
[0369] The selection unit 1302 may also evaluate the relationship between the direct sound and the reflected sound by using the time difference between the direct sound and the reflected sound and a threshold value corresponding to the time difference. For example, the reflected sound that arrives at the listener's position immediately after the direct sound is easily masked by the direct sound and is difficult to perceive. On the other hand, the reflected sound that arrives at the listener's position with a time difference from the direct sound is difficult to be masked by the direct sound and is easy to perceive. The reflected sound may also be given an evaluation score based on such a perceptual model.
[0370] The time difference (T) between the direct sound and the reflected sound may be, for example, the time difference between the time required for the direct sound and the reflected sound to reach the listening position. For example, the time difference (T) between the direct sound and the reflected sound to reach the listening position is obtained by T=tr-td.
[0371] For example, when the relationship between direct sound and reflected sound is used to evaluate reflected sound, a threshold value determined corresponding to the time difference between the direct sound and the reflected sound is used for comparison processing. The threshold value represents a volume pre-set corresponding to the time difference between the direct sound and the reflected sound, and is determined by referring to the threshold value data. In the threshold value data, an indicator indicating a boundary of whether the reflected sound with respect to the direct sound is perceived by the listener may also be used.
[0372] For example, a threshold value refers to a value expressed by a numerical value determined corresponding to a time difference (T), and threshold data refers to table data or a relational expression used to determine or calculate a threshold value under a time difference (T). However, the form and type of threshold data are not limited to table data or a relational expression.
[0373] Fig.14This is a diagram showing an example of a method for setting threshold data based on the temporal masking phenomenon. The threshold data can also be set, for example, by referring to a masking threshold value which is a well-known threshold. As described in Non-Patent Document 1, etc., the temporal masking phenomenon is well known. The shaded portion in the figure represents the time period and amplitude of the masker (blocking signal that blocks the perception of the signal S of the listening object) generated.
[0374] exist Fig.14 In the figure, the masking threshold represents the level (SPL: Sound Pressure Level) at which the signal S can be heard. Of course, the masking threshold is high during the period when the masker is generated. On the other hand, after the masker stops, the masking threshold does not immediately become zero, but gradually decays. That is, the masking threshold is high for a period of time just after the masker stops (during the period of post-masking).
[0375] For example, in Fig.14 The tendency of back masking shown in the area surrounded by the dotted line in FIG. 1 can be used as threshold data for evaluating reflected sound based on the relationship between direct sound and reflected sound. That is, the threshold data can be determined based on the tendency of back masking by considering that the direct sound corresponds to the masker and the reflected sound corresponds to the signal S of the listening object.
[0376] Fig.15 is a diagram showing an example of threshold data. In the above case, Fig.15 The threshold data is determined as shown in the curve. Fig.15 In the graph with the time difference between direct sound and reflected sound on the horizontal axis and the volume of reflected sound on the vertical axis, a curve is used to indicate the boundary (threshold) of whether the reflected sound is perceived. The curve corresponds to the threshold data.
[0377] In addition, the threshold data of this embodiment is stored in the memory 1404 of the sound signal processing device 1001. The form and type of the stored threshold data can be any form and type. For example, the threshold data can also be expressed by an approximate formula having the time difference between the direct sound and the reflected sound as a variable. In addition, the threshold data can also be expressed by the arrangement of the time difference between the direct sound and the reflected sound and the threshold.
[0378] Fig.16 This is a graph showing the relationship between the time difference between direct sound and reflected sound and the threshold. Fig.16 As shown, the threshold data may be stored in an area of the memory 1404 as an array of indices of the time difference between the direct sound and the reflected sound and the thresholds corresponding to the indices.
[0379] certainly, Fig.15 and Fig.16The graph and numerical values shown are examples, and the threshold value data are not limited to these.
[0380] Information related to a relational expression representing the relationship between the time difference (T) and the threshold value may also be stored in the memory 1404. That is, an expression having the time difference (T) as a variable may also be stored. The threshold value of each time difference (T) may also be approximated by a straight line or a curve, and parameters representing the geometric shape of the straight line or the curve may be stored. For example, when the geometric shape is a straight line, the starting point and slope for representing the straight line may also be stored.
[0381] When a plurality of forms and types of threshold values are stored, it may be determined which form and which type of threshold value to use in the reflected sound selection process.
[0382] In addition, the threshold for evaluating reflected sound is not limited to the known masking threshold. Other thresholds may be determined based on the time difference between the direct sound and the reflected sound and the value representing the amplitude value or the volume. For example, the threshold may be determined based on the minimum time difference for detecting the deviation of two sounds based on the listener's perception. The specific numerical value may be derived from known research results or determined through a listening experiment conducted on the premise of applying to the virtual space.
[0383] For example, the threshold is set based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound by referring to the threshold data. The selection unit 1302 may set a high evaluation score when the volume of the reflected sound is greater than the set threshold when the sound arrives.
[0384] The time difference between the arrival time of the direct sound and the arrival time of the reflected sound is, in other words, the difference in time required for the direct sound and the reflected sound to reach the listening position, respectively. Therefore, the difference in distance of the arrival path of the direct sound and the reflected sound can also be used as a value related to the time difference between the arrival time of the direct sound and the reflected sound.
[0385] Alternatively, the time difference between the end of the direct sound and the arrival of the reflected sound at the listening position may be used as the time difference between the direct sound and the reflected sound. The end time of the direct sound may be obtained by, for example, adding the duration of the direct sound to the arrival time of the direct sound.
[0386] Use the position relationship between the listener and the obstacle object Fig. 9 and Fig.10 , and an example of threshold data Fig.15 , illustrating a method for evaluating reflected sound using threshold data.
[0387] Fig.15The graph has the time difference between the direct sound and the reflected sound on the horizontal axis and the volume ratio of the direct sound to the reflected sound on the vertical axis. The curve represents the threshold of the boundary between whether the reflected sound is perceived or not perceived. A, B, and C in the graph represent the reflected sound respectively. In addition, here, the vertical axis uses the volume ratio, that is, the volume of the reflected sound determined relatively to the volume of the direct sound, but the volume of the reflected sound determined absolutely regardless of the volume of the direct sound may also be used.
[0388] In addition, when the volume is expressed by the unit of decibel on the logarithmic axis (when the volume is expressed in the decibel area), the volume ratio of the two signals is of course expressed by the difference in decibel values. Specifically, the volume ratio of the two signals can be the difference when the amplitude values of each signal are expressed in the decibel area. This value can also be calculated based on energy value or power value. In addition, this difference can be called the difference in gain or simply the gain difference in the decibel area.
[0389] That is, the volume ratio in the present disclosure is essentially the ratio of the amplitudes of the signals, so it can also be expressed as Sound volume ratio, Volume ratio, Amplitude ratio, Sound level ratio, Sound intensity ratio or Gain ratio, etc. In addition, when the unit of volume is decibel, the volume ratio in the present disclosure can of course be changed to volume difference.
[0390] In the present disclosure, "volume ratio" typically refers to the gain difference when the volume of two sounds is expressed in decibel units. In the example of the implementation, the threshold data is also typically specified by the gain difference expressed in the decibel area. However, the volume ratio is not limited to the gain difference in the decibel area. When using a volume ratio expressed outside the decibel area, the threshold data specified in the decibel area can also be converted into the unit of the calculated volume ratio for use. Alternatively, the threshold data specified in each unit can also be stored in the memory in advance.
[0391] That is, even if a ratio of energy values or power values is used instead of the volume ratio, it is obvious that the algorithm in the present disclosure can be applied to solving the problem of the present disclosure.
[0392] Fig. 9 Indicates the positional relationship between the listener, the sound source object, and the obstacle object (wall). Fig. 9 In the case of sound source and obstacle, the sound source object and the obstacle object are far away, and the listener hears Fig.15 The reflected sound C. Fig.10 Indicates another positional relationship between the listener, the sound source object, and the obstacle object (wall). Fig.10In the case of sound source object and obstacle object being close, the sound source object is heard by the listener. Fig.15 The reflected sound A or B.
[0393] For example, Fig. 9 As shown, when the listener is relatively far from the obstacle, the arrival time of the reflected sound becomes slower, and the time difference between the direct sound and the reflected sound is larger for the reflected sound C than for the reflected sounds A and B.
[0394] That is, Fig.15 As shown, reflected sound C is located on the right side of the graph compared to reflected sounds A and B. As shown in the curve in the graph, the greater the time difference between the direct sound and the reflected sound, the smaller the threshold. As a result, reflected sound B having the same volume as reflected sound C is smaller than the threshold, and reflected sound C is larger than the threshold. Therefore, the evaluation score of reflected sound C is higher than the evaluation score of reflected sound B.
[0395] In addition, with respect to reflected sounds A and B, although the arrival time is the same, the volume of reflected sound A is greater than the volume of reflected sound B, and the volume of reflected sound B is less than the volume of reflected sound A. In addition, the volume of reflected sound A is greater than the threshold value shown in the curve, and the volume of reflected sound B is less than the threshold value shown in the curve. In this case, reflected sound A is given a higher evaluation score than reflected sound B.
[0396] The reflected sound is evaluated based on a threshold indicating a volume determined in accordance with the time difference between the direct sound and the reflected sound. Thus, the human perception property that the reflected sound that reaches the listener's position with a time difference from the direct sound is not masked by the direct sound and is easily perceived is reflected in the reflected sound evaluation.
[0397] By using a threshold value indicating a volume determined according to the time difference between direct sound and reflected sound, it is possible to more appropriately select a reflected sound having a greater impact on listener perception than using only the time difference between direct sound and reflected sound or only the volume of reflected sound.
[0398] In addition, the calculation of the arrival time and volume of the direct sound and the reflected sound when arriving can be omitted, and the reflected sound can be evaluated based on the path length. In the case of evaluating the reflected sound based on the path lengths of the direct sound and the reflected sound when they reach the listener, the threshold of the path length of the reflected sound can also be set corresponding to the value of the path length difference. In this case, the reflected sound can also be evaluated based on whether the path length of the reflected sound is greater than the threshold determined corresponding to the value of the path length difference.
[0399] In the case of selecting based on the path length, it is possible to select based on information affecting the time difference while reducing the amount of calculation compared to selecting based on the time difference. In addition, in addition to the difference in path length, a parameter representing the propagation speed of sound or a parameter affecting the parameter of the propagation speed of sound may also be used.
[0400] The geometric relationship can also be the relationship between the positions of the sound source, the listener, and the reflection object in the virtual space. Through their relationship, the path lengths of the direct sound and the reflected sound can be geometrically calculated. Therefore, if the relationship that the volume is inversely proportional to the distance is used, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.
[0401] In calculating the reference volume of the reflected sound, the reflection coefficient of the reflecting object may also be used. In addition, a typical value commonly used may also be used as the reflection coefficient. On the other hand, in the case of special conditions such as when the reflecting object is covered with a sound absorbing material, a specially assigned reflection coefficient may also be used as the reflection coefficient of the reflecting object.
[0402] The reflected sound can also be evaluated based on the volume of the reflected sound. The volume of the reflected sound can also be obtained based on the geometric relationship between the direct sound and the reflected sound as described above, and the index given to the reflection object. The reflected sound can also be evaluated by comparing the volume with a predetermined threshold.
[0403] Furthermore, information indicating the temporal transition of the volume of the sound source may be reflected in the evaluation. For example, when the information indicating the temporal transition of the volume of the sound source indicates that the duration of the sound interval is long, the evaluation value of the reflected sound may be maintained as it is when the moment is in the sound interval. On the other hand, when the moment is outside the sound interval, even if the reference volume of the reflected sound exceeds the threshold, the evaluation value of the reflected sound may be reduced or set to zero.
[0404] Alternatively, the information indicating the temporal change in the volume of the sound source may be data that lists in time series a duration during which the amplitude of the sound signal is considered to be substantially constant and a plurality of sets of amplitude values of the signal during the duration. In this case, a process may be performed to evaluate the reflected sound by changing the reference volume of the reflected sound in conjunction with changes in the amplitude value in the data.
[0405] In addition, for one reflected sound, an evaluation score related to all the above-mentioned indicators may be assigned, or an evaluation score related to a part of the indicators may be assigned. In addition, the number of indicators used for evaluation may be different for each reflected sound, or the same indicator may be used for all reflected sounds. The indicator used to assign the evaluation score to the reflected sound may be set based on predetermined information, for example, it may be determined based on information included in the input signal, or it may be determined based on information set by the listener or the administrator.
[0406] In addition, a high evaluation score corresponds to a large evaluation score, and a low evaluation score corresponds to a small evaluation score. Similarly, a high evaluation value corresponds to a large evaluation value, and a low evaluation value corresponds to a small evaluation value. These expressions can be interchangeable.
[0407] (Calculation of evaluation value)
[0408] Next, the selection unit 1302 calculates an evaluation value indicating the importance of the reflected sound based on the evaluation score for the reflected sound to which the evaluation score is assigned using each indicator. For example, the selection unit 1302 determines the total value of multiple evaluation scores as the evaluation value of the reflected sound (S307). The total value of multiple evaluation scores may also be a weighted total value. If there is an unevaluated reflected sound ("Yes" in S308), the selection unit 1302 repeats the above-mentioned processing (S301 to S307). If there is no unevaluated reflected sound ("No" in S308), the evaluation processing ends.
[0409] In addition, the evaluation value of the reflected sound is not limited to the total value of multiple evaluation scores obtained by multiple indicators. For example, the evaluation value of a predetermined benchmark and the already calculated evaluation value can also be corrected by multiple evaluation scores. In addition, the evaluation scores of only a part of the indicators can be used for the evaluation value of the reflected sound, and can also be used to correct the evaluation value of the reflected sound. In addition, in the case where multiple evaluation scores are assigned to a reflected sound using multiple indicators, the highest evaluation score can also be determined as the evaluation value of the reflected sound.
[0410] Which evaluation score of an index is used for calculation or correction of the evaluation value may be determined based on predetermined information, may be determined based on information included in the input signal, or may be determined based on information set by the listener or the administrator.
[0411] In addition, in the above description, for convenience, the evaluation score and the evaluation value are divided into the evaluation score obtained by each indicator and the evaluation value obtained by using multiple evaluation scores obtained by multiple indicators. However, both represent the evaluation results of the reflected sound, so the evaluation score and the evaluation value can also be viewed in the same way. In addition, the sound signal processing device 1001 can directly use the evaluation score obtained by one indicator as the evaluation value in the selection process of the reflected sound, and can use multiple evaluation scores obtained by multiple indicators in the selection process of the reflected sound.
[0412] For example, when a plurality of evaluation scores are used in the selection process of the reflected sound, the sound signal processing device 1001 determines whether to select the reflected sound based on each of the plurality of evaluation scores. Furthermore, the sound signal processing device 1001 may also finally determine that the reflected sound is selected when all of the plurality of determination results based on the plurality of evaluation scores indicate that the reflected sound is selected. Alternatively, the sound signal processing device 1001 may finally determine that the reflected sound is selected when one of the plurality of determination results based on the plurality of evaluation scores indicates that the reflected sound is selected.
[0413] In addition, priorities may be set for a plurality of evaluation scores based on a plurality of indicators. For example, the sound signal processing device 1001 determines whether to select the reflected sound based on each of the first to third evaluation scores based on the first to third indicators.
[0414] In the above case, when the determination result based on the first evaluation score indicates that the reflected sound is not selected, the audio signal processing device 1001 may finally determine that the reflected sound is not selected, regardless of the determination results based on the second and third evaluation scores.
[0415] Furthermore, when the determination results based on the first and second evaluation points indicate that the reflected sound is selected, the sound signal processing device 1001 may finally determine that the reflected sound is selected, regardless of the determination result based on the third evaluation point.
[0416] For example, after determining the evaluation value of the reflected sound, Fig.11 The process is performed as described above in the flowchart shown.
[0417] (Processing order and omission)
[0418] about Fig.11 and Fig.12 The processing included in the flowchart shown may be partially omitted or the order of the processing may be changed.
[0419] For example, in Fig.11 In the flowchart shown, the reflected sound selection process is performed based on both the calculation load and the evaluation value (importance) of the reflected sound. However, the reflected sound selection process may be performed based on only one of them.
[0420] Specifically, the selection unit 1302 may omit the calculation of the evaluation value of each reflected sound, and determine not to select the reflected sound when the computational load of the reflected sound is greater than a threshold value. Alternatively, the selection unit 1302 may omit the acquisition of information indicating the upper limit of the computational load, the calculation of the total computational load of the extracted reflected sounds, and the comparison of the total computational load with the upper limit of the computational load, and perform the selection process of the reflected sound only based on the evaluation value of the reflected sound.
[0421] In addition, the extraction of reflected sound with a volume above the threshold value can be performed after the evaluation value is determined, or after the reflected sound is determined to be selected. For example, even if the reflected sound is determined to be selected based on the evaluation value or the calculation load, if the volume of the reflected sound is below the threshold value, the reflected sound can be determined to be not selected again.
[0422] (Generation of direct sound and reflected sound)
[0423] Next, in the generation process of direct sound and reflected sound ( Figure 8 In S103), the synthesizing unit 1303 generates and synthesizes the sound signal of the direct sound and the sound signal of the reflected sound selected by the selecting unit 1302 as the object reflected sound.
[0424] The sound signal of the direct sound is generated by applying the arrival time (td) and the volume (ld) at the time of arrival calculated by the analysis unit 1301 to the sound data of the sound source object included in the input information. Specifically, the sound data is delayed by the amount of the arrival time (td) and multiplied by the volume (ld) at the time of arrival. The process of delaying the sound data is a process of moving the position of the sound data forward or backward on the time axis. For example, the process of delaying the sound data without deteriorating the sound quality as disclosed in Patent Document 2 may also be applied.
[0425] The sound signal of the reflected sound is generated by applying the arrival time (tr) and the arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object, similarly to the direct sound.
[0426] However, the arrival volume (lr) in the generation of reflected sound is different from the arrival volume of direct sound, and is a value to which the attenuation rate G of the volume in the reflection is applied. G can be an attenuation rate applied to the entire frequency band. Alternatively, the reflectivity can be specified for each specified frequency band in order to reflect the bias of the frequency components generated by the reflection. In this case, the processing of applying the arrival volume (lr) can also be implemented as a processing of multiplying the attenuation rate for each frequency band, that is, a frequency equalizer processing.
[0427] (Pipeline processing)
[0428] The processing performed by the above-mentioned analyzing unit 1301, selecting unit 1302, and synthesizing unit 1303 may be performed as pipeline processing as described in Patent Document 3, for example.
[0429] Fig.17 13 is a block diagram showing a configuration example for the rendering unit 1300 to perform pipeline processing.
[0430] Fig.17The rendering unit 1300 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. The reverberation processing unit 1311, the initial reflection processing unit 1312, and the distance attenuation processing unit 1313 perform reverberation processing, initial reflection processing, and distance attenuation processing, respectively. The selection unit 1314 selects reflected sound, the generation unit 1315 generates direct sound and reflected sound, and the binaural processing unit 1316 applies binaural processing to the direct sound and the reflected sound.
[0431] These multiple components can also be Figure 7 The rendering unit 1300 shown in FIG. 1 may be composed of multiple components, or may be composed of Figure 5 The audio signal processing device 1001 shown in the figure is constituted by at least a part of the multiple components.
[0432] Pipeline processing means dividing the processing for providing an acoustic effect into a plurality of processes and executing the plurality of processes in sequence one by one. In each of the plurality of processes, for example, signal processing of a sound signal or generation of parameters used in the signal processing is executed.
[0433] The rendering unit 1300 may also perform reverberation processing, initial reflection processing, distance attenuation processing, binaural processing, etc. as pipeline processing. However, these processing are examples, and pipeline processing may include processing other than these, or may not include some processing. For example, pipeline processing may also include diffraction processing and occlusion processing. In addition, for example, reverberation processing may be omitted when it is not necessary.
[0434] In addition, each process may be represented as a stage. In addition, the result of each process and the generated sound signal such as reflected sound may be represented as a rendering item. The multiple stages in the pipeline processing and their order are not limited to Fig.17 Example shown.
[0435] Here, the parameters used in the selection process (arrival paths and arrival times related to direct sound and reflected sound) may also be calculated in one of the multiple stages used to generate rendering items. That is, the parameters used in the selection of reflected sound are calculated in a part of the pipeline process used to generate rendering items. In addition, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted or performed outside the rendering unit 1300.
[0436] The following describes the reverberation process, initial reflection process, distance attenuation process, selection process, generation process, and binaural process that can be included as stages in the pipeline process. In each stage, metadata included in the input signal can also be analyzed to calculate parameters used in the generation of reflected sound.
[0437] In the reverberation processing, the reverberation processing unit 1311 generates a sound signal representing a reverberation sound or a parameter used in generating a sound signal. The reverberation sound refers to the sound that reaches the listener as a reverberation after the direct sound. As an example, the reverberation sound is the sound that reaches the listener after being reflected more times (e.g., several dozen times) than the initial reflected sound at a relatively late stage (e.g., from the arrival of the direct sound to about one hundred and several dozen ms) after the initial reflected sound described later reaches the listener.
[0438] The reverberation processing unit 1311 refers to the audio signal and space information included in the input signal, and calculates the reverberation sound using a predetermined function prepared in advance as a function for generating the reverberation sound.
[0439] The reverberation processing unit 1311 may also generate reverberation sound by applying a known reverberation generation method to the sound signal included in the input signal. An example of a known reverberation generation method is the Schroeder method, but the known reverberation generation method is not limited to the Schroeder method. In addition, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by the spatial information in the application of the known reverberation generation method. Thus, the reverberation processing unit 1311 can calculate parameters for generating reverberation sound.
[0440] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating initial reflection sound based on the spatial information. The initial reflection sound is the reflection sound that reaches the listener after one or more reflections at a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct sound) after the direct sound reaches the listener from the sound source object.
[0441] The initial reflection processing unit 1312 calculates the path of the reflected sound from the sound source object to the listener by reflecting from the reflection object, for example, with reference to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflection object such as a structure, and the reflectivity of the reflection object may be used in the calculation of the path.
[0442] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct sound. Information on the path may be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected sound, or may be used as a parameter for the selection unit 1314 to select the reflected sound.
[0443] In the distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct sound and the reflected sound reaching the listener based on the length of the path of the direct sound and the reflected sound. The volume of the direct sound and the reflected sound reaching the listener is attenuated in proportion to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct sound by dividing the volume of the sound source by the length of the path of the direct sound, and can calculate the volume of the reflected sound by dividing the volume of the sound source by the length of the path of the reflected sound.
[0444] In the selection process, the selection unit 1314 selects the generation of the object reflected sound based on the parameters calculated before the selection process. In the selection of the generation of the object reflected sound, a certain selection method of the present disclosure may be used.
[0445] In addition, when the selection process is included in the pipeline process, the process after the selection process may not be performed in the pipeline process for the reflected sound that is not selected in the selection process. By not performing the process after the selection process for the reflected sound that is not selected, the computational load of the sound signal processing device 1001 can be reduced compared to not performing the binaural process alone.
[0446] Furthermore, when the selection process is included in the pipeline process, more processes can be omitted and the amount of calculation can be reduced by assigning an earlier order to the selection process among the plurality of processes in the pipeline process.
[0447] In binaural processing, the binaural processing unit 1316 performs signal processing so that the sound signal of the direct sound is perceived as sound reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing so that the reflected sound selected by the selection unit 1314 is perceived as sound reaching the listener from the reflection object.
[0448] For example, the binaural processing unit 1316 performs processing using the HRIR DB based on the position and orientation of the listener in the sound space so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.
[0449] In addition, HRIR (Head-Related Impulse Responses) is the response characteristic when an impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from the frequency domain to the time domain through Fourier transform. The head-related transfer function expresses the change of the sound generated by the surrounding objects including the auricle, the head and shoulders as a transfer function. HRIR DB is a database containing such information.
[0450] In addition, the position and orientation of the listener in the sound space are, for example, the position and orientation of the virtual listener in the virtual sound space. Alternatively, the position and orientation of the virtual listener in the virtual sound space may change in accordance with the movement of the listener's head. In addition, the position and orientation of the virtual listener in the virtual sound space may also be determined based on information obtained from the sensor 1405.
[0451] The program, spatial information, HRIR DB, threshold data, other parameters, and the like used in the above-mentioned processing are acquired from the memory 1404 included in the sound signal processing device 1001 or from outside the sound signal processing device 1001 .
[0452] In addition, pipeline processing may include other processing. Furthermore, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in the pipeline processing. For example, the rendering unit 1300 may include a diffraction processing unit and an occlusion processing unit.
[0453] The diffraction processing unit performs processing for generating a sound signal representing a sound including diffracted sound caused by an obstacle object between a listener and a sound source object in a three-dimensional sound field (space). The diffracted sound is a sound that, when an obstacle object exists between the sound source object and the listener, bypasses the obstacle object and reaches the listener from the sound source object.
[0454] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object to the listener around the obstacle object, and generates the diffracted sound based on the path. In calculating the path, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), and the shape and size of the obstacle object may also be used.
[0455] When a sound source object exists on the opposite side of the obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on the spatial information and information such as the material of the obstacle object.
[0456] (Example of sound source object)
[0457] In the above description, the position information given to the sound source object indicates a "point" in the virtual space as the position of the sound source object. That is, in the above description, the sound source is defined as a "point sound source".
[0458] On the other hand, the sound source in the virtual space can also be defined as an object with length, size, shape, etc., that is, a sound source that is defined as a non-point sound source and extends in space. In this case, the distance between the listener and the sound source and the direction of arrival of the sound are uncertain. Therefore, the reflected sound caused by such a sound source does not need to be analyzed by the analysis unit 1301, or is limited to being selected by the selection unit 1302 regardless of the analysis result. In this way, it is possible to avoid the degradation of sound quality that may occur due to not selecting the reflected sound.
[0459] Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied assuming that the sound is generated from the representative point. In this case, the threshold may be adjusted based on spatial extension information of the sound source.
[0460] (Examples of direct sound and reflected sound)
[0461] For example, direct sound is sound that is not reflected by a reflective object, and reflected sound is sound that is reflected by a reflective object. Direct sound can also be sound that reaches the listener from the sound source without being reflected by a reflective object, and reflected sound can also be sound that reaches the listener from the sound source after being reflected by a reflective object.
[0462] Furthermore, direct sound and reflected sound are not limited to sounds reaching the listener, but may be sounds before reaching the listener. For example, direct sound may be sound output from a sound source, or in other words, sound of the sound source.
[0463] (Bitstream Structure Example)
[0464] The bitstream includes, for example, an audio signal and metadata. The audio signal is audio data that represents the sound, and indicates information related to the frequency and strength of the sound, etc. In addition, the metadata includes spatial information related to the sound space, that is, the sound field space.
[0465] For example, spatial information is information about the space where a listener who listens to a sound based on a sound signal is located. Specifically, spatial information is information related to a predetermined position (localization position) for localizing a sound image in a sound space (e.g., a three-dimensional sound field), that is, for allowing a listener to perceive a sound coming from a direction corresponding to the predetermined position. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
[0466] The sound source object information is information about a sound source object that generates sound based on a sound signal. That is, the sound source object information is information about an object (sound source object) that reproduces a sound signal, and is information about a virtual sound source object configured in a virtual sound space. Here, the virtual sound space may also correspond to a real space where an object that generates sound is configured, and the sound source object in the virtual sound space may also correspond to an object that generates sound in the real space.
[0467] The sound source object information may also indicate the position of the sound source object arranged in the sound space, the direction of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living thing, whether the sound source object is a moving body, etc. For example, a sound signal is associated with one or more sound source objects indicated by the sound source object information.
[0468] The bit stream has a data structure composed of, for example, metadata (control information) and an audio signal.
[0469] The audio signal and metadata may be included in one bit stream or in multiple bit streams. In addition, the audio signal and metadata may be included in one file or in multiple files.
[0470] The bit stream may exist for each sound source or for each playback time. When the bit stream exists for each playback time, a plurality of bit streams may be processed in parallel at the same time.
[0471] The metadata may be assigned to each bit stream, or may be assigned to multiple bit streams together as information for controlling the multiple bit streams. In this case, the metadata may be shared by multiple bit streams. In addition, the metadata may be assigned for each playback time.
[0472] When there are multiple bitstreams or multiple files, information indicating related bitstreams or related files may be included in more than one bitstream or more than one file. Alternatively, information indicating related bitstreams or related files may be included in all bitstreams or all files.
[0473] Here, the associated bitstream or associated file refers to, for example, a bitstream or file that may be used simultaneously during audio processing. In addition, a bitstream or file in which information indicating the associated bitstream or associated file is recorded together may also be included.
[0474] Here, the information indicating the associated bitstream or the associated file may be, for example, an identifier indicating the associated bitstream or the associated file. In addition, the information indicating the associated bitstream or the associated file may be, for example, a file name, a URL (Uniform Resource Locator) or a URI (Uniform Resource Identifier) indicating the associated bitstream or the associated file.
[0475] In this case, the acquisition unit may also determine and acquire the associated bitstream or associated file based on the information indicating the associated bitstream or associated file. In addition, the information indicating the associated bitstream or associated file may be included in a bitstream or file, and the information indicating the associated bitstream or associated file may be included in another bitstream or other file.
[0476] Here, the file including information indicating the associated bitstream or the associated file may be, for example, a control file such as a declaration file used for content distribution.
[0477] In addition, all or part of the metadata may be obtained from outside the bitstream of the audio signal. For example, metadata for controlling the audio or metadata for controlling the video may be obtained from outside the bitstream, or both metadata may be obtained from outside the bitstream.
[0478] Furthermore, metadata for controlling images may be included in the bit stream obtained by the stereoscopic sound reproduction system 1000. In this case, the stereoscopic sound reproduction system 1000 may output metadata for controlling images to a display device that displays images or a stereoscopic image reproduction device that reproduces stereoscopic images.
[0479] (Example of information included in metadata)
[0480] Metadata may be information used to describe a scene represented by a sound space. Here, a scene is a term that refers to a collection of all elements of three-dimensional images and sound events in a sound space modeled by a sound signal reproduction system using metadata.
[0481] That is, metadata may include not only information for controlling audio processing but also information for controlling video processing. Metadata may include only one of the information for controlling audio processing and the information for controlling video processing, or both.
[0482] The stereo sound reproduction system 1000 generates virtual sound effects by performing sound processing on sound signals using metadata included in a bitstream and interactive listener position information obtained by addition. The sound effects may include initial reflection processing, obstacle processing, diffraction processing, occlusion processing, and reverberation processing, and other sound processing may be performed using metadata. For example, sound effects such as distance attenuation effect, localization, or Doppler effect may be added.
[0483] Furthermore, information on switching on / off of all or part of the additional acoustic effects or priority information on a plurality of processes for the acoustic effects may be added to the metadata.
[0484] In addition, as an example, metadata includes information related to the sound space including the sound source object and the obstacle object, and information related to the localization position used to localize the sound image at a specified position in the sound space (i.e., to make the listener perceive the sound coming from a specified direction).
[0485] Here, an obstacle object is an object that may affect the sound perceived by the listener, such as blocking or reflecting the sound, during the period from the sound emitted by the sound source object to the sound reaching the listener. In addition to stationary objects, obstacle objects may also include moving objects such as animals or machinery. Animals may also be humans, etc.
[0486] Furthermore, when there are multiple sound source objects in the sound space, other sound source objects may become obstacle objects for any sound source object. That is, objects that do not make sound, such as building materials or inanimate objects, i.e., non-sound-generating objects, and sound source objects that make sound may all become obstacle objects.
[0487] The metadata includes all or part of information indicating the shape of the sound space, the shape and position of obstacle objects in the sound space, the shape and position of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
[0488] The sound space may be a closed space or an open space. In addition, the metadata may also include information indicating the reflectivity of an obstacle object that can reflect sound in the sound space. For example, a floor, a wall, or a ceiling that constitutes a boundary of the sound space may also be an obstacle object.
[0489] The reflectivity is the energy ratio of the reflected sound to the incident sound, and can also be set for each frequency band of the sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, for example, uniformly set parameters such as the attenuation rate, diffracted sound, and initial reflected sound can also be used.
[0490] The metadata may also include information other than reflectivity as a parameter related to an obstacle object or a sound source object. For example, the metadata may also include information related to the material of the object as a parameter related to both the sound source object and the non-sound-emitting object. Specifically, the metadata may also include information such as diffusivity, transmittance, and sound absorption coefficient.
[0491] The information about the sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information indicating the sound source area in the object. The reproduction conditions may also be set to be, for example, whether the sound is continuously flowing or the sound is triggered by an event. The sound source area in the object may be set based on the relative relationship between the position of the listener and the position of the object, or may be set using the object as a reference.
[0492] For example, when the sound source area is set based on the relative relationship between the listener's position and the object's position, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.
[0493] Furthermore, when the sound source area is set using the object as a reference, it is possible to use the object as a reference to fix which sound is emitted from which area of the object. For example, when the listener is looking at the object from the front, the listener can perceive high tones from the right side of the object and low tones from the left side of the object. Also, when the listener is looking at the object from the back, the listener can perceive low tones from the right side of the object and high tones from the left side of the object.
[0494] Metadata related to the space may include time until initial reflected sound, reverberation time, and the ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the listener can perceive only direct sound.
[0495] (Replenish)
[0496] In addition, the aspects grasped based on the present disclosure are not limited to the embodiments, and various modifications can be made and implemented.
[0497] For example, the processing executed by a specific component in the embodiment may be executed by another component instead of the specific component. In addition, the order of a plurality of processing may be changed, or a plurality of processing may be executed in parallel.
[0498] In addition, the ordinal numbers such as 1 and 2 used in the description may be replaced, removed, or newly assigned as appropriate. These ordinal numbers do not necessarily correspond to a meaningful order, but may be used to identify elements.
[0499] In addition, for example, in the comparison of the threshold, above the threshold and greater than the threshold can be replaced with each other. Similarly, below the threshold and less than the threshold can be replaced with each other. In addition, for example, time and moment can be replaced with each other.
[0500] Furthermore, in the process of selecting one or more processing target sounds from a plurality of sounds, if there is no sound satisfying the condition, none of the sounds may be selected as the processing target sounds. That is, in the process of selecting one or more processing target sounds from a plurality of sounds, the case of not selecting the processing target sound may also be included.
[0501] Furthermore, for example, an expression such as at least one of the first element, the second element, and the third element may correspond to the first element, the second element, the third element, or any combination thereof.
[0502] In addition, for example, the embodiments describe the case where the form grasped based on the present disclosure is implemented as an audio processing device, an encoding device, or a decoding device. However, the form grasped based on the present disclosure is not limited to these, and can also be implemented as software for executing the audio processing method, encoding method, or decoding method.
[0503] For example, a program for executing the above-mentioned sound processing method, encoding method, or decoding method may be stored in advance in the ROM, and the CPU may operate according to the program.
[0504] In addition, a program for executing the above-mentioned sound processing method, encoding method or decoding method may be stored in a computer-readable recording medium. Furthermore, the computer may record the program stored in the recording medium into the RAM of the computer and operate according to the program.
[0505] Furthermore, each of the above-mentioned components can be typically implemented as an integrated circuit (LSI) having input terminals and output terminals. They can be formed into one chip individually, or can be formed into one chip in a manner that includes all or part of the components of the implementation mode. LSI can also be expressed as IC, system LSI, super LSI or ultra-large-scale LSI according to the difference in integration.
[0506] In addition, it is not limited to LSI, and a dedicated circuit or a general-purpose processor can also be used. In addition, an FPGA that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connection or setting of the circuit unit inside the LSI can also be used. Furthermore, if a technology for integrated circuitization that replaces LSI appears due to the progress of semiconductor technology or other derived technologies, then of course, this technology can also be used to integrate the components. It may be an application of biotechnology, etc.
[0507] In addition, the FPGA or CPU may download all or part of the software for implementing the sound processing method, encoding method or decoding method described in the present disclosure through wireless communication or wired communication. Furthermore, the whole or part of the software for updating may be downloaded through wireless communication or wired communication. Furthermore, the digital signal processing described in the present disclosure may be performed by storing the downloaded software in a memory by the FPGA or CPU and operating based on the stored software.
[0508] At this time, the device with FPGA or CPU etc. can also be connected to the signal processing device wirelessly or by wire, or can be connected to the signal processing server via a network. Furthermore, the device and the signal processing device or the signal processing server can also perform the audio processing method, encoding method or decoding method described in the present disclosure.
[0509] For example, the sound processing device, encoding device, or decoding device of the present disclosure may also include an FPGA or a CPU, etc. Furthermore, the sound processing device, encoding device, or decoding device may also include an interface for obtaining software for operating the FPGA or the CPU, etc. from the outside, and a memory for storing the obtained software. Furthermore, the FPGA or the CPU, etc. may also perform the signal processing described in the present disclosure by operating based on the stored software.
[0510] Alternatively, the server may provide software related to the audio processing, encoding processing, or decoding processing of the present disclosure. Furthermore, the terminal or device may operate as the audio processing device, encoding device, or decoding device described in the present disclosure by installing the software. Alternatively, the terminal or device may be connected to the server via a network to install the software.
[0511] In addition, another device different from the terminal or device may be connected to the server via a network to obtain data for software installation, and the software may be installed in the terminal or device by providing the software installation data to the terminal or device. In addition, an example of software may also be VR software or AR software for causing the terminal or device to execute the audio processing method described in the embodiment.
[0512] In addition, in the above-mentioned embodiments, each component may be formed by dedicated hardware, or implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.
[0513] In the above, the device and the like of one or more forms are described based on the embodiments, but the forms grasped by the present disclosure are not limited to the embodiments. As long as it does not depart from the main purpose of the present disclosure, the forms obtained by applying various modifications to the embodiments that can be thought of by those skilled in the art, and the forms constructed by combining the constituent elements in different modified examples are also included in the scope of one or more forms.
[0514] (Note)
[0515] The following techniques are disclosed through the description of the above embodiments.
[0516] (Technology 1) A sound processing device comprising a circuit and a memory, wherein the circuit uses the memory to obtain sound space information, the sound space information including information about a sound source in the sound space, information about an object in the sound space, and information about a position of a listener in the sound space; and uses the sound space information to calculate an evaluation value of a reflected sound generated corresponding to a sound generated from the sound source.
[0517] (Technique 2) In the sound processing device according to Technique 1, the circuit controls whether to select the reflected sound based on the evaluation value.
[0518] (Technique 3) In the sound processing device according to Technique 2, when the reflected sound is not selected, the circuit does not perform binaural processing on the reflected sound.
[0519] (Technique 4) In the sound processing device according to any one of Techniques 1 to 3, the circuit calculates the volume of the reflected sound, and calculates the evaluation value when the volume exceeds a predetermined threshold value.
[0520] (Technique 5) In the sound processing device described in Technique 2, when the reflected sound is selected based on the evaluation value, the circuit calculates the total computational load of one or more selected reflected sounds including the reflected sound, and when the total computational load exceeds a predetermined upper limit, the selection of the reflected sound is terminated.
[0521] (Technique 6) In the sound processing device according to Technique 5, the total calculation load is defined by the number of the one or more selected reflected sounds or the processing amount of the one or more selected reflected sounds.
[0522] (Technique 7) In the sound processing device as described in any one of Techniques 1 to 6, the circuit calculates the volume of each of the multiple reflected sounds generated as the reflected sound in the sound space, and calculates the evaluation value of the reflected sound for each of one or more reflected sounds among the multiple reflected sounds having a volume above a predetermined threshold.
[0523] (Technique 8) In the sound processing device described in Technique 7, the circuit calculates a total computational load of the one or more reflected sounds, and when the total computational load exceeds a predetermined upper limit, calculates the evaluation value of the one or more reflected sounds for each of the one or more reflected sounds.
[0524] (Technique 9) In the sound processing device as described in any one of Techniques 1 to 8, the circuit calculates the evaluation value of each of the multiple reflected sounds generated as the reflected sound in the sound space, and adds the computational load of the reflected sound to the total computational load for each of the multiple reflected sounds in descending order of the evaluation values, and each time the computational load of the reflected sound is added to the total computational load, the total computational load is compared with a predetermined upper limit, and when the total computational load obtained by adding the computational loads of the reflected sounds does not exceed the predetermined upper limit, the reflected sound is selected, and when the total computational load obtained by adding the computational loads of the reflected sounds exceeds the predetermined upper limit, one or more remaining reflected sounds after the reflected sound among the multiple reflected sounds are not selected.
[0525] (Technique 10) In the sound processing device as described in any one of Techniques 1 to 9, the evaluation value is the total value of at least one of an index value related to volume, a visual index value, an index value related to the object, and an index value representing the relationship between the direct sound corresponding to the reflected sound and the reflected sound.
[0526] (Technique 11) In the sound processing device according to Technique 10, the circuit increases the index value related to the volume as the volume of the sound generated by the sound source increases.
[0527] (Technique 12) In the sound processing device according to Technique 10 or 11, the circuit makes the visibility index value larger when the sound source enters the field of view of the listener than when the sound source does not enter the field of view of the listener.
[0528] (Technique 13) In the sound processing device according to any one of Techniques 10 to 12, the circuit increases the visual index value as the moving speed of the sound source is slower.
[0529] (Technique 14) In the sound processing device according to any one of Techniques 10 to 13, the index value related to the object is given for each object in the sound space and is included in the sound space information.
[0530] (Technique 15) In the sound processing device according to any one of Techniques 10 to 14, the circuit increases the index value representing the relationship between the direct sound and the reflected sound as the angle formed by the direction in which the direct sound arrives and the direction in which the reflected sound arrives increases.
[0531] (Technique 16) In the audio processing device described in any one of Techniques 10 to 15, the greater the difference between the distance from the sound source to the listener of the direct sound and the distance from the sound source to the listener of the reflected sound after reflection, the larger the index value representing the relationship between the direct sound and the reflected sound will be.
[0532] (Technique 17) In the sound processing device described in any one of Techniques 10 to 16, the more the amplitude value of the reflected sound exceeds the temporal masking threshold, the larger the index value representing the relationship between the direct sound and the reflected sound will be. The temporal masking threshold is the threshold of the temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is below the threshold.
[0533] (Technique 18) In the sound processing device described in any one of Techniques 10 to 17, the circuit repeatedly implements the following processing: among the multiple reflected sounds generated as the reflected sounds in the sound space, the index value related to the object related to the selected reflected sound is reduced, the evaluation value is calculated for the reflected sounds that have not yet been selected, and the reflected sounds are selected in descending order of the evaluation values; when the total computational load of one or more selected reflected sounds among the multiple reflected sounds exceeds a predetermined upper limit, the repeatedly implemented processing is terminated.
[0534] (Technology 19) An audio processing device comprises a circuit and a memory, wherein the circuit uses the memory to obtain information on the volume of a sound output from a sound source, uses the volume information to correct an evaluation value of a reflected sound corresponding to the sound, and controls whether to select the reflected sound based on the corrected evaluation value.
[0535] (Technique 20) In the sound processing device as described in Technique 19, the volume has a transition.
[0536] (Technique 21) A sound processing method, comprising: a step of obtaining sound space information, wherein the sound space information includes information about a sound source in the sound space, information about an object in the sound space, and information about a position of a listener in the sound space; and using the sound space information, calculating an evaluation value of a reflected sound generated corresponding to a sound generated from the sound source.
[0537] (Technique 22) A program for causing a computer to execute the sound processing method described in Technique 21.
[0538] Industrial Applicability
[0539] The present disclosure includes, for example, aspects that can be applied to an audio processing device, an encoding device, a decoding device, or a terminal or device including any of these devices.
[0540] Description of symbols
[0541] 1000 stereo sound reproduction system
[0542] 1001 Sound signal processing device (sound processing device)
[0543] 1002 Sound prompt device
[0544] 1100, 1120, 1500 encoding device
[0545] 1101, 1113 Input data
[0546] 1102 Encoder
[0547] 1103 Encoded Data
[0548] 1104, 1114, 1404, 1503 memory
[0549] 1110, 1130 Decoding device
[0550] 1111 Sound signal
[0551] 1112, 1200, 1210 decoders
[0552] 1121 Sending Department
[0553] 1122 Send signal
[0554] 1131 Receiving Department
[0555] 1132 Receiving signal
[0556] 1201, 1211 Space Information Management Department
[0557] 1202 Sound Data Decoder
[0558] 1203, 1213, 1300 Rendering Department
[0559] 1301 Analysis Department
[0560] 1302, 1314 Selection Department
[0561] 1303 Synthesis Department
[0562] 1311 Reverberation Processing Unit
[0563] 1312 Initial reflection processing unit
[0564] 1313 Distance Attenuation Processing Unit
[0565] 1315 Generation Department
[0566] 1316 Binaural Processing Department
[0567] 1401 Speaker
[0568] 1402, 1501 processors
[0569] 1403, 1502 communication IF
[0570] 1405 Sensor
Claims
1. A sound processing device, wherein: A circuit and a memory are provided, The circuit uses the memory, Acquiring sound space information, the sound space information including information of a sound source in the sound space, information of an object in the sound space, and information of a position of a listener in the sound space, An evaluation value of the reflected sound generated in response to the sound generated from the sound source is calculated using the sound space information.
2. The sound processing device according to claim 1, wherein: The circuit controls whether to select the reflected sound based on the evaluation value.
3. The sound processing device according to claim 2, wherein: In a case where the reflected sound is not selected, the circuit does not perform binaural processing on the reflected sound.
4. The sound processing device according to any one of claims 1 to 3, wherein: The circuit, Calculate the volume of the reflected sound, When the volume exceeds a predetermined threshold, the evaluation value is calculated.
5. The sound processing device according to claim 2, wherein: When the reflected sound is selected based on the evaluation value, The circuit, calculating a total computational load of one or more selected reflected sounds including the reflected sound, When the total calculation load exceeds a predetermined upper limit, the selection of the reflected sound is stopped.
6. The sound processing device according to claim 5, wherein: The total calculation load is defined by the number of the one or more selected reflected sounds or the processing amount of the one or more selected reflected sounds.
7. The sound processing device according to any one of claims 1 to 3, wherein: The circuit, calculating the volume of each of a plurality of reflected sounds generated as the reflected sounds in the sound space, For each of one or more reflected sounds having a volume equal to or larger than a predetermined threshold value among the plurality of reflected sounds, the evaluation value of the reflected sound is calculated.
8. The sound processing device according to claim 7, wherein: The circuit, calculating the total computational load of the one or more reflected sounds, When the total calculation load exceeds a predetermined upper limit, the evaluation value of the reflected sound is calculated for each of the one or more reflected sounds.
9. The sound processing device according to any one of claims 1 to 3, wherein: The circuit, calculating the evaluation value of each of a plurality of reflected sounds generated as the reflected sounds in the sound space, In descending order of the evaluation values, for each of the plurality of reflected sounds, the calculation load of the reflected sound is added to the total calculation load, Each time the calculation load of the reflected sound is added to the total calculation load, the total calculation load is compared with a predetermined upper limit. When the total calculation load obtained by adding the calculation loads of the reflected sounds does not exceed the predetermined upper limit, the reflected sound is selected, When the total calculation load obtained by adding the calculation loads of the reflected sounds exceeds the predetermined upper limit, one or more remaining reflected sounds after the reflected sound among the plurality of reflected sounds are not selected.
10. The sound processing device according to any one of claims 1 to 3, wherein: The evaluation value is a total value of at least one of an index value related to volume, a visual index value, an index value related to the object, and an index value indicating a relationship between a direct sound corresponding to the reflected sound and the reflected sound.
11. The sound processing device according to claim 10, wherein: The circuit increases the index value related to the volume as the volume of the sound generated by the sound source increases.
12. The sound processing device according to claim 10, wherein: The circuit increases the visibility index value when the sound source enters the field of vision of the listener compared to when the sound source does not enter the field of vision of the listener.
13. The sound processing device according to claim 10, wherein: The slower the moving speed of the sound source, the larger the value of the visibility index.
14. The sound processing device according to claim 10, wherein: The index value related to the object is assigned to each object in the sound space and is included in the sound space information.
15. The sound processing device according to claim 10, wherein: The circuit increases the index value indicating the relationship between the direct sound and the reflected sound as the angle formed by the direction in which the direct sound arrives and the direction in which the reflected sound arrives increases.
16. The sound processing device according to claim 10, wherein: The circuit increases the index value indicating the relationship between the direct sound and the reflected sound as the difference between the distance from the sound source to the listener and the distance from the sound source to the listener after reflection increases.
17. The sound processing device according to claim 10, wherein: The more the amplitude value of the reflected sound exceeds the temporal masking threshold, the larger the index value representing the relationship between the direct sound and the reflected sound will be. The temporal masking threshold is the threshold of the temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is below the threshold.
18. The sound processing device according to claim 10, wherein: The circuit, The following processing is repeatedly performed: among a plurality of reflected sounds generated as the reflected sounds in the sound space, the index value related to the object related to the selected reflected sound is reduced, the evaluation value is calculated for the reflected sounds that have not been selected, and the reflected sounds are selected in descending order of the evaluation values, When the total calculation load of one or more selected reflected sounds among the plurality of reflected sounds exceeds a predetermined upper limit, the repeatedly performed processing is terminated.
19. A sound processing device, wherein: Having circuit and memory, The circuit uses the memory, Get the volume information of the sound output from the sound source, An evaluation value of the reflected sound corresponding to the sound is corrected using the volume information, and whether or not to select the reflected sound is controlled based on the corrected evaluation value.
20. The sound processing device according to claim 19, wherein: The volume has a transition.
21. A method for processing sound, wherein: include: A step of acquiring sound space information, wherein the sound space information includes information about a sound source in the sound space, information about an object in the sound space, and information about a position of a listener in the sound space; as well as An evaluation value of the reflected sound generated in response to the sound generated from the sound source is calculated using the sound space information.
22. A program for causing a computer to execute the sound processing method according to claim 21.
Citation Information
Patent Citations
Signal processor
JP2019022049A
Apparatus and method for rendering a sound scene using pipeline stages
WO2021180938A1