Signal processing device, signal processing method, and program
Patent Information
- Application Number
- PCT/JP2026/011798
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026011798_01102026_PF_FP_ABST
Abstract
Description
Signal processing apparatus, signal processing method, and program
[0001] The present disclosure relates to a signal processing apparatus, a signal processing method, and a program.
[0002] Conventionally, a technology related to sound reproduction for allowing a user (that is, a listener) to perceive three-dimensional sound in a virtual three-dimensional space is known (see, for example, Patent Document 1). Further, in order to allow a user to perceive sound as coming from a sound source object to the user (more precisely, the user's avatar in the three-dimensional space) in such a three-dimensional space, processing for generating output sound information from original sound information is required. Here, in order to generate an output sound signal, information of the three-dimensional space and processing of a sound signal based on the information of the three-dimensional space are required.
[0003] Japanese Unexamined Patent Application Publication No. 2020-18620
[0004] Here, appropriate information may not be used as the information of the three-dimensional space in some cases. Accordingly, an object of the present disclosure is to provide a signal processing apparatus and the like that facilitate use of appropriate information as information of a three-dimensional space.
[0005] A signal processing apparatus according to an aspect of the present disclosure is a signal processing apparatus that reproduces content representing a virtual sound space, and includes: a first processing unit that updates information of the virtual sound space; a second processing unit that reproduces a sound signal of sound generated in the virtual sound space based on the updated information of the virtual sound space, and presents the sound signal to a user via the user's avatar in the virtual sound space; a first setting unit that sets frequency information indicating a processing frequency of the first processing unit; and a control unit that controls execution of processing by the first processing unit and the second processing unit, the control unit controlling processing of the first processing unit to be executed based on the set frequency information.
[0006] Furthermore, a signal processing method according to one aspect of the present disclosure is a signal processing method performed by a computer for reproducing content that represents a virtual sound space, and includes the steps of: updating information of the virtual sound space; reproducing sound signals of sounds generated in the virtual sound space based on the updated information of the virtual sound space and presenting them to the user via the user's avatar in the virtual sound space; and setting frequency information indicating the processing frequency of the updating step, wherein the processing of the updating step is performed based on the set frequency information.
[0007] Furthermore, one aspect of this disclosure can also be implemented as a program for causing the computer to execute the signal processing method.
[0008] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0009] According to this disclosure, a signal processing device and the like that facilitate the use of appropriate information are provided.
[0010] Figure 1 is a schematic diagram showing an example of use of the sound reproduction system according to the embodiment. Figure 2 is a block diagram showing the functional configuration of the sound reproduction system according to the embodiment. Figure 3 is a diagram illustrating an example of an audio signal according to the embodiment. Figure 4 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. Figure 5 is a block diagram showing the functional configuration of the path calculation unit according to the embodiment. Figure 6 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 7 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 8 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 9 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 10 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 11 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 12 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 13 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 14 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 15 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 16 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 17 is a diagram illustrating another example of the sound reproduction system according to the embodiment. Figure 18 is a flowchart of the signal processing method according to the embodiment. Figure 19 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 20 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 21 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 22 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 23 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 24 is a diagram illustrating the method for acquiring frequency information according to the embodiment. Figure 25 is a diagram illustrating the method for defining processor capability according to the embodiment.
[0011] (Knowledge forming the basis of the disclosure) Conventionally, there is a known technology for sound reproduction that allows users to perceive three-dimensional sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field and virtual sound space, or simply virtual space) (see, for example, Patent Document 1). By using this technology, users can perceive sound as if a sound source object exists at a predetermined location in the virtual space and the sound is coming from that direction. In order to localize a sound image at a predetermined location in a virtual three-dimensional space in this way, it is necessary to perform calculations on the sound signal emitted by the sound source object (also called the sound emitted by the sound source object or the reproduced sound) to generate a difference in the arrival time of the sound between the two ears, and a difference in the sound level (or sound pressure difference) between the two ears, so that the sound is perceived as three-dimensional. Such calculations are performed by applying a stereophonic sound filter. A stereophonic sound filter is an information processing filter that, when the output sound signal after applying the filter to the original sound information is reproduced, causes the position such as the direction and distance of the sound, the size of the sound source, and the size of the space to be perceived with a sense of three-dimensionality.
[0012] One example of the computational process for applying such a spatial sound filter is the process of convolving a head-related transfer function (HRP) onto the target sound signal to make it perceive the sound as coming from a predetermined direction. By performing this HRP convolution process at a sufficiently fine angle with respect to the direction of arrival of the reproduced sound from the sound source object's position to the user's position, the sense of presence experienced by the user is improved.
[0013] Furthermore, in recent years, there has been a great deal of development in virtual reality (VR) technology. In virtual reality, the main focus is on making it possible for the user to experience moving as if they were actually moving within a virtual space, by appropriately changing the position of sound source objects in a virtual three-dimensional space in response to the user's movements. To achieve this, it is necessary to move the localization position of sound images in the virtual sound space (or simply the virtual space) relative to the user's movements. This kind of processing has been done by applying a spatial sound filter, such as the head-related transfer function mentioned above, to the original sound information.
[0014] Here, signal processing for constructing a virtual space requires information such as how the virtual space is structured and what kinds of objects are placed at what locations (hereinafter also referred to as virtual space information), and processing to reproduce and present sound signals so that sounds generated in such a virtual space arrive at the user's avatar in the virtual space (also referred to as the user in the virtual space). Depending on the content, it may be necessary to update the virtual space information when objects in the virtual space move or when the user moves within the virtual space.
[0015] For example, in a virtual space where the user's avatar can move with 6 degrees of freedom, in order to play an audio signal, the avatar must be able to change its position, so the sound path (sound ray) arriving at the avatar changes moment by moment. In order to reflect this acoustic effect in the sound the user hears, it is necessary to know the position information assigned to the avatar moment by moment. Therefore, in a signal processing device that plays such an audio signal, a software module (first processing unit) for knowing the sound path (sound ray) arriving at the avatar moment by moment is required in addition to the software module (second processing unit) for playing the audio signal. Here, the second processing unit is responsible for playing the audio signal that the user will hear, and real-time performance is required because any interruption in the playback sound will be immediately perceived by the user as a quality defect. Therefore, the frequency of performing this processing, or the computing resources required for its performance, must be allocated without excess or deficiency. On the other hand, since the processing of the first processing unit deals with the degree of detail of sound, the frequency of performing the processing of the first processing unit, or the allocation of computing resources required for its performance, can be left to the discretion of the creator, administrator, or user of the virtual space.
[0016] Furthermore, the computer used to realize the virtual space may be a high-performance device such as a large-scale server or supercomputer, depending on the intended use of the virtual space. It may also be a CPU built into a personal mobile device such as a home PC, smartphone, or tablet, or a VR goggle that constitutes the virtual video space. Here, regarding the second processing unit, which requires real-time processing, it is essential to allocate the processing frequency and the necessary computing resources appropriately according to the content, regardless of the type of computer used to realize the virtual space. However, it is desirable to adjust the processing frequency of the first processing unit and the allocation of necessary computing resources as appropriate according to the capabilities of the computer that realizes the virtual space.
[0017] Furthermore, since there may be other sound source objects that can move within the virtual space besides the avatar (for example, cars, airplanes, people, animals, etc.), even if the avatar does not move, the aforementioned sound lines will fluctuate moment by moment as the position information assigned to each of the above sound source objects changes. In this case, the degree of fluctuation of the sound lines will depend heavily on the attributes of the sound source objects that make up the virtual space. For example, the speed at which the sound lines arrive at the user will differ greatly depending on whether the sound source object is a racing car or a cat, so it is desirable to appropriately adjust the frequency of processing performed by the first processing unit, or the allocation of computing resources required for execution, according to the attributes of the sound source objects that make up the virtual space.
[0018] Alternatively, it is desirable to appropriately adjust the frequency of processing performed by the first processing unit, or the allocation of computing resources required for its execution, according to the state of the sound source objects constituting the virtual space. For example, even in a virtual space where a racing car exists, the speed of sound ray fluctuations arriving differs greatly depending on whether the car is stationary, moving slowly, or racing. Therefore, it is desirable to appropriately adjust the frequency of processing performed by the first processing unit, or the allocation of computing resources required for its execution, according to the state of the sound source objects constituting the virtual space.
[0019] Therefore, in view of this perspective, this disclosure provides a signal processing device that can determine the frequency of processing performed by the first processing unit for determining the arrival path (sound ray) of sound arriving at an avatar in a virtual space having 6 degrees of freedom, or the allocation of computing resources necessary for such processing, according to the characteristics of the virtual space and the performance of the computer.
[0020] A more detailed overview of this disclosure is as follows:
[0021] A signal processing device according to a first aspect of the present disclosure is a signal processing device for reproducing content that represents a virtual sound space, comprising: a first processing unit for updating information of the virtual sound space; a second processing unit for reproducing sound signals of sounds generated in the virtual sound space based on the updated information of the virtual sound space and presenting them to the user via the user's avatar in the virtual sound space; a first setting unit for setting frequency information indicating the processing frequency of the first processing unit; and a control unit for controlling the execution of the processing of the first processing unit and the second processing unit, the control unit for controlling the processing of the first processing unit to be executed based on the set frequency information.
[0022] With such a signal processing device, the frequency at which the first processing unit performs its processing, or the allocation of computing resources required for its performance, can be appropriately adjusted according to the attributes of the sound source objects constituting the virtual space, by utilizing the frequency information set in the first setting unit. Therefore, by setting the frequency information set in the first setting unit to be appropriate, this appropriate frequency information will be used. Thus, it is possible to provide a signal processing device that makes it easier to use appropriate frequency information.
[0023] Furthermore, the signal processing device according to the second embodiment is the signal processing device described in the first embodiment, wherein the first setting unit is set with frequency information based on information associated with a virtual sound space.
[0024] According to this, appropriate frequency information based on information linked to the virtual sound space can be more easily used.
[0025] Furthermore, the signal processing device according to the third embodiment is the signal processing device according to the first or second embodiment, wherein the first setting unit is set with frequency information defined in the configuration information that defines the operating conditions of the virtual sound space.
[0026] According to this, appropriate frequency information defined within the configuration information that defines the operating conditions of the virtual sound space can be easily used.
[0027] Furthermore, the signal processing device according to the fourth embodiment is a signal processing device according to any of the first to third embodiments, comprising a second setting unit for setting the sampling frequency of an audio signal, and the first setting unit is set with frequency information corresponding to the set sampling frequency.
[0028] This makes it easier to use appropriate frequency information corresponding to the set sampling frequency.
[0029] Furthermore, the signal processing device according to the fifth embodiment is a signal processing device according to any of the first to fourth embodiments, comprising a third setting unit for setting the processing frame length in the processing of the second processing unit, and the first setting unit is set with frequency information corresponding to the set processing frame length.
[0030] This makes it easier to use appropriate frequency information corresponding to the set processing frame length.
[0031] Furthermore, the signal processing device according to the sixth embodiment is a signal processing device according to any of the first to fifth embodiments, wherein the processing of the first processing unit and the second processing unit is performed by a processor, and frequency information corresponding to the processing capacity of the processor is set in the first setting unit.
[0032] This makes it easier to use appropriate frequency information according to the processor's processing power.
[0033] Furthermore, the signal processing device according to the seventh embodiment is the signal processing device described in the sixth embodiment, wherein the first setting unit is set to frequency information related to the time resolution setting value for the processing of the first processing unit and the second processing unit, according to the processing capability of the processor.
[0034] According to this, appropriate frequency information related to the time resolution settings for the processing of the first and second processing units can be easily used depending on the processing power of the processor.
[0035] Furthermore, the signal processing device according to the eighth embodiment is the signal processing device described in the seventh embodiment, wherein the set value of the time resolution for the processing of the second processing unit includes the sampling frequency of the sound signal and the processing frame length in the processing of the second processing unit, and the signal processing device comprises a second setting unit for setting the sampling frequency and a third setting unit for setting the processing frame length.
[0036] According to this, the sampling frequency and the processing frame length can be set.
[0037] Furthermore, the signal processing device according to the ninth embodiment is a signal processing device according to any of the first to eighth embodiments, comprising: a second setting unit for setting the sampling frequency of an audio signal; and a third setting unit for setting the processing frame length in the processing of the second processing unit, wherein a region for specifying frequency information to be set in the first setting unit, a region for specifying the sampling frequency to be set in the second setting unit, and a region for specifying the processing frame length to be set in the third setting unit are arranged adjacent to each other.
[0038] According to this, the area for specifying frequency information set in the first setting unit, the area for specifying the sampling frequency set in the second setting unit, and the area for specifying the processing frame length set in the third setting unit can be placed adjacent to each other, enabling continuous writing and reading.
[0039] Furthermore, the signal processing device according to the tenth embodiment is a signal processing device according to any of the first to ninth embodiments, comprising a second setting unit for setting the sampling frequency of an audio signal and a third setting unit for setting the processing frame length in the processing of the second processing unit, wherein the first setting unit is set with frequency information corresponding to the set sampling frequency and the set processing frame length.
[0040] This makes it easier to use appropriate frequency information corresponding to the set sampling frequency and set processing frame length.
[0041] Further, the signal processing device according to the eleventh aspect is the signal processing device according to any one of the first to tenth aspects, wherein the first processing unit updates an arrival path of sound arriving at an avatar in a virtual sound space where the avatar is movable with 6DoF.
[0042] According to this, with the processing by the first processing unit, the arrival path of sound arriving at the avatar can be updated in the virtual sound space where the avatar is movable with 6DoF.
[0043] Further, the signal processing device according to the twelfth aspect is the signal processing device according to any one of the first to eleventh aspects, wherein the frequency information includes a fixed reference frequency and a coefficient of a variable value to be multiplied by the reference frequency.
[0044] According to this, frequency information including a fixed reference frequency and a coefficient of a variable value to be multiplied by the reference frequency can be set.
[0045] Further, the signal processing device according to the thirteenth aspect is the signal processing device according to any one of the first to twelfth aspects, wherein the frequency information is obtained from a bit stream that defines a configuration of a virtual sound space and set.
[0046] According to this, appropriate frequency information obtained from a bit stream that defines a configuration of a virtual sound space and set can be easily used.
[0047] Further, the signal processing device according to the fourteenth aspect is the signal processing device according to any one of the first to thirteenth aspects, wherein the frequency information is calculated and set based on a variation in information indicating a position of a sound source object included in the virtual sound space.
[0048] According to this, appropriate frequency information calculated and set based on a variation in information indicating a position of a sound source object included in the virtual sound space can be easily used.
[0049] Furthermore, the signal processing device according to the 15th embodiment is a signal processing device according to any of the 1st to 14th embodiments, wherein frequency information is obtained and set from at least one of a bitstream defining the configuration of the virtual sound space and configuration information defining the operating conditions of the virtual sound space, and if frequency information is obtained from both the bitstream defining the configuration of the virtual sound space and the configuration information defining the operating conditions of the virtual sound space, the one obtained from the configuration information defining the operating conditions of the virtual sound space is set.
[0050] According to this, frequency information can be obtained from at least one of the bitstream defining the structure of the virtual sound space and the configuration information defining the operating conditions of the virtual sound space. Furthermore, if frequency information is obtained from both the bitstream defining the structure of the virtual sound space and the configuration information defining the operating conditions of the virtual sound space, or if it is obtained from the configuration information defining the operating conditions of the virtual sound space, the frequency information obtained from the configuration information defining the operating conditions of the virtual sound space, which is more likely to be appropriate, will be more likely to be used.
[0051] Furthermore, the signal processing method according to the 16th embodiment is a signal processing method performed by a computer for reproducing content that represents a virtual sound space, and includes the steps of: updating information of the virtual sound space; reproducing sound signals of sounds generated in the virtual sound space based on the updated information of the virtual sound space and presenting them to the user via the user's avatar in the virtual sound space; and setting frequency information indicating the processing frequency of the updating step, wherein the processing of the updating step is performed based on the set frequency information.
[0052] According to this, the same effects as those of the signal processing device described above can be achieved.
[0053] Furthermore, the program relating to the 17th embodiment is a program that causes a computer to execute the signal processing method described above.
[0054] According to this, a computer can be used to achieve the same effect as the signal processing device described above.
[0055] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0056] The embodiments will be described in detail below with reference to the drawings. The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the embodiments below are examples only and are not intended to limit the disclosure. Furthermore, components in the embodiments below that are not described in an independent claim will be described as optional components. The figures are schematic diagrams and are not necessarily strictly accurate. Also, in the figures, substantially identical components are denoted by the same reference numerals, and redundant explanations may be omitted or simplified.
[0057] Furthermore, in the following explanation, elements may be assigned ordinal numbers such as 1st, 2nd, and 3rd. These ordinal numbers are assigned to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly assigned, or removed as appropriate.
[0058] Furthermore, in the following explanation, acoustic signals included in sound information may be described, and these acoustic signals may be expressed as speech signals or sound signals. In other words, in this disclosure, acoustic signals have the same meaning as speech signals or sound signals.
[0059] (Embodiment) [Overview] First, an overview of the sound reproduction system according to the embodiment will be described. Figure 1 is a schematic diagram showing an example of use of the sound reproduction system according to the embodiment. In Figure 1, a user 99 using the sound reproduction system 100 is shown.
[0060] The sound reproduction system 100 shown in Figure 1 is used, for example, simultaneously with a stereoscopic video playback device 300. By viewing stereoscopic images and stereoscopic sounds simultaneously, the images enhance the auditory sense of presence, and the sounds enhance the visual sense of presence, allowing the user to experience the scene as if they were actually there where the images and sounds were taken. For example, when an image (moving image) of people talking is displayed, even if the localization of the sound image (sound source object) of the conversation is not aligned with the person's mouth, it is known that the user 99 will still perceive it as conversational sound emanating from that person's mouth. In this way, the sense of presence can be enhanced by combining images and sounds, such as by correcting the position of the sound image based on visual information.
[0061] The stereoscopic image playback device 300 is an image display device worn on the head of the user 99. Therefore, the stereoscopic image playback device 300 moves integrally with the head of the user 99. For example, as shown in the figure, the stereoscopic image playback device 300 is a glasses-type device supported by the ears and nose of the user 99.
[0062] The stereoscopic image playback device 300 changes the displayed image in accordance with the user 99's head movements, thereby making the user 99 perceive that they are moving their head within the three-dimensional image space. In other words, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 turns to the right, the object moves to the left of the user 99, and if the user 99 turns to the left, the object moves to the right of the user 99. In this way, the stereoscopic image playback device 300 moves the three-dimensional image space in the opposite direction to the user 99's movements.
[0063] The stereoscopic image playback device 300 displays two images, each with a disparity in parallax, to the left and right eyes of the user 99. Based on the disparity in parallax of the displayed images, the user 99 can perceive the three-dimensional position of objects in the images. Note that the stereoscopic image playback device 300 does not need to be used simultaneously when the user 99 uses the system with their eyes closed, such as when the sound playback system 100 is used to play healing sounds for sleep induction. In other words, the stereoscopic image playback device 300 is not an essential component of this disclosure. In addition to a dedicated video display device, the stereoscopic image playback device 300 may also be a general-purpose mobile device such as a smartphone or tablet owned by the user 99.
[0064] Such general-purpose mobile devices are equipped with various sensors to detect the device's orientation and movement, in addition to a display for showing images. Furthermore, they are equipped with a processor for information processing and can connect to a network to send and receive information with server devices such as cloud servers. In other words, the stereoscopic video playback device 300 and the sound playback system 100 can also be realized by combining a smartphone with general-purpose headphones or similar devices that do not have information processing capabilities.
[0065] As in this example, a stereoscopic video playback device 300 and an audio playback system 100 may be realized by appropriately arranging a function for detecting head movements, a function for displaying images, a function for processing displaying image information, a function for displaying sound, and a function for processing displaying sound information in one or more devices. If the stereoscopic video playback device 300 is not required, it is sufficient to appropriately arrange a function for detecting head movements, a function for displaying sound, and a function for processing displaying sound information in one or more devices. For example, an audio playback system 100 can be realized by a processing device such as a computer or smartphone having a function for processing displaying sound information, and headphones having a function for detecting head movements and a function for displaying sound.
[0066] The sound reproduction system 100 is a sound presentation device worn on the user 99's head. Therefore, the sound reproduction system 100 moves integrally with the user 99's head. For example, the sound reproduction system 100 in this embodiment is a so-called over-ear headphone type device. However, there are no particular limitations on the form of the sound reproduction system 100; for example, it may consist of two earplug-type devices worn independently on the left and right ears of the user 99.
[0067] The sound reproduction system 100 changes the sound it presents in response to the user 99's head movements, thereby making the user 99 perceive that they are moving their head within a three-dimensional sound field. To this end, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the user 99's movements.
[0068] [Configuration] Next, the configuration of the sound reproduction system 100 according to this embodiment will be described with reference to Figure 2. Figure 2 is a block diagram showing the functional configuration of the sound reproduction system according to this embodiment.
[0069] As shown in Figure 2, the sound reproduction system 100 according to this embodiment includes a signal processing device 101, a communication module 102, a detector 103, a driver 104, and a database 105.
[0070] The signal processing device 101 is a computing device for performing various signal processing in the sound reproduction system 100. The signal processing device 101 is implemented, for example, by comprising a processor and memory, such as a computer, and by executing a program stored in memory by the processor. The execution of this program enables the functions of each functional unit described below.
[0071] The signal processing device 101 includes an acquisition unit 111, a path calculation unit 121, an output sound generation unit 131, and a signal output unit 141. Details of each functional unit of the signal processing device 101, along with details of the other components, will be described below.
[0072] The communication module 102 is an interface device for receiving sound information input to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. More specifically, the communication module 102 receives a wireless signal indicating sound information converted into a format for wireless communication using the antenna, and the signal converter converts the wireless signal back into sound information. As a result, the sound reproduction system 100 acquires sound information from an external device via wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. Thus, the acquisition unit 111 is an example of a sound acquisition unit. The sound information is input to the signal processing device 101 in the manner described above. Note that communication between the sound reproduction system 100 and the external device may be performed by wired communication.
[0073] The sound information acquired by the sound reproduction system 100 is encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound information includes information about the reproduced sound played by the sound reproduction system 100, and information about the localization position when the sound image of the sound is localized to a predetermined position in a three-dimensional sound field (i.e., perceived as a sound coming from a predetermined direction). The sound information can also be interpreted as information about the sound source object. In other words, the sound information includes the position of the sound source object in a three-dimensional sound field and the sound emitted by the sound source object.
[0074] As described above, sound information is obtained as input data and includes an audio signal (acoustic signal), which is information about the reproduced sound, and other information, which is information about the position of the sound source object in the three-dimensional sound field. Other information may also include information for defining the three-dimensional sound field. Therefore, sometimes the other information, including the position information of the sound source object and the information for defining the three-dimensional sound field, is referred to as spatial information. When considering the audio signal as the main component, the input data can be said to be sound information with other information (metadata) attached to the audio signal. When considering the spatial information as the main component, the input data can be said to be information with the audio signal attached to the spatial information. Alternatively, since the input data has both of these aspects, it can also be considered as sound spatial information.
[0075] As a concrete example, sound information includes information about multiple sounds, including a first and a second reproduced sound, and the sound image when each sound is reproduced is localized so that it is perceived as sound arriving from different positions in a three-dimensional sound field. For this reason, the sound source object for the first reproduced sound is localized to a first position in the three-dimensional sound field, and the sound source object for the second reproduced sound is localized to a second position in the three-dimensional sound field. Sound information may thus include multiple sounds. In other words, sound information may include multiple audio signals corresponding to the first and second reproduced sounds, and the positions of multiple sound source objects at a first and second position that correspond one-to-one with these audio signals.
[0076] Figure 3 is a diagram illustrating an example of an audio signal according to an embodiment. For example, as shown in Figure 3(a), the sound information may include an audio signal of a first direct sound that arrives at the user 99's position from a first position (from a first direction) and an audio signal of a second direct sound that arrives at the user 99's position from a second position (from a second direction). Note that the sound information acquired immediately afterward may contain only information about the playback sound. In this case, information regarding a predetermined position may be acquired separately, and subsequent processing may be performed when all of this information is available.
[0077] Alternatively, the sound information may include multiple audio signals and the location of a single sound source object that corresponds to those multiple audio signals in a many-to-one relationship. For example, such sound information is used in situations where multiple sounds are played from a sound source object. For instance, each of the multiple audio signals corresponds to a direct sound that arrives directly from the sound source object's location to the user's location, and a secondary sound (a sound produced by indirect propagation) that is generated along with the direct sound and arrives via a different path than the direct sound.
[0078] For example, as shown in Figure 3(b), the sound information immediately after acquisition contains the audio signal related to the direct sound. This is converted into sound information containing the respective audio signals of reverberation, first-order reflection, diffraction, etc., through a conversion process that calculates secondary sounds. This conversion process that calculates secondary sounds uses information about the spatial environment conditions of the three-dimensional sound field (e.g., the position of objects in the three-dimensional sound field, reflection, diffraction characteristics, etc.). Thus, secondary sounds are computationally generated from the sound information related to one reproduced sound based on the spatial environment conditions of the three-dimensional sound field, and are therefore not included in the sound information immediately after acquisition. Sound information containing these secondary sounds is generated through a conversion process that calculates secondary sounds. From one secondary sound, further secondary sounds may be generated by the propagation of that secondary sound. Note that the spatial environment condition information is part of the spatial information and may be acquired together with the audio signal by the input sound information. Alternatively, the audio signal and spatial information may be acquired separately. In other words, sound information may be acquired from a single file or bitstream, or it may be acquired separately by dividing it into multiple files or bitstreams. For example, the audio signal and spatial information may be obtained from separate files or bitstreams, or each of the audio signal and spatial information may be obtained from multiple files or bitstreams.
[0079] Thus, there are no particular limitations on the form of the input sound information; the sound reproduction system 100 simply needs to be equipped with an acquisition unit 111 that corresponds to various forms of sound information.
[0080] Here, an example of the acquisition unit 111 will be described using Figure 4. Figure 4 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. As shown in Figure 4, the acquisition unit 111 in this embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.
[0081] The encoded sound information input unit 112 is a processing unit that receives encoded sound information acquired by the acquisition unit 111. The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that decodes the sound information output from the encoded sound information input unit 112 to generate the playback sound and the position of the sound source object contained in the sound information in a format to be used in subsequent processing. The sensing information input unit 114 will be described below along with the function of the detector 103.
[0082] The detector 103 is a device for detecting the speed of the user 99's head movement. The detector 103 is composed of a combination of various sensors used for motion detection, such as a gyro sensor and an acceleration sensor. In this embodiment, the detector 103 is built into the sound playback system 100, but it may also be built into an external device, such as a stereoscopic image playback device 300 that operates in accordance with the user 99's head movement, similar to the sound playback system 100. In this case, the detector 103 does not have to be included in the sound playback system 100. Alternatively, the detector 103 may be an external imaging device that captures the user 99's head movement and processes the captured image to detect the user 99's movement.
[0083] The detector 103 is, for example, integrally fixed to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. Since the sound reproduction system 100, including the housing, moves integrally with the user 99's head after the user 99 puts it on, the detector 103 can consequently detect the speed of movement of the user 99's head.
[0084] The detector 103 may, for example, detect the amount of rotation of the user 99's head, with at least one of the three mutually orthogonal axes in three-dimensional space as the rotation axis, or it may detect the amount of displacement with at least one of the three axes as the displacement direction. Alternatively, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of movement of the user 99's head. In other words, the detector 103 may be capable of detecting changes in the user's 6DoF degrees of freedom in real space.
[0085] The sensing information input unit 114 acquires the movement velocity of the user 99's head from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of movement of the user 99's head detected by the detector 103 per unit time as the movement velocity. In this way, the sensing information input unit 114 acquires at least one of the rotational velocity and displacement velocity from the detector 103. The amount of movement of the user 99's head acquired here is used to determine the position and orientation (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. Therefore, the acquisition unit 111 also functions as a position acquisition unit by the sensing information input unit 114. In the sound reproduction system 100, the relative position of the sound source object to the user 99 is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced. Specifically, the above functions are realized by the path calculation unit 121 and the output sound generation unit 131.
[0086] The path calculation unit 121 includes an arrival direction calculation function that calculates the relative arrival direction of the playback sound from the sound source object to the user 99's position based on the determined coordinates and orientation of the user 99, and a conversion process that calculates the secondary sounds described above. Therefore, the path calculation unit 121 includes a function to calculate the propagation path from the sound source object and to calculate the secondary sounds that arrive at the user 99's position due to the indirect propagation of the playback sound according to the calculated propagation path of the playback sound, as well as the arrival direction of said secondary sounds. The arrival direction of secondary sounds includes additional information such as what kind of object the reflected sound is reflected by and the degree of attenuation during that reflection. This additional information is included in the arrival direction of secondary sounds calculated from the input sound information. In other words, the additional information is computationally generated and obtained from the sound information.
[0087] To summarize the spatial information, it includes the spatial position of the sound source object in space (three-dimensional sound field) (information about the position of the sound source object), the reflection and diffraction characteristics of sound at the sound source object (also information about the conditions of the spatial environment), and further information such as the size of the three-dimensional sound field. Based on the spatial information, the path calculation unit 121 generates secondary sounds depending on which sound source object the reproduced sound is reflected or diffracted by, and calculates the direction of arrival of the secondary sounds and the volume after the secondary sounds have been attenuated by reflection or diffraction as additional information. The sound information (input data) includes spatial information in the form of metadata attached to the audio signal, and as described above, this spatial information includes information other than the audio signal that is necessary to make the sound into a three-dimensional sound and position the sound source object in the three-dimensional sound field, and / or information used to calculate the information necessary to make the sound into a three-dimensional sound and position the sound source object in the three-dimensional sound field.
[0088] The path calculation unit 121 can be implemented by any processing method as long as it can calculate the direction of arrival of the reproduced sound when the reproduced sound reaches the user as direct sound, and calculate the direction of arrival of the secondary sound that arrives at the user 99's position due to the secondary propagation of the reproduced sound. Based on the user 99's coordinates and orientation, the path calculation unit 121 determines which direction in the three-dimensional sound field the reproduced sound and secondary sound should be perceived by the user 99 as coming from, and processes the sound information so that when the output sound signal is reproduced, it is perceived as such a sound.
[0089] Here, the route calculation unit 121 will be described in more detail with reference to Figure 5. Figure 5 is a block diagram showing the functional configuration of the route calculation unit according to the embodiment. As shown in Figure 5, the route calculation unit 121 includes an update processing unit 122 that updates spatial information, a route calculation unit 123 that performs the main function of the route calculation unit 121 as described above, and a memory 124.
[0090] The update processing unit 122 performs a process to update spatial information. In other words, the path calculation unit 123 can calculate the arrival path (sound trajectory) of the reproduced sound while updating the spatial information by using the spatial information updated by the update processing unit 122. Therefore, the path calculation unit 121, which includes the update processing unit 122 and the path calculation unit 123, is an example of a first processing unit. Furthermore, the processor of the signal processing device 101 that controls the execution of the process of the path calculation unit 121 is an example of a control unit.
[0091] Memory 124 is a storage device that stores values of information used to update spatial information. Specifically, memory 124 stores frequency information indicating the frequency of processing by the first processing unit. Therefore, memory 124 is an example of a first setting unit that sets frequency information. Memory 124 is provided in the path calculation unit 121, but it may be provided in other locations. For example, the memory of a signal processing device 101 (not shown) may be used to store frequency information. Frequency information is set in memory 124 by one of the following methods: for example, by including it in the content itself when encoding the content and acquiring it from the bitstream together with the content and storing it in memory 124; by determining it in a timely manner based on configuration information that defines the performance of the sound reproduction system 100 or the signal processing device 101 and the operating conditions of the three-dimensional sound field when decoding the content and storing it in memory 124; or by computationally determining it from the content being decoded (for example, fluctuations in information indicating the position of sound source objects included in the three-dimensional sound field) and storing it in memory 124.
[0092] The specific setting values for frequency information will be described later, but in setting the frequency information, the sampling frequency of the sound signal and the processing frame length of the process that reproduces the sound signal of the sound generated in the three-dimensional sound field and presents it to the user 99 in real space (the processing of the second processing unit, described later) may be used. For this reason, the setting values for the sampling frequency of the sound signal and the processing frame length of the processing of the second processing unit may be stored in the memory 124 or the memory of the signal processing unit 101 (not shown). Thus, the memory 124 or the memory of the signal processing unit 101 (not shown) can be said to be an example of a second setting unit and a third setting unit that set the sampling frequency and processing frame length. It is preferable that the memory areas that specify the frequency information, sampling frequency and processing frame length are adjacent areas. That is, it is preferable that the area that specifies the value of the frequency information of the first setting unit, the area that specifies the value of the sampling frequency of the second setting unit and the area that specifies the processing frame length of the third setting unit are continuous. This makes it possible to write the sampling frequency and processing frame length to and read them from memory in a series of processes.
[0093] The output sound generation unit 131 is a processing unit that generates an output sound signal by processing information related to the reproduced sound contained in the sound information.
[0094] In this embodiment, the output sound generation unit 131 obtains head-related transfer functions (HRTFs) used for generating output sound signals from the database 105. The database 105 is an information storage device that combines the functions of a storage device for storing information and a storage controller for reading stored information and outputting it to an external device. The database 105 stores HRTFs for each direction of arrival to the user 99. The HRTFs included in the database 105 are a set of general-purpose HRTFs that can be used by anyone, a set of HRTFs optimized for the individual user 99, or a set of publicly available HRTFs. The database 105 receives a query from the output sound generation unit 131 with the direction of arrival as the query, and outputs the HRTF corresponding to that direction of arrival to the output sound generation unit 131. The output sound generation unit 131 may also output the entire set of HRTFs, or output the characteristics of the HRTF set itself.
[0095] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and generates sound waves in the driver 104 based on the waveform signal, presenting sound to the user 99. The driver 104 has, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism in accordance with the waveform signal, and the drive mechanism vibrates the diaphragm. In this way, the driver 104 generates sound waves by the vibration of the diaphragm in accordance with the output sound signal (meaning "reproducing" the output sound signal, i.e., the user 99's perception is not included in the meaning of "reproduction"), the sound waves propagate through the air and are transmitted to the user 99's ears, and the user 99 perceives the sound.
[0096] Furthermore, the output sound generation unit 131 and the signal output unit 141 together can be considered as a second processing unit that reproduces the sound signals of sounds generated in a three-dimensional sound field and presents them to a user in real space via a user (avatar) in the three-dimensional sound field. In addition, the processor of the signal processing device 101 that controls the execution of the processing of the output sound generation unit 131 and the signal output unit 141 is an example of a control unit.
[0097] [Another Configuration Example] In the above example, the sound reproduction system 100 according to this embodiment was described as a sound presentation device comprising a signal processing device 101, a communication module 102, a detector 103, a database 105, and a driver 104. However, the functions of the sound reproduction system 100 may be realized by multiple devices or by a single device. This will be explained in detail using Figures 6 to 17. Figures 6 to 17 are diagrams illustrating another example of the sound reproduction system according to this embodiment.
[0098] For example, the information processing device 601 may be included in the voice presentation device 602, and the voice presentation device 602 may perform both acoustic processing and sound presentation. Alternatively, the information processing device 601 and the voice presentation device 602 may share the acoustic processing described herein, or a server connected to the information processing device 601 or the voice presentation device 602 via a network may perform some or all of the acoustic processing described herein.
[0099] In the above description, the information processing device is referred to as the 601. However, if the 601 encodes at least a portion of the data of spatial information used for audio signals or sound processing, and then decodes the resulting bitstream to perform sound processing, the 601 may also be called a decoding device, and the sound reproduction system 100 (i.e., the stereophonic sound reproduction system 600 in the figure) may also be called a decoding processing system.
[0100] Here, we will describe an example in which the sound reproduction system 100 functions as a decoding processing system.
[0101] <Example of an encoding device> Figure 7 is a functional block diagram showing the configuration of an encoding device 700, which is an example of an encoding device according to this disclosure.
[0102] The input data 701 is data to be encoded, including spatial information and / or audio signals, which are input to the encoder 702. Details of the spatial information will be explained later.
[0103] The encoder 702 encodes the input data 701 to generate encoded data 703. The encoded data 703 is, for example, a bitstream generated by the encoding process.
[0104] Memory 704 stores the encoded data 703. Memory 704 may be, for example, a hard disk or an SSD (Solid State Drive), or other storage device.
[0105] In the above description, a bitstream generated by the encoding process was given as an example of encoded data 703 stored in memory 704, but other data besides a bitstream may also be used. For example, the encoding device 700 may store converted data generated by converting the bitstream to a predetermined data format in memory 704. The converted data may be, for example, a file or multiplexed stream containing one or more bitstreams. Here, the file is, for example, a file having a file format such as ISOBMFF (ISO Base Media File Format). The encoded data 703 may also be in the form of multiple packets generated by dividing the above bitstream or file. When converting a bitstream generated by encoder 702 to data different from a bitstream, the encoding device 700 may include a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).
[0106] <Example of a Decryption Device> Figure 8 is a functional block diagram showing the configuration of a decoding device 800, which is an example of a decoding device according to the present disclosure.
[0107] Memory 804 stores, for example, the same data as the encoded data 703 generated by the encoding device 700. Memory 804 reads the stored data and inputs it to the decoder 802 as input data 803. Input data 803 is, for example, the bitstream to be decoded. Memory 804 may be, for example, a hard disk or SSD, or other storage device.
[0108] The decoding device 800 may not use the data stored in memory 804 as input data 803 directly, but may use the converted data generated by converting the read data as input data 803. The data before conversion may be, for example, multiplexed data containing one or more bitstreams. Here, the multiplexed data may be a file having a file format such as ISOBMFF. Alternatively, the data before conversion may be in the form of multiple packets generated by dividing the above bitstream or file. When converting data different from the bitstream read from memory 804 into a bitstream, the decoding device 800 may include a conversion unit (not shown), or the conversion process may be performed by the CPU.
[0109] The decoder 802 decodes the input data 803 and generates an audio signal 801 that is presented to the listener.
[0110] <Another Example of an Encoding Device> Figure 9 is a functional block diagram showing the configuration of an encoding device 900, which is another example of an encoding device of the present disclosure. In Figure 9, the same reference numerals are used for components that have the same functions as those in Figure 7, and these components will not be described.
[0111] The encoding device 700 has a memory 704 for storing encoded data 703, whereas the encoding device 900 differs from the encoding device 700 in that it has a transmission unit 901 for transmitting encoded data 703 to the outside.
[0112] The transmitting unit 901 transmits a transmission signal 902 to another device or server based on encoded data 703 or data in another data format generated by converting the encoded data 703. The data used to generate the transmission signal 902 is, for example, a bitstream, multiplexed data, file, or packet as described in the encoding device 700.
[0113] <Another Example of a Combined Device> Figure 10 is a functional block diagram showing the configuration of a decoding device 1000, which is another example of a decoding device of the present disclosure. In Figure 10, the same reference numerals are used for components that have the same functions as those in Figure 8, and these components will not be described.
[0114] Decoding device 800 is equipped with a memory 804 for reading input data 803, whereas decoding device 1000 is equipped with a receiving unit 1001 for receiving input data 803 from an external source, thus differing from decoding device 800.
[0115] The receiving unit 1001 receives the received signal 1002, acquires the received data, and outputs the input data 803 to be input to the decoder 802. The received data may be the same as the input data 803 input to the decoder 802, or it may be data in a different data format than the input data 803. If the received data is in a different data format than the input data 803, the receiving unit 1001 may convert the received data to the input data 803, or a conversion unit or CPU (not shown) in the decoding device 1000 may convert the received data to the input data 803. The received data may be, for example, a bitstream, multiplexed data, file, or packet as described in the encoding device 900.
[0116] <Decoder Function Description> Figure 11 is a functional block diagram showing the configuration of decoder 1100, which is an example of decoder 802 in Figure 8 or Figure 10.
[0117] The input data 803 is an encoded bitstream and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.
[0118] The spatial information management unit 1101 acquires metadata contained in the input data 803 and analyzes the metadata. The metadata includes information describing elements that act on sounds placed in the sound space. The spatial information management unit 1101 manages the spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1103. In this disclosure, the information used for acoustic processing is called spatial information, but it may be called by other names. The information used for acoustic processing may be called, for example, sound space information or scene information. Furthermore, if the information used for acoustic processing changes over time, the spatial information input to the rendering unit 1103 may be called spatial state, sound space state, scene state, etc.
[0119] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, each room may be managed as a scene with a different sound space, or even if it is the same space, the spatial information may be managed as different scenes depending on the scene being represented. In managing spatial information, an identifier may be assigned to identify each piece of spatial information. The spatial information data may be included in a bitstream, which is a form of input data 803, or the bitstream may contain the spatial information identifier, and the spatial information data may be obtained from a source other than the bitstream. If the bitstream contains only the spatial information identifier, the spatial information data stored in the memory of the sound signal processing device or an external server may be obtained as input data using the spatial information identifier during rendering.
[0120] Furthermore, the information managed by the spatial information management unit 1101 is not limited to the information contained in the bitstream. For example, the input data 803 may include data not included in the bitstream, such as data indicating the characteristics or structure of the space obtained from a software application or server providing VR or AR. Also, for example, the input data 803 may include data not included in the bitstream, such as data indicating the characteristics or location of a listener or object. In addition, the input data 803 may include information indicating the location of a listener, such as information obtained by a sensor equipped in a terminal including a decoding device, or information indicating the location of a terminal estimated based on information obtained by the sensor. In other words, the spatial information management unit 1101 may communicate with an external system or server to obtain spatial information and the location of a listener. Furthermore, the spatial information management unit 1101 may obtain clock synchronization information from an external system and execute a process to synchronize with the clock of the rendering unit 1103. Furthermore, the space described above may be a virtually created space, i.e., a VR space, or a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR (Mixed Reality) space. The virtual space may also be called a sound field or sound space. In addition, the position information described above may be information such as coordinate values indicating a position within space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within space.
[0121] The audio data decoder 1102 decodes the encoded audio data contained in the input data 803 to obtain an audio signal.
[0122] The encoded audio data acquired by the 3D audio playback system 600 is a bitstream encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream; the system may also include bitstreams encoded using other encoding methods as encoded audio data. For example, the encoding scheme used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis; or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec); or any other encoding scheme may be used. For example, PCM (Pulse Code Modulation) data may be considered a type of encoded audio data. In this case, the decoding process may, for example, be a process of converting the N-bit binary number of the PCM data into a number format that the rendering unit 1103 can process (e.g., floating-point format) if the number of quantization bits of the PCM data is N.
[0123] The rendering unit 1103 takes an audio signal and spatial information as input, applies acoustic processing to the audio signal using the spatial information, and outputs the acoustically processed audio signal 801.
[0124] Before rendering begins, the spatial information management unit 1101 reads the metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and transmits them to the rendering unit 1103. After rendering begins, the spatial information management unit 1101 grasps the temporal changes in the spatial information and the position of the listener, and updates and manages the spatial information. Then, the spatial information management unit 1101 transmits the updated spatial information to the rendering unit 1103. The rendering unit 1103 generates and outputs an audio signal with added acoustic processing based on the audio signal contained in the input data and the spatial information received from the spatial information management unit 1101.
[0125] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit 1101 and the rendering unit 1103 may be assigned to separate threads. If the spatial information update process and the audio signal output process with added acoustic processing are handled in different threads, the thread startup frequencies may be set individually, or the processes may be executed in parallel.
[0126] By having the spatial information management unit 1101 and the rendering unit 1103 execute processing in separate, independent threads, computing resources can be preferentially allocated to the rendering unit 1103. This allows for safe execution of sound output processing where even slight delays are unacceptable, such as when a delay of even one sample (0.02 msec) would cause a popping noise. In this case, the allocation of computing resources to the spatial information management unit 1101 is limited. However, updating spatial information is a low-frequency process compared to audio signal output processing (for example, updating the direction of the listener's face). Therefore, it does not necessarily need to respond instantaneously like audio signal output processing, and limiting the allocation of computing resources does not significantly affect the acoustic quality provided to the listener.
[0127] Spatial information updates may be performed periodically at predetermined times or periods, or when predetermined conditions are met. Furthermore, spatial information updates may be performed manually by a listener or the sound space administrator, or triggered by changes in an external system. For example, if a listener operates a controller to instantaneously warp their avatar's position, instantly advance or rewind time, or if the virtual space administrator suddenly alters the environment, the thread where the spatial information management unit 1101 is located may be activated as a one-off interrupt in addition to its periodic activation.
[0128] The role of the information update thread that performs spatial information update processing is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. This is handled within a processing thread that is activated at a relatively low frequency of about tens of Hz. It is also possible to have the processing that reflects the properties of direct sound performed in such an infrequently occurring processing thread. This is because the frequency at which the properties of direct sound change is lower than the frequency at which audio processing frames for audio output occur. In fact, doing so can relatively reduce the computational load of the processing, and it also avoids the risk of generating pulsating noise, which can occur if information is updated at an unnecessarily fast frequency.
[0129] Figure 12 is a functional block diagram showing the configuration of decoder 1200, which is another example of decoder 802 in Figure 8 or Figure 10.
[0130] Figure 12 differs from Figure 11 in that the input data 803 includes an unencoded audio signal rather than encoded audio data. The input data 803 includes a bitstream containing metadata and an audio signal.
[0131] The spatial information management unit 1201 is the same as the spatial information management unit 1101 in Figure 11, so its explanation is omitted.
[0132] The rendering unit 1202 is the same as the rendering unit 1103 in Figure 11, so its description is omitted.
[0133] In the above explanation, the configuration shown in Figure 12 is referred to as a decoder, but it may also be called an acoustic processing unit that performs acoustic processing. Furthermore, a device including an acoustic processing unit may be called an acoustic processing unit instead of a decoding device. Also, an acoustic signal processing unit (information processing unit 601) may be called an acoustic processing unit.
[0134] <Physical Configuration of the Encoding Device> Figure 13 shows an example of the physical configuration of the encoding device. The encoding device shown in Figure 13 is an example of the encoding devices 700 and 900 described above.
[0135] The encoding device shown in Figure 13 comprises a processor, memory, and a communication interface.
[0136] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the encoding process of this disclosure may be performed by the CPU, DSP, or GPU executing a program stored in memory. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of this disclosure.
[0137] Memory consists of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). Memory may also include magnetic storage media such as hard disks or semiconductor memory such as SSDs (Solid State Drives). Furthermore, the term "memory" may also include internal memory built into the CPU or GPU.
[0138] A communication interface (IF) is a communication module that supports communication methods such as Bluetooth® or WIGIG®. The encoding device has the function of communicating with other communication devices via the communication interface and transmits the encoded bitstream.
[0139] The communication module consists of, for example, a signal processing circuit and an antenna that correspond to the communication method. In the above example, Bluetooth® or WIGIG® were given as examples of communication methods, but it may also support communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi®. Furthermore, the communication interface may be a wired communication method such as Ethernet®, USB (Universal Serial Bus), or HDMI® (High-Definition Multimedia Interface), rather than the wireless communication methods described above.
[0140] <Physical Configuration of the Acoustic Signal Processing Device> Figure 14 shows an example of the physical configuration of the acoustic signal processing device. Note that the acoustic signal processing device in Figure 14 may be a decoding device. Also, some of the configuration described here may be provided in the voice presentation device 602. Furthermore, the acoustic signal processing device shown in Figure 14 is an example of the acoustic signal processing device 601 described above.
[0141] The acoustic signal processing device shown in Figure 14 comprises a processor, memory, a communication interface, a sensor, and a speaker.
[0142] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may perform the acoustic processing or decoding processing of this disclosure by executing a program stored in memory. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the acoustic processing of this disclosure.
[0143] Memory consists of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). Memory may also include magnetic storage media such as hard disks or semiconductor memory such as SSDs (Solid State Drives). Furthermore, the term "memory" may also include internal memory built into the CPU or GPU.
[0144] A communication interface (IF) is a communication module that supports communication methods such as Bluetooth® or WIGIG®. The acoustic signal processing device shown in Figure 14 has the function of communicating with other communication devices via the communication interface and acquires the bitstream to be decoded. The acquired bitstream is stored in memory, for example.
[0145] The communication module consists of, for example, a signal processing circuit and an antenna that correspond to the communication method. In the above example, Bluetooth® or WIGIG® were given as examples of communication methods, but it may also support communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi®. Furthermore, the communication interface may be a wired communication method such as Ethernet®, USB (Universal Serial Bus), or HDMI® (High-Definition Multimedia Interface), rather than the wireless communication methods described above.
[0146] The sensor performs sensing to estimate the listener's position or orientation. Specifically, the sensor estimates the listener's position and / or orientation based on one or more detection results from the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the whole of the listener's body, such as the head, and generates position information indicating the listener's position and / or orientation. This position information may indicate the listener's position and / or orientation in real space, or it may indicate the displacement of the listener's position and / or orientation relative to the listener's position and / or orientation at a predetermined point in time. Furthermore, the position information may indicate the relative position and / or orientation to the stereophonic sound reproduction system or an external device equipped with a sensor.
[0147] The sensor may be, for example, an imaging device such as a camera or a distance measuring device such as LiDAR (Laser Imaging Detection and Ranging), and may detect the listener's head movement by imaging the listener's head movement and processing the captured image. Alternatively, a device that performs position estimation using wireless communication in an arbitrary frequency band such as millimeter waves may be used as the sensor.
[0148] The acoustic signal processing device shown in Figure 14 may acquire position information via a communication interface from an external device equipped with a sensor. In this case, the acoustic signal processing device does not need to include a sensor. Here, the external device is, for example, the audio presentation device 602 described in Figure 6 or a stereoscopic image playback device attached to the listener's head. In this case, the sensor is composed of a combination of various sensors, such as a gyro sensor and an acceleration sensor.
[0149] The sensor may, for example, detect the angular velocity of rotation with at least one of three mutually orthogonal axes in the sound space as the axis of rotation, or it may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.
[0150] The sensor may, for example, detect the amount of rotation around at least one of three mutually orthogonal axes in the sound space as the axis of rotation, or it may detect the amount of displacement around at least one of the three axes as the direction of displacement, as the amount of movement of the listener's head. Specifically, the sensor detects the listener's position in 6DoF (position (x, y, z) and angle (yaw, pitch, roll)). The sensor is composed of a combination of various sensors used for motion detection, such as a gyroscope and an accelerometer.
[0151] The sensor only needs to be able to detect the listener's position and may be implemented using a camera or a GPS (Global Positioning System) receiver. Alternatively, positional information obtained by self-position estimation using LiDAR (Laser Imaging Detection and Ranging) may be used. For example, if the audio signal playback system is implemented using a smartphone, the sensor can be built into the smartphone.
[0152] The sensors may also include temperature sensors such as thermocouples for detecting the temperature of the acoustic signal processing device shown in Figure 14, and sensors for detecting the remaining battery level of the acoustic signal processing device or a battery connected to the acoustic signal processing device.
[0153] A speaker, for example, has a diaphragm, a drive mechanism such as a magnet or voice coil, and an amplifier, and presents the processed audio signal as sound to the listener. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal showing the waveform of the sound) amplified via the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, allowing the listener to perceive the sound.
[0154] In this explanation, we have used the example of the acoustic signal processing device shown in Figure 14, which includes a speaker and presents the processed audio signal through the speaker. However, the means for presenting the audio signal is not limited to the above configuration. For example, the processed audio signal may be output to an external audio presentation device 602 connected via a communication module. The communication performed by the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in Figure 14 may have a terminal for outputting an analog audio signal, and the audio signal may be presented through an earphone or the like by connecting a cable to the terminal. In the above case, the audio presentation device 602, such as headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, or a surround speaker composed of multiple fixed speakers, reproduces the audio signal.
[0155] <Explanation of Rendering Unit Functions> Figure 15 is a functional block diagram showing an example of the detailed configuration of rendering units 1103 and 1202 in Figures 11 and 12. The rendering unit consists of a pipeline processing unit and a playback unit, and adds acoustic processing to the sound data contained in the input signal and outputs it.
[0156] The input signal consists, for example, of spatial information, sensor information, and sound data. The input signal may also include a bitstream consisting of sound data and metadata (control information), in which case the metadata may include spatial information. In addition, spatial information may be input to the pipeline processing unit separately from the input signal. 22.1ch broadcast data is an example of sound data included in the input signal.
[0157] Spatial information refers to information about the sound space (three-dimensional sound field) created by a 3D sound reproduction system, and consists of information about objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sounding objects that do not emit sound. Non-sounding objects function as obstacle objects that reflect sound emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sound emitted by other sound source objects.
[0158] Information that is attached to both sound source objects and non-sounding objects includes positional information, shape information, and the rate of sound attenuation when the object reflects sound.
[0159] Position information is represented by coordinate values on three axes in Euclidean space, for example, the X, Y, and Z axes, but it does not necessarily have to be three-dimensional information. For example, it may be two-dimensional information represented by coordinate values on two axes, the X and Y axes. The position information of an object is determined by the representative position of the shape represented by a mesh or voxel.
[0160] For example, information representing a virtual studio, which is an example of a virtual space, is an example of information that represents the size and shape of the space included in spatial information. In other words, the virtual space described herein is not limited to a virtual studio. For example, a virtual space may be formed by overlaying it onto the interior of a car in real space, or virtual speakers may be placed inside the car to play content.
[0161] For example, information about a virtual speaker is an example of information about a sound source object included in spatial information. Relative position information is an example of information assigned to a sound source object. The sound source object in this specification is not limited to a virtual speaker.
[0162] For example, information representing a virtual screen placed in a virtual studio is an example of object information included in spatial information. Object information includes the position information of the object in the virtual space. In this specification, based on this position information, the position where an object such as a virtual screen is placed can be set as the reference position, and the direction in which the reference position is viewed from a position directly facing the object can be set as "forward." Shape information included in the object information may include information such as "front," "side," and "back" associated with multiple surface areas of the shape, and this information may be used to identify a position directly facing the object. Alternatively, the position from which the image projected on the virtual screen is viewed from the front may be set as the "position directly facing." In addition, a position on a line segment perpendicular to the virtual screen, where the center (centroid) of the screen is the foot of the perpendicular, may be set as the "position directly facing."
[0163] Shape information may include information about the surface material.
[0164] Furthermore, the information may include whether or not the object belongs to a living organism, or whether or not the object is a moving object. If the object is a moving object, the positional information may move over time, and the changed positional information or the amount of change is transmitted to the rendering unit. In addition to the information provided above for both sound source objects and non-sounding objects, the information regarding the sound source object includes sound data and information necessary to radiate the sound data into the sound space.
[0165] Audio data is data that represents the sound perceived by the listener, including information about the frequency and intensity of the sound. Audio data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In that case, the signal needs to be decoded at least before it reaches the playback unit, so the rendering unit may include a decoding unit (not shown). Alternatively, it may be decoded by the audio data decoder 1102.
[0166] A sound source object only needs to have at least one sound data set, but it may have multiple sound data sets. Furthermore, identification information may be assigned to each sound data, and this identification information may be stored as information about the sound source object.
[0167] Information necessary for radiating sound data into sound space may include, for example, information on the reference volume used as a reference when playing back the sound data, information indicating the properties (also called characteristics) of the sound data, information on the position of the sound source object, information on the orientation of the sound source object, and information on the directivity of the sound emitted by the sound source object. The reference volume information may be, for example, the effective value of the amplitude of the sound data at the sound source position when radiating the sound data into sound space, and may be expressed as a floating-point value in decibels (dB).
[0168] Figure 16 is a flowchart showing the operation of the audio signal processing device according to this embodiment. After acquiring the input signal (S101), the pipeline processing unit performs pipeline processing (S102). Specifically, the pipeline processing unit analyzes the input signal input to the audio signal processing device and calculates the information necessary for generating direct sound and reflected sound, as well as the information necessary for selecting the reflected sound to be generated.
[0169] First, the characteristics of both direct and reflected sound are calculated. Specifically, the arrival time and volume of each sound upon arrival at the listener are calculated. If multiple objects functioning as reflectors exist in the sound space, the characteristics of the reflected sound are calculated for each object.
[0170] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is the path connecting the sound source object's position information S (xs, ys, zs) and the listener's position information A (xa, ya, za). The direct sound arrival time (td) is the value obtained by dividing the length of the path connecting position information S (xs, ys, zs) and position information A (xa, ya, za) by the speed of sound (approximately 340 m / s). For example, the path length (X) can be calculated as (xs - xa)^2 + (ys - ya)^2 + (zs - za)^2)^0.5. Since volume attenuates inversely proportional to distance, the volume at the arrival of the direct sound (ld) can be calculated as ld = N * U / X, where N is the volume at the sound source object's position S (xs, ys, zs) and U is the unit distance.
[0171] The time to arrive of the reflected sound (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is the path connecting the position of the reflected sound image to the position information A (xa, ya, za). The position of the reflected sound image may be derived using methods such as the "image method" or "ray tracing method," or any other arbitrary method for deriving the sound image position. The image method is a technique that simulates the sound image by assuming that a mirror image exists at a position symmetrical to the sound source relative to the wall surface of the reflected wave from the wall surface, and that the sound wave is radiated from the position of that mirror image. Ray tracing is a technique that simulates an image (sound image) observed at a certain point by tracing waves that propagate linearly, such as light rays and sound rays.
[0172] By assuming that the sound image of the reflected sound is formed symmetrically across a wall from the sound source, the position of the reflected sound image can be determined along the x, y, and z axes, and the arrival time of the reflected sound can be calculated in the same way as the arrival time of the direct sound.
[0173] The time to arrival of the reflected sound (tr) is obtained by dividing the length (Y) of the path connecting the position of the reflected sound image and the position information A (xa, ya, za) by the speed of sound (approximately 340 m / sec). The volume of the reflected sound upon arrival (lr) is given by lr = N * G * U / Y, where N is the volume at the sound source, U is the unit distance, and G is the volume attenuation rate in the reflection, because volume attenuates inversely proportional to distance. As explained earlier, the attenuation rate G may be expressed as a real number less than or equal to 1 and greater than or equal to 0, or as a negative decibel value. In this case, the volume of the entire signal will be attenuated by G. Alternatively, the attenuation rate may be set for each frequency band that makes up multiple frequency bands. In this case, the specified attenuation rate is multiplied by each frequency component of the signal. Furthermore, to reduce the amount of computation, the overall attenuation rate may be represented by a representative value or average value among the attenuation rates for each frequency band, and the volume of the entire signal may be attenuated by that amount.
[0174] Next, in the reflected sound selection step, the pipeline processing unit selects whether or not to generate the calculated reflected sound. The choice of whether or not to generate the reflected sound may also be a choice between "playing the direct sound and the reflected sound separately, or combining the direct sound and the reflected sound."
[0175] Next, in the step of generating the output signal (S103), for example, direct sound and reflected sound are generated. The playback unit generates the audio signal of the direct sound and the audio signal of the reflected sound that the selection unit has selected to generate.
[0176] Here, the playback unit may perform volume compensation processing based on the unselected reflected sound. The direct sound audio signal is generated by applying the arrival time (td) and arrival volume (ld) calculated by the pipeline processing unit to the sound data of the sound source object included in the input information. Specifically, the sound data is delayed by td and then multiplied by (ld) to perform a delay process. The delay process is a process that moves the position of the sound data forward or backward on the time axis. For example, existing methods that perform delay processing without degrading sound quality may be applied.
[0177] The reflected sound signal is generated by applying the arrival time (tr) and arrival volume (ld), calculated by the pipeline processing unit, to the sound data of the sound source object. The method for applying the arrival time and arrival volume can be the same as the process for generating the direct sound. However, the arrival volume (lr) in the generation of the reflected sound differs from the arrival volume of the direct sound in that it is a value to which the attenuation rate G of the volume in reflection has been applied. G may be an attenuation rate applied collectively to the entire frequency band, but in order to reflect the bias in frequency components caused by reflection, the reflectance rate may be defined for each predetermined frequency band. In that case, the process of applying the arrival volume (lr) may be carried out as a frequency equalizer process, which is a process of multiplying each band by the attenuation rate. Figure 17 is a block diagram showing an example configuration for the rendering unit 1300 to perform pipeline processing.
[0178] The rendering unit 1300 in Figure 17 includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may be composed of the multiple components of the rendering unit shown in Figure 15, or at least some of the multiple components of the acoustic signal processing device shown in Figure 14.
[0179] Pipelining refers to the process of dividing the process of applying sound effects into multiple processes and executing them one by one in sequence. Each of these processes may perform, for example, signal processing on the audio signal or the generation of parameters used in the signal processing.
[0180] The rendering unit 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, and binaural processing as pipeline processing. However, these processing are just examples, and pipeline processing may include other processing or may omit some of the processing. For example, pipeline processing may include diffraction processing and occlusion processing. Also, for example, reverberation processing may be omitted if it is not necessary. Furthermore, not all sounds are processed in the binaural processing stage.
[0181] Furthermore, each process may be referred to as a stage. Also, the resulting audio signals, such as reflected sound, may be referred to as rendering items. The number of stages in a pipeline process, and their order, are not limited to the example shown in Figure 17.
[0182] [Operation] Here, Figure 18 is a flowchart of the signal processing method according to the embodiment. The operation example shown in the figure shows the operation after the acquisition unit 111 receives information via the communication module. As shown in the figure, the frequency information is set by storing setting values in the memory 124 based on the acquired sound information, configuration information, and the processing performance of the processor constituting the sound reproduction system 100 or the signal processing device 101 (S11).
[0183] Next, the path calculation unit 121 updates the spatial information at a frequency based on the frequency information set in the memory 124 using the update processing unit 122 (S12). As the spatial information is updated, sound will arrive from the sound source object to the user (avatar) in the three-dimensional sound field via an arrival path that changes with each update.
[0184] In this way, based on the updated spatial information, sound signals can be reproduced and presented to the user according to arrival paths that change at a frequency suitable for various conditions (S13).
[0185] [Frequency Information] The frequency information settings will be explained in more detail below. The frequency information indicates the frequency of processing by the first processing unit, that is, the processing performed by the route calculation unit 121. The processing of the first processing unit and the processing of the second processing unit are either executed by the first processing unit interrupting the processing of the second processing unit, or these processes are executed in parallel. At that time, the frequency of how many times the processing of the first processing unit is executed per second is set as the frequency information.
[0186] In acquiring (or determining) frequency information, there is a method of obtaining frequency information from information pre-embedded in the bitstream. Specifically, SOs (Scene Objects: people, objects, and sound sources placed in a three-dimensional sound field) as defined in the MPEG-I standard are assigned information called "isStatic," and this information is transmitted to the renderer (decoder) via the bitstream. For example, objectSources (sound source objects) have a flag called "objectSourceIsStatic" associated with them, which is multiplexed within the bitstream. If this flag is true (i.e., static), the sound source object will not move during the playback time of that scene.
[0187] This applies not only to sound source objects but also to any object that can be the subject of sound reflection or diffraction. For example, meshes (object models made of polygon meshes) also have a flag called "isMeshStatic" associated with them, which is multiplexed within the bitstream. If this flag is true (i.e., static), the object will not move during the playback time of that scene.
[0188] In addition, for sound sources, "isStatic" information is similarly assigned to hoaSource, channelSource, and objects that can be subject to sound reflection or diffraction, such as boxes, spheres, and cylinders, and this information is transmitted to the renderer (decoder) via a bitstream.
[0189] Frequency information may be generated in relation to the isStatic flag mentioned above. That is, the flag for each SO may be extracted from the input bitstream, and its value may be reflected in the frequency at which the first processing unit executes its processing. For example, if the isStatic flag is true for all SOs (except the user's avatar), the only SO that can move is the user's avatar (since it is a 6DoF system, the isStatic flag corresponding to the user's avatar will never be true), and considering that the user's avatar can only move at a maximum speed equivalent to human movement, the frequency at which the first processing unit executes its processing can be small. In this case, for example, the first processing unit may process at a small frequency of about 10 Hz. On the other hand, the second processing unit processes at a frequency that ensures real-time performance without excess or deficiency during the time when the first processing unit is not running. For example, if the unit frame time for processing of the second processing unit is 4 msec, it is controlled to process the audio frame signal at a frequency of at least 250 Hz (1 / 0.004).
[0190] Conversely, if the isStatic flag corresponding to many SOs is false, the scene may be a scene with a lot of motion, so the processing frequency of the first processing unit needs to be increased. In this case, for example, the frequency information may be processed at a high frequency such as 60 Hz.
[0191] As described above, the processing frequency of the second processing unit is logically (automatically) determined by the unit frame time of the second processing unit's processing, whereas the processing frequency of the first processing unit is not determined logically (automatically). Therefore, it is advisable to pre-determine the frequency information based on the information contained in the bitstream. Alternatively, the frequency information itself can be included in the bitstream. In this way, content creators can adopt any frequency during the content encoding stage.
[0192] On the other hand, another method for obtaining frequency information is to computationally obtain it from the fluctuations in information indicating the position of the sound source object. Specifically, the SO mentioned above is assigned "velocity" information, which is managed internally by the renderer (decoder). For example, "velocity" information is included in the attributes associated with the SO.
[0193] Frequency information may be generated in relation to the above-mentioned "velocity" information. That is, the amount of movement per unit time for each SO may be calculated and this value may be stored as the "velocity" information associated with that SO. For example, the largest value among the "velocity" information of all SOs may be detected, and this value may be used to determine the frequency at which the first processing unit executes its processing. Alternatively, the average value of the "velocity" information of all SOs may be detected, and this value may be used to determine the frequency at which the first processing unit executes its processing. If the detected value is small, it means that the scene has little movement, so the frequency at which the first processing unit executes its processing can be small. In this case, for example, the first processing unit may process at a small frequency of about 10 Hz. The second processing unit processes at a frequency that ensures real-time performance without excess or deficiency during the time when the first processing unit is not running. For example, if the unit frame time of the second processing unit is 4 msec, it is controlled to process the audio frame signal at a frequency of at least 250 Hz (1 / 0.004).
[0194] Conversely, if the detected value is large, it indicates that the scene is one with a lot of movement, and therefore the processing frequency of the first processing unit needs to be increased. In this case, for example, the frequency information may be processed at a high frequency such as 60 Hz.
[0195] As described above, the processing frequency of the second processing unit is logically (automatically) determined by the unit frame time of the second processing unit, whereas the processing frequency of the first processing unit is not logically (automatically) determined. Therefore, it is advisable to pre-determine it according to the information of attributes associated with each SO placed in the three-dimensional sound field.
[0196] Furthermore, there are yet another method for obtaining frequency information. Figures 19 to 24 are diagrams illustrating the method for obtaining frequency information according to the embodiment.
[0197] For example, Figure 19 is an example of configuration information that defines the operating conditions of a three-dimensional sound field. As shown in the figure, if "Update frequency of scene information (i.e., frequency information)" is described alongside information such as "global sampling frequency" and "audio frame length" included in the configuration information, which is one of the pieces of information associated with the three-dimensional sound field (virtual sound space), then the frequency information can be obtained by acquiring the configuration information.
[0198] Configuration information defines the operating conditions of a three-dimensional sound field and is therefore more suitable as a source of information for obtaining frequency information. For this reason, even if the frequency information itself is included in the bitstream as explained above, if frequency information can be obtained from the configuration information, the frequency information obtained from the configuration information takes precedence. In other words, if frequency information is obtained from both the bitstream that defines the configuration of the three-dimensional sound field and the configuration information that defines the operating conditions of the three-dimensional sound field, the frequency information obtained from the configuration information may be set as the frequency information used in the processing of the first processing unit, and the information obtained from the bitstream that defines the configuration of the three-dimensional sound field may be discarded.
[0199] Furthermore, frequency information can be said to define the time resolution at which spatial information is updated. On the other hand, the sampling frequency of the sound signal, which relates to the time resolution of the sound signal, is used in the processing of the second processing unit. From the perspective of matching this time resolution, it is desirable to set frequency information with a higher frequency as the sampling frequency of the sound signal increases. Figure 20 is a diagram illustrating the relationship between the sampling frequency of the sound signal and frequency information. From the above perspective, as shown in Figure 20, for example, if the sampling frequency of the sound signal is as high as 48 kHz, frequency information with a high frequency such as 60 Hz should be set, and if the sampling frequency of the sound signal is as low as 32 kHz, frequency information with a low frequency such as 40 Hz should be set. In other words, frequency information can also be obtained by multiplying the sampling frequency of the sound signal by a predetermined correlation coefficient or the like.
[0200] Furthermore, the processing frame length of the second processing unit, which deals with the temporal resolution of the sound signal, is used in the processing of the second processing unit. As with the above, from the viewpoint of matching the temporal resolution, it is desirable to set frequency information of a higher frequency the smaller the processing frame length of the second processing unit. Figure 21 is a diagram illustrating the relationship between the processing frame length of the second processing unit and the frequency information. From the above viewpoint, as shown in Figure 21, for example, if the processing frame length of the second processing unit is small, such as 128 samples, then frequency information of a higher frequency such as 60 Hz should be set, and if the processing frame length of the second processing unit is large, such as 256 samples, then frequency information of a lower frequency such as 30 Hz should be set. In other words, frequency information may be obtained by multiplying the processing frame length of the second processing unit by a predetermined correlation coefficient, etc.
[0201] Figures 22 and 23 illustrate the relationship between the sampling frequency of the sound signal, the processing frame length of the second processing unit, and frequency information. As shown in Figure 22, frequency information may be obtained by multiplying both the sampling frequency of the sound signal and the processing frame length of the second processing unit by a predetermined correlation coefficient. In other words, as shown in Figure 23, when the sampling frequency of the sound signal (A Hz) and the processing frame length of the second processing unit (B samples) are obtained, frequency information may be obtained using the formula P × A / B. Here, P is a predetermined coefficient, and a value such as 0.1 is appropriately set to associate the sampling frequency of the sound signal and the processing frame length of the second processing unit with an appropriate frequency.
[0202] Furthermore, frequency information may be obtained according to the processing performance of the processor used in the processing of the sound reproduction system 100 or the signal processing device 101.
[0203] The term "processor processing power" as used here may be specified, for example, by the processing power (so-called Mips value) of the CPU running the virtual space software, or by its operating frequency, but it is not necessarily limited to being expressed solely by such numerical values. For example, "processor processing power" may be distinguished by the type of processing unit running the virtual space software. For instance, terms such as "supercomputer" and "cloud" may indicate extremely high processor power, "PC" may indicate high processor power, "mobile terminal" may indicate medium processor power, and "CPU built into devices such as HMDs" may indicate low processor power.
[0204] Alternatively, "processor capabilities" may be distinguished by specifying the intended use. For example, the term "wearable use" may indicate low processor capabilities, "mobile use" may indicate medium processor capabilities, "home use" may indicate high processor capabilities, and "server use" may indicate extremely high processor capabilities. Using such expressions, the frequency information may be set to, for example, 80 Hz when the processor capabilities are extremely high (maximum), 60 Hz when the processor capabilities are high, 40 Hz when the processor capabilities are medium, and 20 Hz when the processor capabilities are low.
[0205] Profiles may be defined according to processor capabilities and applications, in which case processing frame length, sampling frequency, etc., may be set for each. The information to be set is not limited to these. Profiles may be defined according to the processor capabilities of the decoder-side signal processing device, such as for wearable devices, mobile devices, PCs, servers, etc., or according to applications such as distribution, broadcasting, professional use, etc.
[0206] Here, Figure 25 is a diagram illustrating a method for defining processor capabilities according to an embodiment. Profile type information may be indicated by an ID, for example, as shown in Figure 25. Profile type information may be included, for example, in the header portion of the bitstream. For example, the encoder selects a profile and generates a bitstream by adding information of the selected profile type. The decoder obtains the profile type information from the bitstream and uses it for rendering. For example, the profile type information is used to determine the processing frequency in rendering.
[0207] By including type information in the bitstream, the encoder can enable the decoder to perform appropriate decoding after checking playback conditions and compatibility.
[0208] Note that profile type information may be included in the configuration information.
[0209] Furthermore, the sampling frequency of the sound signal and the processing frame length of the second processing unit, as described above, may be used in conjunction with the processor's processing power. However, here, the sampling frequency of the sound signal and the processing frame length of the second processing unit are used from the perspective of dividing the limited processing performance between the processing of the first processing unit and the processing of the second processing unit. Therefore, when the sampling frequency of the sound signal, which requires a lot of computational resources for the processing of the second processing unit, is large, and the processing frame length of the processing of the second processing unit is small, the frequency information is reduced in order to reduce the computational resources allocated to the processing of the first processing unit. Conversely, when the sampling frequency of the sound signal, which requires less computational resources for the processing of the second processing unit, is small, and the processing frame length of the processing of the second processing unit is large, the frequency information is increased in order to increase the computational resources allocated to the processing of the first processing unit. As an example, Figure 24 illustrates the relationship between the processor's processing power, the sampling frequency of the sound signal, the processing frame length of the processing of the second processing unit, and the frequency information. As shown in Figure 24, under the condition that the processor's processing power is the same (viewing the table in the row direction), the coarser the time resolution, the larger the frequency information can be. Also, under the condition that the time resolution is the same (viewing the table in the column direction), the greater the processor's processing power, the larger the frequency information can be.
[0210] In the above explanation, frequency information was described simply as the number of operations performed by the first processing unit per second. However, frequency information may also be expressed to include a reference frequency and a coefficient of a variable value multiplied by the reference frequency. The reference frequency is the same as the frequency information described above. The coefficient is a numerical value between 0 and 1 and changes depending on the state of the sound reproduction system 100. For example, when the processor temperature of the sound reproduction system 100 is normal, the coefficient is 1, and when the processor temperature is higher than a threshold, the coefficient is set to be less than 1. This makes it possible to easily change the processing frequency of the first processing unit in accordance with the processor temperature. Also, for example, when the battery level of the sound reproduction system 100 is normal, the coefficient is 1, and when the battery level is lower than a threshold, the coefficient is set to be less than 1. This makes it possible to easily change the processing frequency of the first processing unit in accordance with the battery level. In this way, by representing frequency information with a coefficient, it is possible to easily change the processing frequency of the first processing unit depending on how the coefficient is set.
[0211] (Other Embodiments) Although embodiments have been described above, this disclosure is not limited to the embodiments described above.
[0212] For example, the sound reproduction system described in the above embodiment may be realized as a single device comprising all its components, or it may be realized by assigning each function to multiple devices and having these multiple devices cooperate. In the latter case, an information processing device such as a smartphone, tablet terminal, or PC may be used as the information processing device. For example, in a sound reproduction system 100 that has the function of a renderer that generates an acoustic signal with added sound effects, a server may be responsible for all or part of the renderer's functions. That is, all or part of the acquisition unit 111, the route calculation unit 121, the output sound generation unit 131, and the signal output unit 141 may reside on a server (not shown). In that case, the sound reproduction system 100 is realized by combining, for example, an information processing device such as a computer or smartphone, a sound presentation device such as a head-mounted display (HMD) or earphones worn by a user 99, and a server (not shown). The computer, the sound presentation device, and the server may be connected to communicate on the same network, or they may be connected on different networks. If connected via different networks, communication delays are more likely to occur. Therefore, processing on the server may only be permitted if the computer, sound presentation device, and server are all connected to the same network and capable of communication. Furthermore, depending on the amount of bitstream data received by the sound playback system 100, it may be determined whether the server handles all or part of the renderer's functions.
[0213] Furthermore, the sound reproduction system of this disclosure can also be implemented as an information processing device that is connected to a playback device equipped only with a driver and reproduces an output sound signal generated based on acquired sound information to the playback device. In this case, the information processing device may be implemented as hardware equipped with a dedicated circuit, or as software that causes a general-purpose processor to perform specific processing.
[0214] Furthermore, in the above embodiment, the processing performed by a specific processing unit may be performed by another processing unit. Also, the order of multiple processing units may be changed, or multiple processing units may be executed in parallel.
[0215] Furthermore, in the above embodiment, each component may be realized by executing a software program suitable for that component. Each component may also be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0216] Furthermore, each component may be implemented by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or they may be separate circuits. Also, each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0217] Furthermore, the general or specific embodiments of this disclosure may be implemented as a system, apparatus, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM. Also, the general or specific embodiments of this disclosure may be implemented as any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.
[0218] For example, the present disclosure may be implemented as a method for reproducing an audio signal executed by a computer, or as a program for causing a computer to execute an audio signal reproduction method. The present disclosure may also be implemented as a computer-readable non-temporary recording medium on which such a program is recorded.
[0219] Furthermore, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art could conceive, or forms realized by arbitrarily combining the components and functions of each embodiment without departing from the spirit of this disclosure.
[0220] In this disclosure, the encoded sound information can be rephrased as a bitstream containing a sound signal, which is information about a predetermined sound to be reproduced by the sound reproduction system 100, and metadata, which is information about the localization position when localizing the sound image of the predetermined sound to a predetermined position in a three-dimensional sound field. For example, the sound information may be acquired by the sound reproduction system 100 as a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about a predetermined sound to be reproduced by the sound reproduction system 100. The predetermined sound here is a sound emitted by a sound source object present in the three-dimensional sound field or a natural environmental sound, and may include, for example, machine sounds or animal sounds, including human voices. If there are multiple sound source objects in the three-dimensional sound field, the sound reproduction system 100 will acquire multiple sound signals corresponding to each of the multiple sound source objects.
[0221] On the other hand, metadata is information used, for example, to control acoustic processing of sound signals in the sound reproduction system 100. Metadata may also be information used to describe a scene represented in a virtual space (three-dimensional sound field). Here, "scene" refers to the collection of all elements representing three-dimensional images and acoustic events in a virtual space, which are modeled by the sound reproduction system 100 using metadata. In other words, the metadata referred to here may include not only information that controls acoustic processing, but also information that controls video processing. Of course, metadata may include information that controls only one of acoustic processing or video processing, or it may include information used to control both. The bitstream acquired by the sound reproduction system 100 in this disclosure may contain such metadata. Alternatively, the sound reproduction system 100 may acquire metadata separately from the bitstream, as will be described later.
[0222] The sound reproduction system 100 generates virtual sound effects by performing acoustic processing on the sound signal using metadata included in the bitstream and additionally acquired location information of the interactive user 99. For example, sound effects such as early reflection generation, late reverberation generation, diffraction generation, distance attenuation effect, localization, sound image localization processing, or Doppler effect may be added. Information to switch all or some of the sound effects on or off may also be added as metadata.
[0223] Furthermore, all or some of the metadata may be obtained from sources other than the audio bitstream. For example, either the metadata controlling the sound or the metadata controlling the video may be obtained from a source other than the bitstream, or both metadata may be obtained from a source other than the bitstream.
[0224] Furthermore, if metadata for controlling the video is included in the bitstream acquired by the sound playback system 100, the sound playback system 100 may also have a function to output metadata that can be used to control the video to a display device that displays an image, or a stereoscopic video playback device that plays stereoscopic video.
[0225] As an example, the encoded metadata includes information about a three-dimensional sound field, including sound source objects that emit sound and obstacle objects, and information about the localization position when the sound image of the sound is localized to a predetermined position within the three-dimensional sound field (i.e., perceived as sound arriving from a predetermined direction), i.e., information about the predetermined direction. Here, an obstacle object is an object that can affect the sound perceived by the user 99 by, for example, blocking or reflecting sound until the sound emitted by the sound source object reaches the user 99. Obstacle objects may include stationary objects, animals such as people, or moving objects such as machines. Also, if there are multiple sound source objects in the three-dimensional sound field, any other sound source object can be an obstacle object for any of the sound source objects. Furthermore, non-sounding objects such as building materials or inanimate objects, as well as sound source objects that emit sound, can all be obstacle objects.
[0226] The spatial information constituting the metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects present in the three-dimensional sound field, and the shape and position of sound source objects present in the three-dimensional sound field. The three-dimensional sound field may be either a closed or open space, and the metadata may include information representing the reflectivity of structures that can reflect sound in the three-dimensional sound field, such as floors, walls, or ceilings, and the reflectivity of obstacle objects present in the three-dimensional sound field. Here, the reflectivity is the ratio of the energy of the reflected sound to the energy of the incident sound, and is set for each sound frequency band. Of course, the reflectivity may be set uniformly regardless of the sound frequency band. Furthermore, if the three-dimensional sound field is an open space, parameters such as a uniformly set attenuation rate, diffracted sound, or early reflections may be used.
[0227] In the above explanation, reflectivity was mentioned as a parameter related to obstacle objects or sound source objects included in the metadata, but metadata may include information other than reflectivity. For example, metadata relating to both sound source objects and non-sound source objects may include information about the material of the objects. Specifically, metadata may include parameters such as diffusion, transmittance, or sound absorption.
[0228] Information regarding the sound source object may include volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, or information specifying the sound source area within the object. Playback conditions may, for example, specify whether the sound is continuously playing or triggered by an event. The sound source area within an object may be defined by the relative relationship between the user 99's position and the object's position, or it may be defined relative to the object. When defined relative to the user 99's position, the user 99 can perceive that sound X is emitted from the right side of the object and sound Y from the left side, based on the view from which the user 99 is looking at the object. When defined relative to the object, the area from which each sound is emitted can be fixed regardless of the direction the user 99 is looking. For example, when viewing the object from the front, the user 99 can perceive that a high-pitched sound is emitted from the right side and a low-pitched sound from the left side. In this case, if the user 99 moves behind the object, the user 99 can perceive that a low-pitched sound is emitted from the right side and a high-pitched sound from the left side.
[0229] Metadata related to the space can include the time to the first reflection, reverberation time, or the ratio of direct sound to diffused sound. If the ratio of direct sound to diffused sound is zero, only direct sound can be perceived by the user 99.
[0230] Furthermore, information indicating the position and orientation of the user 99 in the three-dimensional sound field may or may not be included in the bitstream as metadata as an initial setting. If information indicating the position and orientation of the user 99 is not included in the bitstream, this information is obtained from information other than the bitstream. For example, if the user 99's position information is in a VR space, it may be obtained from an application that provides VR content. If the user 99's position information is for presenting sound as AR, for example, position information obtained by a mobile device performing self-position estimation using GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging) may be used. The sound signal and metadata may be stored in a single bitstream or separately in multiple bitstreams. Similarly, the sound signal and metadata may be stored in a single file or separately in multiple files.
[0231] If the audio signal and metadata are stored separately in multiple bitstreams, information indicating the other related bitstreams may be included in one or some of the bitstreams in which the audio signal and metadata are stored. Alternatively, information indicating the other related bitstreams may be included in the metadata or control information of each bitstream in the multiple bitstreams in which the audio signal and metadata are stored. If the audio signal and metadata are stored separately in multiple files, information indicating the other related bitstreams or files may be included in one or some of the files in which the audio signal and metadata are stored. Alternatively, information indicating the other related bitstreams or files may be included in the metadata or control information of each bitstream in the multiple bitstreams in which the audio signal and metadata are stored.
[0232] Here, each related bitstream or file is, for example, a bitstream or file that may be used simultaneously during sound processing. Furthermore, information indicating other related bitstreams may be collectively described in the metadata or control information of one bitstream among multiple bitstreams storing sound signals and metadata, or it may be divided and described in the metadata or control information of two or more bitstreams among multiple bitstreams storing sound signals and metadata. Similarly, information indicating other related bitstreams or files may be collectively described in the metadata or control information of one file among multiple files storing sound signals and metadata, or it may be divided and described in the metadata or control information of two or more files among multiple files storing sound signals and metadata. Additionally, a control file containing the information indicating other related bitstreams or files may be generated separately from the multiple files storing sound signals and metadata. In this case, the control file does not necessarily have to store sound signals and metadata.
[0233] Here, information indicating other related bitstreams or files may include, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit identifies or acquires the bitstream or file based on the information indicating other related bitstreams or files. Furthermore, information indicating other related bitstreams may be included in the metadata or control information of at least some of the bitstreams among a plurality of bitstreams storing sound signals and metadata, and information indicating other related files may be included in the metadata or control information of at least some of the files among a plurality of files storing sound signals and metadata. Here, a file containing information indicating other related bitstreams or files may be, for example, a control file such as a manifest file used for content distribution.
[0234] This disclosure is useful for sound reproduction, such as making users perceive three-dimensional sound.
[0235] 99 User 100 Sound playback system 101 Signal processing unit 102 Communication module 103 Detector 104 Driver 105 Database 111 Acquisition unit 112 Encoded sound information input unit 113 Decode processing unit 114 Sensing information input unit 121 Route calculation unit 122 Update processing unit 123 Route calculation unit 124 Memory 131 Output sound generation unit 141 Signal output unit 300 Stereoscopic video playback device
Claims
1. A signal processing device for playing content that represents a virtual sound space, comprising: a first processing unit for updating information of the virtual sound space; a second processing unit for playing sound signals of sounds generated in the virtual sound space based on the updated information of the virtual sound space and presenting them to the user via the user's avatar in the virtual sound space; a first setting unit for setting frequency information indicating the processing frequency of the first processing unit; and a control unit for controlling the execution of the processing of the first processing unit and the second processing unit, the control unit for controlling the processing of the first processing unit to be executed based on the set frequency information.
2. The signal processing device according to claim 1, wherein the frequency information based on information associated with the virtual sound space is set in the first setting unit.
3. The signal processing device according to claim 1, wherein the frequency information defined in the configuration information defining the operating conditions of the virtual sound space is set in the first setting unit.
4. The signal processing apparatus according to claim 1, further comprising a second setting unit for setting the sampling frequency of the sound signal, wherein the first setting unit is set with the frequency information corresponding to the set sampling frequency.
5. The signal processing apparatus according to claim 1, further comprising a third setting unit for setting the processing frame length in the processing of the second processing unit, wherein the first setting unit is set with the frequency information corresponding to the set processing frame length.
6. The signal processing apparatus according to claim 1, wherein the processing of the first processing unit and the second processing unit is performed by a processor, and the frequency information corresponding to the processing capacity of the processor is set in the first setting unit.
7. The signal processing apparatus according to claim 6, wherein the first setting unit is set to frequency information related to the time resolution setting value for the processing of the first processing unit and the second processing unit, according to the processing capability of the processor.
8. The signal processing apparatus according to claim 7, wherein the set value for the time resolution relating to the processing of the second processing unit includes the sampling frequency of the sound signal and the processing frame length in the processing of the second processing unit, and the signal processing apparatus comprises a second setting unit for setting the sampling frequency and a third setting unit for setting the processing frame length.
9. The signal processing apparatus according to claim 1, comprising: a second setting unit for setting the sampling frequency of the sound signal; and a third setting unit for setting the processing frame length in the processing of the second processing unit, wherein a region for specifying the frequency information set in the first setting unit, a region for specifying the sampling frequency set in the second setting unit, and a region for specifying the processing frame length set in the third setting unit are arranged adjacent to each other.
10. The signal processing apparatus according to claim 1, comprising: a second setting unit for setting the sampling frequency of the sound signal; and a third setting unit for setting the processing frame length in the processing of the second processing unit, wherein the first setting unit is set with the frequency information corresponding to the set sampling frequency and the set processing frame length.
11. The signal processing apparatus according to claim 1, wherein the first processing unit updates the arrival path of sound arriving at the avatar in the virtual sound space in which the avatar can move by 6DoF.
12. The signal processing apparatus according to claim 1, wherein the frequency information includes a fixed reference frequency and a coefficient of a variable value multiplied by the reference frequency.
13. The signal processing apparatus according to claim 1, wherein the frequency information is obtained and set from a bitstream that defines the configuration of the virtual sound space.
14. The signal processing device according to claim 1, wherein the frequency information is calculated and set based on fluctuations in information indicating the position of sound source objects included in the virtual sound space.
15. The signal processing apparatus according to claim 1, wherein the frequency information is obtained and set from at least one of a bitstream defining the configuration of the virtual sound space and configuration information defining the operating conditions of the virtual sound space, and if it is obtained from both the bitstream defining the configuration of the virtual sound space and the configuration information defining the operating conditions of the virtual sound space, the one obtained from the configuration information defining the operating conditions of the virtual sound space is set.
16. A signal processing method performed by a computer for playing content that represents a virtual sound space, comprising: updating information of the virtual sound space; playing sound signals of sounds generated in the virtual sound space based on the updated information of the virtual sound space and presenting them to the user via the user's avatar in the virtual sound space; and setting frequency information indicating the processing frequency of the updating step, wherein the processing of the updating step is performed based on the set frequency information.
17. A program for causing the computer to execute the signal processing method described in claim 16.