Acoustic processing method, acoustic processing device, and program

By repeatedly changing the relative position of the audio processing device and the sound source in the time domain, the problem of insufficient sound positioning in the prior art is solved, and more appropriate audio processing and a stronger sense of presence are achieved.

CN120019673APending Publication Date: 2025-05-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380071658.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-09-28
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art increases the sound positioning sense, it is difficult to perform sound processing appropriately, resulting in a weakening of the sense of presence of sound.

Method used

By obtaining the sound signal and repeatedly changing the relative position of the audio processing device in the time domain, an audio processing that repeatedly changes the relative position of the audio receiving device and the sound source in the time domain is performed.

Benefits of technology

The loss of presence in the sound signal is reconstructed, the appropriateness of audio processing is improved, and the user can better feel the sound position in the three-dimensional space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019673A_ABST
    Figure CN120019673A_ABST
Patent Text Reader

Abstract

The information processing method comprises: a step (S101) for acquiring a sound signal obtained by collecting sound emitted from a sound source using a sound reception device; a step (S103) for executing sound processing for repeatedly changing the relative position of the sound reception device and the sound source in the time domain with respect to the sound signal; and a step (S105) for outputting the output sound signal on which the sound processing has been performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a sound processing method, a sound processing device, and a program. Background Art

[0002] In the past, there are known technologies related to sound reproduction for making users perceive stereoscopic sounds in a virtual three-dimensional space (for example, see Patent Document 1). In addition, in order to make the sound in such a three-dimensional space be perceived as reaching the user from the sound source object, it is necessary to perform processing to generate output sound information based on the original sound information. Here, in order to make the user who listens to the sound better feel the sense of presence in the three-dimensional space, sound processing that increases the sense of localization of the sound is performed. For example, there is known a stereo sound processing device that brings a sense of localization so that the sound is heard from the direction of the sound source coordinates input from the coordinate fluctuation adding device (see Patent Document 1).

[0003] Prior art literature

[0004] Patent Literature

[0005] Patent Document 1: Japanese Patent Application Publication No. 2005-295416 Summary of the invention

[0006] Problems to be solved by the invention

[0007] When a fluctuation is applied to increase the localization sense of a sound, the acoustic processing for applying the fluctuation may not be properly performed. Therefore, in the present disclosure, an acoustic processing method for performing the acoustic processing more properly is described.

[0008] Means used to solve problems

[0009] A sound processing method according to a technical solution of the present disclosure includes: obtaining a sound signal obtained by collecting sound emitted from a sound source using a sound receiving device; performing a sound processing step on the sound signal so that the relative position of the sound receiving device and the sound source repeatedly changes in the time domain; and outputting an output sound signal on which the sound processing has been performed.

[0010] In addition, another technical solution of the present disclosure is a sound processing method for outputting an output sound signal that makes the sound emitted from a sound source object in a virtual sound space be perceived as being heard at a listening point in the virtual sound space, comprising: a step of obtaining a sound signal including the sound emitted from the sound source object; a step of accepting an instruction to change the relative position of the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed by the instruction; a step of performing sound processing on the sound signal to change the relative position by the first change amount and to repeatedly change the relative position by a second change amount in the time domain; and a step of outputting the output sound signal on which the sound processing has been performed.

[0011] In addition, a sound processing device related to a technical solution of the present disclosure includes: an acquisition unit, which acquires a sound signal obtained by using a sound pickup device to collect sound emitted from a sound source; a processing unit, which performs sound processing on the sound signal to repeatedly change the relative position of the sound pickup device and the sound source in the time domain; and an output unit, which outputs an output sound signal that has undergone the sound processing.

[0012] In addition, another technical solution of the present disclosure is a sound processing device, which is a sound processing device for outputting an output sound signal that makes the sound emitted from a sound source object in a virtual sound space be perceived as being heard at a listening point in the virtual sound space, and includes: an acquisition unit, which acquires a sound signal including the sound emitted from the sound source object; a receiving unit, which receives an instruction to change the relative position of the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed by the instruction; a processing unit, which performs sound processing on the sound signal to change the relative position by the first change amount and repeatedly change the relative position by a second change amount in the time domain; and an output unit, which outputs the output sound signal on which the sound processing has been performed.

[0013] Furthermore, one aspect of the present disclosure can also be implemented as a program for causing a computer to execute the above-described sound processing method.

[0014] In addition, these inclusive or specific technical solutions can also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0015] Effects of the Invention

[0016] According to the present disclosure, it is possible to perform audio processing more appropriately. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram showing an example of use of the sound reproduction system according to the embodiment.

[0018] Figure 2A This is a diagram for explaining an example of use of the sound reproduction system according to the embodiment.

[0019] Figure 2B This is a diagram for explaining an example of use of the sound reproduction system according to the embodiment.

[0020] Figure 3 This is a block diagram showing the functional structure of the sound reproduction system according to the embodiment.

[0021] Figure 4 This is a block diagram showing the functional structure of an acquisition unit according to the embodiment.

[0022] Figure 5 It is a block diagram showing the functional structure of a processing unit according to an embodiment.

[0023] Figure 6 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0024] Figure 7 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0025] Figure 8 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0026] Fig. 9 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0027] Fig.10 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0028] Fig.11 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0029] Fig.12 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0030] Fig.13 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0031] Fig.14 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0032] Fig.15 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0033] Fig.16 This is a flowchart showing the operation of the sound processing device according to the embodiment.

[0034] Fig.17 It is a diagram for explaining the frequency characteristics of the sound processing according to the embodiment.

[0035] Fig.18 This is a diagram for explaining the magnitude of fluctuations in the acoustic processing according to the embodiment.

[0036] Fig.19 It is a diagram for explaining the period and angle of the ripple of the sound processing related to the embodiment.

[0037] Fig. 20 This is a block diagram showing a functional structure of a processing unit according to another example of the embodiment.

[0038] Fig.21 This is a flowchart showing the operation of the sound processing device according to another example of the embodiment. DETAILED DESCRIPTION

[0039] (Knowledge as a basis for disclosure)

[0040] In the past, there is a known technology related to sound reproduction for making users perceive stereoscopic sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field or a virtual sound space) (for example, refer to Patent Document 1). By using this technology, the user can perceive the sound as if the sound source object exists at a specified position in the virtual space and the sound comes from its direction. In order to position the sound image at a specified position in the virtual three-dimensional space in this way, for example, for the signal of the sound of the sound source object, it is necessary to perform calculation processing such as the arrival time difference of the sound between the two ears and the level difference (or sound pressure difference) of the sound between the two ears so that it can be perceived as a stereoscopic sound. Such calculation processing is performed by applying a stereo sound filter. The stereo sound filter is an information processing filter that, if the output sound signal after applying the filter to the original sound information is reproduced, the direction of the sound, the distance, etc., the size of the sound source, the size of the space, etc., can be perceived in a three-dimensional sense.

[0041] As an example of calculation processing for applying such a stereo filter, there is known a process of convolving a target sound signal with a head-related transfer function so as to make it perceived as a sound coming from a predetermined direction. By performing the convolution process of the head-related transfer function at a sufficiently small angle with respect to the direction of arrival of the sound from the position of the sound source object to the user's position, the sense of presence felt by the user is improved.

[0042] In addition, in recent years, the development of technologies related to virtual reality (VR) has been actively carried out. In virtual reality, since the sense of positioning of sound in a three-dimensional sound field also brings about a sense of presence of the image, audio processing is performed to increase the sense of positioning. In the case of imparting fluctuations in order to increase the sense of positioning of the sound, from the perspective of the effect, it is not necessary to impart fluctuations to all sounds in the same way. In other words, there are conditions under which the imparting of fluctuations works effectively. By imparting fluctuations only when such conditions are met, there is no need to prepare processing resources unnecessarily, so it can be said to be preferred.

[0043] A more specific summary of the present disclosure is as follows.

[0044] The sound processing method of the first technical solution related to the present disclosure includes: a step of obtaining a sound signal obtained by using a sound receiving device to collect sound emitted from a sound source; a step of performing sound processing on the sound signal so that the relative position of the sound receiving device and the sound source repeatedly changes in the time domain; and a step of outputting an output sound signal that has undergone sound processing.

[0045] According to such an acoustic processing method, in the case of a condition where the sense of presence is lost, such as when the position of the sound pickup device relative to the position of the sound source does not change relative to the sound signal collected by the sound pickup device, the relative position of the sound pickup device and the sound source is repeatedly changed in the time domain by acoustic processing to give fluctuations, thereby recreating the lost sense of presence. In this way, the acoustic processing can be performed more appropriately from the viewpoint of recreating the sense of presence.

[0046] In addition, in the sound processing method related to the second technical solution, in the sound processing method described in the first technical solution, in the step of performing sound processing, it is determined whether the change in the sound pressure of the sound signal in the time domain satisfies a specified condition related to the change; the sound processing is performed when it is determined that the specified condition is satisfied; and the sound processing is not performed when it is determined that the specified condition is not satisfied.

[0047] According to such an acoustic processing method, whether or not to execute acoustic processing can be changed by determining whether or not a predetermined condition regarding a change in the sound pressure of an audio signal in the time domain is satisfied.

[0048] In addition, in the sound processing method related to the third technical solution, in the sound processing method described in the first or second technical solution, in the step of performing sound processing, the positional relationship between the sound receiving device and the sound source is estimated using a sound signal; it is determined whether the estimated positional relationship satisfies a prescribed condition related to the positional relationship; the sound processing is performed when it is determined that the prescribed condition is satisfied; and the sound processing is not performed when it is determined that the prescribed condition is not satisfied.

[0049] According to such a sound processing method, whether or not to execute sound processing can be changed by determining whether or not a predetermined condition regarding the positional relationship between the sound pickup device and the sound source estimated using the sound signal is satisfied.

[0050] In addition, in the sound processing method related to the fourth technical solution, in the sound processing method described in any one of the first to third technical solutions, the sound signal includes sound collection status information related to the conditions when collecting sound; in the step of performing sound processing, it is determined whether the sound collection status information included in the sound signal satisfies a specified condition related to the sound collection status information; the sound processing is performed when it is determined that the specified condition is satisfied; and the sound processing is not performed when it is determined that the specified condition is not satisfied.

[0051] According to such an audio processing method, whether or not to execute audio processing can be changed by determining a predetermined condition related to the sound collection status information included in the audio signal.

[0052] In addition, in the sound processing method related to the fifth technical solution, in the sound processing method described in any one of the first to fourth technical solutions, in the step of performing sound processing, the positional relationship between the sound receiving device and the sound source is estimated using the sound signal; and the sound processing is performed under processing conditions corresponding to the estimated positional relationship.

[0053] According to such a sound processing method, it is possible to execute sound processing under processing conditions corresponding to the positional relationship between the sound pickup device and the sound source estimated using the sound signal.

[0054] In addition, the sound processing method related to the sixth technical solution of the present disclosure is an sound processing method for outputting an output sound signal so that the sound emitted from the sound source object in the virtual sound space is perceived as being heard at a listening point in the virtual sound space, including: a step of obtaining a sound signal including the sound emitted from the sound source object; a step of accepting an instruction to change the relative position of the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed by the instruction; a step of performing sound processing on the sound signal to change the relative position by the first change amount and to repeatedly change the relative position by a second change amount in the time domain; and a step of outputting the output sound signal for which the sound processing has been performed.

[0055] According to such an acoustic processing method, when a sound emitted from a sound source object in a virtual sound space is perceived to be heard at a listening point in the virtual sound space, independently of the change in relative position of the first change amount generated based on an instruction to change the relative position of the listening point and the sound source object, when the sense of presence has been lost in the sound signal, the lost sense of presence can be reconstructed by repeatedly changing the relative position of the listening point and the sound source object by a second change amount in the time domain through acoustic processing to give fluctuations. In this way, the acoustic processing can be performed more appropriately from the viewpoint of reconstructing the sense of presence.

[0056] In addition, in the sound processing method related to the seventh technical solution, in the sound processing method described in the sixth technical solution, the sound source object simulates a user in a real space; the sound processing method also includes a step of obtaining a detection result from a sensor for detecting the user set in the real space; and the second change is calculated based on the detection result.

[0057] According to such a sound processing method, the second change amount can be calculated based on the detection result obtained from the sensor that detects the user in the real space corresponding to the sound source object.

[0058] In addition, in the sound processing method related to the 8th technical solution, in the sound processing method described in the 6th technical solution, the sound source object simulates a user in the real space; the sound processing method also includes a step of obtaining a detection result from a sensor for detecting the user set in the real space; and the second change is calculated independently of the detection result.

[0059] According to such a sound processing method, the second variation amount can be calculated independently from the detection result obtained from the sensor that detects the user in the real space corresponding to the sound source object.

[0060] In addition, in the sound processing method according to a ninth technical solution, in the sound processing method according to the first technical solution, the second change amount is calculated independently of the first change amount.

[0061] According to such a sound processing method, it is possible to calculate the second variation independent of the first variation.

[0062] In addition, in the sound processing method according to a tenth aspect, in the sound processing method according to the sixth aspect, the second change amount is calculated as a larger value as the first change amount is larger.

[0063] According to such a sound processing method, it is possible to calculate the second change amount which becomes larger as the first change amount becomes larger.

[0064] In addition, in the sound processing method according to the eleventh technical solution, in the sound processing method according to the sixth technical solution, the second change amount is calculated as a larger value as the first change amount is smaller.

[0065] According to such a sound processing method, it is possible to calculate the second change amount which becomes larger as the first change amount becomes smaller.

[0066] In addition, the sound processing method related to the 12th technical solution is in the sound processing method described in any one of the technical solutions 1 to 11, and further includes a step of obtaining control information on the sound signal; in the step of performing sound processing, the sound processing is performed when the control information indicates that the sound processing is to be performed.

[0067] According to such a sound processing method, when the acquired control information indicates that the sound processing is to be performed, the sound processing can be performed.

[0068] In addition, the sound processing device related to the 13th technical solution of the present disclosure comprises: an acquisition unit, which acquires a sound signal obtained by using a sound receiving device to collect sound emitted from a sound source; a processing unit, which performs sound processing on the sound signal to repeatedly change the relative position of the sound receiving device and the sound source in the time domain; and an output unit, which outputs an output sound signal that has undergone sound processing.

[0069] According to such a sound processing device, it is possible to achieve the same effects as those of the sound processing method described above.

[0070] In addition, the sound processing device related to the 14th technical solution of the present disclosure is an sound processing device for outputting an output sound signal so that the sound emitted from the sound source object in the virtual sound space is perceived as being heard at a listening point in the virtual sound space, including: an acquisition unit, which acquires a sound signal including the sound emitted from the sound source object; a receiving unit, which receives an instruction to change the relative position of the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed by the instruction; a processing unit, which performs sound processing on the sound signal to change the relative position by the first change amount and repeatedly change the relative position by a second change amount in the time domain; and an output unit, which outputs the output sound signal that has undergone the sound processing.

[0071] According to such a sound processing device, it is possible to achieve the same effects as those of the sound processing method described above.

[0072] Furthermore, these inclusive or specific technical solutions may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0073] Hereinafter, the embodiments are described in detail with reference to the accompanying drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements and connection forms, steps, the order of steps, etc. shown in the following embodiments are examples and are not intended to limit the present disclosure. In addition, regarding the constituent elements of the following embodiments that are not described in the independent claims, they are described as arbitrary constituent elements. In addition, each figure is a schematic diagram and is not necessarily a strict illustration. In addition, in each figure, the same reference numerals are given to substantially the same structure, and there are cases where repeated descriptions are omitted or simplified.

[0074] In addition, in the following description, elements are sometimes given ordinals such as 1st, 2nd, and 3rd. These ordinals are given to elements for the purpose of identifying the elements and do not necessarily correspond to a meaningful order. These ordinals may be replaced appropriately, newly given, or removed.

[0075] (Implementation Method)

[0076] [summary]

[0077] First, an overview of the sound reproduction system according to the embodiment will be described. Figure 1 1 is a schematic diagram showing an example of use of the sound reproduction system according to the embodiment. Figure 1 , a user 99 who uses the audio reproduction system 100 is represented.

[0078] Figure 1 The audio reproduction system 100 shown is used simultaneously with the stereoscopic image reproduction device 200. By simultaneously viewing and listening to stereoscopic images and stereoscopic sounds, the images enhance the auditory sense of presence, and the sounds enhance the visual sense of presence, so that it is possible to experience as if one is at the scene where the images and sounds are shot. For example, it is known that in the case of an image (moving image) showing a person speaking, even if the positioning of the sound image of the speaking sound deviates from the person's mouth, the user 99 will perceive the speaking sound as coming from the person's mouth. In this way, the position of the sound image is corrected through visual information, and the sense of presence is improved by coordinating the image and sound.

[0079] The stereoscopic image reproduction device 200 is an image display device worn on the head of the user 99. Therefore, the stereoscopic image reproduction device 200 moves integrally with the head of the user 99. For example, as shown in the figure, the stereoscopic image reproduction device 200 is a glasses-type device supported by the ears and nose of the user 99.

[0080] The stereoscopic image reproduction device 200 changes the displayed image according to the movement of the head of the user 99, so that the user 99 perceives that the user 99 is moving his head in the three-dimensional image space. That is, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 faces the right, the object moves to the left of the user 99, and if the user 99 faces the left, the object moves to the right of the user 99. In this way, the stereoscopic image reproduction device 200 moves the three-dimensional image space in the opposite direction to the movement of the user 99, relative to the movement of the user 99.

[0081] The stereoscopic image reproduction device 200 displays two images with deviations corresponding to parallax in the left and right eyes of the user 99, respectively. The user 99 can perceive the three-dimensional position of the object on the image based on the deviation of the displayed image corresponding to the parallax. In addition, in the case where the sound reproduction system 100 is used for the reproduction of healing sounds for sleep guidance, etc., and the user 99 closes his eyes, it is not necessary to use the stereoscopic image reproduction device 200 at the same time. That is, the stereoscopic image reproduction device 200 is not an essential component of the present disclosure. As the stereoscopic image reproduction device 200, in addition to a dedicated image display device, there is also a case where a general-purpose portable terminal such as a smart phone or a tablet device owned by the user 99 is used.

[0082] In addition to a display for displaying images, such a general-purpose portable terminal is also equipped with various sensors for detecting the posture and movement of the terminal. Furthermore, a processor for information processing is also equipped, and it can be connected to a network to send and receive information with a server device such as a cloud server. In other words, the stereoscopic image reproduction device 200 and the sound reproduction system 100 can also be realized by combining a smart phone with a general-purpose headset without an information processing function.

[0083] As in this example, the function of detecting the movement of the head, the function of presenting an image, the function of processing image information for presenting, the function of presenting sound information, and the function of processing sound information for presenting may be appropriately arranged in one or more devices to realize the stereoscopic image reproduction device 200 and the sound reproduction system 100. When the stereoscopic image reproduction device 200 is not required, as long as the function of detecting the movement of the head, the function of presenting sound information, and the function of processing sound information for presenting can be appropriately arranged in one or more devices, for example, the sound reproduction system 100 may be realized by a processing device such as a computer or a smart phone having a function of processing sound information for presenting, or a headset having a function of detecting the movement of the head and a function of presenting sound information.

[0084] The sound reproduction system 100 is a sound prompt device worn on the head of the user 99. Therefore, the sound reproduction system 100 moves integrally with the head of the user 99. For example, the sound reproduction system 100 of this embodiment is a so-called headphone type device. In addition, there is no particular limitation on the form of the sound reproduction system 100, and for example, it may be a two-earbud type device that is independently worn on the left and right ears of the user 99.

[0085] The sound reproduction system 100 changes the prompt sound according to the movement of the head of the user 99, so that the user 99 perceives that the user 99 moves his head in the three-dimensional sound field. Therefore, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the movement of the user 99 with respect to the movement of the user 99.

[0086] Here, in order to enhance the sense of presence of the sound listened to by the user 99, an acoustic processing is performed to give fluctuations to the sound. For example, Figure 2A and Figure 2B FIG. 2 is a diagram for explaining an example of using the sound reproduction system according to the embodiment. Figure 2A , which represents a user who is currently in a so-called video call. Figure 2A In the left picture, the sound is collected under the condition that the positions of the mouth (sound source) and the microphone (sound pickup device) of the earphones are almost unchanged, as in headphones. However, at the other party of the call in the right picture, the positions of the sound source and the sound pickup device hardly move relative to the user moving on the image, which creates a sense of disharmony. In such a case, by applying sound fluctuations that match the movements of the user moving on the image, or sound fluctuations that match the general movements of the users in the conversation, the disharmony of the sound can be reduced and the sense of presence can be increased.

[0087] In addition, Figure 2B , which shows a user who is collecting the singing voice for the so-called virtual live broadcast in stereo. The user who is collecting the sound may also be a user different from the user 99 who is the listener. For example, a singer, an artist, etc. may be considered. Figure 2BIn the left figure, singing voices are collected by the user singing toward a fixed microphone. The collected sounds are reproduced on the virtual image in the right figure, and viewed together with the image of the user's avatar dancing and singing in the live broadcast venue in the virtual space, thereby realizing a virtual live broadcast. At this time, if the position of the sound source object (the avatar's head) in the virtual sound space is specified as the reproduction position of the sound following the movement of the avatar, even if the positions match, the tiny movements of the fluctuations that the actual user should have will not be reproduced, resulting in a reduction in the sense of presence of the sound. In the present disclosure, by giving the sound the fluctuations that it should have in this way, audio processing that increases the sense of presence of the sound is performed. In addition, as other situations that give rise to the same problem, there are the following situations: in a case such as Figure 2A Even if a sound pickup device capable of collecting sound including the user's fluctuations is used during a video call, a mechanical sound processing such as AGC (Automatic Volume Control) is applied to make the sound easier for the listener to hear, and the fluctuations in the sound are suppressed, which creates a sense of incongruity. The present disclosure also includes a case where the incongruity of the sound is reduced and the sense of presence is increased by re-applying the fluctuations suppressed by such mechanical sound processing.

[0088] On the other hand, the imparting of fluctuations is performed by filtering the output sound signal to repeatedly move the sound in the time domain. This process is complicated because different filters need to be applied at two consecutive time points in the time domain. It is preferred not to apply sound processing when it is expected that the effect of fluctuations will be difficult to obtain.

[0089] [constitute]

[0090] Next, refer to Figure 3 The configuration of the sound reproduction system 100 according to the present embodiment will be described. Figure 3 This is a block diagram showing the functional structure of the sound reproduction system according to the embodiment.

[0091] like Figure 3 As shown, the sound reproduction system 100 according to the present embodiment includes an information processing device 101 , a communication module 102 , a detector 103 , and a driver 104 .

[0092] The information processing device 101 is an example of an audio processing device, and is a computing device for performing various signal processing in the audio reproduction system 100. The information processing device 101 may be realized by, for example, a computer having a processor and a memory, and the processor executing a program stored in the memory. The execution of the program enables the functions related to each functional unit described below to be exerted.

[0093] The information processing device 101 includes an acquisition unit 111, a processing unit 121, and a signal output unit 141. Hereinafter, details of each functional unit of the information processing device 101 will be described together with details of configurations other than the information processing device 101.

[0094] The communication module 102 is an interface device for accepting input of sound information to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device through wireless communication. More specifically, the communication module 102 uses the antenna to receive a wireless signal representing sound information converted into a format for wireless communication, and reconverts the wireless signal into sound information through the signal converter. Thus, the sound reproduction system 100 obtains sound information from an external device through wireless communication. The sound information obtained by the communication module 102 is obtained by the obtaining unit 111. In this way, the sound information is input to the information processing device 101. In addition, the communication between the sound reproduction system 100 and the external device can also be performed through wired communication.

[0095] The sound information obtained by the sound reproduction system 100 is a sound signal obtained by collecting the sound emitted from the sound source using a sound collecting device. The sound information can also be encoded in a format specified by, for example, MPEG-H 3D Audio (ISO / IEC 23008-3), MPEG-I, etc. As an example, the encoded sound information includes information about the specified sound reproduced by the sound reproduction system 100, and information related to the positioning position when the sound image of the sound is localized at a specified position in the three-dimensional sound field (that is, perceived as sound coming from a specified direction) and other metadata. For example, the sound information includes information related to multiple sounds including a first specified sound and a second specified sound, and the sound image is localized so that the sound image when each sound is reproduced is perceived as sound coming from different positions in the three-dimensional sound field.

[0096] This stereoscopic sound can, for example, enhance the sense of presence of the audiovisual content together with the image recognized by the stereoscopic image reproduction device 200. In addition, the sound information may include only information about the specified sound. In this case, information related to the specified position may also be obtained separately. In addition, as described above, the sound information includes the first sound information related to the first specified sound and the second sound information related to the second specified sound, but a plurality of sound information separately including them may be obtained separately, and the sound image may be positioned at different positions in the three-dimensional sound field by simultaneous reproduction. In this way, there is no particular limitation on the form of the input sound information, as long as the sound reproduction system 100 has an acquisition unit 111 corresponding to various forms of sound information.

[0097] The metadata included in the sound information includes control information for controlling the sound processing for imparting fluctuations. The control information is information for specifying whether to execute the sound processing. For example, when the control information specifies to execute the sound processing, it is further possible to determine whether a predetermined condition is satisfied and execute the sound processing when the predetermined condition is satisfied, or it is possible to execute the sound processing regardless of the determination of whether the predetermined condition is satisfied. On the other hand, when the control information specifies not to execute the sound processing, the sound processing is not executed. In this way, the sound processing can be executed by two opportunities, namely, the determination of whether the predetermined condition is satisfied and whether the sound processing is specified to be executed in the control information, or it can be executed by one opportunity, namely, whether the sound processing is specified to be executed. The control information may not be included in the metadata. For example, the control information may be specified by the action setting of the sound reproduction system 100 and stored in the storage unit. Furthermore, the control information may be acquired at the time of starting the sound reproduction system 100 and used as described above.

[0098] Furthermore, the metadata may include sound collection status information. The sound collection status information is the reverberation level and noise level related to the collection of predetermined sounds included in the sound information. The details of the sound collection status information will be described later.

[0099] Sound information can also be obtained as a bit stream. An example of the construction of a bit stream in the case of obtaining sound information as a bit stream is described. The bit stream includes, for example, a sound signal and metadata. The sound signal is sound data that represents the sound, indicating information related to the frequency and strength of the sound. Metadata can also include spatial information other than the above information. Spatial information is information related to the space in which a listener who hears the sound based on the sound signal is located. Specifically, spatial information is information related to the prescribed position (positioning position) when the sound image of the sound is localized at a prescribed position in the sound space (for example, in a three-dimensional sound field), that is, when the listener perceives the sound as arriving from a prescribed direction. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.

[0100] The sound source object information is information about an object that generates sound based on a sound signal, that is, an object that reproduces the sound signal, and is information about a virtual object (sound source object) arranged in a virtual space corresponding to the real space in which the object is arranged, that is, a sound space. The sound source object information includes, for example, information indicating the position of the sound source object arranged in the sound space, information about the direction of the sound source object, information about the directionality of the sound emitted by the sound source object, information indicating whether the sound source object is a living thing, and information indicating whether the sound source object is a moving body. For example, a sound signal corresponds to one or more sound source objects indicated by the sound source object information.

[0101] As an example of the data structure of a bit stream, the bit stream is composed of, for example, metadata (control information) and an audio signal.

[0102] The sound signal and metadata can be stored in one bitstream or in multiple bitstreams. Similarly, the sound signal and metadata can be stored in one file or in multiple files.

[0103] The bit stream may exist for each sound source or for each playback time. When the bit stream exists for each playback time, a plurality of bit streams may be processed in parallel at the same time.

[0104] The metadata may be assigned to each bit stream, or may be assigned as information for controlling a plurality of bit streams. Furthermore, the metadata may be assigned to each playback time.

[0105] When the sound signal and metadata are stored in multiple bitstreams or multiple files, they may include information indicating other bitstreams or files associated with one or part of the bitstreams or files, or may include information indicating other bitstreams or files associated with all the bitstreams or files. Here, the associated bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. In addition, the associated bitstreams or files may also include bitstreams or files that describe information indicating other associated bitstreams or files.

[0106] Here, the information indicating other associated bitstreams or files is, for example, an identifier indicating the other bitstream, a file name indicating other files, a URL (Uniform Resource Locator) or a URI (Uniform Resource Identifier), etc. In this case, the acquisition unit 111 determines or acquires the bitstream or file based on the information indicating other associated bitstreams or files. In addition, the bitstream may include information indicating other associated bitstreams, and the bitstream may include information indicating bitstreams or files associated with other bitstreams or files. Furthermore, the file containing information indicating associated bitstreams or files may be, for example, a control file such as a manifest file used for content distribution.

[0107] In addition, all or part of the metadata may be obtained from outside the bitstream of the audio signal. For example, either metadata for controlling audio or metadata for controlling video may be obtained from outside the bitstream, or both metadata may be obtained from outside the bitstream. Furthermore, when metadata for controlling video is included in the bitstream obtained by the audio signal reproduction system (corresponding to the audio reproduction system 100), the audio signal reproduction system may also have a function of outputting metadata that can be used for video control to a display device that displays an image or a stereoscopic video reproduction device that reproduces a stereoscopic video (for example, the stereoscopic video reproduction device 200 in the embodiment).

[0108] Furthermore, examples of information included in metadata will be described.

[0109] Metadata may also be information used to describe a scene represented by a sound space. Here, a scene is a term used to represent a collection of all elements of a three-dimensional image and sound event in a sound space modeled by a sound signal reproduction system using metadata. That is, the metadata described here includes not only information for controlling sound processing, but also information for controlling image processing. Of course, metadata may include information for controlling only one of the sound processing and the image processing, or information used for controlling both.

[0110] The sound signal reproduction system performs sound processing on the sound signal by using metadata included in the bitstream and additional interactive listener position information, etc., to generate a virtual sound effect. In this embodiment, the case where initial reflection processing, obstacle processing, diffraction processing, occlusion processing and reverberation processing are performed in the sound effect is described, but other sound processing can also be performed using metadata. For example, it can be considered that the sound signal reproduction system adds sound effects such as distance attenuation effect, positioning, and Doppler effect. In addition, information for switching on and off all or part of the sound effects and priority information can also be added as metadata.

[0111] In addition, as an example, the encoded metadata includes: information related to the sound space containing the sound source object and the obstacle object, and information related to the positioning position when the sound image of the sound is positioned at a specified position in the sound space (that is, so that it is perceived as a sound arriving from a specified direction). Here, the obstacle object is an object that may affect the sound perceived by the listener, such as blocking or reflecting the sound during the period from the sound emitted by the sound source object to the sound reaching the listener. In addition to stationary objects, obstacle objects may also include animals such as humans or moving objects such as machinery. In addition, when there are multiple sound source objects in the sound space, for any sound source object, other sound source objects may become obstacle objects. Non-sound-emitting objects that do not emit sound, such as building materials or inanimate objects, and sound source objects that emit sound may also become obstacle objects.

[0112] The metadata includes all or part of the information representing the shape of the sound space, the shape information and position information of obstacle objects existing in the sound space, the shape information and position information of sound source objects existing in the sound space, and the position and orientation of the listener in the sound space.

[0113] The sound space may be a closed space or an open space. In addition, the metadata includes information indicating the reflectivity of structures such as floors, walls, or ceilings that can reflect sound in the sound space, and the reflectivity of obstacle objects present in the sound space. Here, the reflectivity is the energy ratio of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In the case where the sound space is an open space, for example, uniformly set parameters such as attenuation rate, diffracted sound, or initial reflected sound can also be used.

[0114] In the above description, reflectivity is listed, but the parameter related to the obstacle object or the sound source object included in the metadata may include information other than reflectivity. For example, information other than reflectivity may include information related to the material of the object as metadata related to both the sound source object and the non-sound object. Specifically, information other than reflectivity may include parameters such as diffusivity, transmittance, and sound absorption.

[0115] As information related to the sound source object, it is also possible to include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources emitted from an object, and information specifying the sound source area in the object. For example, the reproduction conditions can also be set to be a sound that flows continuously or a sound triggered by an event. The sound source area in the object can be set by the relative relationship between the position of the listener and the position of the object, or it can be set based on the object. In the case where the sound source area in the object is set by the relative relationship between the position of the listener and the position of the object, the listener can perceive that sound A is emitted from the right side of the object and sound B is emitted from the left side when the listener is looking at it, based on the surface of the object that the listener is looking at. In the case where the sound source area in the object is set based on the object, which sound is emitted from which area of ​​the object regardless of the direction the listener is looking at can be fixed. For example, the listener can perceive that when looking at the object from the front, high sound flows from the right side and low sound flows from the left side. In this case, when the listener goes around the back of the object, the listener can perceive that low sound flows from the right side and high sound flows from the left side when looking from the back.

[0116] Metadata related to the space may include the time until the initial reflected sound, the reverberation time, the ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the listener can perceive only the direct sound.

[0117] use Figure 4 An example of the acquisition unit 111 will be described. Figure 4 is a block diagram showing the functional structure of the acquisition unit according to the embodiment. Figure 4 As shown, the acquisition unit 111 of the present embodiment includes, for example, a coded sound information input unit 112 , a decoding processing unit 113 , and a sensing information input unit 114 .

[0118] The coded sound information input unit 112 is a processing unit to which the coded (in other words, decoded) sound information acquired by the acquisition unit 111 is input. The coded sound information input unit 112 outputs the input sound information to the decoding processing unit 113.

[0119] The decoding processing unit 113 is a processing unit that decodes (in other words, decodes) the sound information output from the coded sound information input unit 112 to generate information on a predetermined sound and information on a predetermined position included in the sound information in a format used for subsequent processing.

[0120] The sensing information input unit 114 will be described below together with the function of the detector 103 .

[0121] The detector 103 is a device for detecting the movement speed of the head of the user 99. The detector 103 is composed of a combination of various sensors for detecting movement, such as a gyro sensor and an acceleration sensor. In the present embodiment, the detector 103 is built into the sound reproduction system 100, but it can also be built into an external device such as a stereoscopic image reproduction device 200 that operates in accordance with the movement of the head of the user 99 in the same way as the sound reproduction system 100. In this case, the detector 103 may not be included in the sound reproduction system 100. In addition, as the detector 103, an external camera device or the like may be used to capture the movement of the head of the user 99, and the movement of the user 99 may be detected by processing the captured image.

[0122] The detector 103 is, for example, integrally fixed to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. Since the sound reproduction system 100 including the housing moves integrally with the head of the user 99 after being worn by the user 99, the detector 103 can detect the speed of movement of the head of the user 99 as a result.

[0123] The detector 103 can detect, for example, the amount of movement of the head of the user 99, the amount of rotation with at least one of three axes orthogonal to each other in the three-dimensional space as the rotation axis, or the amount of displacement with at least one of the three axes as the displacement direction. In addition, the detector 103 can detect both the amount of rotation and the amount of displacement as the amount of movement of the head of the user 99.

[0124] The sensing information input unit 114 obtains the movement speed of the head of the user 99 from the detector 103. More specifically, the sensing information input unit 114 obtains the movement amount of the head of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 obtains at least one of the rotation speed and the displacement speed from the detector 103. The movement amount of the head of the user 99 obtained here is used to determine the position and posture (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. In the sound reproduction system 100, the relative position of the sound image is determined based on the determined coordinates and orientation of the user 99, and the sound is reproduced. Therefore, the listening point in the three-dimensional sound field can be changed according to the movement amount of the head of the user 99. In other words, the sensing information input unit 114 can receive an instruction to change the relative position of the listening point and the sound image (sound source object), and the instruction includes a first change amount to change the relative position according to the instruction. In addition, the relative position is a concept that represents the position of one party relative to the other party, which is expressed by at least one of the relative distance and relative direction between the sound pickup device or the listening point and the sound image (sound source object).

[0125] The processing unit 121 determines, based on the determined coordinates and orientation of the user 99, from which direction in the three-dimensional sound field the user 99 perceives the sound as coming, and processes the sound information so that the reproduced output sound information becomes such a sound. In addition, the processing unit 121 performs an acoustic processing for imparting fluctuations together with the above-mentioned processing. The fluctuations imparted here include fluctuations in the relative distance between the sound source object and the sound pickup device that repeatedly changes in the time domain, and fluctuations in the relative direction between the sound source object and the sound pickup device that repeatedly changes in the time domain.

[0126] Figure 5 1 is a block diagram showing the functional structure of the processing unit according to the embodiment. Figure 5 As shown, the processing unit 121 includes a determination unit 122, a storage unit 123, and an execution unit 124 as a functional unit for executing sound processing. In addition, the processing unit 121 includes other functional units (not shown) as functional units related to the processing of the sound information.

[0127] The determination unit 122 makes a determination to determine whether to execute the audio processing. The determination unit 122 determines whether a predetermined condition is satisfied, for example, and determines to execute the audio processing if the predetermined condition is satisfied, and determines not to execute the audio processing if the predetermined condition is not satisfied. Details of the predetermined condition will be described later. Information indicating the predetermined condition is stored in a storage device, for example, by the storage unit 123.

[0128] The storage unit 123 is a storage controller that performs processing for storing information in a storage device (not shown) storing information and reading information.

[0129] The execution unit 124 executes the audio processing according to the determination result of the determination unit 122 .

[0130] The signal output unit 141 is a functional unit that generates an output sound signal and outputs the generated output sound signal to the driver 104 .

[0131] The signal output unit 141 generates an output sound signal as digital data for the sound information after performing the sound processing according to the judgment result together with the processing for determining the fixed position of the sound and positioning it at the position. In addition, the signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and based on the waveform signal, the driver 104 generates sound waves to prompt the user 99 with sound. The driver 104 has a driving mechanism such as a vibration plate, a magnet, and a voice coil. The driver 104 operates the driving mechanism according to the waveform signal, and vibrates the vibration plate through the driving mechanism. In this way, the driver 104 generates sound waves (refers to "reproducing" the output sound signal, that is, the perception of the user 99 is not included in the meaning of "reproduction") through the vibration of the vibration plate corresponding to the output sound signal, and the sound waves propagate in the air and are transmitted to the ears of the user 99, and the user 99 perceives the sound.

[0132] [Another Example of the Sound Reproduction System According to the Present Embodiment]

[0133] In the above example, the sound reproduction system 100 of the present embodiment is described as a sound prompt device, which includes an information processing device 101, a communication module 102, a detector 103, and a driver 104. However, the functions of the sound reproduction system 100 may be implemented by multiple devices or by a single device. Figure 6 to Figure 15 Provide explanation. Figure 6 to Figure 15 This is a diagram for explaining another example of the sound reproduction system according to the embodiment.

[0134] For example, the information processing device 601 may be included in the sound prompt device 602, and the sound prompt device 602 may perform both sound processing and sound prompting. In addition, the information processing device 601 and the sound prompt device 602 may share the sound processing described in the present disclosure, or a server connected to the information processing device 601 or the sound prompt device 602 via a network may implement part or all of the sound processing described in the present disclosure.

[0135] In the above description, the information processing device 601 is referred to, but when the information processing device 601 decodes a bit stream generated by encoding data of at least a part of the spatial information used in the sound signal or sound processing to perform sound processing, the information processing device 601 may be referred to as a decoding device, and the sound reproduction system 100 (i.e., the stereo sound reproduction system 600 in the figure) may be referred to as a decoding processing system.

[0136] Here, an example in which the sound reproduction system 100 functions as a decoding processing system will be described.

[0137] <Example of Encoding Device>

[0138] Figure 7 Detailed Description of the Invention It is a functional block diagram showing the configuration of an encoding device 700 which is an example of an encoding device according to the present disclosure.

[0139] Input data 701 is data to be encoded and includes spatial information and / or a sound signal input to encoder 702. The details of spatial information will be described later.

[0140] The encoder 702 encodes the input data 701 to generate encoded data 703. The encoded data 703 is, for example, a bit stream generated by encoding processing.

[0141] The memory 704 stores the encoded data 703. The memory 704 may be, for example, a hard disk or an SSD (Solid-State Drive), or may be another storage device.

[0142] In addition, in the above description, a bit stream generated by encoding processing is cited as an example of the encoded data 703 stored in the memory 704, but it can also be data other than a bit stream. For example, the encoding device 700 can also store in the memory 704 the transformed data generated by converting the bit stream into a prescribed data format. The transformed data can also be, for example, a file or a multiplexed stream storing one or more bit streams. Here, the file is a file having a file format such as ISOBMFF (ISOBase Media File Format: ISO base media file format). In addition, the encoded data 703 can also be in the form of multiple packets generated by dividing the above-mentioned bit stream or file. In the case of converting the bit stream generated by the encoder 702 into data different from the bit stream, the encoding device 700 can also have a conversion unit not shown in the figure, and the conversion processing can also be performed by a CPU (Central Processing Unit: central processing unit).

[0143] <Example of Decoding Device>

[0144] Figure 8 2 is a functional block diagram showing the configuration of a decoding device 800 which is an example of a decoding device according to the present disclosure.

[0145] The memory 804 stores, for example, the same data as the coded data 703 generated by the coding device 700. The memory 804 reads out the stored data and inputs it as input data 803 of the decoder 802. The input data 803 is, for example, a bit stream to be decoded. The memory 804 may be, for example, a hard disk or an SSD, or other storage device.

[0146] In addition, the decoding device 800 may not use the data stored in the memory 804 as input data 803 as it is, but may use the transformed data generated by transforming the read data as input data 803. The data before transformation may be, for example, multiplexed data storing one or more bit streams. Here, the multiplexed data may also be a file having a file format such as ISOBMFF. In addition, the data before transformation may also be in the form of a plurality of packets generated by dividing the above-mentioned bit stream or file. In the case of transforming data different from the bit stream read from the memory 804 into a bit stream, the decoding device 800 may also include a transform unit not shown in the figure, or the transform processing may be performed by the CPU.

[0147] The decoder 802 decodes the input data 803 and generates a sound signal 801 for prompting the listener.

[0148] <Another Example of Encoding Device>

[0149] Fig. 9 is a functional block diagram showing the configuration of an encoding device 900 as another example of the encoding device of the present disclosure. Fig. 9 For Figure 7 The same function as the structure of Figure 7 The same reference numerals are used for the components, and description of these components is omitted.

[0150] The coding device 700 includes a memory 704 for storing the coded data 703 , whereas the coding device 900 is different from the coding device 700 in that it includes a transmission unit 901 for transmitting the coded data 703 to the outside.

[0151] Transmitter 901 transmits transmission signal 902 to another device or server based on coded data 703 or data in another data format generated by transforming coded data 703. Data used to generate transmission signal 902 is, for example, the bit stream, multiplexed data, file, or packet described in coding device 700.

[0152] <Another Example of Decoding Device>

[0153] Fig.10 1 is a functional block diagram showing the structure of a decoding device 1000 which is another example of the decoding device of the present disclosure. Fig.10 For Figure 8 The same function as the structure of Figure 8 The same reference numerals are used for the components, and description of these components is omitted.

[0154] The decoding device 800 includes a memory 804 that reads the input data 803 , whereas the decoding device 1000 is different from the decoding device 800 in that it includes a receiving unit 1001 that receives the input data 803 from the outside.

[0155] The receiving unit 1001 receives the received signal 1002 to obtain received data, and outputs input data 803 to be input to the decoder 802. The received data may be the same as the input data 803 to be input to the decoder 802, or may be data in a data format different from that of the input data 803. When the received data is data in a data format different from that of the input data 803, the receiving unit 1001 may convert the received data into the input data 803, or a conversion unit (not shown) or a CPU included in the decoding device 1000 may convert the received data into the input data 803. The received data is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 900.

[0156] <Decoder Functional Description>

[0157] Fig.11 It means as Figure 8 or Fig.10 A functional block diagram of the structure of a decoder 1100 which is an example of the decoder 802 in FIG.

[0158] The input data 803 is a coded bit stream, and includes coded audio data, which is a coded audio signal, and metadata used in audio processing.

[0159] The spatial information management unit 1101 obtains metadata included in the input data 803 and parses the metadata. The metadata includes information describing elements that act on the sound and are configured in the sound space. The spatial information management unit 1101 manages the spatial information required for sound processing obtained by parsing the metadata, and provides the spatial information to the rendering unit 1103. In addition, in the present disclosure, the information used in the sound processing is referred to as spatial information, but it may be referred to by other names. The information used in the sound processing may be referred to as sound spatial information or scene information, for example. In addition, in the case where the information used in the sound processing changes over time, the spatial information input to the rendering unit 1103 may also be referred to as a spatial state, a sound spatial state, a scene state, or the like.

[0160] Furthermore, the spatial information may be managed for each sound space or for each scene. For example, when representing different rooms as virtual spaces, each room may be managed as a scene of a different sound space, or the spatial information may be managed as the same space but as different scenes depending on the occasion of the representation. In the management of the spatial information, an identifier for identifying each piece of spatial information may also be assigned. The spatial information data may be included in a bit stream as a form of input data 803, or the bit stream may include an identifier of the spatial information and the spatial information data may be obtained from outside the bit stream. In the case where only the identifier of the spatial information is included in the bit stream, the identifier of the spatial information may be used during rendering to obtain the spatial information data stored in the memory of the sound signal processing device or in an external server as input data.

[0161] In addition, the information managed by the spatial information management unit 1101 is not limited to the information included in the bitstream. For example, the input data 803 may include data indicating spatial characteristics or structures obtained from a software application or server that provides VR or AR as data not included in the bitstream. In addition, for example, the input data 803 may include data indicating characteristics or positions of listeners or objects as data not included in the bitstream. In addition, the input data 803 may include information obtained by a sensor provided by a terminal including a decoding device, or information indicating the position of the terminal estimated based on the information obtained by the sensor, as information indicating the position of the listener. That is, the spatial information management unit 1101 may communicate with an external system or server to obtain spatial information and the position of the listener. In addition, the spatial information management unit 1101 may obtain clock synchronization information from an external system and perform processing synchronized with the clock of the rendering unit 1103. In addition, the space in the above description may be a virtually formed space, that is, a VR space, or a real space or a virtual space corresponding to a real space, that is, an AR space or an MR (Mixed Reality) space. In addition, the virtual space may also be referred to as a sound field or a sound space. In addition, the information indicating the position in the above description may be information such as coordinate values ​​indicating the position in the space, information indicating the relative position relative to a predetermined reference position, or information indicating the movement or acceleration of the position in the space.

[0162] The audio data decoder 1102 decodes the encoded audio data included in the input data 803 to obtain an audio signal.

[0163] The coded audio data acquired by the stereo sound reproduction system 600 is, for example, a bit stream coded in a prescribed format such as MPEG-H 3D Audio (ISO / IEC 23008-3). In addition, MPEG-H 3D Audio is only an example of a coding method that can be used when generating coded audio data included in a bit stream, and may also include a bit stream and coded audio data coded in other coding methods. For example, the coding method used may be a non-reversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), Vorbis, or a reversible codec such as ALAC (Apple Lossless Audio Codec), FLAC (Free Lossless Audio Codec), or any coding method other than the above may be used. For example, PCM (pulse code modulation) data may be one type of coded audio data. In this case, for example, when the number of quantization bits of the PCM data is N, the decoding process may be a process of converting an N-bit binary number into a number format (eg, floating point format) that can be processed by the rendering unit 1103 .

[0164] The rendering unit 1103 takes the sound signal and the spatial information as input, performs acoustic processing on the sound signal using the spatial information, and outputs the sound signal 801 after the acoustic processing.

[0165] Before starting rendering, the spatial information management unit 1101 reads metadata of the input signal, detects rendering items such as objects or sounds specified by the spatial information, and sends them to the rendering unit 1103. After the rendering starts, the spatial information management unit 1101 grasps the changes over time of the spatial information and the position of the listener, updates the spatial information and manages it. In addition, the spatial information management unit 1101 sends the updated spatial information to the rendering unit 1103. The rendering unit 1103 generates a sound signal with sound processing added based on the sound signal included in the input data and the spatial information received from the spatial information management unit 1101, and outputs it.

[0166] The spatial information update process and the sound signal output process with added acoustic processing may be executed by the same thread, or the spatial information management unit 1101 and the rendering unit 1103 may be assigned to separate threads. When the spatial information update process and the sound signal output process with added acoustic processing are processed by different threads, the thread activation frequency may be set separately, or the processes may be executed in parallel.

[0167] Since the spatial information management unit 1101 and the rendering unit 1103 execute processing in different independent threads, the computing resources can be preferentially allocated to the rendering unit 1103, so in the case of sound processing in which no delay can be allowed, for example, even in the case of a 1-sample (0.02msec) delay, a small noise can be generated. The sound processing can be safely implemented. At this time, the allocation of computing resources to the spatial information management unit 1101 is restricted. However, the update of spatial information is a lower-frequency process than the output processing of the sound signal (for example, a process such as updating the orientation of the listener's face). Therefore, unlike the output processing of the sound signal, it does not require an instantaneous response, so even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality brought to the listener.

[0168] The updating of spatial information can be performed periodically at a predetermined time or period, or when a predetermined condition is satisfied. In addition, the updating of spatial information can be performed manually by the listener or the manager of the sound space, or it can be performed in response to changes in the external system. For example, when the listener operates the controller and the standing position of his or her avatar is instantly tilted, or the time is instantly moved forward or backward, or when the manager of the virtual space suddenly changes the scene, the thread configured with the spatial information management unit 1101 can be started as a one-shot interrupt processing in addition to the regular startup.

[0169] The information update thread that performs the update processing of spatial information is responsible for, for example, updating the position or orientation of the listener's avatar configured in the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of objects moving in the virtual space, etc., which are provided in a processing thread that is started at a relatively low frequency of about tens of Hz. Such a processing thread with a low occurrence frequency can also be used to perform processing that reflects the properties of direct sound. This is because the frequency of changes in the properties of direct sound is lower than the frequency of occurrence of audio processing frames used for audio output. On the contrary, doing so can relatively reduce the computational load of the processing, and if the information is updated at an unnecessarily fast frequency, there is a risk of generating pulsed noise, so this risk can also be avoided.

[0170] Fig.12 It means as Figure 8 or Fig.10 A functional block diagram of the structure of a decoder 1200 which is another example of the decoder 802 in FIG.

[0171] Fig.12 The input data 803 does not include coded audio data but includes an uncoded audio signal. Fig.11 The input data 803 includes a bit stream including metadata and an audio signal.

[0172] Spatial Information Management Department 1201 and Fig.11 Since the spatial information management unit 1101 is the same as that of FIG. 1101 , the description thereof will be omitted.

[0173] Rendering unit 1202 and Fig.11 The rendering unit 1103 is the same as that of FIG. 1 , so the description is omitted.

[0174] In addition, in the above description Fig.12 The structure of is called a decoder, but it can also be called an audio processing unit that performs audio processing. In addition, the device including the audio processing unit can also be called an audio processing device instead of a decoding device. In addition, the audio signal processing device (information processing device 601) can also be called an audio processing device.

[0175] <Physical Structure of Encoding Device>

[0176] Fig.13 is a diagram showing an example of the physical structure of an encoding device. Fig.13 The illustrated encoding device is an example of the encoding devices 700 and 900 described above.

[0177] Fig.13 The encoding device includes a processor, a memory and a communication IF.

[0178] The processor is, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit), and the encoding process of the present disclosure may be implemented by executing a program stored in a memory by the CPU, DSP or GPU. In addition, the processor may also be a dedicated circuit for performing signal processing of sound signals including the encoding process of the present disclosure.

[0179] The memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). In addition, the memory may also include an internal memory embedded in a CPU or a GPU.

[0180] The communication IF (Interface) is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF, and transmits an encoded bit stream.

[0181] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is cited as an example of the communication method, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio) or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.

[0182] <Physical Structure of the Acoustic Signal Processing Device>

[0183] Fig.14 is a diagram showing an example of the physical structure of an audio signal processing device. Fig.14 The audio signal processing device may also be a decoding device. In addition, a part of the structure described here may also be equipped in the audio prompting device 602. In addition, Fig.14 The illustrated sound signal processing device is an example of the sound signal processing device 601 described above.

[0184] Fig.14 The sound signal processing device includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0185] The processor is, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit), and the CPU, DSP or GPU may execute a program stored in a memory to implement the audio processing or decoding processing of the present invention. In addition, the processor may also be a dedicated circuit for performing signal processing of sound signals including the audio processing of the present invention.

[0186] The memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). In addition, the memory may also include an internal memory embedded in a CPU or a GPU.

[0187] The communication IF (Interface) is, for example, a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The sound signal processing device shown in FIG2I has a function of communicating with other communication devices via the communication IF, and obtains a bit stream of a decoding object. The obtained bit stream is stored in, for example, a memory.

[0188] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is cited as an example of the communication method, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio) or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.

[0189] The sensor performs sensing for estimating the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, speed, angular velocity, or acceleration of a part or the whole of the body such as the head of the listener, and generates position information indicating the position and / or orientation of the listener. In addition, the position information may also be information indicating the position and / or orientation of the listener in real space, or information indicating the displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a specified point in time. In addition, the position information may also be information indicating the relative position and / or orientation to a stereo sound reproduction system or an external device having a sensor.

[0190] The sensor may be, for example, a camera or other imaging device or a distance measuring device such as LiDAR (Light Detection And Ranging), and may capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. In addition, as a sensor, for example, a device that uses wireless position estimation using any frequency band such as millimeter waves may be used.

[0191] in addition, Fig.14 The acoustic signal processing device shown in FIG. 1 may also obtain position information from an external device having a sensor via a communication IF. In this case, the acoustic signal processing device may not include a sensor. Here, the external device is, for example, Figure 6 The sound presentation device 602 described in the above or the stereoscopic image reproduction device worn on the head of the listener, etc. In this case, the sensor is composed of a combination of various sensors such as a gyro sensor and an acceleration sensor.

[0192] For example, as the speed of movement of the listener's head, the sensor can detect the angular velocity of rotation about at least one of the three axes orthogonal to each other in the sound space, or the acceleration of displacement in the direction of displacement about at least one of the three axes.

[0193] For example, as the amount of movement of the listener's head, the sensor can detect the amount of rotation with at least one of the three axes orthogonal to each other in the sound space as the rotation axis, or the amount of displacement with at least one of the three axes as the displacement direction. Specifically, the sensor detects 6DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is composed of a combination of various sensors for detecting movement, such as a gyro sensor and an acceleration sensor.

[0194] In addition, the sensor can be implemented as long as it can detect the position of the listener, and can be implemented by a camera or a GPS (Global Positioning System) receiver, etc. It is also possible to use position information obtained by self-position estimation using LiDAR (Laser Imaging Detection and Ranging) etc. For example, when the sound signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.

[0195] In addition, sensors can also include detection Fig.14 A temperature sensor such as a thermocouple for detecting the temperature of the sound signal processing device shown, and a sensor for detecting the remaining amount of a battery included in or connected to the sound signal processing device, etc.

[0196] The speaker has a driving mechanism such as a diaphragm, a magnet or a voice coil, and an amplifier, and presents the sound signal after acoustic processing to the listener as sound. The speaker operates the driving mechanism according to the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), and the driving mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the sound signal generates sound waves, which propagate in the air and are transmitted to the listener's ears, and the listener perceives the sound.

[0197] In addition, here are Fig.14 The example of the case where the audio signal processing device shown in the figure has a speaker and outputs the audio signal after the audio processing through the speaker is described, but the audio signal prompting mechanism is not limited to the above-mentioned structure. For example, the audio signal after the audio processing can also be output to the external audio prompting device 602 connected by the communication module. The communication performed by the communication module can be either wired or wireless. In addition, as another example, it can also be Fig.14 The audio signal processing device shown has a terminal for outputting an analog signal of sound, and a cable of earplugs or the like is connected to the terminal to prompt the sound signal from the earplugs or the like. In the above case, the sound prompting device 602, such as headphones, earplugs, head-mounted displays, neck speakers, wearable speakers, or surround speakers composed of a plurality of fixed speakers, etc., which are worn on the head or a part of the body of the listener, reproduces the sound signal.

[0198] <Rendering Functional Description>

[0199] Fig.15 Yes means Fig.11 and Fig.12 A functional block diagram of an example of the detailed structure of the rendering units 1103 and 1202.

[0200] The rendering unit is composed of an analyzing unit and a synthesizing unit, and performs sound processing on the sound data included in the input signal and outputs the resultant signal.

[0201] The following describes information included in the input signal.

[0202] The input signal is composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bit stream composed of sound data and metadata (control information), and in this case, the metadata may also include spatial information.

[0203] Spatial information is information about the sound space (three-dimensional sound field) formed by the stereo sound reproduction system, and is composed of information about objects contained in the sound space and information about listeners. Objects include sound source objects that emit sound and become sound sources, and non-sound objects that do not emit sound. Non-sound objects function as obstacle objects that reflect the sound emitted by sound source objects, but there are also cases where they function as obstacle objects that reflect the sound emitted by other sound source objects.

[0204] The information given to both the sound source object and the non-sound emitting object includes position information, shape information, and a volume attenuation rate when the object reflects sound.

[0205] The position information is represented by the coordinate values ​​of three axes, such as the X-axis, Y-axis, and Z-axis in Euclidean space, but it does not necessarily have to be three-dimensional information. For example, it can also be two-dimensional information represented by the coordinate values ​​of the X-axis and Y-axis. The position information of the object is determined by the representative position of the shape represented by the grid or voxel.

[0206] The shape information may also include information about the material of the surface.

[0207] In addition, it may include information indicating whether the object is a living being, information indicating whether the object is a moving object, etc. When the object is a moving object, the position information may move over time, and the changed position information or the amount of change may be transmitted to the rendering unit.

[0208] The information on the sound source object includes, in addition to the information commonly given to the sound source object and the non-sound emitting object, sound data and information necessary for emitting the sound data into the sound space.

[0209] The sound data is data that represents the sound perceived by the listener, such as information about the frequency and strength of the sound. The sound data is typically a PCM signal, but it can also be data compressed using an encoding method such as MP3. In this case, it is necessary to decode the signal at least before it reaches the synthesis unit, so a decoding unit (not shown) can also be included in the rendering unit. Alternatively, it can also be decoded by the sound data decoder 1102.

[0210] It is sufficient to set at least one sound data for one sound source object, or a plurality of sound data may be set. In addition, identification information for identifying each sound data may be given, and the identification information of the sound data may be stored as information related to the sound source object.

[0211] The information required for radiating the sound data into the sound space may include, for example, information on a reference volume used as a reference when reproducing the sound data, information indicating the nature (also called characteristics) of the sound data, information on the position of the sound source object, information on the orientation of the sound source object, information on the directivity of the sound emitted by the sound source object, etc. The information on the reference volume may be, for example, the effective value of the amplitude value of the sound data at the sound source position when the sound data is radiated into the sound space, and may be expressed as a decibel (dB) value using a floating point.

[0212] For example, when the reference volume is 0 dB, it may be indicated that the sound is radiated to the sound space from the position indicated by the information related to the position without increasing or decreasing the volume of the signal level indicated by the sound data, and when it is -6 dB, it may be indicated that the sound is radiated to the sound space from the position indicated by the information related to the position with the volume of the signal level indicated by the sound data reduced to about half. Such information may be given to one sound data or to a plurality of sound data at once.

[0213] Information indicating the nature of sound data may be, for example, information about the volume of a sound source, and information indicating its changes in time series. For example, when the sound space is a virtual conference room and the sound source is a speaker, the volume changes intermittently in a short period of time. To express it more simply, it can be said that the sound part and the silent part are generated alternately.

[0214] In addition, when the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain length of time. In addition, when the sound space is a battlefield and the sound source is an explosive, the volume of the explosion sound increases only for a moment, and then becomes silent and continues. In this way, the information on the volume of the sound source includes not only the information on the size of the sound, but also the information on the change of the size of the sound, and such information can also be used as information indicating the nature of the sound data.

[0215] Here, the information on the change in the volume of the sound may be data representing the frequency characteristics in a time series. It may also be data representing the duration of a sound interval. It may also be data representing the duration of a sound interval and the duration of a silent interval. It may also be data listing a plurality of time series for a duration that can be regarded as a constant (regarded as approximately constant) amplitude of a sound signal and the amplitude value of the signal during that duration, etc. It may also be data for a duration that can be regarded as a constant frequency characteristic of a sound signal. It may also be data listing a plurality of time series for a duration that can be regarded as a constant frequency characteristic of a sound signal and the frequency characteristics during that duration, etc.

[0216] As a form of data, for example, data representing the approximate shape of a spectrogram may be used. In addition, the volume that serves as a reference for the above-mentioned frequency characteristics may be set as the above-mentioned reference volume. In addition to being used to calculate the volume of direct sound or reflected sound perceived by the listener, the information representing the reference volume and the information of the properties of the sound data are also used in the selection process for selecting whether to make the listener perceive it. Other examples of the information representing the properties of the sound data or the use thereof in specific selection processes will be described later.

[0217] The information related to the orientation is typically expressed by yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted and the orientation information may be expressed by azimuth (yaw) and elevation (pitch). The orientation information may also change over time and, if changed, is transmitted to the rendering unit.

[0218] Information related to the listener is information related to the position information and orientation of the listener in the sound space. The position information is represented by the position of the XYZ axis of the Euclidean space, but it does not necessarily have to be three-dimensional information, and it can also be two-dimensional information. Information related to the orientation is typically represented by yaw, pitch, and roll. Alternatively, the rotation of roll can be omitted and represented by azimuth (yaw) and elevation (pitch). The position information and orientation information can also change over time and are transmitted to the rendering unit when changes occur.

[0219] The sensor information is information including the amount of rotation or displacement detected by the sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering unit, and the rendering unit updates the information of the position and orientation of the listener based on the sensor information. For example, the sensor information can also use the position information obtained by the portable terminal using GPS, camera or LiDAR (Laser Imaging Detection and Ranging) to estimate its own position. In addition, information obtained from the outside via the communication module other than the sensor can also be detected as sensor information. Information indicating the temperature of the sound signal processing device and information indicating the remaining battery level can also be obtained from the sensor. The computing resources (CPU capacity, memory resources, PC performance) of the sound signal processing device or the sound signal prompting device can also be obtained in real time.

[0220] The analyzing unit has the same function as the acquiring unit 111 in the above example. That is, it analyzes the input signal and acquires necessary information from the processing unit 121.

[0221] The synthesizing unit has the same functions as the processing unit 121 and the signal output unit 141 in the above example. Based on the sound signal of the direct sound and the information of the arrival time of the direct sound and the volume of the direct sound calculated by the analyzing unit, the input sound signal is processed to generate the direct sound. In addition, based on the information of the arrival time of the reflected sound and the volume of the reflected sound calculated by the analyzing unit, the input sound signal is processed to generate the reflected sound. The synthesizing unit synthesizes the generated direct sound and the reflected sound and outputs them.

[0222] [action]

[0223] Next, refer to Figure 16 to Figure 19 , the operation of the sound reproduction system 100 described above will be described. Fig.16 : is a flowchart showing the operation of the sound reproduction system according to the embodiment. Fig.17 It is a diagram for explaining the frequency characteristics of the sound processing according to the embodiment. Fig.18 This is a diagram for explaining the magnitude of fluctuations in the acoustic processing according to the embodiment. Fig.19 It is a diagram for explaining the period and angle of ripple in the acoustic processing according to the embodiment.

[0224] In addition, assuming that Fig.16 The steps shown above are described by setting the audio processing to be performed based on the determination of the control information. Fig.16As shown, first, the sound information (sound signal) is acquired by the acquisition unit 111 (S101). Then, the determination unit 122 determines whether to execute the sound processing. Specifically, the determination unit 122 reads the prescribed conditions stored in the storage unit 123, and determines whether to execute the sound processing by determining whether the prescribed conditions are satisfied (S102).

[0225] Several examples of the prescribed conditions are described below.

[0226] First, when the change in the sound pressure of the predetermined sound in the acquired sound information in the time domain is less than a predetermined threshold value, it is considered that the predetermined sound in the sound information does not contain fluctuations and the addition of fluctuations is appropriate. If a condition related to the change in the sound pressure in the time domain is set as a condition that can be said to be suitable for performing sound processing, it can be determined that the predetermined condition is satisfied when the change in the sound pressure in the time domain is less than the above threshold value.

[0227] Here, in Fig.17 The figure shows the difference in distances reached by the sound at the same sound pressure in each direction in the horizontal plane when the sound of each frequency is emitted from the sound source (the center of each dotted circle). Fig.17 The figures shown show the difference in propagation characteristics of sound in each direction at that frequency. It can be said that the more irregular the shape, the easier it is to reflect the fluctuation of the sound source. In other words, in order to judge the fluctuation of the sound source based on the change of sound pressure in the time domain, the specified sound can be decomposed into each frequency, and the frequency that more easily reflects the fluctuation of the sound source can be used to determine whether it represents the change of sound pressure in the time domain. For example, if the frequency is above 1000Hz as shown in the figure, the shape changes from a circle to an irregular shape, which can be said to be easier to reflect fluctuations. In addition, if the frequency is above 4000Hz as shown in the figure, the shape changes from a circle to a more irregular shape, which can be said to be easier to reflect fluctuations.

[0228] On the contrary, Fig.17 As shown in FIG. 1 , when fluctuations are applied, it can be said that even if the sound processing is performed on a frequency less than 1000 Hz, it is difficult to obtain the effect of fluctuations. Therefore, in the sound processing, the sound processing may be performed only on frequencies above 1000 Hz, or only on frequencies above 4000 Hz. Alternatively, the sound processing may be performed such that the larger the frequency, the greater the fluctuation.

[0229] Furthermore, the positional relationship between the sound pickup device and the sound source is estimated using the sound pressure of the predetermined position or predetermined sound in the acquired sound information. When the estimated positional relationship is below a predetermined threshold, it can be considered that a close-talking sound pickup device such as a microphone of an earphone is being used, so it can be considered that the predetermined sound in the sound information does not contain fluctuations and the addition of fluctuations is appropriate. As a condition that can be said to be suitable for performing sound processing, if a condition related to the estimated positional relationship is set, it can be determined that the predetermined condition is satisfied when the positional relationship is below the above threshold.

[0230] Here, in Fig.18 The figure shows the result of plotting the movement of the human head in the three axes of XYZ. Fig.18 In the figure, the upper part shows the plot of the movement of the head in the Y-axis direction (up and down direction), the middle part shows the plot of the movement of the head in the Z-axis direction (front and back direction), and the lower part shows the plot of the movement of the head in the X-axis direction (left and right direction). As shown in the figure, it can be seen that the human head moves ±0.2m in the X-axis direction (left and right direction), ±0.02m in the Y-axis direction (up and down direction), and ±0.05m in the Z-axis direction (front and back direction).

[0231] That is, if there is no movement of such magnitude, it can be expected that the estimated positional relationship is below a predetermined threshold value such as a close-talking sound pickup device such as a microphone of an earphone.

[0232] On the contrary, Fig.18 As shown, when the fluctuation is applied, the sound processing can be performed by reproducing the motion of ±0.2m in the X-axis direction (left-right direction), ±0.02m in the Y-axis direction (up-down direction), and ±0.05m in the Z-axis direction (front-back direction). In this way, the sound processing can be performed under the processing conditions corresponding to the positional relationship between the sound pickup device and the sound source.

[0233] In addition, Fig.19 The figure shows the movement of a person's head, and the rotation angles are plotted on the three rotation axes of Yaw, Pitch, and Roll. Fig.19 In the figure, the upper part shows the rotation angle of the Yaw angle, the middle part shows the rotation angle of the Pitch angle, and the lower part shows the rotation angle of the Roll angle. As shown in the figure, it can be seen that the human head rotates ±20 degrees in the Yaw angle, ±10 degrees in the Pitch angle, and ±3 degrees in the Yaw angle in a period of 3 to 4 seconds.

[0234] That is, if there is no such periodic and angular movement, it can be expected that the estimated positional relationship is below a predetermined threshold value like a close-talking sound pickup device such as a microphone of an earphone.

[0235] On the contrary, Fig.19 As shown, when the fluctuation is applied, the sound processing can be performed by reproducing a rotation of ±20 degrees in the Yaw angle, a rotation of ±10 degrees in the Pitch angle, and a rotation of ±3 degrees in the Yaw angle in a period of 3 to 4 seconds. In this way, the sound processing can be performed under the processing conditions corresponding to the positional relationship between the sound receiving device and the sound source.

[0236] Furthermore, using the sound collection status information related to the status when collecting sound, when the reverberation level and / or noise level indicated by the sound collection status information is below a predetermined threshold value, it can be assumed that a close-talking sound collection device such as a microphone of an earphone is being used, so it can be considered that the predetermined sound of the sound information does not contain fluctuations and the addition of fluctuations is appropriate. If a condition related to the reverberation level and / or noise level indicated by the sound collection status information is set as a condition that can be said to be suitable for performing sound processing, it can be determined that the predetermined condition is satisfied when the reverberation level and / or noise level indicated by the sound collection status information is below the aforementioned threshold value.

[0237] In addition to this, information related to the sound receiving device (information identifying the device such as a model or information representing the characteristics of the device such as whether waves need to be given) that collects sounds and the like may be used by a close-talking sound receiving device such as a microphone of headphones, and when this information indicates that a close-talking sound receiving device such as a microphone of headphones is used, it is determined that the specified conditions are met.

[0238] Back to Fig.16 If the determination unit 122 determines that the above-mentioned conditions are satisfied ("Yes" in S102), the execution unit 124 executes the audio processing (S103). On the other hand, if the determination unit 122 determines that the above-mentioned conditions are not satisfied ("No" in S102), the execution unit 124 does not execute the audio processing (S104). Next, the signal output unit 141 generates and outputs an output audio signal (S105).

[0239] [Another example]

[0240] Below, use Fig. 20 and Fig.21 A sound reproduction system according to another example of the embodiment will be described. Fig. 20 This is a block diagram showing a functional structure of a processing unit according to another example of the embodiment. Fig.21 1 is a flowchart showing the operation of the sound processing device according to another example of the embodiment. In the following description of another example, the “sound pickup device” described in part of the above embodiment may be referred to as a “listening point” and the description thereof may be omitted.

[0241] The sound reproduction system according to another example of the embodiment is different from the sound reproduction system 100 according to the above-described embodiment in that a processing unit 121 a is provided instead of the processing unit 121 .

[0242] The processing unit 121a includes a calculation unit 125 instead of the determination unit 122. The calculation unit 125 calculates a first change amount and a second change amount. The first change amount is a change amount based on an instruction to change the relative position of the listening point and the sound source object, and corresponds to the movement amount of the so-called movement in the VR space. And, if limited to the virtual sound space, it is a change amount of the change of the relative position of the listening point and the sound source object with the movement of the listening point. Regarding the first change amount, the indication of the change of the relative position at this time, that is, the change amount is obtained by obtaining the detection result from the detector 103 as a sensor. That is, in this example, the acquisition unit 111 (particularly the sensing information input unit 114) accepts the instruction including the first change amount.

[0243] In the present embodiment, in addition to such a change in relative position, a change in the listening point due to fluctuation also occurs, so the first change amount and the second change amount are calculated separately. In addition, by setting the second change amount to 0, it is possible to distinguish whether the sound processing is performed or not performed without the processing performed by the determination unit 122. The second change amount can be calculated based on the detection result or independently of the detection result. For example, the second change amount can also be calculated by using the change speed of the change in the relative position of the sound source object and the listening point indicated by the detection result, or a function of the first change amount as the change amount. Alternatively, the second change amount can also be calculated uniquely based on information such as control information and sound reception status information given to the content when the content is produced, without using (independent of) the change speed of the change in the relative position of the sound source object and the listening point or the first change amount as the change amount.

[0244] In addition, when the first variation is large, there is a case where the sound source object moves greatly relative to the stopped listening point. In such a case, it is natural that the larger the first variation is, the larger the fluctuation of the sound source object is. In other words, the larger the first variation is, the larger the second variation is. Therefore, in the sound processing, the second variation corresponding to the size of the fluctuation can be made larger according to the first variation as long as the first variation is larger.

[0245] On the other hand, in the audio processing, as an example, the second variation corresponding to the size of the fluctuation changes according to the first variation, and conversely, it is appropriate to set the second variation smaller (for example, 0) as the first variation is larger. Specifically, for example, when the first variation is large (or the speed of change of the relative position is fast), even if the fluctuation is given, the effect of increasing the sense of presence is not very visible. This is because the changes caused by the fluctuation and the changes in the relative position overlap or cancel each other synchronously, making it difficult for the listener to perceive that the fluctuation is given. In such a case, it is sufficient to set the second variation smaller (for example, 0) as the first variation is larger.

[0246] The following describes the operation of the audio reproduction system of this example. Fig.21 In the above steps, the following description is given of the setting for executing the audio processing by determining based on the control information. Fig.21 As shown, first, the sound information (sound signal) is acquired by the acquisition unit 111 (S201). Then, the calculation unit 125 calculates the first change amount (S202). In addition, the calculation unit 125 calculates the second change amount (S203). Whether to perform the sound processing (whether to give fluctuations) can be set according to whether the second change amount is calculated to be 0. In addition, the execution unit 124 performs the sound processing as the sound processing, which changes the relative position by the first change amount and repeatedly changes the relative position by the second change amount in the time domain (S204). Then, the signal output unit 141 generates an output sound signal and outputs it (S205).

[0247] (Other embodiments)

[0248] As mentioned above, although embodiment was described, this disclosure is not limited to the said embodiment.

[0249] For example, the sound reproduction system described in the above-mentioned embodiment can be implemented as a single device having all the components, or can be implemented by assigning each function to a plurality of devices and cooperating with the plurality of devices. In the latter case, as a device equivalent to the sound processing device, a sound processing device such as a smartphone, a tablet terminal or a PC can also be used. For example, in the sound reproduction system 100 having a function as a renderer that generates a sound signal with sound effects added, all or part of the functions of the renderer can also be assumed by a server. That is, all or part of the acquisition unit 111, the processing unit 121, and the signal output unit 141 can also exist in a server not shown. In this case, the sound reproduction system 100 is implemented by combining, for example, a sound processing device such as a computer or a smartphone, a sound prompt device such as a head-mounted display (HMD) worn on the user 99, and a server not shown. In addition, the computer, the sound prompt device, and the server can be connected to each other in a communicative manner through the same network or in different networks. When connected through different networks, there is a high possibility that communication delays occur, so the processing in the server can also be permitted only when the computer, the sound prompt device, and the server are connected to each other in a communicative manner through the same network. Furthermore, whether the server should assume all or part of the functions of the renderer may be determined based on the data volume of the bit stream received by the audio reproduction system 100 .

[0250] In addition, the sound reproduction system of the present disclosure can also be connected to a reproduction device having only a driver, and implemented as a sound processing device for the reproduction device that only reproduces the output sound signal generated based on the acquired sound information. In this case, the sound processing device can be implemented as hardware having a dedicated circuit or as software for causing a general-purpose processor to perform specific processing.

[0251] Furthermore, in the above-described embodiment, the processing executed by a specific processing unit may be executed by another processing unit. Furthermore, the order of a plurality of processing may be changed, or a plurality of processing may be executed in parallel.

[0252] In addition, in the above-mentioned embodiments, each component can also be implemented by executing a software program suitable for each component. Each component can also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[0253] In addition, each component may also be implemented by hardware. For example, each component may also be a circuit (or integrated circuit). These circuits may constitute one circuit as a whole, or they may be different circuits. In addition, these circuits may be general circuits or dedicated circuits.

[0254] In addition, the overall or specific form of the present disclosure may also be implemented by an apparatus, device, method, integrated circuit, computer program, or a computer-readable CD-ROM or other recording medium. In addition, the overall or specific form of the present disclosure may also be implemented by any combination of an apparatus, device, method, integrated circuit, computer program, and recording medium.

[0255] For example, the present disclosure may be implemented as a sound signal reproducing method executed by a computer, or as a program for causing a computer to execute the sound signal reproducing method. The present disclosure may also be implemented as a computer-readable non-transitory recording medium having such a program recorded thereon.

[0256] In addition, various modifications that can be conceived by those skilled in the art to the embodiments or embodiments achieved by arbitrarily combining the components and functions of the embodiments without departing from the gist of the present disclosure are also included in the present disclosure.

[0257] In addition, the encoded sound information in the present disclosure can be, in other words, a bit stream including a sound signal and metadata, wherein the sound signal is information about a predetermined sound reproduced by the sound reproduction system 100, and the metadata is information related to the localization position when the sound image of the predetermined sound is localized at a predetermined position in the three-dimensional sound field. The sound reproduction system 100 can also obtain the sound information as a bit stream encoded in a format specified by, for example, MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about the predetermined sound reproduced by the sound reproduction system 100. The predetermined sound mentioned here is a sound or natural environment sound emitted by a sound source object existing in the three-dimensional sound field, and can include, for example, mechanical sound or the sound of animals including humans. In addition, when there are multiple sound source objects in the three-dimensional sound field, the sound reproduction system 100 obtains multiple sound signals corresponding to the multiple sound source objects.

[0258] On the other hand, metadata is, for example, information used in the sound reproduction system 100 to control the sound processing of the sound signal. Metadata can also be information used to describe the scene represented by the virtual space (three-dimensional sound field). The scene mentioned here refers to a term that represents the collection of all elements of the three-dimensional image and sound events in the virtual space modeled by the sound reproduction system 100 using metadata. That is, the metadata mentioned here includes not only information for controlling the sound processing, but also information for controlling the image processing. Of course, the metadata can include information for controlling only one of the sound processing and the image processing, or information used for controlling both. In the present disclosure, the bit stream obtained by the sound reproduction system 100 sometimes includes such metadata. Alternatively, the sound reproduction system 100 can also obtain metadata separately from the bit stream as a single body as described later.

[0259] The sound reproduction system 100 generates virtual sound effects by performing sound processing on the sound signal using metadata included in the bitstream and additionally acquired interactive user 99 position information. For example, it is possible to add sound effects such as initial reflection sound generation, late reverberation sound generation, diffraction sound generation, distance attenuation effect, localization, sound image localization processing, or Doppler effect. In addition, information for switching on / off all or part of the sound effects may be added as metadata.

[0260] In addition, all or part of the metadata may be obtained from outside the bitstream of the audio information. For example, metadata for controlling the audio or metadata for controlling the video may be obtained from outside the bitstream, or both metadata may be obtained from outside the bitstream.

[0261] Furthermore, when the bitstream obtained by the sound reproduction system 100 includes metadata for controlling an image, the sound reproduction system 100 may output the metadata that can be used to control the image to a display device that displays an image or a stereoscopic image reproduction device that reproduces a stereoscopic image.

[0262] In addition, as an example, the encoded metadata includes: information related to a three-dimensional sound field including a sound source object that emits sound and an obstacle object, and information related to the positioning position when the sound image of the sound is positioned at a specified position in the three-dimensional sound field (that is, so that it is perceived as a sound arriving from a specified direction), that is, information related to the specified direction. Here, an obstacle object is an object that may affect the sound perceived by the user 99, such as blocking the sound or reflecting the sound, during the period from the sound emitted by the sound source object to the sound reaching the user 99. In addition to stationary objects, obstacle objects may also include moving objects such as animals such as humans or machinery. In addition, when there are multiple sound source objects in the three-dimensional sound field, for any sound source object, other sound source objects may become obstacle objects. In addition, non-sound source objects such as building materials or inanimate objects and sound source objects that emit sound may also become obstacle objects.

[0263] Spatial information constituting metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects existing in the three-dimensional sound field and the shape and position of sound source objects existing in the three-dimensional sound field. The three-dimensional sound field may be either a closed space or an open space, and the metadata includes information on the reflectivity of structures such as floors, walls, or ceilings that can reflect sound in the three-dimensional sound field and the reflectivity of obstacle objects existing in the three-dimensional sound field. Here, the reflectivity is the energy ratio of the reflected sound to the incident sound, and is set for each frequency band of the sound. Of course, the reflectivity may also be set uniformly regardless of the frequency band of the sound. In the case where the three-dimensional sound field is an open space, for example, uniformly set parameters such as attenuation rate, diffracted sound, or initial reflected sound may also be used.

[0264] In the above description, reflectivity is listed as a parameter related to an obstacle object or a sound source object included in metadata, but metadata may include information other than reflectivity. For example, metadata related to both a sound source object and a non-sound source object may include information related to the material of the object. Specifically, metadata may include parameters such as diffusivity, transmittance, or sound absorption.

[0265] As information related to the sound source object, it may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources emitted from an object, or information on the sound source area in the specified object. In the reproduction conditions, for example, it may also be set whether it is a sound that flows continuously or a sound triggered by an event. The sound source area in the object can be set by the relative relationship between the position of the user 99 and the position of the object, or it can be set based on the object. In the case of setting by the relative relationship between the position of the user 99 and the position of the object, the user 99 can perceive that sound X is emitted from the right side of the object and sound Y is emitted from the left side when the user 99 looks at it, based on the side of the object from which the user 99 looks. In the case of setting based on the object, which sound is emitted from which area of ​​the object regardless of the direction the user 99 looks at can be fixed. For example, the user 99 can perceive that when looking at the object from the front, high sound flows from the right side and low sound flows from the left side. In this case, if the user 99 goes around to the back of the object, the user 99 can perceive that the low sound flows from the right side and the high sound flows from the left side when looking from the back.

[0266] Metadata related to the space may include time until initial reflected sound, reverberation time, ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the user 99 can perceive only direct sound.

[0267] In addition, information indicating the position and orientation of the user 99 in the three-dimensional sound field may be included in the bitstream as metadata as an initial setting, or may not be included in the bitstream. In the case where the information indicating the position and orientation of the user 99 is not included in the bitstream, the information indicating the position and orientation of the user 99 is obtained from information outside the bitstream. For example, if it is the position information of the user 99 in the VR space, it can be obtained from the application that provides VR content. If it is the position information of the user 99 used to prompt the sound as AR, the position information obtained by self-position estimation using GPS, camera or LiDAR (Laser Imaging Detection and Ranging) in a portable terminal is used. In addition, the sound signal and metadata can be stored in one bitstream or in multiple bitstreams respectively. Similarly, the sound signal and metadata can be stored in one file or in multiple files respectively.

[0268] In the case where the sound signal and metadata are stored in multiple bitstreams respectively, information indicating other associated bitstreams may be included in one or a part of the bitstreams storing the sound signal and metadata. In addition, information indicating other associated bitstreams may be included in metadata or control information of each bitstream of the multiple bitstreams storing the sound signal and metadata. In the case where the sound signal and metadata are stored in multiple files respectively, information indicating other associated bitstreams or files may be included in one or a part of the files storing the sound signal and metadata. In addition, information indicating other associated bitstreams or files may be included in metadata or control information of each bitstream of the multiple bitstreams storing the sound signal and metadata.

[0269] Here, the associated bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. In addition, information indicating other associated bitstreams may be recorded together in the metadata or control information of one bitstream among multiple bitstreams storing sound signals and metadata, or may be recorded separately in the metadata or control information of two or more bitstreams among multiple bitstreams storing sound signals and metadata. Similarly, information indicating other associated bitstreams or files may be recorded together in the metadata or control information of one file among multiple files storing sound signals and metadata, or may be recorded separately in the metadata or control information of two or more files among multiple files storing sound signals and metadata. In addition, a control file that records information indicating other associated bitstreams or files may be generated separately from multiple files storing sound signals and metadata. In this case, the control file may not store sound signals and metadata.

[0270] Here, the information indicating the other associated bitstreams or files is, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator) or a URI (Uniform Resource Identifier), etc. In this case, the acquisition unit 111 determines or acquires the bitstream or file based on the information indicating the other associated bitstream or file. In addition, the information indicating the other associated bitstream may be included in the metadata or control information of at least a part of the bitstreams among the multiple bitstreams storing the sound signal and metadata, and the information indicating the other associated files may be included in the metadata or control information of at least a part of the files among the multiple files storing the sound signal and metadata. Here, the file containing the information indicating the associated bitstream or file is, for example, a control file such as a manifest file for distributing content.

[0271] Industrial Applicability

[0272] The present disclosure is useful in enabling users to perceive stereoscopic sound reproduction or the like.

[0273] Description of symbols

[0274] 99 users

[0275] 100 Sound reproduction system

[0276] 101 Information processing device

[0277] 102 Communication module

[0278] 103 Detector

[0279] 104 Driver

[0280] 111 Acquisition Department

[0281] 112 Coded sound information input unit

[0282] 113 Decoding Processing Unit

[0283] 114 sensor information input unit

[0284] 121, 121a Processing unit

[0285] 122 Judgment Department

[0286] 123 Storage

[0287] 124 Executive Department

[0288] 125 Computing Department

[0289] 141 Signal output unit

[0290] 200 Stereoscopic image reproduction device

[0291] 600 Stereo Sound Reproduction System

[0292] 601 Information Processing Device

[0293] 602 Sound prompt equipment

[0294] 700, 900 Encoding Device

[0295] 701, 803 Input data

[0296] 702 Encoder

[0297] 703 Encoded Data

[0298] 704, 804 Memory

[0299] 800, 1000 decoding device

[0300] 801 Sound Signal

[0301] 802, 1100, 1200 decoder

[0302] 901 Sending Department

[0303] 902 Send signal

[0304] 1001 Receiving Department

[0305] 1002 Receive signal

[0306] 1101, 1201 Space Information Management Department

[0307] 1102 Sound Data Decoder

[0308] 1103, 1202 Rendering Department

Claims

1. A sound processing method, wherein: include: The step of obtaining a sound signal obtained by collecting the sound emitted from the sound source using a sound receiving device; A step of performing an acoustic processing on the sound signal so as to repeatedly change the relative position between the sound receiving device and the sound source in the time domain; as well as The step of outputting the output audio signal after the audio processing is performed.

2. The sound processing method according to claim 1, wherein: In the step of performing the audio processing, determining whether a change in the sound pressure of the sound signal in the time domain satisfies a prescribed condition related to the change, If it is determined that the predetermined condition is satisfied, the audio processing is executed. When it is determined that the predetermined condition is not satisfied, the audio processing is not executed.

3. The sound processing method according to claim 1, wherein: In the step of performing the audio processing, Using the sound signal to estimate the positional relationship between the sound receiving device and the sound source, determining whether the inferred positional relationship satisfies a predetermined condition related to the positional relationship, If it is determined that the predetermined condition is satisfied, the audio processing is executed. When it is determined that the predetermined condition is not satisfied, the audio processing is not executed.

4. The sound processing method according to claim 1, wherein: The sound signal includes sound collection status information related to the status when the sound is collected, In the step of performing the audio processing, determining whether the sound collection status information included in the sound signal satisfies a predetermined condition related to the sound collection status information, If it is determined that the predetermined condition is satisfied, the audio processing is executed. When it is determined that the predetermined condition is not satisfied, the audio processing is not executed.

5. The sound processing method according to claim 1, wherein: In the step of performing the audio processing, Using the sound signal to estimate the positional relationship between the sound receiving device and the sound source, The audio processing is performed under a processing condition corresponding to the estimated positional relationship.

6. A sound processing method for outputting an output sound signal that causes a sound emitted from a sound source object in a virtual sound space to be perceived as being heard at a listening point in the virtual sound space, wherein: include: A step of obtaining a sound signal including a sound emitted from the sound source object; a step of accepting an instruction to change a relative position between the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed according to the instruction; A step of performing an acoustic processing on the sound signal to change the relative position by the first change amount and to repeatedly change the relative position by a second change amount in the time domain; as well as The step of outputting the output audio signal after the audio processing has been performed.

7. The sound processing method according to claim 6, wherein: The sound source object simulates a user in real space. The sound processing method further includes the step of obtaining a detection result from a sensor disposed in the real space to detect the user. The second change amount is calculated based on the detection result.

8. The sound processing method according to claim 6, wherein: The sound source object simulates a user in real space. The sound processing method further includes the step of obtaining a detection result from a sensor disposed in the real space to detect the user. The second variation is calculated independently of the detection result.

9. The sound processing method according to claim 6, wherein: The second change amount is calculated independently of the first change amount.

10. The sound processing method according to claim 6, wherein: The second change amount is calculated to be a larger value as the first change amount is larger.

11. The sound processing method according to claim 6, wherein: The second change amount is calculated to be a larger value as the first change amount is smaller.

12. The sound processing method according to claim 1 or 6, wherein: It also includes the step of obtaining control information of the sound signal, In the step of performing the audio processing, When the control information indicates that the audio process is to be executed, the audio process is executed.

13. A sound processing device, wherein: have: An acquisition unit that acquires a sound signal obtained by collecting a sound emitted from a sound source using a sound collecting device; a processing unit that performs, on the sound signal, an acoustic processing that causes the relative position of the sound pickup device and the sound source to repeatedly change in the time domain; as well as The output unit outputs an output audio signal on which the audio processing has been performed.

14. A sound processing device for outputting an output sound signal that causes a sound emitted from a sound source object in a virtual sound space to be perceived as being heard at a listening point in the virtual sound space, wherein: have: an acquisition unit that acquires a sound signal including a sound emitted from the sound source object; a receiving unit receiving an instruction to change a relative position between the listening point and the sound source object, the instruction including a first change amount by which the relative position is changed according to the instruction; a processing unit that performs an acoustic processing on the sound signal to change the relative position by the first change amount and to repeatedly change the relative position by a second change amount in the time domain; as well as The output unit outputs the output audio signal after the audio processing has been performed.

15. A program for causing a computer to execute the sound processing method according to claim 1 or 6.

Citation Information

Patent Citations

  • Stereophonic processing apparatus and method

    JP2005295416A