Acoustic signal processing method, computer program, and acoustic signal processing device
By obtaining object information and prescribed timing in the virtual space and outputting aerodynamic sound data caused by wind, the problem of difficulty in bringing a sense of presence to listeners in the prior art is solved, and a more realistic immersive experience is achieved.
Patent Information
- Application Number
- CN202380071659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-10-03
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to bring a sense of presence to listeners, especially in virtual spaces where the arrival time of aerodynamic sounds cannot be effectively controlled.
By obtaining object information, including object changes that cause wind and related prescribed timing, aerodynamic sound data caused by wind are output. After a certain time has elapsed after the specified timing, aerodynamic sound data is output to simulate the arrival time of the wind.
It realizes the sense of presence to the listener in the virtual space, and by accurately controlling the arrival time of aerodynamic sounds, the sense of incongruity is reduced and the immersion experience is enhanced.
Smart Images

Figure CN120113259A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio signal processing method and the like. Background Art
[0002] Patent Document 1 discloses a technique related to a stereophonic calculation method as a sound signal processing method. In this sound signal processing method, the arrival time of sound to a listener (observer) is controlled so as to vary according to the distance between the sound source and the listener and the speed of sound.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Publication No. 2013-201577
[0006] Patent Document 2: International Publication No. 2021 / 180938 Summary of the invention
[0007] Problems to be solved by the invention
[0008] However, in the technology disclosed in Patent Document 1, it is sometimes difficult to provide a sense of presence to the listener.
[0009] Therefore, an object of the present disclosure is to provide an audio signal processing method and the like that can provide a sense of presence to a listener.
[0010] Means used to solve problems
[0011] An acoustic signal processing method according to a technical solution of the present disclosure includes: an acquisition step of acquiring object information indicating changes in an object causing wind and a specified timing related to the changes in the object; and an output step of outputting aerodynamic sound data indicating aerodynamic sound caused by the wind after a specified time based on the changes in the object has passed since the specified timing indicated by the acquired object information.
[0012] Furthermore, a computer program according to one technical aspect of the present disclosure causes a computer to execute the above-mentioned sound signal processing method.
[0013] In addition, an acoustic signal processing device related to a technical solution of the present disclosure comprises: an acquisition unit, which acquires object information, wherein the object information represents a change in an object causing wind and a specified timing related to the change in the object; and an output unit, which outputs aerodynamic sound data representing aerodynamic sound caused by the wind after a specified time based on the change in the object has passed from the specified timing represented by the acquired object information.
[0014] Furthermore, these inclusive or specific technical solutions may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0015] Effects of the Invention
[0016] According to an audio signal processing method according to a technical solution of the present disclosure, a sense of presence can be brought to the listener. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a diagram showing an immersive audio reproduction system as an example of a system to which the audio processing or decoding processing of the present disclosure can be applied.
[0018] Figure 2 This is a functional block diagram showing the configuration of an encoding device as an example of the encoding device of the present disclosure.
[0019] Figure 3 This is a functional block diagram showing the configuration of a decoding device as an example of the decoding device of the present disclosure.
[0020] Figure 4 This is a functional block diagram showing the configuration of an encoding device as another example of the encoding device of the present disclosure.
[0021] Figure 5 This is a functional block diagram showing the configuration of a decoding device as another example of the decoding device of the present disclosure.
[0022] Figure 6 It means as Figure 3 or Figure 5 A functional block diagram of the structure of a decoder of an example of a decoder in FIG.
[0023] Figure 7 It means as Figure 3 or Figure 5 A functional block diagram of the structure of another example of a decoder in the decoder.
[0024] Figure 8 This is a diagram showing an example of the physical structure of the sound signal processing device.
[0025] Fig. 9 This is a diagram showing an example of the physical structure of an encoding device.
[0026] Fig.10 This is a block diagram showing the functional structure of the sound signal processing device according to the embodiment.
[0027] Fig.11This is a flowchart of operation example 1 of the sound signal processing device according to the embodiment.
[0028] Fig.12 This is a diagram showing fans and listeners who are targets in operation example 1.
[0029] Fig.13A It is explained in Fig.11 FIG. 1 is a diagram showing a process of determining a predetermined time in step S40.
[0030] Fig. 13B It is a diagram for explaining a detailed example of output of pneumatic sound data according to the embodiment.
[0031] Fig. 13C This is a diagram for explaining another example of the details of the output of the pneumatic sound data according to the embodiment.
[0032] Fig.14 This is a flowchart of operation example 2 of the sound signal processing device according to the embodiment.
[0033] Fig.15 This is a diagram showing an ambulance and a listener as objects related to Action Example 2.
[0034] Fig.16 This is a schematic diagram for explaining the prescribed timing of Action Example 2.
[0035] Fig.17 This is a flowchart for explaining the details of step S35 in operation example 2.
[0036] Fig.18 This is a flowchart for explaining the details of step S35 of another first example of operation example 2.
[0037] Fig.19 It is used to explain Figure 6 and Figure 7 A diagram showing an example of a functional block diagram and steps in which a rendering unit performs pipeline processing. DETAILED DESCRIPTION
[0038] (Understanding that is the basis of the present disclosure)
[0039] Conventionally, there is known an acoustic signal processing method for controlling the arrival time of sound to a listener in a virtual space.
[0040] Patent document 1 discloses a technique related to a stereophonic sound calculation method as a sound signal processing method. In this sound signal processing method, the arrival time of sound to a listener is controlled so as to vary according to the distance between the sound source and the listener and the speed of sound. More specifically, the arrival time is controlled so as to become longer as the distance increases and to become longer as the speed of sound slows down. As a result, the listener can recognize the distance between the object emitting the sound, i.e., the sound source, and the listener himself.
[0041] The sound after such control is used in applications such as virtual reality (VR (Virtual Reality)) and augmented reality (AR (Augmented Reality)) to reproduce three-dimensional sound in a space (virtual space) where a user (listener) exists. The sound after such control is particularly used in a virtual space such as one that senses the 6DoF (Degrees of Freedom) information of the listener.
[0042] In addition, the sound reaching the listener disclosed in Patent Document 1 is the driving sound of the vehicle (mobile sound source) as the object in VR or AR, and is the sound emitted by the vehicle itself (engine sound, etc.). However, in real space, for example, if a vehicle is driving, it will cause wind. Aerodynamic sound is generated by the wind caused by the vehicle reaching the ears of the listener. The aerodynamic sound is a sound generated, for example, corresponding to the shape of the ears of the listener L when the wind caused by the object (such as a vehicle) reaches the listener. In addition, the object that causes wind is not limited to an object that drives (moves) like the above-mentioned vehicle, but also includes an object that generates wind like a fan.
[0043] However, Patent Document 1 does not disclose how to make the listener listen to the aerodynamic sound. More specifically, Patent Document 1 does not disclose a technique for controlling the arrival time of the aerodynamic sound to the listener when the object causes wind. In the technique disclosed in Patent Document 1, since the listener cannot listen to the aerodynamic sound at an appropriate timing, the listener will feel uncomfortable and it is difficult for the listener to get a sense of presence. Therefore, an acoustic signal processing method that can give the listener a sense of presence is required.
[0044] Therefore, the sound signal processing method of the first technical solution of the present disclosure includes: an acquisition step of acquiring object information, wherein the object information represents changes in an object causing wind and a specified timing related to the changes in the object; and an output step of outputting aerodynamic sound data representing aerodynamic sound caused by the wind after a specified time based on the changes in the object has passed from the specified timing represented by the acquired object information.
[0045] Thus, the aerodynamic sound data can be output at a timing after a predetermined time has passed from a predetermined timing. Therefore, the listener can listen to the aerodynamic sound at an appropriate timing, so the listener is less likely to feel a sense of disharmony and can get a sense of presence. That is, an acoustic signal processing method that can give the listener a sense of presence is realized.
[0046] In addition, for example, in the sound signal processing method related to the second technical solution of the present disclosure, in the sound signal processing method related to the first technical solution, the object information represents: the change of the wind caused by the change of the object; and the specified timing is the timing of the change of the wind; the sound signal processing method includes a determination step, in which the specified time is determined based on the wind represented by the obtained object information.
[0047] Thus, since the aerodynamic sound data can be output at a timing after a predetermined time determined based on the wind has passed from the timing at which the wind has changed, the listener can hear the aerodynamic sound at a more appropriate timing.
[0048] In addition, for example, in the sound signal processing method related to the third technical solution of the present disclosure, in the sound signal processing method related to the second technical solution, the change in the wind represented by the object information represents a change in the wind speed of the wind; and in the determination step, the specified time is determined based on the wind speed.
[0049] Thus, the predetermined time is determined based on the wind speed, so the listener can listen to the aerodynamic sound at a more appropriate timing.
[0050] Furthermore, for example, in the sound signal processing method according to the fourth technical aspect of the present disclosure, in the sound signal processing method according to the third technical aspect, the aerodynamic sound is a sound generated at the changed wind speed.
[0051] This makes it possible to make the aerodynamic sound heard by the listener in the virtual space closer to the aerodynamic sound heard by the listener in the real space.
[0052] In addition, for example, in the sound signal processing method related to the fifth technical solution of the present disclosure, in the sound signal processing method related to the first technical solution, the object information represents the position of the object; the sound signal processing method includes a determination step, in which the specified time is determined based on the distance between the position of the listener of the aerodynamic sound and the position of the object represented by the obtained object information.
[0053] Thus, the predetermined time is determined based on the distance, so the listener can listen to the aerodynamic sound at a more appropriate timing.
[0054] In addition, for example, in the sound signal processing method related to the 6th technical solution of the present disclosure, in the sound signal processing method related to the 3rd or 4th technical solution, the object information represents the position of the object; in the determination step, the specified time is determined based on the wind speed and the distance between the position of the listener of the aerodynamic sound and the position of the object represented by the obtained object information.
[0055] Thus, the predetermined time is determined based on the wind speed and the distance, so the listener can listen to the aerodynamic sound at a more appropriate timing.
[0056] In addition, for example, in the sound signal processing method related to the 7th technical solution of the present disclosure, in the sound signal processing method related to any one of the 1st to 6th technical solutions, the object information indicates that the specified timing is the first timing for outputting the sound data corresponding to the object; and in the output step, the pneumatic sound data is output after the specified time has passed from the first timing indicated by the obtained object information.
[0057] Thus, for example, when a subject produces a sound, the pneumatic sound data can be output at a timing when a predetermined time has elapsed from a first timing when the sound is output, so the listener can hear the pneumatic sound at a more appropriate timing.
[0058] In addition, for example, in the sound signal processing method related to the 8th technical solution of the present disclosure, in the sound signal processing method related to any one of the 1st to 6th technical solutions, the object information represents: the position of the object; and the specified timing is a second timing when the distance between the position of the listener of the pneumatic sound and the position of the object becomes shorter than a specified distance; in the output step, the pneumatic sound data is output after the specified time has passed from the second timing represented by the obtained object information.
[0059] Thus, since the aerodynamic sound data can be output at the timing when the predetermined time has elapsed from the second timing when the distance becomes shorter than the predetermined distance, that is, the second timing when the object approaches the listener, the listener can hear the aerodynamic sound at a more appropriate timing.
[0060] In addition, for example, in the sound signal processing method related to the 9th technical solution of the present disclosure, in the sound signal processing method related to any one of the 1st to 6th technical solutions, the object information indicates: the change in the wind caused by the change in the object is a change in the direction of the wind; and the specified timing is the third timing at which the change in the direction of the wind occurs; in the output step, the aerodynamic sound data is output after the specified time has passed from the third timing represented by the obtained object information.
[0061] Thus, since the aerodynamic sound data can be output at the timing when a predetermined time has elapsed from the third timing when the wind direction has changed, the listener can listen to the aerodynamic sound at a more appropriate timing.
[0062] In addition, for example, in the sound signal processing method related to the 10th technical solution of the present disclosure, in the sound signal processing method related to the 6th technical solution, the object is an object that generates sound and wind represented by sound data corresponding to the object; and the aerodynamic sound is aerodynamic sound generated by the wind generated by the object reaching the listener.
[0063] Thereby, a fan or the like that generates sound and wind can be targeted, and aerodynamic sound caused by the wind blown out from the target can be realized.
[0064] In addition, for example, in the sound signal processing method related to the 11th technical solution of the present disclosure, in the sound signal processing method related to the 10th technical solution, when the distance is set to D, the distance from the position of the object with the wind speed So is set to U, and the specified time is set to t, t satisfies the following formula.
[0065] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0066] Thus, in the determination step, the time from the predetermined timing to the wind generated by the object reaching the listener can be determined as the predetermined time. Therefore, the aerodynamic sound data can be output at a timing after the predetermined time has passed from the predetermined timing, so the listener can hear the aerodynamic sound at a more appropriate timing.
[0067] In addition, for example, in the sound signal processing method related to the 12th technical solution of the present disclosure, in the sound signal processing method related to the 6th technical solution, the object is an object that generates the wind by moving the position of the object; and the aerodynamic sound is the aerodynamic sound generated by the wind generated by the movement reaching the listener.
[0068] Thereby, a vehicle or the like that generates wind by movement can be targeted, and aerodynamic sound caused by the wind generated by the movement can be realized.
[0069] Furthermore, for example, in the sound signal processing method according to the 13th technical solution of the present disclosure, in the sound signal processing method according to the 12th technical solution, the predetermined timing indicated by the object information is the timing at which the change in the distance with the passage of time changes from negative to positive.
[0070] Thus, since the pneumatic sound data can be output at a timing when a predetermined time has elapsed from a timing when the distance between the position of the listener and the position of the object becomes the shortest, the listener can hear the pneumatic sound at a more appropriate timing.
[0071] In addition, for example, in the sound signal processing method related to the 14th technical solution of the present disclosure, in the sound signal processing method related to the 12th or 13th technical solution, when the distance is set to D, the distance from the position of the object where the wind speed of the wind generated by the movement is So is set to U, and the specified time is set to t, t satisfies the following formula.
[0072] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0073] Thus, in the determination step, the time from the predetermined timing to the wind generated by the object reaching the listener can be determined as the predetermined time. Therefore, the aerodynamic sound data can be output at a timing after the predetermined time has passed from the predetermined timing, so the listener can hear the aerodynamic sound at a more appropriate timing.
[0074] Furthermore, for example, a computer program according to a fifteenth aspect of the present disclosure is a computer program for causing a computer to execute the sound signal processing method according to any one of the first to fourteenth aspects.
[0075] Thereby, the computer can execute the acoustic signal processing method according to the computer program.
[0076] In addition, for example, the sound signal processing device related to the 16th technical solution of the present disclosure includes: an acquisition unit, which acquires object information, wherein the object information represents a change in an object causing wind and a specified timing related to the change in the object; and an output unit, which outputs aerodynamic sound data representing aerodynamic sound caused by the wind after a specified time based on the change in the object has passed from the specified timing represented by the acquired object information.
[0077] Thus, the aerodynamic sound data can be output at a timing after a predetermined time has passed from a predetermined timing. Therefore, the listener can listen to the aerodynamic sound at an appropriate timing, so the listener is less likely to feel a sense of disharmony and can get a sense of presence. That is, an acoustic signal processing device that can give the listener a sense of presence is realized.
[0078] Furthermore, these inclusive or specific technical solutions may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0079] Hereinafter, the embodiments will be described in detail with reference to the drawings.
[0080] In addition, the embodiments described below are all inclusive or specific examples. The numerical values, shapes, materials, components, configuration positions and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the present disclosure.
[0081] In addition, in the following description, elements are sometimes given ordinals such as 1st and 2nd. These ordinals are given to elements for the purpose of identifying the elements and do not necessarily correspond to a meaningful order. These ordinals may be replaced as appropriate, may be newly given, or may be removed.
[0082] In addition, each figure is a schematic diagram and does not necessarily illustrate the exact figure. Therefore, the scales and the like in each figure are not necessarily the same. In each figure, the same reference numerals are given to substantially the same structure, and repeated descriptions are omitted or simplified.
[0083] In this specification, terms such as vertical and other terms indicating the relationship between elements and numerical ranges do not indicate only strict expressions, but mean substantially equivalent ranges, for example, expressions including differences of about several percentage points.
[0084] (Implementation Method 1)
[0085] [Examples of Devices to Which the Sound Processing Technology or Encoding / Decoding Technology of the Present Disclosure Can Be Applied]
[0086] <Stereo sound reproduction system>
[0087] Figure 1 1 is a diagram showing an immersive audio reproduction system A0000 as an example of a system to which the audio processing or decoding processing of the present disclosure can be applied. The immersive audio reproduction system A0000 includes an audio signal processing device A0001 and an audio presentation device A0002.
[0088] The sound signal processing device A0001 performs sound processing on the sound signal emitted by the virtual sound source, and generates a sound signal after sound processing to prompt the listener (i.e., the receiver). The sound signal is not limited to the human voice, as long as it is an audible sound. For example, sound processing is a signal processing performed on the sound signal in order to reproduce one or more sound-related effects received from the time when the sound is emitted to the time when the listener hears it. The sound signal processing device A0001 performs sound processing based on information that records the causes of the above-mentioned sound-related effects. Spatial information includes, for example, information indicating the position of the sound source, the listener, and the surrounding objects, information indicating the shape of the space, parameters related to the propagation of sound, etc. The sound signal processing device A0001 is, for example, a PC (Personal Computer), a smart phone, a tablet computer, or a game console.
[0089] The sound-processed signal is prompted to the listener (user) from the sound prompting device A0002. The sound prompting device A0002 is connected to the sound signal processing device A0001 via wireless or wired communication. The sound signal generated by the sound signal processing device A0001 after sound processing is transmitted to the sound prompting device A0002 via wireless or wired communication. In the case where the sound prompting device A0002 is composed of multiple devices such as a device for the right ear and a device for the left ear, the multiple devices synchronously prompt the sound by communicating between the multiple devices or each of the multiple devices communicates with the sound signal processing device A0001. The sound prompting device A0002 is, for example, headphones, earplugs, a head-mounted display worn on the listener's head, or a surround speaker composed of multiple fixed speakers.
[0090] Furthermore, the stereophonic sound reproduction system A0000 may be used in combination with an image presentation device or a stereoscopic image presentation device that visually provides an ER (Extended Reality) experience including VR or AR.
[0091] in addition, Figure 1 The system configuration example in which the sound signal processing device A0001 and the sound prompting device A0002 are different devices is shown, but the stereo sound reproduction system A0000 to which the sound signal processing method or decoding method disclosed in the present invention can be applied is not limited to Figure 1For example, the sound signal processing device A0001 may be included in the sound prompting device A0002, and the sound prompting device A0002 may perform both sound processing and sound prompting. In addition, the sound signal processing device A0001 and the sound prompting device A0002 may share the sound processing described in the present disclosure, or a server connected to the sound signal processing device A0001 or the sound prompting device A0002 via a network may implement part or all of the sound processing described in the present disclosure.
[0092] In addition, although the above description refers to the sound signal processing device A0001, when the sound signal processing device A0001 performs sound processing by decoding a bit stream generated by encoding data of at least a part of spatial information used in a sound signal or sound processing, the sound signal processing device A0001 may also be referred to as a decoding device.
[0093] <Example of Encoding Device>
[0094] Figure 2 It is a functional block diagram showing the structure of an encoding device A0100 which is an example of the encoding device of the present disclosure.
[0095] Input data A0101 is data to be encoded including spatial information and / or a sound signal input to encoder A0102. The details of spatial information will be described later.
[0096] The encoder A0102 encodes the input data A0101 to generate encoded data A0103. The encoded data A0103 is, for example, a bit stream generated by encoding processing.
[0097] The memory A0104 stores the encoded data A0103. The memory A0104 may be, for example, a hard disk or an SSD (Solid-State Drive), or may be other memory.
[0098] In addition, in the above description, a bit stream generated by encoding processing is cited as an example of the encoded data A0103 stored in the memory A0104, but it can also be data other than a bit stream. For example, the encoding device A0100 can also store the converted data generated by converting the bit stream into a specified data format in the memory A0104. The converted data can also be, for example, a file or a multiplexed stream storing one or more bit streams. Here, the file is a file having a file format such as ISOBMFF (ISO Base Media File Format). In addition, the encoded data A0103 can also be in the form of multiple packets generated by dividing the above-mentioned bit stream or file. In the case of converting the bit stream generated by the encoder A0102 into data different from the bit stream, the encoding device A0100 can also have a conversion unit not shown in the figure, and the conversion process can also be performed by a CPU (Central Processing Unit).
[0099] <Example of Decoding Device>
[0100] Figure 3 It is a functional block diagram showing the structure of a decoding device A0110 which is an example of a decoding device according to the present disclosure.
[0101] The memory A0114 stores, for example, the same data as the coded data A0103 generated by the coding device A0100. The memory A0114 reads the stored data and inputs it as input data A0113 of the decoder A0112. The input data A0113 is, for example, a bit stream to be decoded. The memory A0114 may be, for example, a hard disk or an SSD, or other memory.
[0102] In addition, the decoding device A0110 may use the transformed data generated by transforming the read data instead of the data stored in the memory A0114 as input data A0113. The data before the transformation may be, for example, multiplexed data storing one or more bit streams. Here, the multiplexed data may also be a file having a file format such as ISOBMFF. In addition, the data before the transformation may also be in the form of a plurality of packets generated by dividing the above-mentioned bit stream or file. In the case of transforming data different from the bit stream read from the memory A0114 into a bit stream, the decoding device A0110 may also include a transformation unit not shown in the figure, or the transformation process may be performed by the CPU.
[0103] Decoder A0112 decodes input data A0113 to generate a sound signal A0111 for prompting the listener.
[0104] <Another Example of Encoding Device>
[0105] Figure 4 1 is a functional block diagram showing the structure of an encoding device A0120 which is another example of the encoding device of the present disclosure. Figure 4 For Figure 2 The same function as the structure of Figure 2 The same reference numerals are used for the components, and description of these components is omitted.
[0106] The coding device A0100 stores the coded data A0103 in the memory A0104 , but the coding device A0120 is different from the coding device A0100 in that it includes a transmission unit A0121 that transmits the coded data A0103 to the outside.
[0107] The transmission unit A0121 transmits a transmission signal A0122 to another device or server based on the coded data A0103 or data in another data format generated by transforming the coded data A0103. The data used to generate the transmission signal A0122 is, for example, the bit stream, multiplexed data, file, or packet described in the coding device A0100.
[0108] <Another Example of Decoding Device>
[0109] Figure 5 1 is a functional block diagram showing the structure of a decoding device A0130 which is another example of the decoding device of the present disclosure. Figure 5 For Figure 3 The same function as the structure of Figure 3 The same reference numerals are used for the components, and description of these components is omitted.
[0110] The decoding device A0110 reads the input data A0113 from the memory A0114 , whereas the decoding device A0130 is different from the decoding device A0110 in that it includes a receiving unit A0131 that receives the input data A0113 from the outside.
[0111] The receiving unit A0131 receives the received signal A0132 to obtain received data, and outputs input data A0113 to be input to the decoder A0112. The received data may be the same as the input data A0113 to be input to the decoder A0112, or may be data in a data format different from that of the input data A0113. When the received data is data in a data format different from that of the input data A0113, the receiving unit A0131 may convert the received data into the input data A0113, or a conversion unit or CPU (not shown) included in the decoding device A0130 may convert the received data into the input data A0113. The received data is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device A0120.
[0112] <Decoder Functional Description>
[0113] Figure 6 It means as Figure 3 or Figure 5 A functional block diagram of the structure of a decoder A0200 which is an example of the decoder A0112 in FIG.
[0114] The input data A0113 is a coded bit stream, and includes coded audio data which is a coded audio signal and metadata used in audio processing.
[0115] The spatial information management unit A0201 obtains metadata included in the input data A0113 and parses the metadata. The metadata includes information describing elements that act on the sound and are configured in the sound space. The spatial information management unit A0201 manages the spatial information required for sound processing obtained by parsing the metadata, and provides the spatial information to the rendering unit A0203. In addition, in the present disclosure, the information used in the sound processing is referred to as spatial information, but it may be called something else. The information used in the sound processing may be called, for example, sound spatial information, or may be called scene information. In addition, in the case where the information used in the sound processing changes over time, the spatial information input to the rendering unit A0203 may also be called a spatial state, a sound spatial state, a scene state, or the like.
[0116] In addition, the spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, the spatial information may be managed as scenes of different sound spaces for each room, or the spatial information may be managed as different scenes depending on the occasion of representation despite being the same space. In the management of spatial information, an identifier may be assigned to identify each piece of spatial information. The spatial information data may be included in a bit stream as a form of input data, or the bit stream may include an identifier of the spatial information and the spatial information data may be obtained from outside the bit stream. In the case where only the identifier of the spatial information is included in the bit stream, the identifier of the spatial information may be used during rendering to obtain the spatial information data stored in the memory of the sound signal processing device A0001 or in an external server as input data.
[0117] In addition, the information managed by the spatial information management unit A0201 is not limited to the information included in the bitstream. For example, the input data A0113 may include data representing spatial characteristics or structures obtained from a software application or server that provides VR or AR as data not included in the bitstream. In addition, for example, the input data A0113 may include data representing characteristics or positions of listeners or objects as data not included in the bitstream. In addition, the input data A0113 may include information obtained by a sensor provided by a terminal including a decoding device, or information representing the position of the terminal inferred based on the information obtained by the sensor, as information representing the position of the listener. That is, the spatial information management unit A0201 may also communicate with an external system or server to obtain spatial information and the position of the listener. In addition, the spatial information management unit A0201 may also obtain clock synchronization information from an external system and perform processing synchronized with the clock of the rendering unit A0203. In addition, the space in the above description may be a virtually formed space, i.e., a VR space, or a real space (i.e., a real space) or a virtual space corresponding to the real space, i.e., AR or MR (Mixed Reality). In addition, the virtual space may also be referred to as a sound field or a sound space. In addition, the information indicating the position in the above description may be information such as coordinate values indicating the position in the space, information indicating the relative position relative to a prescribed reference position, or information indicating the movement or acceleration of the position in the space.
[0118] The sound data decoder A0202 decodes the encoded sound data contained in the input data A0113 to obtain a sound signal.
[0119] The coded audio data acquired by the stereo sound reproduction system A0000 is, for example, a bit stream coded in a prescribed format such as MPEG-H 3D Audio (ISO / IEC 23008-3). MPEG-H 3D Audio is merely an example of a coding method that can be used when generating coded audio data included in a bit stream, and may include a bit stream and coded audio data coded in other coding methods. For example, the coding method used may be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), Vorbis, or a reversible codec such as ALAC (Apple Lossless Audio Codec), FLAC (Free Lossless Audio Codec), or any coding method other than the above may be used. For example, PCM (pulse code modulation) data may be one type of coded audio data. In this case, the decoding process may be a process of converting an N-bit binary number into a number format (eg, floating point format) that can be processed by the rendering unit A0203, for example, when the number of quantization bits of the PCM data is N.
[0120] The rendering unit A0203 takes the sound signal and spatial information as input, performs audio processing on the sound signal using the spatial information, and outputs the sound signal A0111 after the audio processing.
[0121] Before starting rendering, the spatial information management unit A0201 reads metadata of the input signal, detects rendering items such as objects or sounds specified by the spatial information, and sends them to the rendering unit A0203. After the rendering starts, the spatial information management unit A0201 grasps the changes in the spatial information and the listener's position over time, updates the spatial information and manages it. In addition, the spatial information management unit A0201 sends the updated spatial information to the rendering unit A0203. The rendering unit A0203 generates a sound signal with sound processing added based on the sound signal included in the input data A0113 and the spatial information received from the spatial information management unit A0201, and outputs it.
[0122] The spatial information update process and the sound signal output process with added sound processing may be executed by the same thread, or the spatial information management unit A0201 and the rendering unit A0203 may be assigned to separate threads. When the spatial information update process and the sound signal output process with added sound processing are processed by different threads, the thread startup frequency may be set separately, or the processes may be executed in parallel.
[0123] Since the spatial information management unit A0201 and the rendering unit A0203 execute processing in different independent threads, the computing resources can be preferentially allocated to the rendering unit A0203. Therefore, in the case of sound processing in which no delay can be allowed, for example, even in the case of a 1-sample (0.02msec) delay, a small noise can be generated. The sound processing can be safely implemented. At this time, the allocation of computing resources to the spatial information management unit A0201 is restricted. However, the update of spatial information is a lower-frequency process than the output processing of the sound signal (for example, a process such as updating the orientation of the listener's face). Therefore, unlike the output processing of the sound signal, it does not require an instantaneous response, so even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality brought to the listener.
[0124] The updating of spatial information can be performed periodically at a predetermined time or period, or when a predetermined condition is met. In addition, the updating of spatial information can be performed manually by the listener or the manager of the sound space, or it can be performed in response to changes in the external system. For example, when the listener operates the controller and the standing position of his or her avatar is instantly tilted, or the time is instantly moved forward or backward, or when the manager of the virtual space suddenly changes the scene, the thread configured with the spatial information management unit A0201 can be started as a one-shot interrupt processing in addition to the regular startup.
[0125] The information update thread that performs the update processing of spatial information is responsible for, for example, updating the position or orientation of the listener's avatar configured in the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of objects moving in the virtual space, etc., which are provided in a processing thread that is started at a relatively low frequency of about tens of Hz. Such a processing thread with a low occurrence frequency can also be used to perform processing that reflects the properties of direct sound. This is because the frequency of changes in the properties of direct sound is lower than the frequency of occurrence of audio processing frames used for audio output. On the contrary, doing so can relatively reduce the computational load of the processing, and if the information is updated at an unnecessarily fast frequency, there is a risk of generating pulsed noise, so this risk can also be avoided.
[0126] Figure 7 It means as Figure 3 or Figure 5 A functional block diagram of the structure of a decoder A0210 which is another example of the decoder A0112 in FIG.
[0127] Figure 7 The decoder A0210 shown is different from the decoder A0210 in that the input data A0113 does not contain coded sound data but contains uncoded sound signals. Figure 6 The decoder A0200 shown is different. Input data A0113 includes a bit stream including metadata and an audio signal.
[0128] Space Information Management Department A0211 and Figure 6 The spatial information management unit A0201 is the same as that of the spatial information management unit A0201, so the description is omitted.
[0129] Rendering Department A0213 and Figure 6 The rendering part A0203 is the same, so the description is omitted.
[0130] In addition, in the above description Figure 7 The structure of is called decoder A0210, but it can also be called an audio processing unit that performs audio processing. In addition, the device including the audio processing unit can also be called an audio processing device instead of a decoding device. In addition, the audio signal processing device A0001 can also be called an audio processing device.
[0131] <Physical Structure of the Acoustic Signal Processing Device>
[0132] Figure 8 is a diagram showing an example of the physical structure of an audio signal processing device. Figure 8 The audio signal processing device may also be a decoding device. In addition, a part of the structure described here may also be equipped in the sound prompting device A0002. In addition, Figure 8 The illustrated sound signal processing device is an example of the sound signal processing device A0001 described above.
[0133] Figure 8 The sound signal processing device includes a processor, a memory, a communication IF, a sensor, and a speaker.
[0134] The processor is, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit), and the CPU, DSP or GPU may execute a program stored in a memory to implement the audio processing or decoding processing of the present invention. In addition, the processor may also be a dedicated circuit for performing signal processing of sound signals including the audio processing of the present invention.
[0135] The memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). In addition, the memory may also include an internal memory embedded in a CPU or a GPU.
[0136] The communication IF (Interface) is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGI (registered trademark). Figure 8 The acoustic signal processing device shown has a function of communicating with other communication devices via a communication IF, and acquires a bit stream to be decoded, and stores the acquired bit stream in, for example, a memory.
[0137] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is cited as an example of the communication method, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio) or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.
[0138] The sensor performs sensing for estimating the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, speed, angular velocity, or acceleration of a part or the whole of the body such as the head of the listener, and generates position information indicating the position and / or orientation of the listener. In addition, the position information may also be information indicating the position and / or orientation of the listener in real space, or information indicating the displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a specified point in time. In addition, the position information may also be information indicating the relative position and / or orientation to the stereo sound reproduction system A0000 or an external device having a sensor.
[0139] The sensor may be, for example, a camera or other imaging device or a distance measuring device such as LiDAR (Light Detection And Ranging), and may capture the movement of the listener's head and detect the movement of the listener's head by processing the captured image. In addition, as a sensor, for example, a device that uses wireless position estimation using any frequency band such as millimeter waves may be used.
[0140] in addition, Figure 8 The acoustic signal processing device shown in FIG. 1 may also obtain position information from an external device having a sensor via a communication IF. In this case, the acoustic signal processing device may not include a sensor. Here, the external device is, for example, Figure 1 The sound presentation device A0002 described in , or a stereoscopic image reproduction device worn on the head of the listener, etc. In this case, the sensor is composed of a combination of various sensors such as a gyro sensor and an acceleration sensor.
[0141] For example, as the speed of movement of the listener's head, the sensor can detect the angular velocity of rotation about at least one of the three axes orthogonal to each other in the sound space, or the acceleration of displacement in the direction of displacement about at least one of the three axes.
[0142] For example, as the amount of movement of the listener's head, the sensor can detect the amount of rotation with at least one of the three axes orthogonal to each other in the sound space as the rotation axis, or the amount of displacement with at least one of the three axes as the displacement direction. Specifically, the sensor detects 6DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is composed of a combination of various sensors for detecting movement, such as a gyro sensor and an acceleration sensor.
[0143] In addition, the sensor can be implemented as long as it can detect the position of the listener, and can be implemented by a camera or a GPS (Global Positioning System) receiver, etc. It is also possible to use position information obtained by self-position estimation using LiDAR (Laser Imaging Detection and Ranging) etc. For example, when the sound signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.
[0144] In addition, sensors can also include detection Figure 8 A temperature sensor such as a thermocouple for detecting the temperature of the sound signal processing device shown, and a sensor for detecting the remaining amount of a battery included in or connected to the sound signal processing device, etc.
[0145] The speaker has a driving mechanism such as a diaphragm, a magnet or a voice coil, and an amplifier, and presents the sound signal after acoustic processing to the listener as sound. The speaker operates the driving mechanism according to the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), and the driving mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the sound signal generates sound waves, which propagate in the air and are transmitted to the listener's ears, and the listener perceives the sound.
[0146] In addition, here are Figure 8 The example of the case where the audio signal processing device shown in the figure has a speaker and outputs the audio signal after the audio processing through the speaker is described, but the audio signal prompting mechanism is not limited to the above-mentioned structure. For example, the audio signal after the audio processing can also be output to the external audio prompting device A0002 connected by the communication module. The communication performed by the communication module can be either wired or wireless. In addition, as another example, it can also be Figure 8 The audio signal processing device shown has a terminal for outputting an analog signal of sound, and a cable of an earplug or the like is connected to the terminal to prompt a sound signal from the earplug or the like. In the above case, the earphone, earplug, head-mounted display, neck speaker, wearable speaker, or surround speaker composed of a plurality of fixed speakers worn on the head or a part of the body of the listener as the sound prompting device A0002 reproduces the sound signal.
[0147] <Physical Structure of Encoding Device>
[0148] Fig. 9 is a diagram showing an example of the physical structure of an encoding device. Fig. 9 The illustrated encoding device is an example of the encoding devices A0100 and A0120 described above.
[0149] Fig. 9 The encoding device includes a processor, a memory and a communication IF.
[0150] The processor is, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor), and the encoding process of the present disclosure may be implemented by executing a program stored in a memory by the CPU or DSP. In addition, the processor may also be a dedicated circuit for performing signal processing on a sound signal including the encoding process of the present disclosure.
[0151] The memory is composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). In addition, the memory may also include an internal memory embedded in a CPU or a GPU.
[0152] The communication IF (Interface) is a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGI (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF, and transmits an encoded bit stream.
[0153] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is cited as the communication method, but it can also correspond to communication methods such as LTE (LongTerm Evolution), NR (New Radio) or Wi-Fi (registered trademark). In addition, the communication IF may not be a wireless communication method as described above, but a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), etc.
[0154] [constitute]
[0155] Next, the configuration of the sound signal processing device 100 according to the embodiment will be described. Fig.10 : is a block diagram showing the functional configuration of the sound signal processing device 100 according to the present embodiment.
[0156] The sound signal processing device 100 according to the present embodiment is a device for outputting aerodynamic sound data representing aerodynamic sound caused by wind caused by an object in a virtual space (sound reproduction space). As an example, the sound signal processing device 100 according to the present embodiment is a device applied to various applications in a virtual space such as virtual reality or augmented reality (VR or AR).
[0157] The object in the virtual space is included in the content (here, as an example, a video) displayed on the display unit 300 that displays the content executed in the virtual space. The object is not particularly limited as long as it is an object that causes wind.
[0158] An object is, for example, a moving body that generates wind by moving the position of the object. Moving bodies include, for example, objects representing animals, plants, artificial objects, or natural objects. Examples of objects representing artificial objects include vehicles, bicycles, and airplanes. In addition, examples of objects representing artificial objects include sports equipment such as baseball bats and tennis rackets; and furniture such as tables, chairs, and clocks. In addition, as an example, an object is at least one of an object that can move within the content and an object that is moved.
[0159] In addition, for example, the object may be an object capable of blowing air, such as a fan, a circulator, a fan, an air conditioner, and the like.
[0160] The following describes aerodynamic sound according to the present embodiment. Aerodynamic sound is sound produced by wind caused by an object in a virtual space and reaching the ears of a listener.
[0161] When the object is an object capable of blowing air such as a fan, aerodynamic sound is generated when the wind generated by the object reaches the listener. More specifically, aerodynamic sound is sound generated when the wind blown from the fan reaches the listener, for example, in accordance with the shape of the listener's ears.
[0162] When the object is a moving body (such as a vehicle), aerodynamic sound is aerodynamic sound generated by wind generated by the movement of the object's position reaching the listener, and more specifically, it is sound generated by the wind reaching the listener and corresponding to the shape of the listener's ear, for example.
[0163] In addition, the object may also be an object that causes wind and thus generates sound. The sound generated by the object is a sound represented by sound data corresponding to the object (hereinafter sometimes referred to as object sound data). For example, in the case where the object is a fan, the sound generated by the object is the motor sound generated by the motor of the fan. In addition, in the case where the object is an ambulance, for example, the sound generated by the object is the whistle sound emitted from the ambulance.
[0164] In addition, in the present embodiment, the object is a fan as an example of an object that can send air.
[0165] The acoustic signal processing device 100 outputs aerodynamic sound data indicating aerodynamic sound in the virtual space to the earphone 200 .
[0166] Next, the earphone 200 will be described.
[0167] The earphone 200 is a device for reproducing pneumatic sound, and is a sound output device for presenting the pneumatic sound to the listener. More specifically, the earphone 200 reproduces the pneumatic sound based on the pneumatic sound data output by the sound signal processing device 100. Thus, the listener can listen to the pneumatic sound. In addition, other output channels such as speakers may be used in addition to the earphone 200.
[0168] like Fig.10 As shown, the headset 200 includes a head sensor unit 201 and an output unit 202 .
[0169] The head sensor unit 201 senses the position of the listener in the virtual space determined by the horizontal coordinates and the vertical height, and outputs second position information indicating the position of the listener of the aerodynamic sound in the virtual space to the sound signal processing device 100 .
[0170] The head sensor unit 201 can sense 6DoF information of the listener's head. For example, the head sensor unit 201 can be an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetic sensor, or a combination thereof.
[0171] The output unit 202 is a device that reproduces the sound reaching the listener in the sound reproduction space. More specifically, the output unit 202 reproduces the aerodynamic sound based on the aerodynamic sound data indicating the aerodynamic sound output from the acoustic signal processing device 100.
[0172] Furthermore, when the object is a fan, sound data representing a motor sound is output from the sound signal processing device 100, and the output unit 202 reproduces the motor sound based on the output sound data. Similarly, when the object is an ambulance, sound data representing a whistle sound is output from the sound signal processing device 100, and the output unit 202 reproduces the whistle sound based on the output sound data.
[0173] Next, the display unit 300 will be described.
[0174] The display unit 300 is a display device that displays content (images) including objects in a virtual space. The processing for the display unit 300 to display content will be described later. The display unit 300 is implemented by a display panel such as a liquid crystal panel or an organic EL (ElectroLuminescence) panel.
[0175] Furthermore, Fig.10 In the present embodiment, the acoustic signal processing device 100 outputs the pneumatic sound data to the earphone 200 after a predetermined time has passed since a predetermined timing.
[0176] like Fig.10 As shown, the sound signal processing device 100 includes an acquisition unit 110 , a determination unit 120 , an output unit 130 , and a storage unit 140 .
[0177] The acquisition unit 110 acquires object information. The object information is information indicating a change in an object causing wind, a predetermined timing related to the change in the object, a change in the wind caused by the change in the object, and a position of the object. In addition, the object information is hereinafter referred to as information including first change information indicating a change in an object causing wind, timing information indicating a predetermined timing related to the change in the object, second change information indicating a change in the wind caused by the change in the object, and first position information indicating the position of the object.
[0178] When the object is an object that generates sound, the object information includes sound data (object sound data) indicating the sound. In addition, the object information may include shape information indicating the shape of the object.
[0179] The acquisition unit 110 acquires the second position information. As described above, the second position information is information indicating the position of the listener in the virtual space. The acquisition unit 110 acquires aerodynamic sound data indicating aerodynamic sound. The aerodynamic sound data is stored in the storage unit 140, and the acquisition unit 110 acquires the aerodynamic sound data stored in the storage unit 140.
[0180] The acquisition unit 110 may acquire the object information, the second position information, and the pneumatic sound data from, for example, the input signal, or may acquire the object information, the second position information, and the pneumatic sound data from other sources. The input signal will be described below. In addition, the object sound data and the pneumatic sound data may be recorded together as sound data below.
[0181] The input signal is composed of, for example, spatial information, sensor information, and sound data (sound signal). In addition, the above information and sound data may be included in one input signal, or in a plurality of different signals. The input signal may also include a bit stream composed of sound data and metadata (control information), in which case the metadata may also include spatial information and information for identifying sound data.
[0182] The first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and aerodynamic sound data described above may also be included in the input signal. More specifically, the first change information, timing information, second change information, first position information, and shape information may also be included in the spatial information, and the second position information may also be generated based on information obtained from sensor information. The sensor information may also be obtained from the head sensor unit 201, or may be obtained from other external devices.
[0183] Spatial information is information about the sound space (three-dimensional sound field) formed by the stereo sound reproduction system A0000, and is composed of information about objects contained in the sound space and information about listeners. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound objects that do not emit sound. Non-sound objects function as obstacle objects that reflect the sound emitted by sound source objects, but there are also cases where they function as obstacle objects that reflect the sound emitted by other sound source objects as sound source objects. Obstacle objects can also be called reflection objects.
[0184] The information given to both the sound source object and the non-sound emitting object includes position information, shape information, and a volume attenuation rate when the object reflects sound.
[0185] The position information is represented by the coordinate values of three axes, such as the X-axis, Y-axis, and Z-axis in Euclidean space, but it does not necessarily have to be three-dimensional information. The position information may also be two-dimensional information represented by the coordinate values of two axes, such as the X-axis and the Y-axis. The position information of the object is determined by the representative position of the shape represented by the grid or voxel.
[0186] The shape information may also include information about the material of the surface.
[0187] The attenuation rate can be expressed by a real number below 1 or above 0, or by a negative decibel value. Since the volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value, but for example, in order to express the horror of an unreal space, an attenuation rate of 1 or above, that is, a positive decibel value, can be deliberately set. In addition, the attenuation rate can be set to a different value for each frequency band constituting a plurality of frequency bands, or a value can be set independently for each frequency band. In addition, when the attenuation rate is set for each type of material on the surface of the object, the corresponding attenuation rate value can also be used based on information related to the material of the surface.
[0188] In addition, the information assigned to the sound source object and the non-sound emitting object may also include information indicating whether the object is a living thing, or information indicating whether the object is a moving object, etc. In the case where the object is a moving object, the position information may also move over time, and the changed position information or the amount of change is transmitted to the rendering units A0203 and A0213.
[0189] The information related to the sound source object includes the object sound data and the information required to radiate the object sound data into the sound space, in addition to the information assigned to the sound source object and the non-sound-generating object. The object sound data is data that represents the sound perceived by the listener, such as information related to the frequency and strength of the sound. The object sound data is typically a PCM signal, but it can also be data compressed using an encoding method such as MP3. In this case, at least the signal needs to be transmitted to the generator (in the Fig.19 Since the signal is decoded before the generation unit 907 described later in the figure, a decoding unit not shown in the figure may be included in the rendering unit A0203 and A0213. Alternatively, the signal may be decoded by the sound data decoder A0202.
[0190] At least one target sound data may be set for one sound source object, or a plurality of target sound data may be set. In addition, identification information for identifying each target sound data may be given, and the identification information of the target sound data may be stored as metadata as information related to the sound source object.
[0191] The information required for radiating the object sound data into the sound space may include, for example, information on a reference volume serving as a reference when reproducing the object sound data, information related to the position of the sound source object, information related to the orientation of the sound source object, and information related to the directionality of the sound emitted by the sound source object.
[0192] The reference volume information is, for example, the effective value of the amplitude value of the object sound data at the sound source position when the object sound data is radiated into the sound space, and may also be expressed as a decibel (db) value in a floating point. For example, when the reference volume is 0db, the reference volume information may also indicate that the sound is radiated to the sound space from the position indicated by the information related to the above position without increasing or decreasing the volume of the signal level represented by the object sound data. In the case of -6db, the reference volume information may also indicate that the sound is radiated to the sound space from the position indicated by the information related to the above position with the volume of the signal level represented by the object sound data reduced to approximately half. The reference volume information may also be assigned to one object sound data or to a plurality of object sound data at once.
[0193] The volume information included in the information required to radiate the object sound data into the sound space may include, for example, information indicating the time series change of the volume of the sound source. For example, when the sound space is a virtual conference room and the sound source is a speaker, the volume changes intermittently in a short period of time. If it is expressed more simply, it can be said that the sound part and the silent part are produced alternately. In addition, when the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain length of time. In addition, when the sound space is a battlefield and the sound source is an explosive, the volume of the explosion only increases for a moment, and then it is silent and continues. In this way, the information on the volume of the sound source not only includes information on the size of the sound, but also includes information on the change in the size of the sound, and such information can also be set as information indicating the nature of the object sound data.
[0194] Here, the information on the change in the volume of the sound may also be data representing the frequency characteristics in a time series. The information on the change in the volume of the sound may also be data representing the duration of the sound interval. The information on the change in the volume of the sound may also be data representing the duration of the sound interval and the time length of the silent interval. The information on the change in the volume of the sound may also be data listing the duration of the amplitude of the sound signal as constant (regarded as approximately constant) and the amplitude value of the signal during the duration in a plurality of time series. The information on the change in the volume of the sound may also be data listing the duration of the frequency characteristics of the sound signal as constant. The information on the change in the volume of the sound may also be data listing the duration of the frequency characteristics of the sound signal as constant and the frequency characteristics during the duration in a plurality of time series. The information on the change in the volume of the sound may also be data representing the approximate shape of a spectrogram as data. In addition, the volume as a reference for the above-mentioned frequency characteristics may also be set as the above-mentioned reference volume. The information indicating the reference volume and the information indicating the properties of the target sound data are used not only for calculating the volume of the direct sound or reflected sound to be perceived by the listener, but also for selection processing for selecting whether to make the listener perceive the sound.
[0195] The information related to the orientation is typically expressed by yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted and expressed by azimuth (yaw) and elevation (pitch). The orientation information may also change over time and, if changed, is transmitted to the rendering units A0203 and A0213.
[0196] Information related to the listener is information related to the position information and orientation of the listener in the sound space. The position information is represented by the position of the X-axis, Y-axis and Z-axis of the Euclidean space, but it does not necessarily have to be three-dimensional information and can also be two-dimensional information. Information related to the orientation is typically represented by yaw, pitch, and roll. Alternatively, the information related to the orientation can also omit the rotation of roll and be represented by azimuth (yaw) and elevation (pitch). The position information and orientation information can also change over time and are transmitted to the rendering units A0203 and A0213 when changes occur.
[0197] The sensor information is information including the amount of rotation or displacement detected by the sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to the rendering units A0203 and A0213, and the rendering units A0203 and A0213 update the information of the position and orientation of the listener based on the sensor information. For example, the sensor information can also use the position information obtained by the portable terminal using GPS, camera or LiDAR (Laser Imaging Detection and Ranging) to estimate its own position. In addition, information obtained from the outside via the communication module other than the sensor can also be detected as sensor information. Information indicating the temperature of the sound signal processing device 100 and information indicating the remaining battery level can also be obtained from the sensor as sensor information. Information indicating the computing resources (CPU capacity, memory resources, PC performance) of the sound signal processing device 100 or the sound prompt device A0002 can also be obtained in real time as sensor information.
[0198] In this embodiment, the acquisition unit 110 acquires the object information from the storage unit 140, but the present invention is not limited thereto, and the object information may be acquired from a device other than the sound signal processing device 100 (for example, a server device 500 such as a cloud server). In addition, the acquisition unit 110 acquires the second position information from the earphone 200 (more specifically, the head sensor unit 201), but the present invention is not limited thereto.
[0199] Here, information included in the object information is described.
[0200] First, the first change information will be described.
[0201] The first change information is information indicating a change in an object causing wind. In the present embodiment, the change in the object refers to a change in the state of the object. Here, since the object is a fan, the change in the state of the object can be exemplified by the following examples.
[0202] For example, the change in the state of the object is that the fan is switched on and off (hereinafter sometimes referred to as "on / off switching"). In addition, for example, the change in the state of the object is that the switch indicating the wind speed of the fan is switched from weak to strong (hereinafter sometimes referred to as "wind speed switching"). In addition, for example, the change in the state of the object is that the switch indicating the swing head of the fan is switched from no swing head to swing head (hereinafter sometimes referred to as "wind direction switching").
[0203] Next, the second change information will be described.
[0204] The second change information is information indicating a change in wind caused by a change in the object. The second change information indicates a change in wind speed or a change in wind direction as a change in wind caused by a change in the object. In this embodiment, the content of the information indicated by the second change information changes in response to a change in the state of the object indicated by the first change information.
[0205] In the case where the change in the state of the object represented by the first change information is "on / off switching", the second change information, for example, indicates that the wind speed is switched from 0 m / s to V1 m / s (V1>0). In addition, in the case where the change in the state of the object represented by the first change information is "wind speed switching", the second change information, for example, indicates that the wind speed is switched from V2 m / s to, for example, V3 m / s (V3>V2). In addition, in the case where the change in the state of the object represented by the first change information is "wind direction switching", the second change information, for example, indicates that the wind direction is switched from a constant state to a changing state. In this way, the second change information can be information that depends on the first change information.
[0206] In addition, the above-mentioned V1, V2, and V3 which show the wind speed are, for example, the wind speed at the position where the target fan is arranged.
[0207] Next, the timing information will be described.
[0208] The timing information is information indicating a predetermined timing related to a change in an object. As described above, the acoustic signal processing device 100 outputs the pneumatic sound data to the earphone 200 after a predetermined time has passed from the predetermined timing. The predetermined timing indicates the timing at which the predetermined time for outputting the pneumatic sound data starts.
[0209] The predetermined timing indicated by the timing information is the timing of wind change, more specifically, the timing of wind change caused by the change of the object. For example, the predetermined timing is the timing of wind speed change or wind direction change due to the change of the object.
[0210] Furthermore, a case where the predetermined timing is the timing of wind speed change will be described.
[0211] As an example of a change in wind speed, an example of a fan being the object being switched from off to on can be cited. In this case, for example, the wind speed changes from 0 m / s to V1 m / s, and the prescribed timing is the timing of the wind speed change, that is, the timing of the wind speed changing from 0 m / s to V1 m / s. In addition, when the fan is switched from off to on, as described above, the fan generates a motor sound. Therefore, in this case, the prescribed timing is the timing of the wind speed change, and is the timing (first timing) of outputting the sound data (object sound data) corresponding to the fan being the object. In other words, the sound signal processing device 100 (more specifically, the output unit 130) of the present embodiment outputs the sound data (object sound data) corresponding to the fan at the prescribed timing (first timing). In addition, the timing information included in the object information shows that the prescribed timing is the timing of the wind change and is the first timing.
[0212] Furthermore, the predetermined timing may be, for example, a timing designated by a manager of the acoustic signal processing device 100 .
[0213] Next, the first position information will be described.
[0214] As described above, the object in the virtual space is included in the content (image) displayed by the display unit 300 , and is a fan in the present embodiment.
[0215] The first position information is information indicating where the fan in the virtual space is located at a certain point in time. In addition, in the virtual space, the fan may be moved, for example, by the user holding the fan in his hand and moving it. Therefore, the acquisition unit 110 continuously acquires the first position information. The acquisition unit 110 acquires the first position information every time the space information is updated by the space information management units A0201 and A0211, for example.
[0216] Furthermore, the sound data including the target sound data and the pneumatic sound data associated with the object will be described.
[0217] The sound data including the target sound data and the pneumatic sound data described in this specification may be a sound signal such as PCM (Pulse Code Modulation) data, but is not limited thereto and may be any information for indicating the properties of the sound.
[0218] For example, if the sound signal is a noise signal with a volume of X decibels, the sound data related to the sound signal may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating that the volume is X decibels. For another example, if the sound signal is a noise signal with a Peak / Dip of a frequency component having a predetermined characteristic, the sound data related to the sound data may be the PCM data itself representing the sound signal, or may be data consisting of information indicating that the component is a noise signal and information indicating the Peak / Dip of the frequency component.
[0219] In this specification, a sound signal based on sound data refers to PCM data representing the sound data.
[0220] In addition, the pneumatic sound data is stored in advance in the storage unit 140 as described above. The pneumatic sound data refers to data obtained by collecting the sound generated by the wind reaching the human ear or a model imitating the human ear. In the present embodiment, the pneumatic sound data is data obtained by collecting the sound generated by the wind reaching the model imitating the human ear. The pneumatic sound data is collected using a dummy head-shaped microphone or the like as the model imitating the human ear.
[0221] In addition, as described above, in the present embodiment, the wind changes due to the change of the object. The aerodynamic sound is the aerodynamic sound caused by the wind before the change or the wind after the change. In addition, the aerodynamic sound may be the aerodynamic sound caused by the wind after the change, for example, the aerodynamic sound caused by the wind of the changed wind speed or the aerodynamic sound caused by the wind of the changed wind direction.
[0222] Next, shape information will be described.
[0223] Shape information is information indicating the shape of an object in a virtual space. Shape information indicates the shape of an object, and more specifically, indicates a three-dimensional shape as a rigid body of the object. The shape of the object is represented by, for example, a sphere, a cuboid, a cube, a polyhedron, a cone, a pyramid, a cylinder, a prism, or a combination thereof. In addition, shape information may also be represented as mesh data, or a collection of multiple faces composed of voxels, a three-dimensional point group, or vertices having three-dimensional coordinates.
[0224] In addition, the first change information includes object identification information for identifying the object. In addition, the timing information also includes object identification information, the second change information also includes object identification information, the first position information also includes object identification information, the object sound data also includes object identification information, and the shape information also includes object identification information.
[0225] Therefore, even if the acquisition unit 110 acquires the first change information, timing information, second change information, first position information, object sound data, and shape information, respectively, it is possible to identify the objects represented by the first change information, timing information, second change information, first position information, object sound data, and shape information by referring to the object identification information contained in each of the first change information, timing information, second change information, first position information, object sound data, and shape information. For example, it is easy to identify here that the objects represented by the first change information, timing information, second change information, first position information, object sound data, and shape information are the same fan. That is, the first change information, timing information, second change information, first position information, object sound data, and shape information acquired by the acquisition unit 110 are clarified as information related to the fan by referring to the six object identification information. Therefore, the first change information, the timing information, the second change information, the first position information, the target sound data, and the shape information are associated as information indicating the fan.
[0226] Next, the second position information will be described.
[0227] The listener can move in the virtual space. The second position information is information indicating where the listener in the virtual space is at a certain point in time. In addition, since the listener can move in the virtual space, the acquisition unit 110 continuously acquires the second position information. The acquisition unit 110 acquires the second position information each time the spatial information management units A0201 and A0211 execute an update of the spatial information, for example.
[0228] In addition, the first change information, timing information, second change information, first position information, shape information, target sound data, second position information, and aerodynamic sound data mentioned above may be included in metadata, control information, or header information included in the input signal. In the case where the sound data including the target sound data and the aerodynamic sound data is a sound signal (PCM data), information for identifying the sound signal may be included in metadata, control information, or header information, and the sound signal may be included in other than metadata, control information, or header information. That is, the sound signal processing device 100 (more specifically, the acquisition unit 110) may also acquire metadata, control information, or header information included in the input signal, and perform sound processing based on the metadata, control information, or header information. In addition, the sound signal processing device 100 (more specifically, the acquisition unit 110) only needs to acquire the first change information, timing information, second change information, first position information, shape information, target sound data, second position information, and aerodynamic sound data mentioned above, and the acquisition location is not limited to the input signal. The sound data including the target sound data and the pneumatic sound data and the metadata may be stored in one input signal or may be stored separately in a plurality of input signals.
[0229] In addition, a sound signal other than the sound data including the object sound data and the pneumatic sound data may be stored in the input signal as audio content information. The audio content information may also be coded using MPEG-H 3D Audio (ISO / IEC23008-3) (hereinafter referred to as MPEG-H 3D Audio) or the like. In addition, the technology used in the coding process is not limited to MPEG-H 3D Audio, and other well-known technologies may also be used. In addition, the above-mentioned first change information, timing information, second change information, first position information, shape information, object sound data, second position information, and pneumatic sound data may also be used as coding processing objects.
[0230] That is, the sound signal processing device 100 obtains the sound signal and metadata included in the encoded bit stream. In the sound signal processing device 100, the audio content information is obtained and decoded. In the present embodiment, the sound signal processing device 100 functions as a decoder (for example, decoders A0200 and A0210) provided by a decoding device (for example, decoding devices A0110 and A0130), and more specifically, functions as rendering units A0203 and A0213 provided by the decoder. In addition, the term "audio content information" in the present disclosure is interpreted as information that can be replaced by the sound signal itself, the first change information, the timing information, the second change information, the first position information, the shape information, the target sound data, the second position information, and the aerodynamic sound data.
[0231] The acquisition unit 110 outputs the acquired object information and second position information to the determination unit 120 and the output unit 130 .
[0232] The determination unit 120 determines the predetermined time based on the wind indicated by the object information acquired by the acquisition unit 110. That is, the determination unit 120 determines the predetermined time based on the wind caused by the object.
[0233] For example, the determination unit 120 determines the prescribed time based on the wind speed indicated by the second change information included in the acquired object information and the distance between the position of the listener and the position of the object. If the prescribed time is t seconds, t>0 is satisfied as an example, but it is not limited to this. The prescribed time may be, for example, not less than 0.1 seconds and not more than 5 seconds. The determination unit 120 may also determine the prescribed time as a time specified by the administrator of the sound signal processing device 100. In addition, the determination unit 120 calculates the distance as follows.
[0234] The determination unit 120 calculates the distance between the position of the listener and the position of the object based on the first position information included in the object information obtained by the acquisition unit 110 and the second position information obtained. As described above, the acquisition unit 110 acquires the first position information and the second position information in the virtual space each time the spatial information is updated by the spatial information management units A0201 and A0211. The determination unit 120 calculates the distance between the position of the listener and the position of the object in the virtual space based on the plurality of first position information and the plurality of second position information acquired each time the spatial information is updated.
[0235] The determination unit 120 determines the predetermined time and outputs it to the output unit 130 .
[0236] The output unit 130 outputs the pneumatic sound data acquired by the acquisition unit 110 after a predetermined time determined by the determination unit 120 has passed from the predetermined timing indicated by the object information acquired by the acquisition unit 110. Here, the output unit 130 outputs the pneumatic sound data to the earphone 200. Thus, the earphone 200 can reproduce the pneumatic sound indicated by the output pneumatic sound data. That is, the listener can listen to the pneumatic sound after a predetermined time has passed from the predetermined timing.
[0237] The storage unit 140 is a storage device that stores computer programs executed by the acquisition unit 110 , the determination unit 120 , and the output unit 130 , object information, and pneumatic sound data.
[0238] Here, the shape information related to this embodiment is described again. The shape information is information used to generate an image of an object in a virtual space, and is information indicating the shape of the object (fan). That is, the shape information is also information used to generate content (image) displayed by the display unit 300.
[0239] The acquisition unit 110 also outputs the acquired shape information to the display unit 300. The display unit 300 acquires the shape information output by the acquisition unit 110. The display unit 300 also acquires attribute information (color, etc.) other than the shape of the object (fan) in the virtual space. The display unit 300 may directly acquire the attribute information from a device other than the sound signal processing device 100 (server device 500), or may acquire it from the sound signal processing device 100. The display unit 300 generates content (image) based on the acquired shape information and attribute information and displays it.
[0240] Hereinafter, operation example 1 of the sound signal processing method performed by the sound signal processing device 100 will be described.
[0241] [Action Example 1]
[0242] Fig.11 This is a flowchart of operation example 1 of the sound signal processing device 100 according to the present embodiment. Fig.12 This is a diagram showing a fan F and a listener L as targets in operation example 1.
[0243] like Fig.11 As shown, first, the acquisition unit 110 acquires object information (S10). As described above, the object information includes first change information indicating a change in an object causing the wind W, timing information indicating a predetermined timing related to the change in the object, second change information indicating a change in the wind W caused by the change in the object, and first position information indicating the position of the object. In addition, the object information includes object sound data indicating motor sound and shape information. This step S10 is equivalent to the acquisition step.
[0244] Here, the second change information indicates a change in wind speed of wind W as a change in wind W caused by a change in the object. The predetermined timing indicated by the timing information is the timing of the change in wind W, more specifically, the timing of the change in wind W caused by a change in the object.
[0245] Next, the acquisition unit 110 acquires second position information indicating the position of the listener L in the virtual space from the earphone 200 (S20). Furthermore, the acquisition unit 110 acquires aerodynamic sound data indicating aerodynamic sound stored in the storage unit 140 (S30).
[0246] Next, the determination unit 120 determines a predetermined time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (the fan F) ( S40 ). This step S40 corresponds to a determination step.
[0247] Furthermore, the output unit 130 outputs the sound data (target sound data) corresponding to the fan F at a predetermined timing (S50). Next, the output unit 130 outputs the aerodynamic sound data representing the aerodynamic sound caused by the wind W after a predetermined time has passed since the predetermined timing (S60). This step S60 corresponds to an output step.
[0248] Here, the prescribed timing and prescribed time in this operation example are described.
[0249] Here, the predetermined timing is the timing of the change of the wind W and the timing of the change of the wind speed due to the change of the object. As an example, when the listener L views the content of the fan F displayed on the display unit 300, the predetermined timing is the timing when the fan F is switched from off to on.
[0250] In the real space, the listener L hears the aerodynamic sound at a timing when the wind W caused by the fan F reaches the listener L after the time when the fan F is switched from off to on (i.e., the predetermined timing). Therefore, the determination unit 120 may determine the time from the predetermined timing to the time when the wind W caused by the fan F reaches the listener L as the predetermined time.
[0251] Fig.13A It is explained in Fig.11 FIG. 1 is a diagram showing a process of determining a predetermined time in step S40.
[0252] Let the distance between the position of the listener L and the position of the object (fan F) be D. More specifically, let the distance between the position of the ear of the listener L and the position of the object (fan F) be D. In addition, the distance D is calculated by the determination unit 120 based on the first position information included in the object information acquired by the acquisition unit 110 and the acquired second position information.
[0253] Let the distance from the position of the object (fan F) where the wind speed So of the wind W generated by the fan F is U. In addition, let the direction from the fan F toward the listener L be the x-axis direction, and let the distance from the fan F to the x-axis direction be x. Since the wind speed V of the wind W is inversely proportional to the distance x, the wind speed V and the distance x satisfy the following formula.
[0254] V=So×(U / x)
[0255] The average wind speed to the position at distance D satisfies the following formula.
[0256] [Formula 1]
[0257]
[0258] The time (predetermined time) t from the timing when the fan F switches from off to on (ie, the predetermined timing) until the wind W caused by the target fan F reaches the listener L is the value obtained by dividing the distance by the average wind speed, and satisfies the following equation.
[0259] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0260] In addition, “^” in the above formula represents an exponentiation operator.
[0261] Then, as described above, in step S60 , the pneumatic sound data is output at the timing when the predetermined time t has elapsed from the predetermined timing.
[0262] Thus, the listener L can listen to the aerodynamic sound output from the headphones 200 at a time (predetermined time t) after the time (i.e., predetermined time) when the wind W caused by the fan F reaches the listener L from the time (i.e., predetermined time) when the fan F is switched from off to on. Therefore, the listener L can listen to the aerodynamic sound at the same timing as in the real space, i.e., at an appropriate timing, so the listener L is less likely to feel a sense of disharmony and can obtain a sense of presence.
[0263] Furthermore, in the present operation example, the predetermined timing is the timing when the fan F is switched from off to on, and is the first timing when the target sound data associated with the target fan F is output.
[0264] In addition, the above operation naturally includes the following meaning. That is, the meaning is "from the specified timing to the timing after the specified time t, the aerodynamic sound represented by the aerodynamic sound data is output as a sound with an amplitude that can be perceived by the listener L". This is achieved, for example, by a filter with the specified time t as a time constant when the aerodynamic sound data is output. Specifically, it can also be achieved as follows.
[0265] Fig. 13B This is a diagram for explaining an example of the details of the output of the pneumatic sound data according to the present embodiment. Fig. 13C This is a diagram for explaining another example of the details of the output of the pneumatic sound data according to the present embodiment.
[0266] Fig. 13B (a) is a diagram showing a trigger signal indicating a change in the on / off state of the fan F. Fig. 13B In (a), a trigger signal is shown whose value is "0" when the fan F is turned off and whose value is "1" when the fan F is turned on. Fig. 13B (b) is a diagram showing the trigger signal after being multiplied by the time constant t. That is, a Low Pass filter having a time constant of a predetermined time t is applied to the trigger signal. Fig. 13B(c) is a diagram showing aerodynamic sound data after the amplitude is amplified according to the magnitude of the output signal of the LowPass filter.
[0267] This makes it possible to easily simulate the operation of outputting pneumatic sound data at a timing when the predetermined time t has elapsed. In addition, this makes it possible to automatically simulate the operation when the cause of the generation of pneumatic sound disappears (operation when the fan F changes from on to off).
[0268] Here, t does not necessarily have to be a value accurately calculated based on the following formula, and may be a value simply approximated so that t becomes larger as the distance D becomes larger.
[0269] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0270] In addition, “^” in the above formula represents an exponentiation operator.
[0271] Fig. 13C (a) and Fig. 13B (a) is a diagram showing a trigger signal indicating a change in the on / off state of the fan F. Fig. 13C (b) and Fig. 13B (b) is a diagram showing the trigger signal after multiplication by the time constant t. Fig. 13B (b) shows the trigger signal after being multiplied by the time constant t which is smaller than the time constant t. Fig. 13C (c) shows that according to Fig. 13C (b) is a graph of pneumatic sound data in which the value of the trigger signal multiplied by the time constant t is controlled.
[0272] As described above, the predetermined timing is the timing when the fan F is switched from off to on, and is the first timing when the target sound data associated with the target fan F is output.
[0273] Therefore, through the processing of step S50, at the timing when the fan F is switched from off to on, the listener L can hear the motor sound of the fan F output from the headphones 200. Furthermore, through the processing of step S60, at the timing when the wind W caused by the fan F switching from off to on reaches the listener L after the listener L hears the motor sound, the listener L can hear the aerodynamic sound output from the headphones 200.
[0274] In the real space, the motor sound reaches the listener L at the speed of sound and is heard by the listener L, and the aerodynamic sound is heard by the listener L when the wind W reaches the listener L. In the real space, the speed of sound is generally faster than the wind speed. In this action example, as in the real space, the listener L first hears the motor sound and then the aerodynamic sound. Therefore, the listener L can hear the motor sound (the sound represented by the sound data corresponding to the object) and the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disharmony and can get a sense of presence.
[0275] In Operation Example 1, the timing of wind speed change and the timing (first timing) of outputting sound data (target sound data) associated with the target fan F are used as the predetermined timing, but the present invention is not limited to this.
[0276] For example, there is a case where the object information indicates a change in the direction of the wind W caused by a change in the object (fan F). More specifically, there is a case where the object information indicates a change in the direction (wind direction) of the wind W as a change in the wind W caused by a change in the object (fan F). This case is, for example, a case where the change in the state of the object indicated by the first change information is "wind direction switching", and the second change information indicates that the wind direction switches from a certain state to a changing state.
[0277] In this case, the timing information included in the target information indicates that the predetermined timing is the third timing at which the direction (wind direction) of the wind W changes.
[0278] In this way, if the wind direction of the fan F changes, the state of the wind W reaching the listener L changes, so the aerodynamic sound heard by the listener L also changes. Fig.11 In step S60 shown in FIG. 1 , the output unit 130 may output the aerodynamic sound data indicating the aerodynamic sound caused by the wind W after a predetermined time has passed from the third timing (predetermined timing) indicated by the object information.
[0279] Furthermore, the prescribed timing and prescribed time are not limited to the timing and time shown in Action Example 1. Alternatively, the prescribed timing may be a timing (specified timing) specified by a user (e.g., an administrator of the sound signal processing device 100), and the prescribed time may be a time (specified time) specified by the administrator. The decision unit 120 may also determine the timing and time specified by the user as the prescribed timing and prescribed time. For example, the sound signal processing device 100 may include an acceptance unit that receives the timing and time specified by the user, and the decision unit 120 determines the timing and time accepted by the acceptance unit as the prescribed timing and prescribed time. In this case, the administrator specifies the specified timing and time so that the listener L can hear the aerodynamic sound at the same timing as in the real space.
[0280] In this case, the listener L can also hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of discomfort and can obtain a sense of presence.
[0281] In addition, in the operation example 1 of the embodiment, the pneumatic sound data is pre-stored in the storage unit 140, but the present invention is not limited thereto. For example, the determination unit 120 may generate the pneumatic sound data. For example, the determination unit 120 may acquire a noise signal and generate the pneumatic sound data by processing the acquired noise signal with a plurality of frequency band emphasis filters.
[0282] Furthermore, in the action example 1 of the embodiment, the determination unit 120 determines the prescribed time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (fan F), but the present invention is not limited thereto. For example, the object information may include the first position information indicating the position of the object, and the determination unit 120 may determine the prescribed time based on the distance between the position of the listener L of the aerodynamic sound and the position of the object indicated by the first position information included in the acquired object information. For example, the prescribed time corresponding to the distance serving as a reference may be set, and the prescribed time may be determined in such a manner that the longer the distance between the position of the listener L of the aerodynamic sound and the position of the object is than the distance serving as a reference, the longer the prescribed time is, and the shorter the distance between the position of the listener L of the aerodynamic sound and the position of the object is than the distance serving as a reference, the shorter the prescribed time is.
[0283] (Variation of the embodiment)
[0284] Hereinafter, a modification of the embodiment will be described. Hereinafter, the description will be centered on the differences from the embodiment, and the description of the common points will be omitted or simplified.
[0285] [constitute]
[0286] In the modification, the sound signal processing device 100 of the embodiment is used, but the object in the virtual space is different. The object of this modification is a vehicle as a moving body. More specifically, the object is an ambulance. In this case, the aerodynamic sound is the sound generated by the wind W generated by the movement of the position of the object reaching the listener L. In addition, the ambulance as the object is an object that generates sound, and generates a whistle sound.
[0287] The object information related to this variant is information indicating the change of the object causing the wind W, the specified timing related to the change of the object, the change of the wind W caused by the change of the object, and the position of the object. In addition, as in the embodiment, the object information is regarded as information including the first change information indicating the change of the object causing the wind W, the timing information indicating the specified timing related to the change of the object, the second change information indicating the change of the wind W caused by the change of the object, and the first position information indicating the position of the object.
[0288] The first change information is information indicating a change in an object causing the wind W. In the present modification, the change in the object refers to a change in the position of the object.
[0289] The first position information is information indicating where the ambulance in the virtual space is located at a certain point in time. In addition, in the virtual space, for example, the ambulance may travel and move its position by being operated by a driver. Therefore, the acquisition unit 110 continuously acquires the first position information.
[0290] The second change information is information indicating a change in the wind W caused by a change in the object. In the present embodiment, the content of the information indicated by the second change information changes in accordance with a change in the position of the object indicated by the first change information.
[0291] For example, when the first change information indicates a change in the position of the object, the second change information indicates that the wind speed of the wind W generated by the movement of the object changes from the first specified value to the second specified value, or the wind direction changes from the first specified direction to the second specified direction. In addition, the first and second specified values are, for example, the wind speed at the position where the ambulance is deployed, and the first and second specified directions are, for example, the wind direction at the position where the ambulance is deployed.
[0292] As a more specific example, a case where the first change information indicates that the ambulance approaches the listener L and then moves away from the listener L will be described. In this case, the wind W generated by the movement of the ambulance blows strongly toward the listener L while the ambulance approaches the listener L, and blows weakly toward the listener L while the ambulance moves away from the listener L. Therefore, the wind speed of the wind W toward the listener L is a high value while the ambulance approaches the listener L, and is a low value toward the listener L while the ambulance moves away from the listener L. In this way, the wind W (more specifically, the wind speed of the wind W) changes.
[0293] In this modification, the wind speed W caused by the target ambulance can be regarded as the same as the moving speed of the ambulance. The moving speed of the ambulance is calculated by differentiating the position of the ambulance in virtual space with time based on the first position information.
[0294] Next, the timing information will be described.
[0295] The timing information is information indicating a prescribed timing related to a change in an object. The prescribed timing indicated by the timing information is the timing of a change in the wind W, and more specifically, the timing of a change in the wind W caused by a change in the position of the object. For example, the prescribed timing is the timing of a change in wind speed due to a change in the position of the object, and as an example, is the timing of an ambulance approaching the listener L and then moving away from the listener L. In this case, the prescribed timing is the timing at which the amount of change in the distance between the position of the listener L in the virtual space and the position of the object changes from negative to positive as time passes. In other words, the prescribed timing is the timing at which the object is closest to the listener L in the virtual space. In addition, for example, the prescribed timing may also be the timing of a change in wind direction due to a change in the position of the object.
[0296] Next, operation example 2 of the sound signal processing method performed by the sound signal processing device 100 will be described.
[0297] [Action Example 2]
[0298] Fig.14 This is a flowchart of operation example 2 of the sound signal processing device 100 according to the present embodiment. Fig.15 This is a diagram showing the ambulance A and the listener L as the objects in operation example 2.
[0299] like Fig.14 As shown, first, the acquisition unit 110 acquires object information (S10). As described above, the object information includes first change information indicating a change in an object causing the wind W, timing information indicating a predetermined timing related to the change in the object, second change information indicating a change in the wind W caused by the change in the object, and first position information indicating the position of the object. In addition, the object information includes object sound data and shape information indicating the whistle sound.
[0300] Here, the second change information indicates a change in wind speed of wind W as a change in wind W caused by a change in the object. The predetermined timing indicated by the timing information is the timing of the change in wind W, more specifically, the timing of the change in wind W caused by a change in the object.
[0301] Next, the acquisition unit 110 acquires second position information indicating the position of the listener L in the virtual space from the earphone 200 (S20). Furthermore, the acquisition unit 110 acquires aerodynamic sound data indicating aerodynamic sound stored in the storage unit 140 (S30).
[0302] Furthermore, the output unit 130 determines whether or not a predetermined timing has been reached (S35). If the predetermined timing has not been reached (No in step S35), the process of step S35 is repeated.
[0303] When the predetermined timing has come (YES in step S35 ), the determination unit 120 determines the predetermined time based on the wind speed indicated by the second change information and the distance between the position of the listener L and the position of the object (ambulance A) ( S40 ).
[0304] Next, the output unit 130 outputs aerodynamic sound data indicating aerodynamic sound caused by the wind W after a predetermined time has passed from the predetermined timing ( S60 ).
[0305] Furthermore, the predetermined timing and the processing of step S35 in this operation example will be described in more detail.
[0306] In this operation example, the predetermined timing is the timing of the change in wind W. More specifically, the predetermined timing is the timing of the change in wind speed due to the change in the position of the object, and is the timing when the change in the distance between the position of the listener L in the virtual space and the position of the object changes from negative to positive as time passes.
[0307] Fig.16 This is a schematic diagram for explaining the prescribed timing of Action Example 2.
[0308] Ambulance A Fig.16 The position of the listener L is fixed while the ambulance A is moving from (a) to (c). The change in the distance between the position of the listener L in the virtual space and the position of the object is negative while the ambulance A is moving from (a) to (b). The change in the distance between the position of the listener L in the virtual space and the position of the object is positive while the ambulance A is moving from (b) to (c). Therefore, the timing when the change in the distance changes from negative to positive is when the ambulance A is in Fig.16 The timing of the position of (b) is shown.
[0309] Therefore, in step S35, the following is performed: Fig.17 Processing shown. Fig.17 This is a flowchart for explaining the details of step S35 in operation example 2.
[0310] After the processing of step S30, the decision unit 120 determines whether it is the timing (predetermined timing) at which the change amount of the distance between the position of the listener L and the position of the object (ambulance A) in the virtual space changes from negative to positive (S35a). In addition, the decision unit 120 calculates the distance between the position of the listener L and the position of the object (ambulance A), and calculates the change amount of the distance by differentiating the calculated distance. In the case of "yes" in step S35a, the processing of step S40 is performed, and in the case of "no" in step S35a, the processing of step S35 is repeated.
[0311] Furthermore, the prescribed time in this operation example will be described in more detail.
[0312] In the real space, the listener L hears the aerodynamic sound at the timing when the wind W caused by the ambulance A reaches the listener L after the timing when the change amount of the distance between the position of the listener L and the position of the object changes from negative to positive. In addition, as described above, the timing when the change amount of the distance changes from negative to positive is the timing when the object is closest to the listener L, which is the predetermined timing. Therefore, the decision unit 120 can determine the time from the predetermined timing to the time when the wind W caused by the ambulance A reaches the listener L as the predetermined time.
[0313] In this operation example, the same operation as described in operation example 1 is used. Fig.13A The same idea determines the time. Fig.15 As shown, let the distance between the position of the listener L and the position of the object (ambulance A) be D. More specifically, let Fig.16 The distance between the position of the ambulance A in the position (b) shown and the position of the listener L is D.
[0314] Let the distance from the position of the object (ambulance A) where the wind speed of wind W generated by the object ambulance A is So be U. In addition, let the direction from the ambulance A toward the listener L be the x-axis direction, and let the distance from the ambulance A to the x-axis direction be x. Since the wind speed V of the wind W is inversely proportional to the distance x, the wind speed V and the distance x satisfy the following formula.
[0315] V=So×(U / x)
[0316] The average wind speed to the position at distance D satisfies the following formula.
[0317] [Formula 2]
[0318]
[0319] t, which is the time from the timing when the change in the distance between the position of the listener L and the position of the object changes from negative to positive (i.e., the specified timing) to the time when the wind W caused by the ambulance A as the object reaches the listener L (specified time), is the value obtained by dividing the distance by the average wind speed, and satisfies the following formula.
[0320] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0321] Next, as described above, in step S60 , the pneumatic sound data is output at a timing when a predetermined time t has elapsed from a predetermined timing.
[0322] Thus, the listener L can hear the aerodynamic sound output from the earphone 200 at the timing (predetermined time t) after the time when the change amount of the distance between the position of the listener L and the position of the object changes from negative to positive (i.e., the predetermined timing) and the time when the wind W caused by the ambulance A reaches the listener L. Therefore, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at the appropriate timing, so the listener L is less likely to feel a sense of disharmony and can obtain a sense of presence.
[0323] Further explanation is as follows. In the real space, the listener L hears the pneumatic sound after the ambulance A or other vehicle approaches the listener L. Therefore, in the virtual space, if the listener L hears the pneumatic sound before the ambulance A approaches the listener L, the listener L will feel a sense of disharmony. In Action Example 2, the timing at which the change in the distance between the position of the listener L and the position of the object changes from negative to positive (that is, the timing at which the object approaches the listener L) is set to a prescribed timing. As a result, the listener L can hear the pneumatic sound after the ambulance A or other vehicle as the object approaches the listener L, that is, the listener L can hear the pneumatic sound at an appropriate timing, so the listener L is not likely to feel a sense of disharmony, and the listener L can obtain a sense of presence.
[0324] In addition, ambulance A is an object that generates sound, and generates whistle sound. Fig.16 As shown, when the position of the ambulance A changes, that is, when the ambulance A moves, the output unit 130 outputs the target sound signal representing the whistle sound so that the listener L hears the whistle sound accompanied by the Doppler effect.
[0325] In the above-mentioned Action Example 2, the predetermined timing is the timing when the change in the distance between the position of the listener L and the position of the object changes from negative to positive, but it is not limited to this. For example, in another first example of Action Example 2, the predetermined timing may be the timing (second timing) when the distance between the position of the listener L and the position of the object becomes shorter than the predetermined distance. The predetermined distance is, for example, several meters to several tens of meters, and is a distance indicating that the distance between the position of the listener L and the position of the object is sufficiently close. The predetermined distance may also be, for example, a value specified by the administrator of the sound signal processing device 100.
[0326] In this case, in step S35, the following is performed: Fig.18 Processing shown. Fig.18 This is a flowchart for explaining the details of step S35 of another first example of operation example 2.
[0327] After the process of step S30 is performed, the determination unit 120 determines whether it is the timing (second timing) at which the distance between the position of the listener L in the virtual space and the position of the object (ambulance A) becomes shorter than a predetermined distance (S35b). As described above, if the answer is "yes" in step S35b, the process of step S40 is performed, and if the answer is "no" in step S35b, the process of step S35 is repeated.
[0328] Thus, in another first example of action example 2, also at the timing when the wind W caused by the ambulance A reaches the listener L after the time has passed from the second timing when the distance between the position of the listener L and the position of the object (ambulance A) is sufficiently close, the listener L is able to hear the aerodynamic sound output from the earphone 200.
[0329] Next, another second example of the operation example 2 is described. In another second example of the operation example 2, in step S35, Fig.17 and Fig.18 If both steps S35a and S35b are "yes", the process of step S40 is performed, and if at least one of steps S35a and S35b is "no", the process of step S35 is repeated. The process shown in another second example of such action example 2 can also be performed.
[0330] Next, pipeline processing will be described.
[0331] Part or all of the processing performed by the above-described sound signal processing device 100 may be performed as part of pipeline processing as described in Patent Document 2, for example. Fig.19 It is used to illustrate Figure 6 and Figure 7 A functional block diagram of the case where the rendering units A0203 and A0213 perform pipeline processing and a diagram showing an example of the steps. Fig.19 In the description, use Figure 6 and Figure 7 The rendering unit 900 is described as an example of the rendering units A0203 and A0213.
[0332] Pipeline processing means dividing the processing for providing the acoustic effect into a plurality of processes and executing each process in sequence. In each of the divided processes, for example, signal processing of the sound signal or generation of parameters used in the signal processing is performed.
[0333] The rendering unit 900 of this embodiment includes, for example, implementing reverberation effects, initial reflection processing, distance attenuation effects, binaural processing and other processing as pipeline processing. However, the above-mentioned processing is an example, and other processing may also be included, or some processing may not be included. For example, the rendering unit 900 may include diffraction processing or occlusion processing as pipeline processing, and it may be omitted when reverberation processing is not required. In addition, each processing may be represented as a stage, and the sound signal such as the reflected sound generated by the result of each processing may be represented as a rendering item. The order of each stage in the pipeline processing and the stages included in the pipeline processing are not limited to Fig.19 Example shown.
[0334] In addition, the rendering unit 900 may not include Fig.19 Of all the stages shown, some stages may be omitted, or other stages may exist in addition to the rendering unit 900 .
[0335] As an example of pipeline processing, the processing performed in each of reverberation processing, initial reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing is described. In each process, metadata included in the input signal is analyzed and parameters required for generating reflected sound are calculated.
[0336] In addition, Fig.19 In the embodiment, the rendering unit 900 includes a reverberation processing unit 901, an initial reflection processing unit 902, a distance attenuation processing unit 903, a selection unit 904, a calculation unit 906, a generation unit 907, and a binaural processing unit 905. Here, an example is described in which the reverberation processing unit 901 performs a reverberation processing step, the initial reflection processing unit 902 performs an initial reflection processing step, the distance attenuation processing unit 903 performs a distance attenuation processing step, the selection unit 904 performs a selection processing step, and the binaural processing unit 905 performs a binaural processing step.
[0337] In the reverberation processing step, the reverberation processing unit 901 generates a sound signal representing a reverberation sound or a parameter required for generating a sound signal. The reverberation sound is a sound that is included in the reverberation sound that reaches the listener as a reverberation after the direct sound. As an example, the reverberation sound is a reverberation sound that reaches the listener after a relatively late period (for example, about a hundred and several dozen ms from the arrival of the direct sound) after the initial reflected sound described later reaches the listener, after being reflected more times (for example, several dozen times) than the initial reflected sound. The reverberation processing unit 901 refers to the sound signal and spatial information included in the input signal, and uses a predetermined function prepared in advance for generating the reverberation sound to calculate.
[0338] The reverberation processing unit 901 may also apply a known reverberation generation method to the sound signal to generate reverberation. As an example, the known reverberation generation method is the Schroeder method, but is not limited thereto. In addition, when applying the known reverberation generation process, the reverberation processing unit 901 uses the shape and acoustic characteristics of the sound reproduction space represented by the spatial information. Thus, the reverberation processing unit 901 can calculate parameters for generating a sound signal representing the reverberation.
[0339] In the initial reflection processing step, the initial reflection processing unit 902 calculates the parameters used to generate the initial reflected sound based on the spatial information. The initial reflected sound is the reflected sound that reaches the listener after more than one reflection during the relatively early period (for example, about tens of ms from the arrival of the direct sound) after the direct sound reaches the listener from the sound source object. The initial reflection processing unit 902, for example, refers to the sound signal and metadata, and uses the shape, size, position of objects such as structures and the reflectivity of the objects in the three-dimensional sound field (space) to calculate the path (the length of the path) of the reflected sound that is reflected from the sound source object and reaches the listener. In addition, the initial reflection processing unit 902 can also calculate the path (the length of the path) of the direct sound. The information representing the path can also be used as a parameter for generating the initial reflected sound, and can also be used as a parameter for the selection processing of the reflected sound in the selection unit 904.
[0340] In the distance attenuation processing step, the distance attenuation processing unit 903 calculates the volume reaching the listener based on the difference between the length of the direct sound path and the length of the reflected sound path calculated by the initial reflection processing unit 902. The volume reaching the listener is attenuated in proportion to the distance to the listener (inversely proportional to the distance) relative to the volume of the sound source, so the volume of the direct sound can be obtained by dividing the volume of the sound source by the length of the direct sound path, and the volume of the reflected sound can be calculated by dividing the volume of the sound source by the length of the reflected sound path.
[0341] In the selection process step, the selection unit 904 selects a sound to be generated. The selection process may be performed based on the parameters calculated in the previous step.
[0342] When the selection process is performed as part of the pipeline process, the sound that is not selected in the selection process may not be the subject of the process after the selection process in the pipeline process. By not performing the process after the selection process on the sound that is not selected, the calculation load of the sound signal processing device 100 can be reduced compared to the case where it is decided not to perform the binaural process on the sound that is not selected.
[0343] Furthermore, when the selection process described in this embodiment is executed by a part of the pipeline process, if the order of the selection process is set to be executed in an earlier order among the orders of the plurality of processes in the pipeline process, more processes after the selection process can be omitted, so that the amount of calculation can be further reduced. For example, if the calculation unit 906 and the generation unit 907 execute the selection process in an earlier order than the process, the process related to the aerodynamic sound related to the object determined not to be selected can be omitted, and the amount of calculation in the acoustic signal processing device 100 can be further reduced.
[0344] Furthermore, the parameters calculated in a part of the pipeline process for generating the rendering item may be used in the selection unit 904 or the calculation unit 906 .
[0345] In the binaural processing step, the binaural processing unit 905 performs signal processing on the sound signal of the direct sound so that it is perceived as sound reaching the listener from the direction of the sound source object. Furthermore, the binaural processing unit 905 performs signal processing so that the reflected sound is perceived as sound reaching the listener from an obstacle object related to the reflection. Based on the coordinates and orientation of the listener in the sound space (i.e., the position and orientation of the listening point), processing of applying HRIR (Head-Related Impulse Responses) DB (Data base) is performed so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object. In addition, the position and direction of the listening point can change, for example, to match the movement of the listener's head. In addition, information indicating the position of the listener can also be obtained from the sensor.
[0346] The programs used in pipeline processing and binaural processing, spatial information required for sound processing, HRIR DB, threshold data and other parameters are obtained from the memory of the sound signal processing device 100 or from the outside. HRIR (Head-Related Impulse Responses) is the response characteristic when one impulse is generated. In other words, HRIR is the response characteristic obtained by Fourier transforming the head-related transfer function that expresses the change of sound generated by the surrounding objects including the ear shell, the human head and shoulders as a transfer function, thereby transforming the expression from the frequency domain to the time domain. HRIR DB is a database containing such information.
[0347] In addition, as an example of pipeline processing, the rendering unit 900 may include a processing unit not shown in the figure, for example, a diffraction processing unit or an occlusion processing unit.
[0348] The diffraction processing unit performs processing to generate a sound signal representing a sound including diffracted sound caused by an obstacle between a listener and a sound source object in a three-dimensional sound field (space). The diffracted sound is a sound that, when there is an obstacle between the sound source object and the listener, bypasses the obstacle and reaches the listener from the sound source object.
[0349] The diffraction processing unit, for example, refers to the sound signal and metadata, and uses the position of the sound source object in the three-dimensional sound field (space), the position of the listener, and the position, shape and size of the obstacle to calculate the path from the sound source object to bypass the obstacle and reach the listener, and generates diffracted sound based on the path.
[0350] The occlusion processing unit generates a sound signal that can be faintly heard when there is a sound source object on the opposite side of the obstacle object, based on the space information acquired in a certain step and information such as the material of the obstacle object.
[0351] In addition, in the above-mentioned embodiment, the position information assigned to the sound source object is defined as the information of a "point" in the virtual space, and the details of the invention are described assuming a so-called "point sound source". On the other hand, as a method of defining a sound source in a virtual space, there is also a case where a sound source that is not a point sound source and is spatially extended is defined as an object having length, size or shape. In such a case, since the distance between the listener and the sound source or the direction of arrival of the sound is uncertain, the reflected sound generated thereby does not need to be analyzed, or regardless of the analysis result, it can be limited to the processing of "selection" in the above-mentioned selection unit 904. This is because, by doing so, it is possible to avoid the degradation of sound quality that may be caused by not selecting the reflected sound. Alternatively, a representative point such as the center of gravity of the object can be set, and the processing of the present disclosure can be applied assuming that the sound is generated from the representative point. In this case, the processing of the present disclosure can also be applied after adjusting the threshold according to the information of the spatial extension of the sound source.
[0352] Next, an example of the construction of a bit stream is described.
[0353] The bitstream includes, for example, a sound signal and metadata. The sound signal is sound data that represents the sound, indicating information related to the frequency and strength of the sound. The spatial information included in the metadata is information related to the space in which the listener who hears the sound based on the sound signal is located. Specifically, the spatial information is information related to the predetermined position (localization position) when the sound image of the sound is localized at a predetermined position in the sound space (for example, in a three-dimensional sound field), that is, when the listener perceives the sound as arriving from a predetermined direction. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
[0354] The sound source object information is information about an object that generates sound based on a sound signal, that is, an object that reproduces the sound signal, and is information about a virtual object (sound source object) arranged in a virtual space corresponding to the real space in which the object is arranged, that is, a sound space. The sound source object information includes, for example, information indicating the position of the sound source object arranged in the sound space, information about the direction of the sound source object, information about the directionality of the sound emitted by the sound source object, information indicating whether the sound source object is a living thing, and information indicating whether the sound source object is a moving body. For example, a sound signal corresponds to one or more sound source objects indicated by the sound source object information.
[0355] As an example of the data structure of a bit stream, the bit stream is composed of, for example, metadata (control information) and an audio signal.
[0356] The sound signal and metadata can be stored in one bitstream or in multiple bitstreams. Similarly, the sound signal and metadata can be stored in one file or in multiple files.
[0357] The bit stream may exist for each sound source or for each playback time. When the bit stream exists for each playback time, a plurality of bit streams may be processed in parallel at the same time.
[0358] The metadata may be assigned to each bit stream, or may be assigned as information for controlling a plurality of bit streams. Furthermore, the metadata may be assigned to each playback time.
[0359] When the sound signal and metadata are stored in a plurality of bitstreams or files, the sound signal and metadata may be included in information indicating other bitstreams or files associated with one or a part of the bitstreams or files, and the sound signal and metadata may be included in information indicating other bitstreams or files associated with all the bitstreams or files. Here, the associated bitstream or file is, for example, a bitstream or file that may be used simultaneously during audio processing. In addition, the associated bitstream or file may also include a bitstream or file in which information indicating other associated bitstreams or files is described together. Here, the information indicating other associated bitstreams or files may be, for example, an identifier indicating the other bitstream, a file name indicating other files, a URL (Uniform Resource Locator) or a URI (Uniform Resource Identifier), etc. In this case, the acquisition unit 110 determines or acquires the bitstream or file based on the information indicating other associated bitstreams or files. In addition, information indicating other associated bitstreams may be included in the bitstream, and information indicating bitstreams or files associated with other bitstreams or files may be included in the bitstream. Here, the file including information indicating the associated bitstream or file may be, for example, a control file such as a manifest file used for content distribution.
[0360] In addition, all or part of the metadata may be obtained from outside the bitstream of the audio signal. For example, either metadata for controlling the audio or metadata for controlling the video may be obtained from outside the bitstream, or both metadata may be obtained from outside the bitstream. Furthermore, when metadata for controlling the video is included in the bitstream obtained by the audio signal reproduction system, the audio signal reproduction system may also have a function of outputting metadata that can be used for controlling the video to a display device that displays an image or a stereoscopic video reproduction device that reproduces a stereoscopic video.
[0361] Furthermore, examples of information included in metadata will be described.
[0362] Metadata may also be information used to describe a scene represented by a sound space. Here, a scene is a term used to represent a collection of all elements of a three-dimensional image and sound event in a sound space modeled by a sound signal reproduction system using metadata. That is, the metadata described here includes not only information for controlling sound processing, but also information for controlling image processing. Of course, metadata may include information for controlling only one of the sound processing and the image processing, or information used for controlling both.
[0363] The sound signal reproduction system performs sound processing on the sound signal by using metadata included in the bitstream and additional interactive listener position information to generate a virtual sound effect. Here, the case where initial reflection processing, obstacle processing, diffraction processing, occlusion processing and reverberation processing are performed in the sound effect is described, but other sound processing can also be performed using metadata. For example, it can be considered that the sound signal reproduction system adds sound effects such as distance attenuation effect, positioning, and Doppler effect. In addition, information for switching on and off all or part of the sound effects and priority information can also be added as metadata.
[0364] In addition, as an example, the encoded metadata includes: information related to the sound space containing the sound source object and the obstacle object, and information related to the positioning position when the sound image of the sound is positioned at a specified position in the sound space (that is, so that it is perceived as a sound arriving from a specified direction). Here, the obstacle object is an object that may affect the sound perceived by the listener, such as blocking or reflecting the sound during the period from the sound emitted by the sound source object to the sound reaching the listener. In addition to stationary objects, obstacle objects may also include animals such as humans or moving objects such as machinery. In addition, when there are multiple sound source objects in the sound space, for any sound source object, other sound source objects may become obstacle objects. Non-sound-emitting objects that do not emit sound, such as building materials or inanimate objects, and sound source objects that emit sound may also become obstacle objects.
[0365] The metadata includes all or part of the information representing the shape of the sound space, the shape information and position information of obstacle objects existing in the sound space, the shape information and position information of sound source objects existing in the sound space, and the position and orientation of the listener in the sound space.
[0366] The sound space may be a closed space or an open space. In addition, the metadata includes information indicating the reflectivity of structures such as floors, walls, or ceilings that can reflect sound in the sound space, and the reflectivity of obstacle objects present in the sound space. Here, the reflectivity is the energy ratio of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectivity can also be set uniformly regardless of the frequency band of the sound. In the case where the sound space is an open space, for example, uniformly set parameters such as attenuation rate, diffracted sound, or initial reflected sound can also be used.
[0367] In the above description, reflectivity is listed as a parameter related to an obstacle object or a sound source object included in metadata, but information other than reflectivity may also be included. For example, information other than reflectivity may also include information related to the material of the object as metadata related to both the sound source object and the non-sound-generating object. Specifically, information other than reflectivity may also include parameters such as diffusivity, transmittance, and sound absorption.
[0368] As information related to the sound source object, it is also possible to include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources emitted from an object, and information specifying the sound source area in the object. For example, the reproduction conditions can also be set to be a sound that flows continuously or a sound triggered by an event. The sound source area in the object can be set by the relative relationship between the position of the listener and the position of the object, or it can be set based on the object. In the case where the sound source area in the object is set by the relative relationship between the position of the listener and the position of the object, the listener can perceive that sound C is emitted from the right side of the object and sound E is emitted from the left side when the listener is looking at it, based on the surface of the object that the listener is looking at. In the case where the sound source area in the object is set based on the object, which sound is emitted from which area of the object regardless of the direction the listener is looking at can be fixed. For example, the listener can perceive that when looking at the object from the front, high sound flows from the right side and low sound flows from the left side. In this case, when the listener goes around the back of the object, the listener can perceive that low sound flows from the right side and high sound flows from the left side when looking from the back.
[0369] Metadata related to the space may include the time until the initial reflected sound, the reverberation time, the ratio of direct sound to diffuse sound, etc. When the ratio of direct sound to diffuse sound is zero, the listener can perceive only the direct sound.
[0370] (Effects, etc.)
[0371] The sound signal processing method of the relevant embodiment includes: an acquisition step, acquiring object information, wherein the object information represents the change of the object causing the wind W and a specified timing related to the change of the object; and an output step, outputting aerodynamic sound data representing the aerodynamic sound caused by the wind W after a specified time based on the change of the object has passed from the specified timing represented by the acquired object information.
[0372] Thus, the pneumatic sound data can be output at a timing after a predetermined time has passed from a predetermined timing. Therefore, the listener L can listen to the pneumatic sound at an appropriate timing, so the listener L is less likely to feel uncomfortable and can get a sense of presence. That is, an acoustic signal processing method that can give the listener L a sense of presence is realized.
[0373] For example, as shown in Operation Example 1, the predetermined timing is, for example, the timing of the change in the wind W, and the predetermined time is, for example, the time when the wind W generated by the fan F reaches the listener L.
[0374] For example, as shown in Operation Example 2, the predetermined timing is, for example, the timing of the change in the wind W. Furthermore, the predetermined time is, for example, the time when the wind W caused by the ambulance A reaches the listener L.
[0375] In the cases shown in the operation examples 1 and 2, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disharmony and can get a sense of presence. In this way, the sound signal processing method of the embodiment can give the listener L a sense of presence.
[0376] In addition, for example, the prescribed timing may be a timing specified by the user (specified timing), and the time specified by the user may be the prescribed time. In this case, the user may specify the specified timing and time so that the listener L can hear the aerodynamic sound at the same timing as in the real space, and the specified specified timing and time may be set as the prescribed timing and prescribed time. In this case, the listener L can also hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disharmony, and the listener L can get a sense of presence.
[0377] In the sound signal processing method according to the embodiment, the object information indicates a change in wind W caused by a change in the object, and the predetermined timing is the timing of the change in wind W. The sound signal processing method includes a determination step of determining the predetermined time based on the wind W indicated by the acquired object information.
[0378] Thus, since the aerodynamic sound data can be output at a timing after a predetermined time determined based on the wind W has elapsed from the timing at which the wind W changes, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0379] Furthermore, in the acoustic signal processing method according to the embodiment, the change in wind W indicated by the object information indicates a change in the wind speed of wind W; and in the determination step, the predetermined time is determined based on the wind speed.
[0380] Thus, since the predetermined time is determined based on the wind speed, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0381] Furthermore, in the acoustic signal processing method according to the embodiment, the aerodynamic sound is a sound generated at a changed wind speed.
[0382] Thereby, the aerodynamic sound heard by the listener L in the virtual space can be made closer to the aerodynamic sound heard by the listener L in the real space.
[0383] In the acoustic signal processing method according to the embodiment, the object information indicates the position of the object. The acoustic signal processing method includes a determination step of determining the predetermined time based on the distance between the position of the listener L of the aerodynamic sound and the position of the object indicated by the acquired object information.
[0384] Thus, since the predetermined time is determined based on the distance, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0385] In the acoustic signal processing method according to the embodiment, the object information indicates the position of the object. In the determination step, the predetermined time is determined based on the wind speed and the distance between the position of the listener L of the aerodynamic sound and the position of the object indicated by the acquired object information.
[0386] Thus, the predetermined time is determined based on the wind speed and the distance, so the listener L can hear the aerodynamic sound at a more appropriate timing.
[0387] In the acoustic signal processing method of the embodiment, the object information indicates that the predetermined timing is a first timing for outputting the sound data associated with the object. In the output step, the aerodynamic sound data is output after a predetermined time has passed from the first timing indicated by the acquired object information.
[0388] Thus, for example, when the subject produces a sound, the pneumatic sound data can be output at a timing when a predetermined time has elapsed from a first timing when the sound is output, so the listener L can listen to the pneumatic sound at a more appropriate timing.
[0389] For example, as shown in Action Example 1, when the object is a fan F and a motor sound is generated, the predetermined timing is, for example, the timing when the fan F is switched from off to on. The listener L can hear the aerodynamic sound output from the headphones 200 at a timing after the time (predetermined time) when the wind W caused by the fan F reaches the listener L from the predetermined timing. Therefore, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disobedience, and the listener L can get a sense of presence. In this way, the sound signal processing method of the embodiment can give the listener L a sense of presence.
[0390] In the acoustic signal processing method according to the modified example of the embodiment, the object information indicates the position of the object, and the predetermined timing is a second timing when the distance between the position of the listener L of the aerodynamic sound and the position of the object becomes shorter than a predetermined distance. In the output step, the aerodynamic sound data is output after a predetermined time has passed from the second timing indicated by the acquired object information.
[0391] Thus, since the aerodynamic sound data can be output at the timing when the predetermined time has elapsed from the second timing when the distance becomes shorter than the predetermined distance, that is, the second timing when the object approaches the listener L, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0392] For example, as shown in Action Example 2, the prescribed timing is, for example, the timing when the change in the distance between the position of the listener L and the position of the object changes from negative to positive. The listener L can listen to the aerodynamic sound output from the earphone 200 at the timing after the wind W caused by the ambulance A reaches the listener L (predetermined time) from the prescribed timing. Therefore, the listener L can listen to the aerodynamic sound at the same timing as the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disobedience, and the listener L can get a sense of presence. In this way, the sound signal processing method of the modified example of the embodiment can give the listener L a sense of presence.
[0393] Furthermore, in the acoustic signal processing method of the embodiment, the object information indicates that the change in wind W caused by the change in the object is a change in the direction of wind W, and the predetermined timing is a third timing at which a change in the direction of wind W occurs. In the output step, the aerodynamic sound data is output after a predetermined time has passed from the third timing indicated by the acquired object information.
[0394] Thus, since the aerodynamic sound data can be output at the timing when a predetermined time has elapsed from the third timing when the direction of the wind W has changed, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0395] In the acoustic signal processing method of the embodiment, the object is an object that generates sound and wind W represented by sound data associated with the object, and the aerodynamic sound is aerodynamic sound generated when the wind W generated by the object reaches the listener L.
[0396] Thereby, the fan F or the like that generates sound and wind W can be targeted, and the aerodynamic sound caused by the wind W blown out from the target can be realized.
[0397] In the acoustic signal processing method according to the embodiment, let D be the distance, and let U be the distance from the position of the object with a wind speed of So. When t is the predetermined time, t satisfies the following equation.
[0398] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0399] Thus, in the determination step, the time from the predetermined timing to the time when the wind W generated by the object reaches the listener L can be determined as the predetermined time. Therefore, the aerodynamic sound data can be output at the timing after the predetermined time has passed from the predetermined timing, so the listener L can hear the aerodynamic sound at a more appropriate timing.
[0400] For example, as shown in Example 1, in the determination step, the time when the wind W generated by the fan F reaches the listener L can be determined as a predetermined time. Therefore, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disharmony, and the listener L can get a sense of presence. In this way, the sound signal processing method of the embodiment can give the listener L a sense of presence.
[0401] In the acoustic signal processing method according to the modified example of the embodiment, the object is an object that generates wind W by the movement of the position of the object, and the aerodynamic sound is aerodynamic sound generated when the wind W generated by the movement reaches the listener L.
[0402] Thereby, a vehicle or the like that generates the wind W by moving can be targeted, and the aerodynamic sound caused by the wind W generated by the movement can be realized.
[0403] Furthermore, in the acoustic signal processing method according to the modified example of the embodiment, the predetermined timing indicated by the object information is the timing at which the amount of change in the distance with the passage of time changes from negative to positive.
[0404] Thus, since the aerodynamic sound data can be output at a timing when a predetermined time has elapsed from a timing when the distance between the position of the listener L and the position of the object is closest, the listener L can hear the aerodynamic sound at a more appropriate timing.
[0405] In the acoustic signal processing method according to the modification of the embodiment, let D be the distance, and let U be the distance from the position of the object where the wind W generated by movement has a wind speed of So. When t is the predetermined time, t satisfies the following equation.
[0406] t={(D-U)^2} / {So×U×(log(D)-log(U))}
[0407] Thus, in the determination step, the time from the predetermined timing to the time when the wind W generated by the object reaches the listener L can be determined as the predetermined time. Therefore, the aerodynamic sound data can be output at the timing after the predetermined time has passed from the predetermined timing, so the listener L can hear the aerodynamic sound at a more appropriate timing.
[0408] For example, as shown in Example 2, in the determination step, the time when the wind W caused by the ambulance A reaches the listener L can be determined as a predetermined time. Therefore, the listener L can hear the aerodynamic sound at the same timing as in the real space, that is, at an appropriate timing, so the listener L is less likely to feel a sense of disharmony, and the listener L can get a sense of presence. In this way, the sound signal processing method of the embodiment can give the listener L a sense of presence.
[0409] Furthermore, the computer program according to the embodiment is a computer program for causing a computer to execute the above-described sound signal processing method.
[0410] Thereby, the computer can execute the above-mentioned sound signal processing method according to the computer program.
[0411] In addition, the sound signal processing device 100 of the relevant embodiment includes: an acquisition unit 110, which acquires object information, wherein the object information represents a change in an object causing the wind W and a specified timing related to the change in the object; and an output unit 130, which outputs aerodynamic sound data representing aerodynamic sound caused by the wind W after a specified time based on the change in the object has passed from the specified timing represented by the acquired object information.
[0412] Thus, the aerodynamic sound data can be output at a timing after a predetermined time has passed from a predetermined timing. Therefore, the listener L can listen to the aerodynamic sound at an appropriate timing, so the listener L is less likely to feel a sense of disharmony and can obtain a sense of presence. That is, the sound signal processing device 100 that can provide the listener L with a sense of presence is realized.
[0413] (Other embodiments)
[0414] The above is a description of the sound signal processing method and the sound signal processing device of the technical solution of the present disclosure based on the implementation mode, but the present disclosure is not limited to the implementation mode and the modified example. For example, other implementation modes that are realized by arbitrarily combining the constituent elements described in this specification or excluding some of the constituent elements may be used as the implementation modes of the present disclosure. In addition, the modified examples obtained by implementing various modifications that can be conceived by those skilled in the art to the above-mentioned implementation modes without departing from the main purpose of the present disclosure, that is, without departing from the meaning of the sentences described in the claims are also included in the present disclosure.
[0415] In the above embodiment, an example is shown in which the object is a fan F, but the present invention is not limited to this. Here, an object that causes wind W is exemplified.
[0416] The object causing the wind W may be, for example, a window or door into which the wind W blows. In the virtual space, in an example where the listener L is in a building and the wind W blows outside the building, the wind W blows into the building due to the opening of the window or door, and thus the listener L hears the aerodynamic sound. In this example, the timing of opening the window or door is equivalent to the predetermined timing, and the technology of the present disclosure is applied by assuming that the wind W is generated at the position of the window or door.
[0417] The object causing the wind W may be, for example, an object from which the wind W is blown out of a vent or an exhaust hole. In the wind W blown out of the vent or the exhaust hole, it is meaningless to correctly define the position where the wind W is generated in the virtual space, and the technology of the present disclosure can be applied by assuming that the wind W is generated at the position of the outlet of the vent or the exhaust hole. In this case, the prescribed timing may be determined by the administrator of the virtual space or the administrator of the sound signal processing device 100. For example, the receiving unit of the sound signal processing device 100 may receive the timing specified by the administrator, and the determining unit 120 may determine the timing received by the receiving unit as the prescribed timing.
[0418] In addition, the following aspects are also included in the scope of one or more technical aspects of the present disclosure.
[0419] (1) A part of the components constituting the above-mentioned sound signal processing device may also be a computer system consisting of a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. A computer program is stored in the RAM or hard disk unit. The above-mentioned microprocessor operates according to the above-mentioned computer program to achieve its function. Here, the computer program is composed of a plurality of command codes representing instructions to the computer in order to achieve a predetermined function.
[0420] (2) A part of the components constituting the above-mentioned sound signal processing device may also be constituted by a system LSI (Large Scale Integration). The system LSI is a super-multifunctional LSI manufactured by integrating multiple components into one chip. Specifically, it is a computer system composed of a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. The microprocessor operates according to the computer program, and the system LSI achieves its function.
[0421] (3) A part of the components constituting the above-mentioned sound signal processing device may also be constituted by an IC card or a single module that is detachable from each device. The above-mentioned IC card or the above-mentioned module is a computer system composed of a microprocessor, ROM, RAM, etc. The above-mentioned IC card or the above-mentioned module may also include the above-mentioned super-multifunctional LSI. The above-mentioned IC card or the above-mentioned module achieves its function through the microprocessor operating according to the computer program. The IC card or the module may also have tamper-resistant properties.
[0422] (4) In addition, a part of the components constituting the above-mentioned sound signal processing device may be a product in which the above-mentioned computer program or the above-mentioned digital signal is recorded on a recording medium that can be read by a computer, such as a floppy disk, a hard disk, a CD-ROM, an MO, a DVD, a DVD-ROM, a DVD-RAM, a BD (Blu-ray (registered trademark) Disc), a semiconductor memory, etc. In addition, it may be a digital signal recorded on these recording media.
[0423] Furthermore, part of the components constituting the sound signal processing device may be obtained by transmitting the computer program or the digital signal via an electric communication line, a wireless or wired communication line, a network represented by the Internet, data broadcasting, or the like.
[0424] (5) The present disclosure may be the methods described above. In addition, the present disclosure may be a computer program that implements the methods on a computer, or a digital signal composed of the computer program.
[0425] (6) Furthermore, the present disclosure may be a computer system including a microprocessor and a memory, wherein the memory stores the computer program and the microprocessor operates according to the computer program.
[0426] (7) In addition, the program or the digital signal may be recorded in the recording medium and transferred, or the program or the digital signal may be transferred via the network or the like, so that the program or the digital signal can be implemented by another independent computer system.
[0427] Industrial Applicability
[0428] The present disclosure can be utilized in an acoustic signal processing method and an acoustic signal processing device, and can be particularly applied to an acoustic system and the like.
[0429] Description of symbols
[0430] 100 Audio signal processing device
[0431] 110 Acquisition Department
[0432] 120 Decision Department
[0433] 130 Output section
[0434] 140 Storage
[0435] 200 Headphones
[0436] 201 Head sensor unit
[0437] 202 Output
[0438] 300 Display unit
[0439] 900 Rendering Department
[0440] 901 Reverberation Processing Unit
[0441] 902 Initial reflection processing unit
[0442] 903 Distance Attenuation Processing Unit
[0443] 904 Selection Department
[0444] 905 Binaural Processing Unit
[0445] 906 Computing Department
[0446] 907 Generation Department
[0447] A. Ambulance
[0448] A0000 Stereo sound reproduction system
[0449] A0001 Audio signal processing device
[0450] A0002 Sound prompt device
[0451] A0100 Encoding Device
[0452] A0101 Input data
[0453] A0102 Encoder
[0454] A0103 Encoded data
[0455] A0104 Memory
[0456] A0110 Decoding device
[0457] A0111 Sound signal
[0458] A0112 Decoder
[0459] A0113 Input data
[0460] A0114 Memory
[0461] A0120 Encoding device
[0462] A0121 Sending Department
[0463] A0122 Send signal
[0464] A0130 Decoding Device
[0465] A0131 Receiving Department
[0466] A0132 Receive signal
[0467] A0200 Decoder
[0468] A0201 Space Information Management Department
[0469] A0202 Sound Data Decoder
[0470] A0203 Rendering Department
[0471] A0210 Decoder
[0472] A0211 Space Information Management Department
[0473] A0213 Rendering Department
[0474] F Fan
[0475] L Listener
Claims
1. A method for processing an acoustic signal, in, include: an acquiring step of acquiring object information indicating a change in an object causing wind and a predetermined timing related to the change in the object; as well as The output step is to output aerodynamic sound data indicating aerodynamic sound caused by the wind after a predetermined time based on a change of the object has elapsed from the predetermined timing indicated by the acquired object information.
2. The sound signal processing method according to claim 1, in, The object information indicates: changes in the wind caused by changes in the object; and The predetermined timing is the timing of the wind change. The acoustic signal processing method includes a determination step of determining the predetermined time based on the wind indicated by the acquired object information.
3. The sound signal processing method according to claim 2, in, The change in the wind represented by the object information represents a change in the wind speed of the wind. In the determination step, the predetermined time is determined based on the wind speed.
4. The sound signal processing method according to claim 3, in, The aerodynamic sound is the sound generated by the changed wind speed.
5. The sound signal processing method according to claim 1, in, The object information indicates the location of the object, The acoustic signal processing method includes a determination step of determining the predetermined time based on a distance between a position of a listener of the aerodynamic sound and a position of the object indicated by the acquired object information.
6. The sound signal processing method according to claim 3, in, The object information indicates the location of the object, In the determination step, the predetermined time is determined based on the wind speed and the distance between the position of a listener of the aerodynamic sound and the position of the object indicated by the acquired object information.
7. The sound signal processing method according to claim 1, in, The object information indicates that the predetermined timing is a first timing for outputting the audio data associated with the object, In the output step, the pneumatic sound data is output after the predetermined time has elapsed from the first timing indicated by the acquired object information.
8. The sound signal processing method according to claim 1, in, The object information indicates: The location of the object; and The predetermined timing is a second timing at which the distance between the position of the listener of the pneumatic sound and the position of the object becomes shorter than a predetermined distance, In the output step, the pneumatic sound data is output after the predetermined time has elapsed from the second timing indicated by the acquired object information.
9. The sound signal processing method according to claim 1, in, The object information indicates: The change of the wind caused by the change of the object is a change of the direction of the wind; and The predetermined timing is a third timing at which the wind direction changes. In the output step, the pneumatic sound data is output after the predetermined time has elapsed from a third timing indicated by the acquired object information.
10. The sound signal processing method according to claim 6, in, The object is an object that generates the sound represented by the sound data associated with the object and the wind, The aerodynamic sound is aerodynamic sound generated by the wind generated by the object reaching the listener.
11. The sound signal processing method according to claim 10, in, When the distance is set to D, Let U be the distance from the position of the object where the wind speed is So, When the prescribed time is set to t, t satisfies the following formula: t={(D-U)^2} / {So×U×(log(D)-log(U))}.
12. The sound signal processing method according to claim 6, in, The object is an object that generates the wind by movement of the position of the object, The aerodynamic sound is aerodynamic sound generated when the wind generated by the movement reaches the listener.
13. The sound signal processing method according to claim 12, in, The predetermined timing indicated by the object information is a timing at which the amount of change in the distance with the passage of time changes from negative to positive.
14. The sound signal processing method according to claim 12, in, When the distance is set to D, Let U be the distance from the position of the object at which the wind speed of the wind generated by the movement is So, When the prescribed time is set to t, t satisfies the following formula: t={(D-U)^2} / {So×U×(log(D)-log(U))}.
15. A computer program for causing a computer to execute the sound signal processing method according to any one of claims 1 to 14.
16. An audio signal processing device, in, have: an acquisition unit that acquires object information indicating a change in an object causing wind and a predetermined timing related to the change in the object; and The output unit outputs aerodynamic sound data indicating aerodynamic sound caused by the wind after a predetermined time based on a change of the object has elapsed from the predetermined timing indicated by the acquired object information.
Citation Information
Patent Citations
Stereophonic sound calculation method, apparatus, program, recording medium, stereophonic sound presentation system, and virtual reality space presentation system
JP2013201577A
Apparatus and method for rendering a sound scene using pipeline stages
WO2021180938A1