ACOUSTIC ROOM CONSTRUCTION FACILITY, ACOUSTIC ROOM CONSTRUCTION SYSTEM, PROGRAM AND ACOUSTIC ROOM CONSTRUCTION METHOD

The sound chamber construction device and system address the challenge of reproducing a sound field at a fixed sound detection device by adapting to multiple sound sources through audio processing and superposition, ensuring accurate sound reproduction at varying listener positions.

DE112022007568B4Active Publication Date: 2026-05-07MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2022-09-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional technologies fail to reproduce a sound field accurately at a fixed sound detection device when multiple sound sources are present, especially with the Ambisonics method, as they cannot adapt to changes in the viewing/listening position.

Method used

A sound chamber construction device and system that includes audio acquisition, sound source determination, audio extraction, format conversion, position acquisition, motion processing, angle-distance adjustment, and superposition units to generate and superimpose stereophonic tones at a listening position, adapting to the movement of the listener relative to multiple sound sources.

Benefits of technology

Enables the reproduction of a sound field at a free position within a virtual space even with multiple sound sources, allowing for accurate sound reproduction regardless of the listener's movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Soundproof chamber construction equipment (100, 200, 300), comprising: an audio procurement unit (101, 201) that procures audio data that includes audio from a variety of sound sources; a sound source identification unit (102, 202) which, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit (103) that generates a multitude of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit (104) that generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; a position procurement unit (105) that procures a listening position as a position at which audio is listened to; a motion processing unit (106) that calculates an angle and distance between the listening position and each of the multitude of sound source positions; an angle-distance adjustment unit (107) that adjusts each of the plurality of stereophonic tones using the angle and distance according to each of the plurality of sound source positions, thereby producing a plurality of set stereophonic tones as a plurality of stereophonic tones at the listening position; and a superposition unit (108, 308) that superimposes the multitude of set stereophonic tones.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates to a soundproof enclosure design device, a soundproof enclosure design system, a program and a soundproof enclosure design method. TECHNICAL BACKGROUND

[0002] The development of stereophonic technology is currently underway. Using the Ambisonics method, for example, a sound field can be reproduced in 360 degrees from a single microphone position. An Ambisonics microphone is typically used to implement the Ambisonics method. If the Ambisonics microphone is fixed in place, the sound field cannot be reproduced at the location where the listener moves freely within the virtual space.

[0003] With regard to this question, patent reference 1 discloses a device suitable for correcting the directional characteristics of recorded directional audio in response to spatial data from a microphone system recording the directional audio. This device allows the directional characteristics of the directional audio to be corrected depending on the movement of a viewing / listening position.

[0004] WO 2021 / 111 030 A1 describes devices and methods for signaling changes in audio scenes with respect to audio objects within an audio scene.

[0005] US 2017 / 0 127 035 A1 describes an information reproduction procedure as well as an information recording procedure. REFERENCES ON THE STATE OF THE TECHNOLOGY PATENT REFERENCE

[0006] Patent reference 1: Publication of Japanese patent application no. 2022-509761 SUMMARY OF THE INVENTION TASK TO BE SOLVED BY THE INVENTION

[0007] However, with conventional technology, room tracking in the Ambisonics-B format with respect to the movement of the viewing / listening position cannot be performed when there are two or more sound sources.

[0008] Therefore, one objective of one or a multitude of aspects of the present disclosure is to enable the reproduction of the sound field at a free position in the state in which a sound detection device is fixed. MEANS TO SOLVE THE PROBLEM

[0009] A sound chamber construction device according to one aspect of the present disclosure comprises an audio acquisition unit that acquires audio data containing audio from a plurality of sound sources, a sound source determination unit that, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources, an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio, a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format, and a position acquisition unit.which procures a listening position as a position at which audio is listened to, a motion processing unit that calculates an angle and distance between the listening position and each of the multitude of sound source positions, an angle-distance setting unit that sets each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby generating a multitude of set stereophonic tones as a multitude of stereophonic tones at the listening position, and a superposition unit that superimposes the multitude of set stereophonic tones together.

[0010] A sound chamber construction system according to one aspect of the present disclosure is a sound chamber construction system comprising a sound chamber construction device and a sound acquisition device connected to the sound chamber construction device by a network and generating audio data comprising audio from a plurality of sound sources, wherein the sound chamber construction device comprises a communication unit that performs communication with the sound acquisition device, an audio acquisition unit that acquires the audio data via the communication unit, a sound source determination unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data, and an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data.with respect to each sound source, and the extraction audio data representing the extracted audio are generated; a format conversion unit that generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; a position acquisition unit that acquires a listening position as a position at which audio is listened to; a motion processing unit that calculates an angle and distance between the listening position and each of the multitude of sound source positions; an angle-distance adjustment unit that adjusts each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby generating a multitude of adjusted stereophonic tones as a multitude of stereophonic tones at the listening position.and a superposition unit that superimposes the multitude of set stereophonic tones.

[0011] A program according to one aspect of the present disclosure is a program that causes a computer to operate as an audio acquisition unit that acquires audio data comprising audio from a plurality of sound sources; a sound source determination unit that, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit that generates a plurality of stereophonic tones corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquisition unit;which procures a listening position as a position at which audio is listened to, a motion processing unit that calculates an angle and distance between the listening position and each of the multitude of sound source positions, an angle-distance setting unit that sets each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby generating a multitude of set stereophonic tones as a multitude of stereophonic tones at the listening position, and a superposition unit that superimposes the multitude of set stereophonic tones together.

[0012] A sound chamber construction method according to one aspect of the present disclosure comprises obtaining audio data containing audio from a plurality of sound sources, determining a plurality of sound source positions as positions of the plurality of sound sources based on the audio data, generating a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio, generating a plurality of stereophonic tones corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format, obtaining a listening position as a position at which audio is listened to, and calculating an angle and distance between the listening position and each of the plurality of sound source positions.Adjusting each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby generating a multitude of adjusted stereophonic tones as a multitude of stereophonic tones at the listening position, and superimposing the multitude of adjusted stereophonic tones with each other. IMPACT OF THE INVENTION

[0013] According to one or more aspects of the present disclosure, the sound field can be reproduced at a free position in the state in which a sound detection device is fixed. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a block diagram that schematically shows the configuration of a soundproof enclosure according to a first embodiment. Fig. Figure 2 is a block diagram that schematically shows the configuration of an audio extraction unit. Fig. Figure 3 is a block diagram that schematically shows the configuration of a computer. Fig. Figure 4 shows a first example to illustrate a processing example that accompanies a movement of a listening position. Fig. Figure 5 shows a second example to illustrate the processing example, which accompanies the movement of the listening position. Fig. Figure 6 shows a third example to illustrate the processing example, which accompanies the movement of the listening position. Fig. Figure 7 is a block diagram that schematically shows the configuration of a soundproof enclosure system according to a second embodiment. Fig. Figure 8 is a block diagram that schematically shows the configuration of a sound detection device in the second embodiment. Fig. Figure 9 is a block diagram that schematically shows the configuration of a soundproof enclosure in the second embodiment. Fig. Figure 10 is a block diagram that schematically shows the configuration of a soundproof enclosure construction device according to a third embodiment. MODE FOR EXECUTING THE INVENTION First embodiment

[0014] Fig. Figure 1 is a block diagram that schematically shows the configuration of a soundproof enclosure 100 according to a first embodiment.

[0015] The sound chamber construction device 100 comprises an audio acquisition unit 101, a sound source determination unit 102, an audio extraction unit 103, a format conversion unit 104, a position acquisition unit 105, a motion processing unit 106, an angle-distance adjustment unit 107, a superposition unit 108 and an output processing unit 109.

[0016] Audio Procurement Unit 101 procures audio data that includes audio from a variety of sound sources.

[0017] The Audio Acquisition Unit 101, for example, acquires audio data generated by a sound acquisition device (not shown), such as a microphone. While the audio in the audio data is intended to be recorded by an Ambisonics microphone (i.e., a microphone that supports the Ambisonics method), the audio in the audio data can also be recorded by a variety of omnidirectional microphones. Furthermore, the Audio Acquisition Unit 101 can acquire the audio data from a sound acquisition device via a connection interface (InterFace, not shown) or from a network, such as the internet, via a communication interface (Communication I / F, not shown). The acquired audio data is then made available to the Sound Source Identification Unit 102.

[0018] The sound source determination unit 102 determines a multitude of sound source positions based on the audio data as the positions of the multitude of sound sources.

[0019] The sound source determination unit 102, for example, performs a sound source number determination of determining the number of sound sources contained in the audio data and a sound source position determination of determining the sound source positions as the positions of the sound sources contained in the audio data.

[0020] A publicly known technology can be used to determine the number of sound sources. For example, Reference 1, listed below, describes a method for determining the number of sound sources using independent component analysis.

[0021] Furthermore, the Sound Source Identification Unit 102 can identify sound sources by analyzing an image represented by image data obtained from an image acquisition device, such as a camera (not shown), and determine the number of sound sources. In other words, the Sound Source Identification Unit 102 can determine the multiple positions of sound sources using an image obtained by photographing a room containing multiple sound sources. For example, the position of an object as a sound source can be determined based on the object's direction and size.

[0022] A publicly known technology can also be used for determining the position of a sound source. For example, Reference 2, listed below, describes a method for determining the position of a sound source using a beamforming method and a MUSIC method.

[0023] The audio data and the sound source count data, which indicate the number of sound sources obtained by performing the sound source count determination on the audio data, are provided to the audio extraction unit 103.

[0024] Sound source position data, which specify the sound source positions obtained through sound source position determination, are provided to the motion processing unit 106.

[0025] The audio extraction unit 103 generates a multitude of extraction audio data pieces by extracting audio, represented by the audio data, with respect to each sound source and generating the extraction audio data representing the extracted audio. The multitude of extraction audio data pieces corresponds to the multitude of sound sources.

[0026] For example, the audio extraction unit 103 extracts the extraction audio data from the audio data as audio data with respect to each sound source. Specifically, the audio extraction unit 103 generates the extraction audio data corresponding to a sound source contained within the multitude of sound sources from the multitude of extraction audio data pieces by subtracting from the audio data the data remaining after separating the audio from that one sound source. The extraction audio data is then provided to the format conversion unit 104.

[0027] Fig. Figure 2 is a block diagram that schematically shows the configuration of the audio extraction unit 103.

[0028] The audio extraction unit 103 comprises a noise reduction unit 110 and an extraction processing unit 111.

[0029] The noise reduction unit 110 reduces the noise in the audio data. Any publicly known technology can be used for the noise reduction process. For example, the noise reduction unit 110 can reduce the noise using a GSC (Global Sidelobe Canceller), which is described in Reference 5, listed later. The processed audio data obtained by reducing the noise in the audio data is provided to the extraction processing unit 111.

[0030] For example, the extraction processing unit 111 extracts the extraction audio data from the processed audio data as the audio data relating to each sound source.

[0031] The extraction processing unit 111 comprises a sound source separation unit 112, a phase adjustment unit 113 and a subtraction unit 114.

[0032] The sound source separation unit 112 generates separation audio data by separating the audio data with respect to each sound source from the processed audio data. A generally known technology can be used as the method for separating the audio data with respect to each sound source. For example, the sound source separation unit 112 performs the separation using a technology called ILRMA (Independent Low-Rank Matrix Analysis), which is described in Reference 3, cited below.

[0033] The phase adjustment unit 113 generates phase-aligned audio data by extracting a phase shift relative to each sound source in the signal processing used for sound source separation in the sound source separation unit 112, and applying a phase shift on the opposite side to the processed audio data to cancel the extracted phase shift. The phase-aligned audio data is then provided to the subtraction unit 114.

[0034] The subtraction unit 114 extracts the extraction audio data as the audio data with respect to each sound source by subtracting the phase-aligned audio data from the processed audio data with respect to each sound source.

[0035] Again with reference to Fig. 1. The format conversion unit 104 generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting the format of the multitude of pieces of extraction audio data into the format of stereophonic audio.

[0036] For example, the format conversion unit 104 converts the extraction audio data into a stereophonic audio format. In this example, the format conversion unit 104 generates stereophonic audio data representing the stereophonic tones by converting the format of the extraction audio data into the Ambisonics B format as a stereophonic audio format.

[0037] If the audio was recorded with an Ambisonics microphone, the format conversion unit 104 can convert the Ambisonics A format of the extracted audio data to the Ambisonics B format. A generally known technology can be used as the conversion method from Ambisonics A to Ambisonics B format. For example, a conversion method from Ambisonics A to Ambisonics B format is described in Reference 4, which is listed later.

[0038] In contrast, the format conversion unit 104 can convert the format of the extracted audio data into the Ambisonics-B format if the audio data was recorded by a variety of omnidirectional microphones, using a generally known technology. For example, a method for generating Ambisonics-B format audio data by creating bidirectionality through beamforming of the sound capture result by an omnidirectional microphone is described in Reference 5, cited below.

[0039] The position acquisition unit 105 acquires a listening position, defined as the position where audio is listened to. For example, the position acquisition unit 105 acquires the listening position by receiving the name of the listening position, where a user in a virtual space listens to the audio, from the user via an unseen input device such as a mouse or keyboard. In this example, it is assumed that the user can move within the virtual space, so the position acquisition unit 105 acquires the listening position periodically or whenever the user's movement is detected.

[0040] Subsequently, the position acquisition unit 105 provides the motion processing unit 106 with position data specifying the acquired listening position.

[0041] The motion processing unit 106 calculates an angle and a distance between the listening position and each of the multitude of sound source positions.

[0042] For example, the motion processing unit 106 calculates the angle and distance between the listening position and each sound source position based on the listening position specified by the position data and the sound source position specified by the sound source position data. The motion processing unit 106 then provides angle-distance data, indicating the calculated angle and distance relative to each sound source, to the angle-distance adjustment unit 107.

[0043] The angle-distance adjustment unit 107 adjusts each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby producing a multitude of set stereophonic tones as a multitude of stereophonic tones at the listening position.

[0044] For example, the angle-distance adjustment unit 107 adjusts the stereophonic sound data with respect to each sound source so that the angle and distance specified by the angle-distance data are satisfied.

[0045] The angle-distance adjustment unit 107, for example, is able to easily change the angle corresponding to the direction of arrival of sound from the sound source in the Ambisonics B format, according to Ambisonics specifications.

[0046] Furthermore, the angle-distance adjustment unit 107 adjusts the amplitude in the stereophonic sound data according to the distance specified by the angle-distance data. For example, if the distance between the listening position and the sound source is half the distance between the sound source and a recording position at the time the audio data was acquired, the angle-distance adjustment unit 107 increases the amplitude by 6 dB. In other words, the angle-distance adjustment unit 107 can adjust the relationship between distance and amplitude, for example, according to the square law.

[0047] The angle-distance adjustment unit 107 provides the superposition unit 108 with set stereophonic sound data, which represents the set stereophonic tones as the stereophonic tones where the angle and distance are set in relation to each sound source.

[0048] The superposition unit 108 superimposes the multitude of set stereophonic tones on each other.

[0049] The superposition unit 108, for example, superimposes the configured stereophonic sound data with respect to the respective sound sources. Specifically, the superposition unit 108 adds the sound signals that are represented by the configured stereophonic sound data with respect to the respective sound sources. In this process, the superposition unit 108 generates synthetic sound data that specifies the summed sound signals. The synthetic sound data is provided to the output processing unit 109.

[0050] The output processing unit 109 generates output sound data representing output tones by converting channel-based tones, represented by the synthetic sound data, into binaural tones for hearing with both ears. A generally known technology can be used as the method for converting the channel-based tones into binaural tones. For example, a method for converting channel-based tones into binaural tones is described in Reference 6, which is listed below.

[0051] The output processing unit 109 then outputs the audio data to an audio output device, such as a loudspeaker, via a connection I / F (not shown). Alternatively, the output processing unit 109 outputs the audio data to an audio output device, such as a loudspeaker, via a communication I / F (not shown).

[0052] The sound chamber construction device 100 described above can be implemented by a computer 10, as in Fig. 3 shown.

[0053] The computer 10 includes, for example, an auxiliary storage device 11 such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), a memory 12, a processor 13 such as a CPU (Central Processing Unit), an input I / F 14 such as a keyboard or a mouse, a connection I / F 15 such as USB (Universal Serial Bus) or similar, and a communication I / F 16 such as a NIC (Network Interface Card).

[0054] In particular, the audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superposition unit 108 and the output processing unit 109 can be implemented by the processor 13, which loads a program stored in the auxiliary memory device 11 into the memory 12 and executes the program.

[0055] The program can be downloaded from a recording medium via a read / write device (not shown) or from a network via the communication interface 16 to the auxiliary storage device 11, and then loaded into memory 12 and executed by the processor 13. The program can also be loaded directly from a recording medium via a read / write device or from a network via the communication interface 16 into memory 12 and then executed by the processor 13.

[0056] In the Ambisonics method, the direction of incidence of the sound from the sound source can be changed according to the user's viewing direction.

[0057] However, if there are multiple sound sources, such as a first sound source 20 and a second sound source 21, as in Fig. As shown in Figure 4, when a user 22 moves from a first listening position 23 to a second listening position 24, the angle between the user 22 and the first sound source 20 changes from angle θ1 to angle θ2 and the angle between the user 22 and the second sound source 21 changes from angle θ3 to angle θ4.

[0058] With the conventional Ambisonics method, it is not possible to change the angle relative to each sound source, as in Fig. 4 is shown, although a uniform change in angle, such as a change in the user's direction, is possible.

[0059] Therefore, in the first embodiment, the process is carried out by extracting the extraction audio data from the first sound source 20 and the extraction audio data from the second sound source 21 from the audio data, as for example in Fig. 5 and Fig. 6 shown.

[0060] More specifically, as in Fig. As shown in Figure 5, the first embodiment changes the angle between the user 22 and the first sound source 20 from a first angle θ1 to a second angle θ2 when the user 22 moves from the first listening position 23 to the second listening position 24. In the first embodiment, the intensity of the sound from the first sound source 20 also changes depending on the change from a first distance d1 between the first listening position 23 and the first sound source 20 to a second distance d2 between the second listening position 24 and the first sound source 20.

[0061] Furthermore, as in Fig. As shown in Figure 6, the first embodiment changes the angle between the user 22 and the second sound source 21 from a third angle θ3 to a fourth angle θ4 when the user 22 moves from the first listening position 23 to the second listening position 24. In the first embodiment, the intensity of the sound from the second sound source 21 also changes depending on the change in a third distance d3 between the first listening position 23 and the second sound source 21 to a fourth distance d4 between the second listening position 24 and the second sound source 21.

[0062] Then, the first embodiment modifies the sound that accompanies the user's movement by superimposing the data processed in relation to the respective sound sources as described above.

[0063] Therefore, according to the first embodiment, the sound field can be reproduced at a free position in the virtual space even if a large number of sound sources are present. Second embodiment

[0064] Fig. Figure 7 is a block diagram that schematically shows the configuration of a soundproof enclosure system 230 according to a second embodiment.

[0065] The sound chamber construction system 230 comprises a sound chamber construction device 200 and a sound recording device 240.

[0066] The sound chamber construction device 200 and the sound recording device 240 are connected to each other via a network 231, such as the Internet.

[0067] The sound recording device 240 records audio in a room separate from the sound chamber construction device 200 and transmits audio data representing the audio via the network 231 to the sound chamber construction device 200.

[0068] Fig. Figure 8 is a block diagram that schematically shows the configuration of the sound detection device 240.

[0069] The sound detection device 240 comprises a sound detection unit 241, a control unit 242 and a communication unit 243.

[0070] The sound recording unit 241 records audio in a room where the sound recording device 240 is installed. The sound recording unit 241 can consist, for example, of an Ambisonics microphone or a number of omnidirectional microphones.

[0071] The control unit 242 controls the processing in the sound detection device 240.

[0072] For example, the control unit 242 generates audio data representing the audio recorded by the sound detection unit 241 and transmits the audio data via the communication unit 243 to the sound chamber construction device 200.

[0073] When a direction for audio reception is instructed by the sound chamber design device 200 via the communication unit 243, the control unit 242 generates audio data representing audio signals from that direction by controlling the sound detection unit 241 and transmits the audio data to the sound chamber design device 200. This is a process by which beam shaping is performed by the sound chamber design device 200.

[0074] Part or all of the control unit 242 described above can consist of a memory and a processor such as a CPU (Central Processing Unit) that executes a program stored in memory, although this is not shown in the drawing. Such a program can be provided over a network or in the form of storage on a recording medium. Specifically, such a program can, for example, be provided as a program product.

[0075] Furthermore, part or all of the control unit 242 may also consist of a processing circuit, such as a single circuit, a combined circuit, a program-controlled processor, a program-controlled parallel processor, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), although this is not shown in the drawing.

[0076] As described above, the control unit 242 can be implemented by a processing circuit network.

[0077] The communication unit 243 communicates with the soundproof chamber construction unit 200 via the network 231.

[0078] For example, the communication unit 243 transmits the audio data via the network 231 to the sound chamber construction device 200.

[0079] Furthermore, the communication unit 243 receives an instruction from the sound chamber construction device 200 via the network 231 and provides the instruction to the control unit 242.

[0080] Here, the communication unit 243 can be implemented by a communication I / F, such as a NIC, although it is not shown in the drawing.

[0081] Fig. Figure 9 is a block diagram that schematically shows the configuration of the sound chamber construction device 200 in the second embodiment.

[0082] The sound room construction device 200 comprises an audio acquisition unit 201, a sound source determination unit 202, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superposition unit 108, the output processing unit 109 and a communication unit 220.

[0083] The audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superposition unit 108 and the output processing unit 109 in the sound chamber construction device 200 in the second embodiment are the same as the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superposition unit 108 and the output processing unit 109 in the sound chamber construction device 100 in the first embodiment.

[0084] The communication unit 220 communicates with the sound detection device 240 via the network 231.

[0085] For example, the communication unit 220 receives the audio data via the network 231 from the sound detection device 240.

[0086] The communication unit 220 also transmits an instruction to the sound detection device 240 via the network 231.

[0087] The communication unit 220 can, incidentally, be used by the in Fig. The 3 communication I / F 16 shown will be implemented.

[0088] The audio acquisition unit 201 acquires the audio data from the sound detection device 240 via the communication unit 220. The acquired audio data is made available to the sound source identification unit 202. In the second embodiment, the audio data represents the audio recorded by the sound detection device 240, which is connected to the sound chamber construction device 200 via the network 231.

[0089] The sound source identification unit 202 performs sound source count determination (determining the number of sound sources contained in the audio data) and sound source position determination (determining the positions of the sound sources contained in the audio data). Sound source count determination and sound source position determination can be performed using the same processes as in the first embodiment.

[0090] Furthermore, if the sound source determination unit 202 performs the sound source position determination, for example by means of the beam shaping method and the MUSIC method, the sound source determination unit 202 transmits an instruction, which specifies the direction for recording audio, to the sound detection device 240 via the communication unit 220.

[0091] As described above, according to the second embodiment, a virtual space can be constructed using audio transmitted from a remote location by installing the sound detection device 240 at the remote location. Third embodiment

[0092] Fig. Figure 10 is a block diagram that schematically shows the configuration of a soundproof enclosure 300 according to a third embodiment.

[0093] The sound room construction device 300 comprises an audio acquisition unit 101, a sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superposition unit 308, the output processing unit 109, a other audio acquisition unit 321 and an angle-distance adjustment unit 322.

[0094] The audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107 and the output processing unit 109 in the sound chamber construction device 300 in the third embodiment are the same as the audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107 and the output processing unit 109 in the sound chamber construction device 100 in the first embodiment.

[0095] However, the motion processing unit 106 provides the angular distance data to the angle-distance setting unit 322.

[0096] Other Audio Procurement Unit 321 procures audio data generated by a sound-collecting device (not shown), such as a microphone. The audio data procured by Other Audio Procurement Unit 321 is assumed to differ from the audio data procured by Audio Procurement Unit 101 in at least one aspect, namely the time and location of the recording. The audio data procured by Other Audio Procurement Unit 321 is also referred to as overlay-specific audio data.

[0097] Here it is assumed that the superposition-specific audio data are data that have undergone the separation with respect to the respective sound sources and the conversion into the Ambisonics-B format by the same processing as the processing by the sound source determination unit 102, the audio extraction unit 103 and the format conversion unit 104 in the first embodiment.

[0098] In other words, an Other Audio Procurement Unit 321 procures the overlay-specific audio data representing overlay-specific stereophonic sound as stereophonic sound produced by converting the audio data of the audio that differs from the audio contained in the audio data procured by the Audio Procurement Unit 101 at at least one of the time and position of the recording into the stereophonic audio format.

[0099] While the audio in the overlay-specific audio data is to be recorded by an Ambisonics microphone (i.e., a microphone that supports the Ambisonics method), the audio in the overlay-specific audio data can also be recorded by a variety of omnidirectional microphones. The Other Audio Procurement Unit 321 can also procure the audio data from a sound detection device via an I / F (not shown) or from a network such as the Internet via a Communication I / F (not shown). Furthermore, the Other Audio Procurement Unit 321 can also procure the overlay-specific audio data from a storage unit (not shown). The procured overlay-specific audio data is provided to the Angle-Distance Adjustment Unit 322.

[0100] The angle-distance adjustment unit 322 operates as a superposition-specific angle-distance adjustment unit, which generates a stereophonic tone at the listening position from the superposition-specific stereophonic tone.

[0101] The angle-distance adjustment unit 322 adjusts the superposition-specific audio data with respect to each sound source so that the angle and distance specified by the angle-distance data are satisfied. For example, if the superposition-specific audio data represents audio in the past at the same location as the audio in the audio data acquired by the audio acquisition unit 101, the angle-distance adjustment unit 322 can adjust the angle and amplitude according to the angle-distance data. The method for adjusting the angle and amplitude is the same as the adjustment method of the angle-distance adjustment unit 107 in the first embodiment.

[0102] In contrast, if the overlay-specific audio data represents audio at a location that differs from the location of the audio in the audio data procured by the audio procurement unit 101, a standard for setting the angle and amplitude with respect to each sound source was previously established according to the angle and distance specified by the angle-distance data, and the angle-distance setting unit 322 can set the angle and amplitude in the overlay-specific audio data according to the standard.

[0103] The angle-distance adjustment unit 322 provides the superposition unit 308 with superposition-specific set stereophonic audio data, which represent the superposition-specific set stereophonic tone as the superposition-specific stereophonic tone after setting the angle and distance with respect to each sound source.

[0104] The superposition unit 308 superimposes the multitude of set stereophonic tones and the superposition-specific set stereophonic tone together.

[0105] The superposition unit 308, for example, superimposes the configured stereophonic sound data for each sound source with the superposition-specific audio data. Specifically, the superposition unit 308 adds the sound signals represented by the configured stereophonic sound data for each sound source and a sound signal represented by the superposition-specific audio data. In this process, the superposition unit 308 generates the synthetic sound data that specifies the summed sound signals. The synthetic sound data is then provided to the output processing unit 109.

[0106] The Other Audio Procurement Unit 321 and the Angle Distance Adjustment Unit 322, described above, can also be replaced by the one described in Fig. 3 shown processor 13 is implemented, which loads a program stored in the auxiliary memory device 11 into memory 12 and executes the program.

[0107] As described above, according to the third embodiment, other audio that does not exist in reality can also be inserted into the virtual space, thereby increasing, for example, the perceived value of long-distance travel. Specifically, the user at the listening position in the virtual space can listen to audio from the past or to audio in a different space than the virtual space itself. For example, the user can listen to audio recordings of Shuri Castle, which no longer exists, in the virtual space. Referenz 1: Sawada et al., „Sound Source Number Estimation Method by Using Independent Component Analysis“, Proceedings of the Autumn Meeting of the Acoustical Society of Japan, 2004 Referenz 2: Futoshi Asano, „Array Signal Processing of Sound - Localization / Tracking and Separation of Sound Source“, Kapitel 4 und 5, Corona Publishing Co. Itd., 2011 Referenz 3: Kitamura et al., „Blind Source Separation Based on Independent Low-rank Matrix Analysis“, IEICE Technical Report, EA2017-56, vol.117, No.255, pp.73-80, Toyama, October 2017 Referenz 4: Ryouichi Nishimura „Ambisonics“, The Journal of the Institute of Image Information and Television Engineers, Vol. 68, No. 8, pp.616-620, 2014 Referenz 5: Japanisches Patent Nr. 6742535 Referenz 6: Japanisches Patent Nr. 4969978 BEZUGSZEICHENLISTE

[0108] 100, 200, 300: Soundproof enclosure construction unit, 101, 201: Audio acquisition unit, 102, 202: Sound source identification unit, 103: Audio extraction unit, 104: Format conversion unit, 105: Position acquisition unit, 106: Motion processing unit, 107: Angle-distance adjustment unit, 108, 308: Superposition unit, 109: Output processing unit, 110: Noise reduction unit, 111: Extraction processing unit, 112: Sound source separation unit, 113: Phase adjustment unit, 114: Subtraction unit, 220: Communication unit, 321: Other audio acquisition unit, 322: Angle-distance adjustment unit, 230: Soundproof enclosure construction system, 231: Network, 240: Sound detection device, 241: Sound detection unit, 242: Control unit, 243: Communication unit.

Claims

[1] Soundproof enclosure construction equipment (100, 200, 300), comprising: an audio procurement unit (101, 201) that procures audio data that includes audio from a variety of sound sources; a sound source identification unit (102, 202) which, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit (103) that generates a multitude of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit (104) that generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; a position procurement unit (105) that procures a listening position as a position at which audio is listened to; a motion processing unit (106) that calculates an angle and distance between the listening position and each of the multitude of sound source positions; an angle-distance adjustment unit (107) that adjusts each of the plurality of stereophonic tones using the angle and distance according to each of the plurality of sound source positions, thereby producing a plurality of set stereophonic tones as a plurality of stereophonic tones at the listening position; and a superposition unit (108, 308) that superimposes the multitude of set stereophonic tones. [2] Sound chamber construction device (100, 200, 300) according to claim 1, wherein the audio extraction unit (103) generates the extraction audio data corresponding to a sound source contained in the plurality of sound sources from the plurality of pieces of extraction audio data by subtracting from the audio data data that remain after separating the audio from the one sound source. [3] Sound room construction device (100, 200, 300) according to claim 1 or 2, wherein the sound source determination unit (102, 202) determines the plurality of sound source positions using an image obtained by photographing a room having the plurality of sound sources. [4] Sound chamber construction device (200) according to one of claims 1 to 3, wherein the audio data are data representing audio data recorded by a sound detection device (240) connected to the sound chamber construction device (200) via a network (231). [5] Sound chamber construction device (300) according to one of claims 1 to 4, further comprising: an other audio procurement unit (321) that procures overlay-specific audio data representing overlay-specific stereophonic sound as stereophonic sound produced by converting audio data of audio that differs from the audio contained in the audio data procured by the audio procurement unit (101) at at least one time and position of the recording into the format of stereophonic audio; and a superposition-specific angle-distance adjustment unit (322) that generates a stereophonic tone at the listening position from the superposition-specific stereophonic tone, wherein the superposition unit (308) superimposes the multitude of set stereophonic tones and the superposition-specific set stereophonic tones together. [6] Sound chamber construction system comprising a sound chamber construction device (100, 200, 300) and a sound acquisition device (240) which is connected to the sound chamber construction device (100, 200, 300) by a network (231) and generates audio data which includes audio from a plurality of sound sources, wherein the sound chamber construction device (100, 200, 300) comprises: a communication unit (243) that communicates with the sound detection device (240); an audio procurement unit (101, 201) that procures the audio data via the communication unit (240); a sound source identification unit (102, 202) which, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit (103) that generates a multitude of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit (104) that generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; a position procurement unit (105) that procures a listening position as a position at which audio is listened to; a motion processing unit (106) that calculates an angle and distance between the listening position and each of the multitude of sound source positions; an angle-distance adjustment unit (107) that adjusts each of the plurality of stereophonic tones using the angle and distance according to each of the plurality of sound source positions, thereby producing a plurality of set stereophonic tones as a plurality of stereophonic tones at the listening position; and a superposition unit (108, 308) that superimposes the multitude of set stereophonic tones. [7] Program that causes a computer (10) to operate as: an audio procurement unit (101, 201) that procures audio data which includes audio from a multitude of sound sources; a sound source identification unit (102, 202) which, based on the audio data, determines a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit (103) that generates a multitude of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit (104) that generates a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; a position procurement unit (105) that procures a listening position as a position at which audio is listened to; a motion processing unit (106) that calculates an angle and distance between the listening position and each of the multitude of sound source positions; an angle-distance adjustment unit (107) that adjusts each of the plurality of stereophonic tones using the angle and distance according to each of the plurality of sound source positions, thereby producing a plurality of set stereophonic tones as a plurality of stereophonic tones at the listening position; and a superposition unit (108, 308) that superimposes the multitude of set stereophonic tones. [8] Soundproof room design methods, including: Obtaining audio data that includes audio from a variety of sound sources; Determining a multitude of sound source positions as the positions of the multitude of sound sources based on the audio data; Generating a multitude of pieces of extraction audio data by extracting audio represented by the audio data in relation to each sound source and generating the extraction audio data representing the extracted audio; Generating a multitude of stereophonic tones corresponding to the multitude of sound sources by converting a format of the multitude of pieces of extraction audio data into a stereophonic audio format; Providing a listening position as a position from which audio is listened to; Calculating an angle and a distance between the listening position and each of the multiple sound source positions; Adjusting each of the multitude of stereophonic tones using the angle and distance according to each of the multitude of sound source positions, thereby producing a multitude of adjusted stereophonic tones as a multitude of stereophonic tones at the listening position; and Overlaying the multitude of set stereophonic tones with each other.

Citation Information

Patent Citations

  • JP002022509761A

  • Information reproducing apparatus and information reproducing method, and information recording apparatus and information recording method

    US20170127035A1

  • Audio scene change signaling

    WO2021111030A1