ACOUSTIC ROOM CONSTRUCTION FACILITY, ACOUSTIC ROOM CONSTRUCTION SYSTEM, PROGRAM AND ACOUSTIC ROOM CONSTRUCTION METHOD
The sound space construction apparatus and system address the challenge of space tracking multiple sound sources in the Ambisonics B format by determining sound source positions, extracting and converting audio data, and superimposing stereo sounds, resulting in accurate sound field reproduction as the listener moves within a virtual space.
Patent Information
- Application Number
- DE112022007568
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-09-28
AI Technical Summary
Conventional technologies are unable to perform space tracking of multiple sound sources in the Ambisonics B format as a listener moves within a virtual space.
A sound space construction apparatus and system that include an audio acquisition unit, sound source determination unit, audio extraction unit, format conversion unit, position acquisition unit, motion processing unit, angle-distance setting unit, and superimposing unit, which work together to determine sound source positions, extract and convert audio data, calculate angles and distances, and superimpose stereo sounds to maintain accurate sound field reproduction as the listener moves.
Enables the reproduction of a sound field at a free position within a virtual space even when the sound sensing device is fixed, effectively addressing the limitations of conventional Ambisonics B format technologies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a sound space design apparatus, a sound space design system, a program, and a sound space design method. TECHNICAL BACKGROUND
[0002] The development of stereophonic technology is currently underway. For example, using the Ambisonics technique, a sound field can be reproduced in 360-degree directions at a single microphone position. An Ambisonics microphone is typically used to implement the Ambisonics technique. If the Ambisonics microphone is fixed, the sound field cannot be reproduced at the position after the movement if a listener moves freely in virtual space.
[0003] With regard to this question, Patent Reference 1 discloses a device suitable for correcting the directional characteristics of recorded directional audio in response to spatial data from a microphone system that records the directional audio. With this device, the directional characteristics of the directional audio can be corrected depending on the movement of a viewing / listening position. PRIOR ART PATENT REFERENCES
[0004] Patent Reference 1: Publication of Japanese Patent Application No. 2022-509761 SUMMARY OF THE INVENTION TASK TO BE SOLVED BY THE INVENTION
[0005] However, with conventional technology, when there are two or more sound sources, Ambisonics B format room tracking cannot be performed with respect to the movement of the viewing / listening position.
[0006] An object of one or more aspects of the present disclosure is therefore to enable reproduction of the sound field at a free position in the state where a sound detecting device is fixed. MEANS TO SOLVE THE PROBLEM
[0007] A sound space construction device according to one aspect of the present disclosure includes an audio acquisition unit that acquires audio data including audio from a plurality of sound sources; a sound source determination unit that determines, based on the audio data, a plurality of sound source positions as positions of the plurality of sound sources; an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquisition unit;which acquires a listening position as a position at which audio is listened to, a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions, an angle-distance adjusting unit that adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position, and a superimposing unit that superimposes the plurality of adjusted stereophonic sounds on each other.
[0008] A sound space construction system according to one aspect of the present disclosure is a sound space construction system comprising a sound space construction device and a sound acquisition device connected to the sound space construction device through a network and generating audio data comprising audio from a plurality of sound sources, wherein the sound space construction device comprises a communication unit that carries out communication with the sound acquisition device, an audio acquisition unit that acquires the audio data via the communication unit, a sound source determination unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data, an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data,with respect to each sound source is extracted and the extraction audio data representing the extracted audio is generated; a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquisition unit that acquires a listening position as a position at which audio is listened to; a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions; an angle-distance adjustment unit that adjusts each of the plurality of stereophonic sounds using the angle and distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position;and a superposition unit that superimposes the plurality of set stereophonic tones with each other.
[0009] A program according to one aspect of the present disclosure is a program that causes a computer to operate as an audio acquiring unit that acquires audio data including audio from a plurality of sound sources, a sound source determining unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data, an audio extracting unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio, a format converting unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format, a position acquiring unit,which acquires a listening position as a position at which audio is listened to, a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions, an angle-distance adjusting unit that adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position, and a superimposing unit that superimposes the plurality of adjusted stereophonic sounds on each other.
[0010] A sound space construction method according to one aspect of the present disclosure includes acquiring audio data including audio from a plurality of sound sources, determining a plurality of sound source positions as positions of the plurality of sound sources based on the audio data, generating a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio, generating a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format, acquiring a listening position as a position at which audio is listened to, calculating an angle and a distance between the listening position and each of the plurality of sound source positions,Adjusting each of the plurality of stereophonic tones using the angle and distance corresponding to each of the plurality of sound source positions and thereby generating a plurality of adjusted stereophonic tones as a plurality of stereophonic tones at the listening position, and superimposing the plurality of adjusted stereophonic tones on each other. EFFECT OF THE INVENTION
[0011] According to one or a plurality of aspects of the present disclosure, the sound field can be reproduced at a free position in the state where a sound detecting device is fixed. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram schematically showing the configuration of a sound room construction device according to a first embodiment. Fig. Figure 2 is a block diagram schematically showing the configuration of an audio extraction unit. Fig. Figure 3 is a block diagram schematically showing the configuration of a computer. Fig. 4 shows a first example for explaining a processing example that accompanies a movement of a listening position. Fig. Figure 5 shows a second example to explain the processing example that accompanies the movement of the listening position. Fig. Figure 6 shows a third example to explain the processing example that accompanies the movement of the listening position. Fig. 7 is a block diagram schematically showing the configuration of a sound room construction system according to a second embodiment. Fig. 8 is a block diagram schematically showing the configuration of a sound detecting device in the second embodiment. Fig. 9 is a block diagram schematically showing the configuration of a sound room constructing device in the second embodiment. Fig. 10 is a block diagram schematically showing the configuration of a sound room construction device according to a third embodiment. MODE FOR CARRYING OUT THE INVENTIONFirst Embodiment
[0012] Fig. 1 is a block diagram schematically showing the configuration of a sound room construction device 100 according to a first embodiment.
[0013] The sound space construction device 100 includes an audio acquisition unit 101, a sound source determination unit 102, an audio extraction unit 103, a format conversion unit 104, a position acquisition unit 105, a motion processing unit 106, an angle-distance adjustment unit 107, a superimposition unit 108, and an output processing unit 109.
[0014] The audio acquiring unit 101 acquires audio data including audio from a plurality of sound sources.
[0015] The audio acquisition unit 101 acquires, for example, audio data generated by a sound detection device (not shown) such as a microphone. While the audio in the audio data is to be captured by an Ambisonics microphone as a microphone that supports the Ambisonics method, the audio in the audio data may also be captured by a plurality of omnidirectional microphones. Furthermore, the audio acquisition unit 101 may acquire the audio data from a sound detection device via a connection interface (connection I / F: Interface) not shown, or acquire the audio data from a network such as the Internet via a communication interface (communication I / F) not shown. The acquired audio data is provided to the sound source determination unit 102.
[0016] The sound source determining unit 102 determines a plurality of sound source positions as the positions of the plurality of sound sources based on the audio data.
[0017] The sound source determination unit 102 performs, for example, a sound source number determination of determining the number of sound sources included in the audio data and a sound source position determination of determining the sound source positions as the positions of the sound sources included in the audio data.
[0018] A publicly known technology can be used to determine the number of sound sources. For example, Reference 1, listed below, describes a method for determining the number of sound sources using independent component analysis.
[0019] Furthermore, the sound source determination unit 102 can identify the sound sources by analyzing an image represented by image data acquired by an image pickup device (not shown), such as a camera, and determine the number of sound sources. In other words, the sound source determination unit 102 can determine the plurality of sound source positions using an image obtained by photographing a space containing the plurality of sound sources. For example, the position of an object as a sound source can be determined based on a direction and size of the object.
[0020] Publicly known technology can also be used for sound source position detection. For example, Reference 2, listed below, describes a sound source position detection method using a beamforming method and a MUSIC method.
[0021] The audio data and the sound source number data indicating the sound source number obtained by performing the sound source number determination on the audio data are provided to the audio extraction unit 103.
[0022] Sound source position data indicating the sound source positions obtained by the sound source position detection is provided to the motion processing unit 106.
[0023] The audio extraction unit 103 generates a plurality of pieces of extraction audio data by extracting the audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio. The plurality of pieces of extraction audio data each correspond to the plurality of sound sources.
[0024] For example, the audio extraction unit 103 extracts the extraction audio data from the audio data as audio data related to each sound source. Specifically, the audio extraction unit 103 generates the extraction audio data corresponding to one sound source included in the plurality of sound sources from the plurality of pieces of extraction audio data by subtracting data remaining after the audio of one sound source is separated from the audio data. The extraction audio data is provided to the format conversion unit 104.
[0025] Fig. Figure 2 is a block diagram schematically showing the configuration of the audio extraction unit 103. The audio extraction unit 103 includes a noise reduction unit 110 and an extraction processing unit 111.
[0026] The noise reduction unit 110 reduces the noise in the audio data. Publicly known technology can be used for the noise reduction process. For example, the noise reduction unit 110 can reduce the noise using a GSC (Global Sidelobe Canceller) described in Reference 5 listed later. Processed audio data obtained by reducing the noise in the audio data is provided to the extraction processing unit 111.
[0027] For example, the extraction processing unit 111 extracts the extraction audio data from the processed audio data as the audio data related to each sound source.
[0028] The extraction processing unit 111 includes a sound source separation unit 112, a phase adjustment unit 113, and a subtraction unit 114.
[0029] The sound source separation unit 112 generates separation audio data by separating the audio data related to each sound source from the processed audio data. A well-known technology can be used as a method for separating the audio data related to each sound source. For example, the sound source separation unit 112 performs the separation using a technology called ILRMA (Independent Low-Rank Matrix Analysis), which is described in Reference 4 listed later.
[0030] The phase adjustment unit 113 generates phase-adjusted audio data by extracting a phase rotation given with respect to each sound source in the signal processing used for sound source separation in the sound source separation unit 112, and applying a phase rotation on the opposite side to the processed audio data to cancel the extracted phase rotation. The phase-adjusted audio data is provided to the subtraction unit 114.
[0031] The subtraction unit 114 extracts the extraction audio data as the audio data related to each sound source by subtracting the phase-adjusted audio data from the processed audio data related to each sound source.
[0032] Again with reference to Fig. 1, the format conversion unit 104 generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting the format of the plurality of pieces of extraction audio data into the format of stereophonic audio.
[0033] For example, the format conversion unit 104 converts the extracted audio data into a stereophonic audio format. In this example, the format conversion unit 104 generates stereophonic audio data representing the stereophonic sounds by converting the format of the extracted audio data into the Ambisonics-B format as the stereophonic audio format.
[0034] If the audio was captured by an Ambisonics microphone, the format conversion unit 104 can convert the Ambisonics A format of the extracted audio data to the Ambisonics B format. A well-known technology can be used as the conversion method from the Ambisonics A format to the Ambisonics B format. For example, a conversion method from the Ambisonics A format to the Ambisonics B format is described in Reference 5 listed later.
[0035] In contrast, when the audio data was captured by a plurality of omnidirectional microphones, the format conversion unit 104 can convert the format of the extracted audio data into the Ambisonics-B format using a well-known technology. A method for generating the Ambisonics-B format by generating bidirectionality by performing beamforming on the result of sound capture by an omnidirectional microphone is described, for example, in Reference 6 listed later.
[0036] The position acquisition unit 105 acquires a listening position as the position where audio is listened to. For example, the position acquisition unit 105 acquires the listening position by receiving the designation of the listening position where a user listens to the audio in a virtual space from the user via an input I / F (not shown), such as a mouse or keyboard. In this example, it is assumed that the user can move in the virtual space, so the position acquisition unit 105 acquires the listening position periodically or whenever a user's movement is detected.
[0037] Subsequently, the position acquisition unit 105 provides the motion processing unit 106 with position data indicating the acquired listening position.
[0038] The motion processing unit 106 calculates an angle and a distance between the listening position and each of the plurality of sound source positions.
[0039] For example, the motion processing unit 106 calculates the angle and distance between the listening position and each sound source position based on the listening position indicated by the position data and the sound source position indicated by the sound source position data. Then, the motion processing unit 106 provides angle-distance data indicating the calculated angle and distance with respect to each sound source to the angle-distance setting unit 107.
[0040] The angle-distance adjusting unit 107 adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, and thereby generates a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position.
[0041] For example, the angle-distance setting unit 107 sets the stereophonic sound data with respect to each sound source so that the angle and distance indicated by the angle-distance data are satisfied.
[0042] For example, the angle-distance adjustment unit 107 is capable of slightly changing the angle corresponding to an arrival direction of the sound from the sound source in the Ambisonics B format according to the specifications of Ambisonics.
[0043] Furthermore, the angle-distance adjustment unit 107 adjusts the amplitude in the stereophonic sound data according to the distance specified by the angle-distance data. For example, if the distance between the listening position and the sound source is 1 / 2 of the distance between the sound source and a recording position at the time the audio data was acquired, the angle-distance adjustment unit 107 increases the amplitude by 6 dB. In other words, the angle-distance adjustment unit 107 can adjust the relationship between the distance and the amplitude, for example, according to the square law.
[0044] The angle-distance setting unit 107 provides the superimposing unit 108 with adjusted stereophonic sound data representing the adjusted stereophonic sounds as the stereophonic sounds in which the angle and distance are adjusted with respect to each sound source.
[0045] The superimposing unit 108 superimposes the plurality of set stereophonic tones with each other.
[0046] For example, the superposition unit 108 superimposes the adjusted stereophonic sound data relating to the respective sound sources. Specifically, the superposition unit 108 adds the sound signals each represented by the adjusted stereophonic sound data relating to the respective sound sources. In this process, the superposition unit 108 generates synthetic sound data indicating the added sound signals. The synthetic sound data is provided to the output processing unit 109.
[0047] The output processing unit 109 generates output sound data representing output tones by converting channel-based tones represented by the synthesized sound data into binaural tones for listening with both ears. A well-known technology can be used as a method for converting the channel-based tones into binaural tones. For example, a method for converting the channel-based tones into binaural tones is described in Reference 7 below.
[0048] Subsequently, the output processing unit 109 outputs the output sound data to an audio output device such as a speaker via a connection I / F (not shown). Alternatively, the output processing unit 109 outputs the output sound data to an audio output device such as a speaker via a communication I / F (not shown).
[0049] The sound room construction device 100 described above can be implemented by a computer 10 such as the one shown in Fig. 3 shown.
[0050] The computer 10 includes, for example, an auxiliary storage device 11 such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), a memory 12, a processor 13 such as a CPU (Central Processing Unit), an input I / F 14 such as a keyboard or a mouse, a connection I / F 15 according to USB (Universal Serial Bus) or the like, and a communication I / F 16 such as a NIC (Network Interface Card).
[0051] Specifically, the audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superimposition unit 108, and the output processing unit 109 may be implemented by the processor 13, which loads a program stored in the auxiliary storage device 11 into the memory 12 and executes the program.
[0052] The program can be downloaded to the auxiliary storage device 11 from a recording medium via a reader / writer (not shown) or from a network via the communication I / F 16, and then loaded into the memory 12 and executed by the processor 13. The program can also be loaded directly from a recording medium via a reader / writer or from a network via the communication I / F 16 into the memory 12 and then executed by the processor 13.
[0053] With the Ambisonics process, the direction of incidence of the sound from the sound source can be changed according to the user's line of sight.
[0054] However, if a plurality of sound sources are present, such as a first sound source 20 and a second sound source 21, as in Fig. 4, when a user 22 moves from a first listening position 23 to a second listening position 24, the angle between the user 22 and the first sound source 20 changes from angle θ1 to angle θ2 and the angle between the user 22 and the second sound source 21 changes from angle θ3 to angle θ4.
[0055] With the conventional Ambisonics method, it is not possible to change the angle with respect to each sound source, as in Fig. 4, although a smooth change in angle, such as a change in the user's direction, is possible.
[0056] Therefore, in the first embodiment, the process is carried out by extracting the extraction audio data from the first sound source 20 and the extraction audio data from the second sound source 21 from the audio data, for example, as shown in Fig. 5 and Fig. 6 shown.
[0057] More specifically, as in Fig. As shown in Figure 5, the first embodiment changes the angle between the user 22 and the first sound source 20 from a first angle θ1 to a second angle θ2 as the user 22 moves from the first listening position 23 to the second listening position 24. In the first embodiment, the intensity of the sound from the first sound source 20 also changes depending on the change from a first distance d1 between the first listening position 23 and the first sound source 20 to a second distance d2 between the second listening position 24 and the first sound source 20.
[0058] Furthermore, as in Fig. As shown in Figure 6, the first embodiment changes the angle between the user 22 and the second sound source 21 from a third angle θ3 to a fourth angle θ4 as the user 22 moves from the first listening position 23 to the second listening position 24. In the first embodiment, the intensity of the sound from the second sound source 21 also changes depending on the change from a third distance d3 between the first listening position 23 and the second sound source 21 to a fourth distance d4 between the second listening position 24 and the second sound source 21.
[0059] Then, the first embodiment changes the sound accompanying the user's movement by superimposing the data processed with respect to the respective sound sources as described above.
[0060] Therefore, according to the first embodiment, the sound field can be reproduced at a free position in the virtual space even if a plurality of sound sources are present. Second embodiment
[0061] Fig. 7 is a block diagram schematically showing the configuration of a sound room construction system 230 according to a second embodiment.
[0062] The sound chamber construction system 230 includes a sound chamber construction device 200 and a sound detection device 240.
[0063] The sound space construction device 200 and the sound detection device 240 are connected to each other via a network 231, such as the Internet.
[0064] The sound acquisition device 240 records audio in a space separate from the sound space construction device 200 and transmits audio data representing the audio to the sound space construction device 200 via the network 231.
[0065] Fig. Fig. 8 is a block diagram schematically showing the configuration of the sound detecting device 240.
[0066] The sound detection device 240 comprises a sound detection unit 241, a control unit 242 and a communication unit 243.
[0067] The sound acquisition unit 241 records audio in a room in which the sound acquisition device 240 is installed. The sound acquisition unit 241 may consist of, for example, an Ambisonics microphone or a plurality of omnidirectional microphones.
[0068] The control unit 242 controls the processing in the sound detection device 240.
[0069] For example, the control unit 242 generates audio data representing the audio captured by the sound acquisition unit 241 and transmits the audio data to the sound room construction device 200 via the communication unit 243.
[0070] When a direction for recording audio is instructed from the sound chamber construction device 200 via the communication unit 243, the control unit 242 generates audio data representing audio signals from that direction by controlling the sound acquisition unit 241 and transmits the audio data to the sound chamber construction device 200. This is a process in which beamforming is performed by the sound chamber construction device 200.
[0071] Although not shown in the drawing, part or all of the control unit 242 described above may be formed of a memory and a processor such as a CPU (Central Processing Unit) that executes a program stored in the memory. Such a program may be provided via a network or stored on a recording medium. Specifically, such a program may be provided, for example, as a program product.
[0072] In addition, part or all of the control unit 242 may also be composed of a processing circuit, such as a single circuit, a combined circuit, a program-controlled processor, a program-controlled parallel processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array), although this is not shown in the drawing.
[0073] As described above, the control unit 242 may be implemented by a processing circuit network.
[0074] The communication unit 243 carries out communication with the sound room construction device 200 via the network 231.
[0075] For example, the communication unit 243 transmits the audio data to the sound room construction device 200 via the network 231.
[0076] Furthermore, the communication unit 243 receives an instruction from the sound room construction device 200 via the network 231 and provides the instruction to the control unit 242.
[0077] Here, the communication unit 243 may be implemented by a communication I / F such as a NIC, although it is not shown in the drawing.
[0078] Fig. 9 is a block diagram schematically showing the configuration of the sound space constructing device 200 in the second embodiment.
[0079] The sound space construction device 200 includes an audio acquisition unit 201, a sound source determination unit 202, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superimposition unit 108, the output processing unit 109, and a communication unit 220.
[0080] The audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superimposition unit 108, and the output processing unit 109 in the sound space construction device 200 in the second embodiment are the same as the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superimposition unit 108, and the output processing unit 109 in the sound space construction device 100 in the first embodiment.
[0081] The communication unit 220 communicates with the sound detection device 240 via the network 231.
[0082] For example, the communication unit 220 receives the audio data via the network 231 from the sound detection device 240.
[0083] The communication unit 220 further transmits an instruction to the sound detection device 240 via the network 231.
[0084] The communication unit 220 can be controlled by the Fig. 3 shown communication I / F 16 can be implemented.
[0085] The audio acquisition unit 201 acquires the audio data from the sound acquisition device 240 via the communication unit 220. The acquired audio data is provided to the sound source determination unit 202. In the second embodiment, the audio data is data representing the audio captured by the sound acquisition device 240, which is connected to the sound space construction device 200 via the network 231.
[0086] The sound source determination unit 202 performs sound source number determination of determining the number of sound sources included in the audio data and sound source position determination of determining the sound source positions as the positions of the sound sources included in the audio data. The sound source number determination and the sound source position determination can be performed according to the same processes as in the first embodiment.
[0087] Incidentally, when the sound source determination unit 202 performs the sound source position determination using, for example, the beamforming method and the MUSIC method, the sound source determination unit 202 transmits an instruction indicating the direction for recording audio to the sound detection device 240 via the communication unit 220.
[0088] As described above, according to the second embodiment, a virtual space can be constructed using audio transmitted from a remote location by installing the sound detection device 240 at the remote location. Third embodiment
[0089] Fig. 10 is a block diagram schematically showing the configuration of a sound space constructing device 300 according to a third embodiment.
[0090] The sound space construction device 300 includes an audio acquisition unit 101, a sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, the superimposition unit 308, the output processing unit 109, an other audio acquisition unit 321, and an angle-distance adjustment unit 322.
[0091] The audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, and the output processing unit 109 in the sound space construction device 300 in the third embodiment are the same as the audio acquisition unit 101, the sound source determination unit 102, the audio extraction unit 103, the format conversion unit 104, the position acquisition unit 105, the motion processing unit 106, the angle-distance adjustment unit 107, and the output processing unit 109 in the sound space construction device 100 in the first embodiment.
[0092] However, the motion processing unit 106 provides the angular distance data to the angle distance setting unit 322.
[0093] The other audio acquisition unit 321 acquires audio data generated by a sound pickup device (not shown), such as a microphone. The audio data acquired by the other audio acquisition unit 321 is assumed to differ from the audio data acquired by the audio acquisition unit 101 in at least one of the recording time and position. The audio data acquired by the other audio acquisition unit 321 is also referred to as overlay-specific audio data.
[0094] Here, it is assumed that the overlay-specific audio data is data that has undergone separation with respect to the respective sound sources and conversion into the Ambisonics B format by the same processing as the processing by the sound source determination unit 102, the audio extraction unit 103, and the format conversion unit 104 in the first embodiment.
[0095] In other words, a different audio acquiring unit 321 acquires the overlay-specific audio data representing overlay-specific stereophonic sound as stereophonic sound generated by converting the audio data of the audio different from the audio included in the audio data acquired by the audio acquiring unit 101 in at least one of the time and position of recording into the stereophonic audio format.
[0096] While the audio in the overlay-specific audio data is to be captured by an Ambisonics microphone as a microphone that supports the Ambisonics method, the audio in the overlay-specific audio data may also be captured by a plurality of omnidirectional microphones. The other audio acquisition unit 321 may also acquire the audio data from a sound detection device via an I / F (not shown) or acquire the audio data from a network such as the Internet via a communication I / F (not shown). Further, the other audio acquisition unit 321 may also acquire the overlay-specific audio data from a storage unit (not shown). The acquired overlay-specific audio data is provided to the angle-distance adjustment unit 322.
[0097] The angle-distance adjusting unit 322 functions as a superposition-specific angle-distance adjusting unit that generates superposition-specific adjusted stereophonic sound from the superposition-specific stereophonic sound as stereophonic sound at the listening position.
[0098] The angle-distance adjustment unit 322 adjusts the overlay-specific audio data with respect to each sound source so that the angle and distance specified by the angle-distance data are satisfied. For example, if the overlay-specific audio data represents audio in the past at the same location as the audio in the audio data acquired by the audio acquisition unit 101, the angle-distance adjustment unit 322 can adjust the angle and amplitude according to the angle-distance data. The method for adjusting the angle and amplitude is the same as the method of the angle-distance adjustment unit 107 in the first embodiment.
[0099] In contrast, when the overlay-specific audio data represents audio at a location different from the location of the audio in the audio data acquired by the audio acquisition unit 101, a standard for setting the angle and amplitude with respect to each sound source has been previously set according to the angle and distance indicated by the angle-distance data, and the angle-distance setting unit 322 can set the angle and amplitude in the overlay-specific audio data according to the standard.
[0100] The angle-distance adjustment unit 322 provides the superposition unit 308 with superposition-specific adjusted stereophonic audio data representing the superposition-specific adjusted stereophonic sound as the superposition-specific stereophonic sound after adjusting the angle and distance with respect to each sound source.
[0101] The superimposing unit 308 superimposes the plurality of set stereophonic tones and the superimposition-specific set stereophonic tone with each other.
[0102] The superimposition unit 308, for example, superimposes the adjusted stereophonic sound data related to the respective sound sources and the superimposition-specific adjusted audio data. Specifically, the superimposition unit 308 adds the sound signals represented by the adjusted stereophonic sound data related to the respective sound sources and a sound signal represented by the superimposition-specific adjusted audio data. In this process, the superimposition unit 308 generates the synthetic sound data indicating the added sound signals. The synthetic sound data is provided to the output processing unit 109.
[0103] The other audio acquiring unit 321 and the angle distance adjusting unit 322 described above can also be implemented by the Fig. 3, which loads a program stored in the auxiliary storage device 11 into the memory 12 and executes the program.
[0104] As described above, according to the third embodiment, other audio that doesn't exist in reality can also be inserted into the virtual space, thereby increasing the value of, for example, long-distance travel or similar experiences. Specifically, the user can listen to audio from the past at the listening position in the virtual space or listen to audio in a space other than the virtual space. For example, the user can listen to audio recordings of Shuri Castle, which no longer exists, in the virtual space. Referenz 1: Sawada et al., „Sound Source Number Estimation Method by Using Independent Component Analysis“, Proceedings of the Autumn Meeting of the Acoustical Society of Japan, 2004 Referenz 2: Futoshi Asano, „Array Signal Processing of Sound - Localization / Tracking and Separation of Sound Source“, Kapitel 4 und 5, Corona Publishing Co. Itd., 2011 Referenz 3: Futoshi Asano, „Array Signal Processing of Sound - Localization / Tracking and Separation of Sound Source“, Kapitel 4 und 5, Corona Publishing Co. Itd., 2011 Referenz 4: Kitamura et al., „Blind Source Separation Based on Independent Low-rank Matrix Analysis“, IEICE Technical Report, EA2017-56, vol.117, No.255, pp.73-80, Toyama, October 2017 Referenz 5: Ryouichi Nishimura „Ambisonics“, The Journal of the Institute of Image Information and Television Engineers, Vol. 68, No. 8, pp.616-620, 2014 Referenz 6: Japanisches Patent Nr. 6742535 Reference 7: Japanese Patent No. 4969978 LIST OF REFERENCE SYMBOLS
[0105] 100, 200, 300: Sound space construction device, 101, 201: Audio acquisition unit, 102, 202: Sound source determination unit, 103: Audio extraction unit, 104: Format conversion unit, 105: Position acquisition unit, 106: Motion processing unit, 107: Angle-distance adjustment unit, 108, 308: Superposition unit, 109: Output processing unit, 110: Noise reduction unit, 111: Extraction processing unit, 112: Sound source separation unit, 113: Phase adjustment unit, 114: Subtraction unit, 220: Communication unit, 321: Other audio acquisition unit, 322: Angle-distance adjustment unit, 230: Sound space construction system, 231: Network, 240: Sound detection device, 241: Sound detection unit, 242: Control unit, 243: Communication unit. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] JP 2022-509761
[0004] JP 6742535
[0104] JP 4969978
[0104] Cited non-patent literature
[0000] Sawada et al., “Sound Source Number Estimation Method by Using Independent Component Analysis,” Proceedings of the Autumn Meeting of the Acoustical Society of Japan, 2004
[0104] Futoshi Asano, “Array Signal Processing of Sound - Localization / Tracking and Separation of Sound Source”, Chapters 4 and 5, Corona Publishing Co. Itd., 2011
[0104] Kitamura et al., „Blind Source Separation Based on Independent Low-rank Matrix Analysis“, IEICE Technical Report, EA2017-56, vol.117, No.255, pp.73-80, Toyama, October 2017
[0104] Ryouichi Nishimura „Ambisonics“, The Journal of the Institute of Image Information and Television Engineers, Vol. 68, No. 8, pp.616-620, 2014
[0104]
Claims
[1] Acoustic chamber construction device, comprising: an audio acquisition unit that acquires audio data comprising audio from a plurality of sound sources; a sound source determining unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data; an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquiring unit that acquires a listening position as a position at which audio is listened to; a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions; an angle-distance adjusting unit that adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position; and a superposition unit that superimposes the multitude of set stereophonic tones. [2] The sound space constructing device according to claim 1, wherein the audio extraction unit generates the extraction audio data corresponding to one sound source included in the plurality of sound sources from the plurality of pieces of extraction audio data by subtracting data remaining after separating the audio from the one sound source from the audio data from the audio data. [3] The sound room constructing device according to claim 1 or 2, wherein the sound source determining unit determines the plurality of sound source positions using an image obtained by photographing a room having the plurality of sound sources. [4] The sound space construction device according to any one of claims 1 to 3, wherein the audio data is data representing audio data acquired by a sound detection device connected to the sound space construction device via a network. [5] Sound chamber construction device according to one of claims 1 to 4, further comprising: a different audio acquisition unit that acquires overlay-specific audio data representing overlay-specific stereophonic sound as stereophonic sound generated by converting audio data of audio different from the audio included in the audio data acquired by the audio acquisition unit in at least one of a time and a position of recording into the format of stereophonic audio; and a superposition-specific angle-distance adjustment unit that generates superposition-specific adjusted stereophonic sound from the superposition-specific stereophonic sound as stereophonic sound at the listening position, wherein the superimposing unit superimposes the plurality of set stereophonic tones and the superimposition-specific set stereophonic tones on each other. [6] A sound space construction system comprising a sound space construction device and a sound acquisition device connected to the sound space construction device by a network and generating audio data comprising audio from a plurality of sound sources, the sound space construction device comprising: a communication unit that carries out communication with the sound detection device; an audio acquisition unit that acquires the audio data via the communication unit; a sound source determining unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data; an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquiring unit that acquires a listening position as a position at which audio is listened to; a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions; an angle-distance adjusting unit that adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position; and a superposition unit that superimposes the multitude of set stereophonic tones. [7] Program that causes a computer to act as: an audio acquisition unit that acquires audio data comprising audio from a plurality of sound sources; a sound source determining unit that determines a plurality of sound source positions as positions of the plurality of sound sources based on the audio data; an audio extraction unit that generates a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; a format conversion unit that generates a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; a position acquiring unit that acquires a listening position as a position at which audio is listened to; a motion processing unit that calculates an angle and a distance between the listening position and each of the plurality of sound source positions; an angle-distance adjusting unit that adjusts each of the plurality of stereophonic sounds using the angle and the distance corresponding to each of the plurality of sound source positions, thereby generating a plurality of adjusted stereophonic sounds as a plurality of stereophonic sounds at the listening position; and a superposition unit that superimposes the multitude of set stereophonic tones. [8] Acoustic chamber design method, comprising: Obtaining audio data comprising audio from a variety of sound sources; Determining a plurality of sound source positions as positions of the plurality of sound sources based on the audio data; Generating a plurality of pieces of extraction audio data by extracting audio represented by the audio data with respect to each sound source and generating the extraction audio data representing the extracted audio; generating a plurality of stereophonic sounds corresponding to the plurality of sound sources by converting a format of the plurality of pieces of extraction audio data into a stereophonic audio format; Providing a listening position as a position at which audio is listened to; Calculating an angle and a distance between the listening position and each of the plurality of sound source positions; Adjusting each of the plurality of stereophonic tones using the angle and distance corresponding to each of the plurality of sound source positions and thereby generating a plurality of adjusted stereophonic tones as a plurality of stereophonic tones at the listening position; and Superimposing the multitude of set stereophonic tones with each other.
Citation Information
Patent Citations
JP002022509761A
Information reproducing apparatus and information reproducing method, and information recording apparatus and information recording method
US20170127035A1
Audio scene change signaling
WO2021111030A1