Live data distribution method, live data distribution system, live data distribution device, live data playback device, and live data playback method
By distributing sound source information and performing signal processing to recreate venue-specific reverberations, the system enhances the immersive experience by providing a realistic sense of presence at the live venue.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-04
AI Technical Summary
Existing live data distribution systems fail to provide the same sense of presence as a live venue when distributing live data.
The system distributes sound source information related to specific locations within a venue, along with positional information, and performs signal processing to localize and reproduce acoustic characteristics based on virtual space information, incorporating early and late reverberation effects to create a realistic sound field.
This approach provides a venue with a sense of realism equivalent to being present at the live venue, enhancing the immersive experience by accurately localizing sounds and reproducing venue-specific reverberations.
Smart Images

Figure 2026035902000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to a live data distribution method, a live data distribution system, a live data distribution device, a live data playback device, and a live data playback method. [Background technology]
[0002] Patent Document 1 discloses a game watching method that enables a user to effectively experience the excitement of a game as if they were in the stadium, using a terminal for watching sports games.
[0003] In the game watching method of Patent Document 1, reaction information indicating the user's reaction is transmitted from each user's terminal, and each user's terminal displays icon information based on the reaction information. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-024157 Summary of the Invention [Problem to be solved by the invention]
[0005] The system of Patent Document 1 only displays icon information, and does not provide the venue of the distribution destination with the same sense of presence as the live venue when distributing live data.
[0006] An embodiment of the present invention aims to provide a live data distribution method, a live data distribution system, a live data distribution device, a live data playback device, and a live data playback method that can provide the same sense of realism as a live venue at a distribution destination venue when live data is distributed. [Means for solving the problem]
[0007] The live data distribution method distributes, as distribution data, first sound source information related to sound of a first sound source occurring at a first location in a first venue and positional information of the first sound source, and second sound source information related to a second sound source occurring at a second location in the first venue, and renders the distribution data to provide, to a second venue, the sound of the first sound source that has been localized based on the positional information of the first sound source and the sound of the second sound source.The live data distribution method receives virtual space information at the second venue, receives the position of a sound-receiving point in the virtual space information, and performs signal processing on the sound of the first sound source based on a first signal processing model that reproduces the acoustic characteristics of the first sound source, thereby generating sound related to the reverberation of the space based on the virtual space information, the positional information of the first sound source, and the position of the sound-receiving point. [Effects of the Invention]
[0008] When live data is distributed, the live data distribution method can provide the venue at the distribution destination with the same sense of realism as the live venue. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing the configuration of a live data distribution system 1. FIG. [Figure 2] FIG. 2 is a schematic plan view of the first venue 10. [Figure 3] FIG. 2 is a schematic plan view of the second venue 20. [Figure 4] FIG. 2 is a block diagram showing the configuration of a mixer 11. [Figure 5] FIG. 2 is a block diagram showing the configuration of a distribution device 12. [Figure 6] 10 is a flowchart showing the operation of the distribution device 12. [Figure 7] FIG. 2 is a block diagram showing the configuration of a playback device 22. [Figure 8] 10 is a flowchart showing the operation of the playback device 22. [Figure 9] FIG. 10 is a block diagram showing the configuration of a live data distribution system 1A according to a first modification. [Figure 10]10 is a schematic plan view of a second venue 20 in a live data distribution system 1A according to a first modification. FIG. [Figure 11] FIG. 10 is a block diagram showing the configuration of a live data distribution system 1B according to a second modification. [Figure 12] FIG. 2 is a block diagram showing the configuration of an AV receiver 32. [Figure 13] FIG. 11 is a block diagram showing the configuration of a live data distribution system 1C according to a third modification. [Figure 14] FIG. 4 is a block diagram showing the configuration of a terminal 42. [Figure 15] FIG. 10 is a block diagram showing the configuration of a live data distribution system 1D according to a fourth modification. [Figure 16] FIG. 7 shows an example of live video 700 displayed on a playback device at each venue. [Figure 17] FIG. 10 is a block diagram showing an application example of signal processing performed in a playback device. [Figure 18] 1 is a schematic diagram showing the path of sound that travels from a sound source 70, reflects off a wall surface, and arrives at a sound receiving point 75. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0010] 1 is a block diagram showing the configuration of a live data distribution system 1. The live data distribution system 1 comprises a plurality of audio devices and information processing devices installed in a first venue 10 and a second venue 20, respectively.
[0011] Fig. 2 is a schematic plan view of the first venue 10, and Fig. 3 is a schematic plan view of the second venue 20. In this example, the first venue 10 is a live venue where performers perform. The second venue 20 is a public viewing venue where listeners in remote locations can watch the performers' performance.
[0012] A mixer 11, a distribution device 12, multiple microphones 13A-13F, multiple speakers 14A-14G, multiple trackers 15A-15C, and a camera 16 are installed in a first venue 10. A mixer 21, a playback device 22, a display 23, and multiple speakers 24A-24F are installed in a second venue 20. The distribution device 12 and the playback device 22 are connected via the Internet 5. Note that the numbers of microphones, speakers, and trackers are not limited to those shown in this embodiment. Furthermore, the installation manner of the microphones and speakers is not limited to the example shown in this embodiment.
[0013] The mixer 11 is connected to a distribution device 12, a plurality of microphones 13A to 13F, a plurality of speakers 14A to 14G, and a plurality of trackers 15A to 15C. The mixer 11, the plurality of microphones 13A to 13F, and the plurality of speakers 14A to 14G are connected via network cables or audio cables. The plurality of trackers 15A to 15C are connected to the mixer 11 via wireless communication. The mixer 11 and the distribution device 12 are connected via a network cable. The distribution device 12 is also connected to a camera 16 via a video cable. The camera 16 captures live video including the performers.
[0014] Multiple speakers 14A to 14G are installed along the walls of the first venue 10. In this example, the first venue 10 is rectangular in plan view. A stage is located in front of the first venue 10. Performers sing or play music on the stage. Speaker 14A is installed on the left side of the stage, speaker 14B is installed in the center of the stage, and speaker 14C is installed on the right side of the stage. Speaker 14D is installed on the left side of the center between the front and rear of the first venue 10, and speaker 14E is installed on the right side of the center between the front and rear of the first venue 10. Speaker 14F is installed on the left side of the rear of the first venue 10, and speaker 14G is installed on the right side of the rear of the first venue 10.
[0015] Microphone 13A is installed on the left side of the stage, microphone 13B is installed in the center of the stage, and microphone 13C is installed on the right side of the stage. Microphone 13D is installed on the left side of the center between the front and back of first venue 10, and microphone 13E is installed in the center behind first venue 10. Microphone 13F is installed on the right side of the center between the front and back of first venue 10.
[0016] Mixer 11 receives sound signals from microphones 13A to 13F. Mixer 11 also outputs sound signals to speakers 14A to 14G. In this embodiment, speakers and microphones are shown as examples of audio devices connected to mixer 11, but in reality, a large number of audio devices are connected to mixer 11. Mixer 11 receives sound signals from multiple audio devices such as microphones, performs signal processing such as mixing, and outputs sound signals to multiple audio devices such as speakers.
[0017] Microphones 13A to 13F each capture the singing or playing sounds of performers as sounds generated in first venue 10. Alternatively, microphones 13A to 13F capture environmental sounds of first venue 10. In the example of FIG. 2, microphones 13A to 13C capture the sounds of performers, and microphones 13D to 13F capture environmental sounds. Environmental sounds include sounds such as cheers, applause, calls, cheers, choruses, and murmurs from listeners. However, performer sounds may also be input via line input. Line input refers to inputting sound signals from an audio cable or the like connected to a sound source, rather than using a microphone to pick up and input sound output from a sound source such as an instrument. It is preferable that performer sounds be captured with a high signal-to-noise ratio and that no other sounds are included.
[0018] The speakers 14A to 14G output the sounds of the performers to the first venue 10. The speakers 14A to 14G may also output early reflected sounds or late reverberation sounds to control the sound field in the first venue 10.
[0019] The mixer 21 in the second venue 20 is connected to a playback device 22 and multiple speakers 24A to 24F. These audio devices are connected via a network cable or an audio cable. The playback device 22 is also connected to a display 23 via a video cable.
[0020] Multiple speakers 24A to 24F are installed along the walls of the second venue 20. In this example, the second venue 20 is rectangular in plan view. A display 23 is disposed in front of the second venue 20. The display 23 displays live video captured at the first venue 10. Speaker 24A is installed on the left side of the display 23, and speaker 24B is installed on the right side of the display 23. Speaker 24C is installed on the left side of the center between the front and rear of the second venue 20, and speaker 24D is installed on the right side of the center between the front and rear of the second venue 20. Speaker 24E is installed on the left side at the rear of the second venue 20, and speaker 24F is installed on the right side at the rear of the second venue 20.
[0021] The mixer 21 outputs sound signals to the speakers 24A to 24F. The mixer 21 receives sound signals from the playback device 22, performs signal processing such as mixing, and outputs the sound signals to a plurality of audio devices such as speakers.
[0022] The speakers 24A to 24F output the sounds of the performers to the second venue 20. The speakers 24A to 24F also output early reflection sounds or late reverberation sounds to reproduce the sound field of the first venue 10. The speakers 24A to 24F also output environmental sounds, such as cheers from the listeners in the first venue 10, to the second venue 20.
[0023] Fig. 4 is a block diagram showing the configuration of mixer 11. Since mixer 21 has the same configuration and functions as mixer 11, Fig. 4 shows the configuration of mixer 11 as a representative. Mixer 11 includes a display 101, a user I / F 102, an audio I / O (Input / Output) 103, a signal processing unit (DSP) 104, a network I / F 105, a CPU 106, a flash memory 107, and a RAM 108.
[0024] The CPU 106 is a control unit that controls the operation of the mixer 11. The CPU 106 performs various operations by reading out a predetermined program stored in a flash memory 107, which is a storage medium, into a RAM 108 and executing the program.
[0025] The program read by CPU 106 does not need to be stored in flash memory 107 within the device itself. For example, the program may be stored in a storage medium of an external device such as a server. In this case, CPU 106 simply reads the program from the server into RAM 108 and executes it each time.
[0026] The signal processing unit 104 is configured with a DSP for performing various signal processing. The signal processing unit 104 performs signal processing such as mixing and filtering on audio signals input from audio equipment such as a microphone via the audio I / O 103 or the network I / F 105. The signal processing unit 104 outputs the processed audio signals to audio equipment such as a speaker via the audio I / O 103 or the network I / F 105.
[0027] The signal processor 104 may also perform panning, early reflection sound generation, and late reverberation sound generation. Panning controls the volume of sound signals distributed to multiple speakers 14A-14G so that a sound image is localized at the performer's position. To perform panning, the CPU 106 acquires performer position information via trackers 15A-15C. The position information indicates two-dimensional or three-dimensional coordinates based on a certain position in the first venue 10. The trackers 15A-15C are tags that transmit and receive radio waves, such as Bluetooth (registered trademark). The performers or musical instruments are equipped with the trackers 15A-15C. At least three beacons are installed in the first venue 10 in advance. Each beacon measures its distance from the tracker 15A-15C based on the time difference between transmitting and receiving radio waves. The CPU 106 acquires the position information of the beacons in advance and measures the distances from at least three beacons to the tag, thereby being able to uniquely determine the positions of the trackers 15A to 15C.
[0028] In this way, CPU 106 acquires position information of each performer, i.e., position information of sounds generated in first venue 10, via trackers 15A-15C. Based on the acquired position information and the positions of speakers 14A-14G, CPU 106 determines the volume of each sound signal to be output to speakers 14A-14G so that sound images are localized at the performer's position. Signal processor 104 controls the volume of each sound signal to be output to speakers 14A-14G under the control of CPU 106. For example, signal processor 104 increases the volume of sound signals output to speakers closer to the performer's position and decreases the volume of sound signals output to speakers farther from the performer's position. In this way, signal processor 104 can localize the sound images of the performer's playing sounds and singing sounds at predetermined positions.
[0029] The early reflection sound generation process and late reverberation sound generation process involve convolving an impulse response with the performer's sound using an FIR filter. The signal processing unit 104 convolves the performer's sound with an impulse response, for example, acquired in advance at a predetermined venue (a venue other than the first venue 10). In this way, the signal processing unit 104 controls the sound field of the first venue 10. Alternatively, the signal processing unit 104 may control the sound field of the first venue 10 by further feeding back sounds acquired by microphones installed near the ceiling or walls of the first venue 10 to the speakers 14A to 14G.
[0030] The signal processing unit 104 outputs the sound of the performer and the information about the performer's position to the distribution device 12. The distribution device 12 acquires the sound of the performer and the information about the performer's position from the mixer 11.
[0031] The distribution device 12 also acquires a video signal from the camera 16. The camera 16 captures images of each performer or the entire first venue 10, and outputs a video signal relating to the live video to the distribution device 12.
[0032] Furthermore, the distribution device 12 acquires spatial reverberation information for the first venue 10. The spatial reverberation information is information for generating indirect sound. Indirect sound is sound that is generated when sound from a sound source reflects within the venue and reaches a listener, and includes at least early reflections and late reverberation. The spatial reverberation information includes, for example, information indicating the size, shape, and wall materials of the first venue 10, as well as an impulse response related to late reverberation. The information indicating the size, shape, and wall materials of the space is information for generating early reflections. The information for generating early reflections may be an impulse response. The impulse response is measured in advance, for example, in the first venue 10. The spatial reverberation information may also be information that changes depending on the position of a performer. The information that changes depending on the position of a performer is, for example, an impulse response measured in advance for each position of a performer in the first venue 10. For example, the distribution device 12 acquires a first impulse response when a performer's sound occurs in front of the stage of the first venue 10, a second impulse response when a performer's sound occurs on the left side of the stage, and a third impulse response when a performer's sound occurs on the right side of the stage. However, the number of impulse responses is not limited to three. Furthermore, the impulse responses do not need to be actually measured in the first venue 10; for example, they may be obtained by simulation based on the size, shape, and wall materials of the space of the first venue 10.
[0033] Note that early reflections are reflections from a fixed direction, while late reverberation is reflections from an unfixed direction. Late reverberation changes less with changes in the performer's position than does the early reflections. Therefore, the spatial reverberation information may be comprised of an impulse response of the early reflections that changes depending on the performer's position, and an impulse response of the late reverberation that is constant regardless of the performer's position.
[0034] The signal processing unit 104 may also acquire ambience information related to environmental sounds and output it to the distribution device 12. Environmental sounds are sounds acquired by the microphones 13D to 13F as described above, and include sounds such as background noise, listeners' cheers, applause, calls, cheers, choruses, and murmurs. However, environmental sounds may also be acquired by the microphones 13A to 13C on the stage. The signal processing unit 104 outputs sound signals related to environmental sounds as ambience information to the distribution device 12. The ambience information may also include positional information of the environmental sounds. Among the environmental sounds, cheers such as "Go for it!" from individual listeners, calls on the names of individual performers, and exclamations such as "Bravo" are sounds that can be recognized as the voices of individual listeners without being drowned out by the audience. The signal processing unit 104 may also acquire positional information of these individual sounds. The positional information of the environmental sounds can be obtained, for example, from sounds acquired by the microphones 13D to 13F. When the signal processing unit 104 recognizes the individual sounds by processing such as voice recognition, it finds the correlation between the sound signals of the microphones 13D to 13F and finds the difference in timing at which the individual sounds are picked up by the microphones 13D to 13F. Based on the difference in timing at which the sounds are picked up by the microphones 13D to 13F, the signal processing unit 104 can uniquely find the position in the first venue 10 at which the sound originated. Furthermore, the position information of the environmental sound may be considered as the position information of each of the microphones 13D to 13F.
[0035] The distribution device 12 encodes and distributes sound source information related to sounds generated in the first venue 10 and spatial reverberation information as distribution data. The sound source information includes at least the sounds of the performers, but may also include positional information of the performers' sounds. The distribution device 12 may also distribute ambience information related to environmental sounds included in the distribution data. The distribution device 12 may also distribute video signals related to the video of the performers included in the distribution data.
[0036] Alternatively, the distribution device 12 may distribute at least sound source information relating to the sound of the performer and the performer's position information, and ambience information relating to environmental sounds, as distribution data.
[0037] Fig. 5 is a block diagram showing the configuration of the distribution device 12. Fig. 6 is a flowchart showing the operation of the distribution device 12.
[0038] The distribution device 12 is an information processing device such as a general personal computer, etc. The distribution device 12 includes a display 201, a user I / F 202, a CPU 203, a RAM 204, a network I / F 205, a flash memory 206, and a general-purpose communication I / F 207.
[0039] CPU 203 reads a program stored in flash memory 206, which is a storage medium, into RAM 204 to implement a predetermined function. Note that the program read by CPU 203 does not need to be stored in flash memory 206 within its own device. For example, the program may be stored in a storage medium of an external device such as a server. In this case, CPU 203 simply reads the program from the server into RAM 204 and executes it each time.
[0040] The CPU 203 acquires the sound of the performers and information about the performers' positions (sound source information) from the mixer 11 via the network I / F 205 (S11). The CPU 203 also acquires information about the reverberation of the space in the first venue 10 (S12). The CPU 203 also acquires ambience information related to environmental sounds (S13). The CPU 203 may also acquire a video signal from the camera 16 via the general-purpose communication I / F 207.
[0041] The CPU 203 encodes and distributes data relating to the performer's sound and sound position information (sound source information), data relating to spatial reverberation information, data relating to ambience information, and data relating to the video signal as distribution data (S14).
[0042] The playback device 22 receives distribution data from the distribution device 12 via the Internet 5. The playback device 22 renders the distribution data and provides the second venue 20 with sounds related to the performers' sounds and the reverberation of the space. Alternatively, the playback device 22 provides the second venue 20 with sounds related to the performers' sounds and the reverberation of the space that corresponds to the ambience information. The playback device 22 may also provide the second venue 20 with sounds related to the reverberation of the space that corresponds to the ambience information.
[0043] Fig. 7 is a block diagram showing the configuration of the playback device 22. Fig. 8 is a flowchart showing the operation of the playback device 22.
[0044] The playback device 22 is an information processing device such as a general personal computer, etc. The playback device 22 includes a display 301, a user I / F 302, a CPU 303, a RAM 304, a network I / F 305, a flash memory 306, and a video I / F 307.
[0045] CPU 303 reads a program stored in flash memory 306, which is a storage medium, into RAM 304 to implement a predetermined function. Note that the program read by CPU 303 does not have to be stored in flash memory 306 within its own device. For example, the program may be stored in a storage medium of an external device such as a server. In this case, CPU 303 simply reads the program from the server into RAM 304 and executes it each time.
[0046] The CPU 303 receives distribution data from the distribution device 12 via the network I / F 305 (S21). The CPU 303 decodes the distribution data into sound source information, spatial reverberation information, ambience information, video signals, etc. (S22), and renders the sound source information, spatial reverberation information, ambience information, video signals, etc.
[0047] As an example of rendering the sound source information, the CPU 303 causes the mixer 21 to perform panning processing of the performer's sound (S23). As described above, panning processing is processing for localizing the performer's sound to the performer's position. The CPU 303 determines the volume of the sound signals to be distributed to the speakers 24A-24F so that the performer's sound is localized to the position indicated by the position information included in the sound source information. The CPU 303 causes the mixer 21 to perform panning processing by outputting the sound signals related to the performer's sound and information indicating the output amount of the sound signals related to the performer's sound to the speakers 24A-24F to the mixer 21.
[0048] This allows listeners in the second venue 20 to perceive sound as coming from the performer's position. For example, listeners in the second venue 20 can hear the sound of a performer on the right side of the stage in the first venue 10 from the front right side in the second venue 20. The CPU 303 may also render the video signal and display live video on the display 23 via the video I / F 307. This allows listeners in the second venue 20 to hear the panned sound of the performer while watching the video of the performer displayed on the display 23. This allows listeners in the second venue 20 to feel more immersed in the live performance because visual information and auditory information match.
[0049] Furthermore, the CPU 303 causes the mixer 21 to perform indirect sound generation processing as an example of rendering the spatial reverberation information (S24). The indirect sound generation processing includes an early reflection sound generation processing and a late reverberation sound generation processing. The early reflection sound is generated based on the performer's sound contained in the sound source information and information indicating the size, shape, wall material, etc. of the first venue 10 contained in the spatial reverberation information. The CPU 303 determines the arrival timing of the early reflection sound based on the size and shape of the space, and determines the level of the early reflection sound based on the wall material. More specifically, the CPU 303 calculates the coordinates of the wall surface on which the sound of the sound source is reflected based on the information on the size and shape of the space. Then, the CPU 303 calculates the position of a virtual sound source (imaginary sound source) that exists with the wall surface as a mirror surface relative to the position of the sound source based on the position of the sound source, the position of the wall surface, and the position of the sound receiving point. The CPU 303 calculates the delay amount of the imaginary sound source based on the distance from the position of the imaginary sound source to the sound receiving point. The CPU 303 also calculates the level of the imaginary sound source based on information about the material of the wall. The material information corresponds to the energy loss when the sound is reflected from the wall. Therefore, the CPU 303 calculates the level of the imaginary sound source by taking this energy loss into account in the sound signal of the sound source. By repeating this process, the CPU 303 can calculate the delay and level of the sound related to the spatial reverberation. The CPU 303 outputs the calculated delay and level to the mixer 21. The mixer 21 convolves the performer's sound with a level tap coefficient corresponding to the delay and level. In this way, the mixer 21 reproduces the spatial reverberation of the first venue 10 in the second venue 20. Furthermore, if the spatial reverberation information includes an impulse response of early reflection sounds, the CPU 303 causes the mixer 11 to execute a process of convolving the impulse response with the performer's sound using an FIR filter. The CPU 303 outputs the spatial reverberation information (impulse response) included in the distribution data to the mixer 21. The mixer 21 convolves the sound of the performer with the spatial reverberation information (impulse response) received from the playback device 22. In this way, the mixer 21 reproduces the spatial reverberation of the first venue 10 in the second venue 20.
[0050] Furthermore, if the spatial reverberation information changes depending on the performer's position, the playback device 22 outputs the spatial reverberation information corresponding to the performer's position to the mixer 21 based on the position information included in the sound source information. For example, if a performer who was in front of the stage in the first venue 10 moves to the left side of the stage, the impulse response convolved with the performer's sound changes from the first impulse response to the second impulse response. Alternatively, if a virtual sound source is reproduced based on information about the size and shape of the space, the delay amount and level are recalculated depending on the performer's position after movement. This allows the appropriate spatial reverberation corresponding to the performer's position to be reproduced in the second venue 20 as well.
[0051] The playback device 22 may also cause the mixer 21 to generate spatial reverberation corresponding to the environmental sound based on the ambience information and spatial reverberation information. In other words, the spatial reverberation may include a first reverberation corresponding to the performer's sound (sound from the first sound source) and a second reverberation corresponding to the environmental sound (sound from the second sound source). In this way, the mixer 21 reproduces the reverberation of the environmental sound in the first venue 10 in the second venue 20. Furthermore, if the ambience information includes position information, the playback device 22 may output spatial reverberation information corresponding to the position of the environmental sound to the mixer 11 based on the position information included in the ambience information. The mixer 21 reproduces the reverberation of the environmental sound based on the position of the environmental sound. For example, if an audience member who was located at the rear left of the first venue 10 moves to the rear right, the impulse response to be convolved with the cheers of that audience member is changed. Alternatively, when recreating a virtual sound source based on information about the size and shape of the space, the delay amount and level are recalculated depending on the position of the audience after they have moved. In this way, the reverb information of the space may include first reverb information that changes depending on the position of the performer's sound (first sound source) and second reverb information that changes depending on the position of the environmental sound (second sound source), and the rendering may include a process of generating the first reverb based on the first reverb information and a process of generating the second reverb based on the second reverb information.
[0052] Late reverberation sounds are reflected sounds that do not have a fixed direction of arrival. Late reverberation sounds change less with changes in sound position than early reflection sounds. Therefore, the playback device 22 may change only the impulse response of the early reflection sounds, which change depending on the performer's position, and keep the impulse response of the late reverberation sounds fixed.
[0053] The playback device 22 may omit the indirect sound generation process and use the reverberation of the second venue 20 as is. The indirect sound generation process may also be limited to the early reflection sound generation process. The late reverberation may use the reverberation of the second venue 20 as is. Alternatively, the mixer 21 may further reinforce the control of the second venue 20 by feeding back to the speakers 24A to 24F the sounds picked up by microphones (not shown) installed near the ceiling or walls of the second venue 20.
[0054] Then, the CPU 303 of the playback device 22 performs playback processing of the environmental sounds based on the ambience information (S25). The ambience information includes sound signals of sounds such as background noise, cheers from listeners, applause, calls, cheers, chorus, or murmurs. The CPU 303 outputs these sound signals to the mixer 21. The mixer 21 outputs the sound signals received from the playback device 22 to the speakers 24A to 24F.
[0055] When the ambience information includes position information of the environmental sound, the CPU 303 causes the mixer 21 to perform a panning process to localize the environmental sound. In this case, the CPU 303 determines the volume of the sound signals to be distributed to the speakers 24A-24F so that the environmental sound is localized at the position of the position information included in the ambience information. The CPU 303 causes the mixer 21 to perform a panning process by outputting the sound signals of the environmental sound and information indicating the output amount of the sound signals related to the environmental sound to the speakers 24A-24F to the mixer 21. The same applies when the position information of the environmental sound is the position information of each of the microphones 13D-13F. The CPU 303 determines the volume of the sound signals to be distributed to the speakers 24A-24F so that the environmental sound is localized at the microphone positions. Each of the microphones 13D-13F collects a plurality of environmental sounds (second sound sources), such as background noise, applause, chorus, cheers such as "Wow!", and murmurs. The sound of each sound source reaches the microphone with a predetermined delay and level. That is, background noise, applause, chorus, cheers such as "Wow!", and murmurs also reach the microphone as individual sound sources with a predetermined delay and level (information for localizing the sound source). The CPU 303 performs panning processing so that the sound picked up by the microphone is localized at the position of the microphone, thereby easily reproducing the localization of each sound source.
[0056] For sounds emitted simultaneously by many listeners that cannot be recognized as the voices of individual listeners, the CPU 303 may perform effects such as reverb on the mixer 21 to create a sense of spatial expansion. For example, background noise, applause, singing together, cheers such as "Wow!", and murmurs are sounds that resonate throughout a live music venue. The CPU 303 causes the mixer 21 to perform effects that create a sense of spatial expansion on these sounds.
[0057] The playback device 22 may provide environmental sounds based on the above-described ambience information to the second venue 20. This allows listeners at the second venue 20 to enjoy a more realistic live performance, as if they were watching a live performance at the first venue 10.
[0058] As described above, the live data distribution system 1 of this embodiment distributes sound source information related to sounds generated in the first venue 10 and spatial reverberation information as distribution data, renders the distribution data, and provides the sounds related to the sound source information and spatial reverberation to the second venue 20. This makes it possible to provide the sense of presence of the live venue to the distribution destination venue as well.
[0059] Furthermore, the live data distribution system 1 distributes, as distribution data, sound from a first sound source (e.g., sound from a performer) generated at a first location (e.g., a stage) in the first venue 10, first sound source information relating to position information of the first sound source, and second sound source information relating to a second sound source (e.g., environmental sound) generated at a second location (e.g., a location where listeners are) in the first venue 10, and renders the distribution data to provide the sound of the first sound source, which has been localized based on the position information of the first sound source, and the sound of the second sound source, to the second venue. This makes it possible to provide the venue to which the sound is distributed with the same sense of presence as at the live venue.
[0060] Next, Fig. 9 is a block diagram showing the configuration of a live data distribution system 1A according to Modification 1. Fig. 10 is a schematic plan view of a second venue 20 in the live data distribution system 1A according to Modification 1. Components common to Figs. 1 and 3 are given the same reference numerals, and descriptions thereof will be omitted.
[0061] A plurality of microphones 25A to 25C are installed in the second venue 20 of the live data distribution system 1A. When facing the stage 80 of the second venue 20, the microphone 25A is installed on the left side of the center between the front and rear of the second venue 20, and the microphone 25B is installed in the center behind the second venue 20. The microphone 25C is installed on the right side of the center between the front and rear of the second venue 20.
[0062] The microphones 25A to 25C acquire environmental sounds in the second venue 20. The mixer 21 outputs sound signals of the environmental sounds as ambience information to the playback device 22. The ambience information may include position information of the environmental sounds. As described above, the position information of the environmental sounds can be obtained from the sounds acquired by the microphones 25A to 25C, for example.
[0063] The playback device 22 transmits ambience information related to the environmental sounds generated in the second venue 20 to another venue as a third sound source. For example, the playback device 22 feeds back the environmental sounds generated in the second venue 20 to the first venue 10. This allows the performers on the stage in the first venue 10 to hear the voices, applause, cheers, etc. of people other than the listeners in the first venue 10, allowing them to perform a live performance in a more realistic environment. Furthermore, the listeners in the first venue 10 can also hear the voices, applause, cheers, etc. of listeners in the other venues, allowing them to watch and listen to the live performance in a more realistic environment.
[0064] Furthermore, if a playback device at another venue renders the distribution data and provides the sound from the first venue to that other venue, and also provides the environmental sounds generated at the second venue 20 to that other venue, listeners at that other venue will also be able to hear the voices, applause, cheers, etc. of many listeners, and will be able to watch the live performance in a realistic environment.
[0065] Next, Fig. 11 is a block diagram showing the configuration of a live data distribution system 1B according to Modification 2. Components common to Fig. 1 are given the same reference numerals, and description thereof will be omitted.
[0066] In the live data distribution system 1B, the distribution device 12 is connected to an AV receiver 32 in a third venue 20A via the Internet 5. The AV receiver 32 is connected to a display 33, multiple speakers 34A-34F, and a microphone 35. The third venue 20A is, for example, the home of a listener. The AV receiver 32 is an example of a playback device. The user of the AV receiver 32 is a listener who remotely watches the live performance in the first venue 10.
[0067] 12 is a block diagram showing the configuration of the AV receiver 32. The AV receiver 32 includes a display 401, a user I / F 402, an audio I / O (Input / Output) 403, a signal processing unit (DSP) 404, a network I / F 405, a CPU 406, a flash memory 407, a RAM 408, and a video I / F 409.
[0068] The CPU 406 is a control unit that controls the operation of the AV receiver 32. The CPU 406 performs various operations by loading predetermined programs stored in a flash memory 407, which is a storage medium, into a RAM 408 and executing the programs.
[0069] The programs read by CPU 406 do not need to be stored in the flash memory 407 of the device itself. For example, the programs may be stored in a storage medium of an external device such as a server. In this case, CPU 406 simply reads the programs from the server into RAM 408 and executes them each time.
[0070] The signal processing unit 404 is configured with a DSP for performing various signal processing. The signal processing unit 404 performs signal processing on an audio signal input via the audio I / O 403 or the network I / F 405. The signal processing unit 404 outputs the processed audio signal to an acoustic device such as a speaker via the audio I / O 403 or the network I / F 405.
[0071] The AV receiver 32 performs processing similar to that performed by the mixer 21 and the playback device 22. The CPU 406 receives distribution data from the distribution device 12 via the network I / F 405. The CPU 406 renders the distribution data and provides the sound of the performers and the sound related to the reverberation of the space to the third venue 20A. Alternatively, the CPU 406 may render the distribution data and provide the environmental sound generated in the first venue 10 to the third venue 20A. Alternatively, the CPU 406 may render the distribution data and display live video on the display 33 via the video I / F 307.
[0072] The signal processing unit 404 performs panning processing of the performer's sound. The signal processing unit 404 also performs processing to generate indirect sound. Alternatively, the signal processing unit 404 may perform panning processing of environmental sound.
[0073] This allows the AV receiver 32 to provide the same sense of realism as in the first venue 10 to the third venue 20A.
[0074] The AV receiver 32 also acquires environmental sounds from the third venue 20A (such as cheers, applause, and calls from listeners) via the microphone 35. The AV receiver 32 transmits the environmental sounds from the third venue 20A to other devices. For example, the AV receiver 32 feeds back the environmental sounds from the third venue 20A to the first venue 10.
[0075] In this way, by feeding back sounds from multiple listeners to the first venue 10, the performers on the stage at the first venue 10 can hear the cheers, applause, cheers, etc. of many listeners other than the listeners at the first venue 10, allowing for a live performance in an environment that feels more realistic. Additionally, listeners at the first venue 10 can hear the cheers, applause, cheers, etc. of many listeners in remote locations, allowing for a live performance to be viewed in an environment that feels more realistic.
[0076] Alternatively, the AV receiver 32 may receive reactions from listeners by displaying icon images such as "cheering," "applause," "calling out," and "bustling" on the display 401 and receiving a selection operation for these icon images from the listener via the user I / F 402. When the AV receiver 32 receives a selection operation for these reactions, it may generate sound signals corresponding to the respective reactions and transmit them to another device as ambience information.
[0077] Alternatively, the AV receiver 32 may transmit, as ambience information, information indicating the type of environmental sound, such as listeners' cheers, applause, or calls. In this case, the receiving device (e.g., the distribution device 12 and the mixer 11) generates a corresponding sound signal based on the ambience information and provides the sounds of listeners' cheers, applause, or calls in the venue. In this way, the ambience information is not a sound signal of the environmental sound, but information indicating the sound to be generated, and the distribution device 12 and the mixer 11 may play back pre-recorded environmental sounds.
[0078] Furthermore, the ambience information for the first venue 10 may be pre-recorded environmental sounds rather than environmental sounds generated in the first venue 10. In this case, the distribution device 12 distributes information indicating the sounds to be generated as the ambience information. The playback device 22 or the AV receiver 32 plays the corresponding environmental sounds based on the ambience information. Furthermore, among the ambience information, background noise and commotion may be recorded sounds, while other environmental sounds (e.g., cheers, applause, calls from listeners) may be sounds generated in the first venue 10.
[0079] The AV receiver 32 may also receive position information of the listener via the user I / F 402. The AV receiver 32 displays an image simulating a plan view or a perspective view of the first venue 10 on the display 401 or the display 33, and receives position information from the listener via the user I / F 402 (see, for example, FIG. 16 ). The position information is information specifying an arbitrary position within the first venue 10. The AV receiver 32 transmits the received listener position information to the first venue 10. The distribution device 12 and mixer 11 at the first venue perform processing to localize the environmental sound of the third venue 20A to the specified position based on the environmental sound of the third venue 20A and the listener position information received from the AV receiver 32.
[0080] The AV receiver 32 may also change the content of the panning process based on position information received from the user. For example, if the listener specifies a position immediately in front of the stage in the first venue 10, the AV receiver 32 will set the localization position of the performer's sound to a position immediately in front of the listener and perform panning. This allows the listener in the third venue 20A to experience the same sense of presence as if they were directly in front of the stage in the first venue 10.
[0081] The sound of the listener at the third venue 20A may be transmitted to the second venue 20 instead of the first venue 10, or may be transmitted to other venues. For example, the sound of the listener at the third venue 20A may be transmitted only to a friend's home (venue 4). The listener at the fourth venue can watch the live performance at the first venue 10 while listening to the sound of the listener at the third venue 20A. Furthermore, a playback device (not shown) at the fourth venue may transmit the sound of the listener at the fourth venue to the third venue 20A. In this case, the listener at the third venue 20A can watch the live performance at the first venue 10 while listening to the sound of the listener at the fourth venue. This allows the listener at the third venue 20A and the listener at the fourth venue to watch the live performance at the first venue 10 while conversing with each other.
[0082] 13 is a block diagram showing the configuration of a live data distribution system 1C according to Modification 3. Components common to those in FIG. 1 are given the same reference numerals, and descriptions thereof will be omitted.
[0083] In the live data distribution system 1C, the distribution device 12 is connected to a terminal 42 at a fifth venue 20B via the Internet 5. The terminal 42 is connected to headphones 43. The fifth venue 20B is, for example, a listener's home. However, if the terminal 42 is portable, the fifth venue 20B may be any location, such as a cafe, a car, or public transportation. In this case, any location can become the fifth venue 20B. The terminal 42 is an example of a playback device. The user of the terminal 42 is a listener who remotely watches the live performance at the first venue 10. In this case, the terminal 42 renders the distribution data and provides sounds related to the sound source information and sounds related to the reverberation of the space to the second venue (the fifth venue 20B in this example) via the headphones 43.
[0084] 14 is a block diagram showing the configuration of the terminal 42. The terminal 42 is an information processing device such as a personal computer, a smartphone, or a tablet computer. The terminal 42 includes a display 501, a user I / F 502, a CPU 503, a RAM 504, a network I / F 505, a flash memory 506, an audio I / O (Input / Output) 507, and a microphone 508.
[0085] The CPU 503 is a control unit that controls the operation of the terminal 42. The CPU 503 performs various operations by reading out a predetermined program stored in a flash memory 506, which is a storage medium, into the RAM 504 and executing the program.
[0086] The programs read by CPU 503 do not need to be stored in the flash memory 506 of the device itself. For example, the programs may be stored in a storage medium of an external device such as a server. In this case, CPU 503 simply reads the programs from the server into RAM 504 and executes them each time.
[0087] The CPU 503 performs signal processing on the audio signal input via the network I / F 505. The CPU 503 outputs the processed audio signal to the headphones 43 via the audio I / O 507.
[0088] The CPU 503 receives distribution data from the distribution device 12 via the network I / F 505. The CPU 503 renders the distribution data and provides the sounds of the performers and the sounds related to the reverberation of the space to the listeners in the fifth venue 20B.
[0089] Specifically, CPU 503 convolves a head-related transfer function (hereinafter referred to as HRTF) into the sound signal related to the sound of the performer, and performs sound image localization processing (binaural processing) so that the sound of the performer is localized to the position of the performer. HRTF corresponds to the transfer function between a predetermined position and the ears of the listener. HRTF is a transfer function that expresses the volume, arrival time, frequency characteristics, etc. of sound from a sound source at a certain position to each of the left and right ears. CPU 503 convolves the HRTF into the sound signal of the performer's sound based on the performer's position. As a result, the performer's sound is localized to a position corresponding to the position information.
[0090] The CPU 503 also performs indirect sound generation processing by binaural processing, which convolves the sound signal of the performer's sound with HRTFs corresponding to the reverberation information of the space. The CPU 503 localizes the early reflection sounds and the late reverberation sounds by convolving the HRTFs from the positions of virtual sound sources corresponding to each early reflection sound included in the reverberation information of the space to the left and right ears, respectively. However, the late reverberation sounds are reflected sounds whose arrival direction is not determined. Therefore, the CPU 503 may perform effect processing such as reverb on the late reverberation sounds without performing localization processing. The CPU 503 may also perform digital filtering processing (headphone inverse characteristic processing) that reproduces the inverse acoustic characteristics of the headphones 43 used by the listener.
[0091] Furthermore, CPU 503 renders the ambience information of the distribution data and provides the environmental sounds generated in first venue 10 to the listener in fifth venue 20B. If the ambience information includes positional information of the environmental sounds, CPU 503 performs localization processing using HRTF, and performs effect processing on sounds whose arrival direction is uncertain.
[0092] Furthermore, the CPU 503 may render the video signal of the distribution data and display live video on the display device 501.
[0093] This allows the terminal 42 to provide the listeners in the fifth venue 20B with the same sense of presence as in the first venue 10.
[0094] The terminal 42 also acquires the sounds of the listeners in the fifth venue 20B via the microphone 508. The terminal 42 transmits the sounds of the listeners to another device. For example, the terminal 42 feeds back the sounds of the listeners to the first venue 10. Alternatively, the terminal 42 may display icon images such as "cheering," "applause," "calling out," and "bustling noise" on the display 501 and receive reactions from listeners by receiving selection operations on these icon images via the user I / F 502. The terminal 42 generates sounds corresponding to the received reactions and transmits the generated sounds as ambience information to another device. Alternatively, the terminal 42 may transmit information indicating the type of ambient sound, such as cheering, applause, or calling out, as ambience information. In this case, a receiving device (e.g., the distribution device 12 and the mixer 11) generates a corresponding sound signal based on the ambience information and provides sounds such as cheering, applause, or calling out from listeners in the venue.
[0095] Terminal 42 may also receive the listener's position information via user I / F 502. Terminal 42 transmits the received listener's position information to first venue 10. Based on the sound and position information of the listener in third venue 20A received from AV receiver 32, distribution device 12 and mixer 11 in the first venue perform processing to localize the listener's sound to a specified position.
[0096] The terminal 42 may also change the HRTF based on location information received from the user. For example, if the listener specifies a position immediately in front of the stage in the first venue 10, the terminal 42 sets the localization position of the performer's sound to a position immediately in front of the listener and convolutes an HRTF that localizes the performer's sound to that position. This allows the listener in the fifth venue 20B to experience the same sense of presence as if they were directly in front of the stage in the first venue 10.
[0097] The sound of the listener at the fifth venue 20B may be transmitted to the second venue 20 instead of the first venue 10, or may be transmitted to another venue. As described above, the sound of the listener at the fifth venue 20B may be transmitted only to the friend's home (the fourth venue). This allows the listener at the fifth venue 20B and the listener at the fourth venue to watch the live performance at the first venue 10 while conversing with each other.
[0098] Furthermore, in the live data distribution system of this embodiment, multiple users can specify the same location. For example, multiple users may each specify a location immediately in front of the stage in the first venue 10. In this case, each listener can experience the same sense of realism as if they were located immediately in front of the stage. This allows multiple listeners to watch a performer's performance with the same sense of realism from one location (seat in the venue). In this case, the live event organizer can provide services that exceed the capacity of the actual space.
[0099] 15 is a block diagram showing the configuration of a live data distribution system 1D according to Modification 4. Components common to those in FIG. 1 are given the same reference numerals, and descriptions thereof will be omitted.
[0100] The live data distribution system 1D further includes a server 50 and a terminal 55. The terminal 55 is installed in the sixth venue 10A. The server 50 is an example of a distribution device, and the hardware configuration of the server 50 is similar to that of the distribution device 12. The hardware configuration of the terminal 55 is similar to that of the terminal 42 shown in FIG. 14.
[0101] The sixth venue 10A is the home or the like of a performer who is performing a performance such as a musical instrument remotely. The performer in the sixth venue 10A performs a performance such as playing or singing in sync with the performance or singing in the first venue. The terminal 55 transmits the sound of the performer in the sixth venue 10A to the server 50. The terminal 55 may also capture an image of the performer in the sixth venue 10A with a camera (not shown) and transmit the video signal to the server 50.
[0102] The server 50 distributes distribution data including the sound of the performers at the first venue 10, the sound of the performers at the sixth venue 10A, information on the reverberation of the space at the first venue 10, ambience information at the first venue 10, live footage at the first venue 10, and footage of the performers at the sixth venue 10A.
[0103] In this case, the playback device 22 renders the distribution data and provides the sounds of the performers at the first venue 10, the sounds of the performers at the sixth venue 10A, the reverberation of the space at the first venue 10, the ambient sounds at the first venue 10, the live video of the first venue 10, and the video of the performers at the sixth venue 10A to the second venue 20. For example, the playback device 22 displays the video of the performers at the sixth venue 10A superimposed on the live video of the first venue 10.
[0104] The sound of the performer in the sixth venue 10A does not need to be localized, but may be localized to a position that matches the video displayed on the display. For example, if the performer in the sixth venue 10A is displayed on the right side of the live video, the sound of the performer in the sixth venue 10A is localized to the right.
[0105] Alternatively, the performer at the sixth venue 10A or the distributor of the distribution data may specify the performer's position. In this case, the distribution data includes position information of the performer at the sixth venue 10A. The playback device 22 localizes the sound of the performer at the sixth venue 10A based on the position information of the performer at the sixth venue 10A.
[0106] The video of the performers in the sixth venue 10A is not limited to video captured by a camera. For example, a character image (virtual video) made up of a two-dimensional image or 3D modeling may be distributed as the video of the performers in the sixth venue 10A.
[0107] The distribution data may include audio recording data. The distribution data may also include video recording data. For example, the distribution device may distribute distribution data including the sound of the performers at the first venue 10, the audio recording data, reverberation information about the space at the first venue 10, ambience information about the first venue 10, live video from the first venue 10, and video recording data. In this case, the playback device renders the distribution data and provides the sound of the performers at the first venue 10, the audio related to the audio recording data, the reverberation information about the space at the first venue 10, the ambient sound at the first venue 10, the live video from the first venue 10, and the video related to the video recording data to another venue. The playback device 22 displays the video of the performers corresponding to the video recording data superimposed on the live video from the first venue 10.
[0108] The distribution device may also determine the type of instrument when recording the sound associated with the recording data. In this case, the distribution device distributes the recording data including information indicating the type of instrument determined to be associated with the recording data. The playback device generates a video of the corresponding instrument based on the information indicating the type of instrument. The playback device may display the video of the instrument superimposed on the live video from the first venue 10.
[0109] Furthermore, the distribution data does not need to superimpose the video of the performer at the sixth venue 10A on the live video from the first venue 10. For example, the distribution data may distribute the video of each performer at the first venue 10 and the sixth venue 10A and the background video as separate data. In this case, the distribution data includes information indicating the display position of each video. The playback device renders the video of each performer based on the information indicating the display position.
[0110] Furthermore, the background image is not limited to an image of a venue where a live performance is actually taking place, such as first venue 10. The background image may be an image of a venue other than the venue where the live performance is taking place.
[0111] Furthermore, the spatial reverberation information included in the distribution data does not need to correspond to the spatial reverberation of first venue 10. For example, the spatial reverberation information may be virtual space information for virtually recreating the spatial reverberation of the venue corresponding to the background video (information indicating the size, shape, wall material, etc. of each venue, or an impulse response indicating the transfer function of each venue). The impulse response of each venue may be measured in advance, or may be calculated by simulation based on the size, shape, wall material, etc. of each venue.
[0112] Furthermore, the ambience information may also be changed to match the background video. For example, in the case of a background video of a large venue, the ambience information includes sounds such as cheers, applause, and cheers from many listeners. Furthermore, an outdoor venue includes different background noises than an indoor venue. The reverberation of environmental sounds may also change according to the reverberation information of the space. Furthermore, the ambience information may include information indicating the number of spectators and information indicating the level of congestion (how densely packed the venue is). The playback device increases or decreases the number of sounds such as cheers, applause, and cheers from listeners based on the information indicating the number of spectators. Furthermore, the playback device increases or decreases the volume of the cheers, applause, and cheers from listeners based on the information indicating the level of congestion.
[0113] Alternatively, the ambience information may be changed depending on the performer. For example, if a performer with many female fans is performing live, the sounds of the listeners' cheers, applause, cheers, etc. included in the ambience information may be changed to female voices. The ambience information may include audio signals of the listeners' voices, but may also include information indicating audience attributes such as the gender ratio or age ratio. The playback device changes the voice quality of the listeners' cheers, applause, cheers, etc. based on the information indicating the attributes.
[0114] Furthermore, listeners at each venue may specify background video and spatial reverberation information using the user I / F of the playback device.
[0115] FIG. 16 shows an example of live video 700 displayed on a playback device at each venue. The live video 700 may consist of footage captured of the first venue 10 or another venue, or virtual video (computer graphics) corresponding to each venue. The live video 700 is displayed on the display of the playback device. The live video 700 displays the background of the venue, the stage, performers including their instruments, and video of listeners within the venue. The video of the background of the venue, the stage, performers including their instruments, and listeners within the venue may all be actual video footage or virtual video footage. Alternatively, only the background video may be actual video footage, while the other video footage may be virtual video footage. The live video 700 also displays icon images 751 and 752 for specifying a space. Icon image 751 is an image for specifying the space of a certain venue, Stage A (e.g., first venue 10), and icon image 752 is an image for specifying the space of another venue, Stage B (e.g., another concert hall). Furthermore, the live video 700 displays a listener image 753 for specifying the listener's position.
[0116] A listener using the playback device specifies a desired space by specifying either icon image 751 or icon image 752 using the playback device's user I / F. The distribution device distributes distribution data that includes background video and spatial reverb information corresponding to the specified space. Alternatively, the distribution device may distribute distribution data that includes multiple background videos and spatial reverb information. In this case, the playback device renders the background video and spatial reverb information corresponding to the space specified by the listener from the received distribution data.
[0117] In the example of Fig. 16, icon image 751 is specified. The playback device displays a background video (for example, a video of first venue 10) corresponding to Stage A of icon image 751, and plays back sounds related to the reverberation of the space corresponding to the specified Stage A. When the listener specifies icon image 752, the playback device switches to and displays a background video of Stage B, which is another space corresponding to icon image 752, and plays back sounds related to the reverberation of the corresponding other space based on the virtual space information corresponding to Stage B.
[0118] This allows the listeners of each playback device to experience the same sense of presence as if they were watching a live performance in a desired space.
[0119] Furthermore, the listener of each playback device can specify a desired position within the venue by moving the listener image 753 in the live video 700. The playback device performs localization processing based on the position specified by the user. For example, if the listener moves the listener image 753 to a position just in front of the stage, the playback device sets the localization position of the performer's sound to a position just in front of the listener and performs localization processing so that the performer's sound is localized at that position. This allows the listener of each playback device to experience the same sense of presence as if they were right in front of the stage.
[0120] Furthermore, as described above, when the position of the sound source and the position of the listener (position of the sound receiving point) change, the sound related to the reverberation of the space also changes. The playback device can calculate the early reflection sound even when the space, the position of the sound source, or the position of the sound receiving point changes. Therefore, even if measurements such as impulse responses are not performed in the actual space, the playback device can calculate the sound related to the reverberation of the space based on virtual space information. Therefore, the playback device can accurately reproduce the reverberation that occurs in spaces, including real spaces.
[0121] For example, mixer 11 may function as a distribution device, and mixer 21 may function as a playback device. Furthermore, playback devices do not need to be installed at each venue. For example, server 50 shown in FIG. 15 may render the distribution data and distribute the processed audio signals to terminals at each venue. In this case, server 50 functions as a playback device.
[0122] The sound source information may include information indicating the posture of the performer (e.g., the performer's orientation). The playback device may adjust the volume or frequency characteristics based on the performer's posture information. For example, the playback device may use a case where the performer is facing directly ahead as a reference and perform a process of reducing the volume as the performer's orientation increases to the left or right. The playback device may also perform a process of attenuating high frequencies more than low frequencies as the performer's orientation increases to the left or right. This allows the sound to change depending on the performer's posture, allowing listeners to enjoy a more realistic live performance.
[0123] Next, Fig. 17 is a block diagram showing an application example of signal processing performed by a playback device. In this example, rendering is performed using the terminal 42 and headphones 43 shown in Fig. 13. The playback device (terminal 42 in the example of Fig. 13) functionally comprises an instrument model processing unit 551, an amplifier model processing unit 552, a speaker model processing unit 553, a space model processing unit 554, a binaural processing unit 555, and a headphone inverse characteristics processing unit 556.
[0124] The instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 perform signal processing to impart the acoustic characteristics of the audio equipment to the sound signal related to the performance sound. A first digital signal processing model for performing this signal processing is included, for example, in the sound source information distributed by the distribution device 12. The first digital signal processing model is a digital filter that simulates the acoustic characteristics of the instrument, the amplifier, and the speaker, respectively. The first digital signal processing model is created in advance by the instrument manufacturer, the amplifier manufacturer, and the speaker manufacturer through simulation or the like. The instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 perform digital filter processing that simulates the acoustic characteristics of the instrument, the amplifier, and the speaker, respectively. Note that if the instrument is an electronic instrument such as a synthesizer, the instrument model processing unit 551 inputs note event data (information indicating the timing, pitch, etc. of the sound to be produced) instead of a sound signal, and generates a sound signal having the acoustic characteristics of the electronic instrument such as a synthesizer.
[0125] This allows the playback device to reproduce the acoustic characteristics of any musical instrument, etc. For example, in FIG. 16, a virtual (computer graphic) live video 700 is displayed. Here, a listener using the playback device may change the video to another virtual musical instrument using the user I / F of the playback device. When the listener changes the instrument displayed in the live video 700 to another instrument, the musical instrument model processing unit 551 of the playback device performs signal processing according to the first digital signal processing model corresponding to the changed instrument. This allows the playback device to output sound that reproduces the acoustic characteristics of the instrument displayed in the live video 700.
[0126] Similarly, a listener using the playback device may change the amplifier type and speaker type to different types using the user I / F of the playback device. The amplifier model processing unit 552 and the speaker model processing unit 553 perform digital filter processing that simulates the acoustic characteristics of the changed amplifier type and speaker type. Note that the speaker model processing unit 553 may simulate the acoustic characteristics for each speaker direction. In this case, a listener using the playback device may change the speaker orientation using the user I / F of the playback device. The speaker model processing unit 553 performs digital filter processing according to the changed speaker orientation.
[0127] The spatial model processing unit 554 is a second digital signal processing model that reproduces the acoustic characteristics of a room in a live venue (for example, the reverberation of the space described above). The second digital signal processing model may be obtained, for example, by using a test sound or the like in an actual live venue. Alternatively, the second digital signal processing model may determine the delay amount and level of an imaginary sound source by calculation from virtual space information (information indicating the size, shape, and wall materials of each venue) as described above.
[0128] When the position of the sound source and the position of the listener (position of the sound receiving point) change, the sound related to the reverberation of the space also changes. The playback device can calculate the delay amount and level of the imaginary sound source even when the space, the position of the sound source, or the position of the sound receiving point changes. Therefore, even if measurements such as impulse responses are not performed in the actual space, the playback device can calculate the sound related to the reverberation of the space based on virtual space information. Therefore, the playback device can accurately reproduce the reverberation that occurs in spaces, including real spaces.
[0129] The virtual space information may also include information on the position and material of structures (acoustic obstacles) such as pillars. In the sound source localization and indirect sound generation process, if an obstacle exists in the path of the direct sound and indirect sound arriving from the sound source, the playback device reproduces the phenomena of reflection, obstruction, and diffraction caused by the obstacle.
[0130] FIG. 18 is a schematic diagram showing the path of sound that travels from a sound source 70 to a sound receiving point 75 after being reflected by a wall. The sound source 70 shown in FIG. 18 may be either a performance sound (first sound source) or an environmental sound (second sound source). The playback device determines the position of an imaginary sound source 70A, which exists with the wall as a mirror surface relative to the position of the sound source 70, based on the positions of the sound source 70, the wall, and the sound receiving point 75. The playback device then determines the delay amount of the imaginary sound source 70A based on the distance from the imaginary sound source 70A to the sound receiving point 75. The playback device also determines the level of the imaginary sound source 70A based on information about the material of the wall. Furthermore, if an obstacle 77 is present on the path from the position of the imaginary sound source 70A to the sound receiving point 75, as shown in FIG. 18, the playback device determines the frequency characteristics resulting from diffraction by the obstacle 77. Diffraction, for example, attenuates high-frequency sounds. 18, if an obstacle 77 is present on the path from the position of the imaginary sound source 70A to the sound receiving point 75, the playback device performs equalizer processing to reduce the high frequency level. The frequency characteristics caused by diffraction may be included in the virtual space information.
[0131] The playback device may also set new second imaginary sound source 77A and third imaginary sound source 77B to the left and right of obstacle 77. The second imaginary sound source 77A and third imaginary sound source 77B correspond to new sound sources generated by diffraction. The second imaginary sound source 77A and third imaginary sound source 77B are sounds obtained by adding frequency characteristics generated by diffraction to the sound of imaginary sound source 70A. The playback device recalculates the delay amount and level based on the positions of second imaginary sound source 77A and third imaginary sound source 77B and the position of sound receiving point 75. This allows the diffraction phenomenon of obstacle 77 to be reproduced.
[0132] The playback device may calculate the delay and level of the sound from the imaginary sound source 70A that is reflected by the obstacle 77, then reflected on a wall, and reaches the sound receiving point 75. Furthermore, the playback device may erase the imaginary sound source 70A when it determines that the imaginary sound source 70A is blocked by the obstacle 77. Information for determining whether to block the imaginary sound source 70A may be included in the virtual space information.
[0133] By performing the above processing, the playback device performs first digital signal processing that represents the acoustic characteristics of the audio equipment and second digital signal processing that represents the acoustic characteristics of the room, thereby generating sound related to the sound of the sound source and the reverberation of the space.
[0134] The binaural processing unit 555 then convolves a head-related transfer function (hereinafter referred to as HRTF) with the sound signal to perform sound image localization processing for the sound source and various indirect sounds. The headphone inverse characteristic processing unit 556 performs digital filtering to reproduce the inverse characteristics of the acoustic characteristics of the headphones used by the listener.
[0135] Through the above processing, the user can experience the same sense of realism as if they were watching a live performance in a desired space with desired audio equipment.
[0136] It should be noted that the playback device does not need to include all of the instrument model processing unit 551, amplifier model processing unit 552, speaker model processing unit 553, and space model processing unit 554 shown in FIG. 17. The playback device only needs to perform signal processing using at least one digital signal processing model. Furthermore, the playback device may perform signal processing using one digital signal processing model on a single sound signal (e.g., the sound of a certain performer), or may perform signal processing using one digital signal processing model on each of multiple sound signals. The playback device may perform signal processing using multiple digital signal processing models on a single sound signal (e.g., the sound of a certain performer), or may perform signal processing using multiple digital signal processing models on multiple sound signals. The playback device may perform signal processing using a digital signal processing model on environmental sound.
[0137] The description of the present embodiment is illustrative in all respects and is not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to the claims and fall within the scope thereof. [Explanation of symbols]
[0138] 1, 1A, 1B, 1C, 1D...Live data distribution system 5. Internet 10...1st Venue 10A...6th venue 11...Mixer 12...Distribution device 13A~13F...Microphone 14A~14G...Speakers 15A~15C...Tracker 16...Camera 20...Second Venue 20A...3rd Venue 20B...5th venue 21...Mixer 22...Playback device 23...Indicator 24A~24F...Speakers 25A~25C...Microphone 32...AV receiver 33...Indicator 34A...Speaker 35…Mike 42...Terminal 43...Headphones 50...Server 55...Terminal 101...Indicator 102...User I / F 103...Audio I / O 104...signal processing unit 105...Network I / F 106...CPU 107...Flash memory 108...RAM 201…Display unit 202...User I / F 203...CPU 204...RAM 205...Network I / F 206...Flash memory 207...General-purpose communication I / F 301...Indicator 302...User I / F 303...CPU 304...RAM 305...Network I / F 306...Flash memory 307…Video I / F 401...Indicator 402...User I / F 403...Audio I / O 404...Signal processing unit 405...Network I / F 406...CPU 407...Flash memory 408...RAM 409…Video I / F 501...Display unit 503...CPU 504...RAM 505...Network I / F 506...Flash memory 507...Audio I / O 508...Mike 700...Live footage
Claims
1. Distributing, as distribution data, first sound source information relating to a sound of a first sound source occurring at a first location in a first venue and position information of the first sound source, and second sound source information relating to a second sound source including environmental sounds occurring at a second location in the first venue; Rendering the distribution data and providing the sound of the first sound source, which has been localized based on the position information of the first sound source, and the sound of the second sound source to a second venue. A live data distribution method, comprising: receiving virtual space information at the second venue, receiving the position of the sound receiving point in the virtual space information; Accepts the specification of any type of audio equipment, performing signal processing on the sound of the first sound source based on a first signal processing model that is a digital filter that reproduces the acoustic characteristics of the acoustic device of the accepted type; generating a sound relating to the reverberation of the space based on the virtual space information, the position information of the first sound source, and the position of the sound receiving point; outputting the sound of the first sound source that has been subjected to signal processing based on the first signal processing model and the sound related to the reverberation of the space; Live data delivery method.
2. The distribution data includes audio recording data. The live data distribution method according to claim 1 .
3. the first signal processing model includes a digital filter that simulates the acoustic characteristics of a musical instrument, an amplifier, or a speaker; The live data distribution method according to claim 1 .
4. the first signal processing model is created in advance by simulation; The live data distribution method according to claim 3 .
5. the first signal processing model includes a digital filter that simulates acoustic characteristics of a speaker; receiving an instruction from a listener to change the orientation of the speaker; performing digital filtering using the first signal processing model in accordance with the changed speaker orientation; 3. The live data distribution method according to claim 1 or 2.
6. A second digital signal processing model that reproduces the acoustic characteristics of the room is obtained based on the virtual space information, generating a sound relating to the reverberation of the space based on the second digital signal processing model, the position information of the first sound source, and the position of the sound receiving point; The live data distribution method according to any one of claims 1 to 5.
7. the virtual space information includes information on the position and material of an acoustic obstacle; In the process of generating the sound related to the reverberation of the space, a process of reproducing the phenomena of reflection, obstruction, and diffraction caused by the obstacle is performed. The live data distribution method according to any one of claims 1 to 6.
8. Transmitting ambience information related to the environmental sound of the second venue to a venue other than the second venue. The live data distribution method according to any one of claims 1 to 7.
9. feeding back the ambience information to the first venue; providing a sound related to the ambience information to users of the first venue; The live data distribution method according to claim 8.
10. the ambience information includes information corresponding to a user's reaction; providing a sound corresponding to the reaction to the users of the first venue; The live data distribution method according to claim 9.
11. The ambience information includes sounds picked up by a microphone installed in the second venue. The live data distribution method according to any one of claims 8 to 10.
12. the ambience information includes pre-created sounds; The live data distribution method according to any one of claims 8 to 11.
13. The pre-created sounds are different for each venue, The live data distribution method according to claim 12.
14. the ambience information includes information related to an attribute of a user corresponding to the second sound source, the rendering includes providing a sound based on the attribute. The live data distribution method according to any one of claims 8 to 13.
15. the second sound source information includes position information of the second sound source, the rendering includes a process of providing a sound of the second sound source that has been subjected to localization processing based on position information of the second sound source; The live data distribution method according to any one of claims 1 to 14.
16. receiving posture information of a performer corresponding to the sound of the first sound source; performing a process of adjusting the volume or frequency characteristics of the sound of the first sound source based on the posture information; The live data distribution method according to any one of claims 1 to 15.
17. The sound related to the reverberation in the space includes a first reverberation corresponding to the sound of the first sound source and a second reverberation corresponding to the sound of the second sound source. The live data distribution method according to any one of claims 1 to 16.
18. the spatial reverberation information includes first reverberation information that changes depending on the position of the first sound source and second reverberation information that changes depending on the position of the second sound source; the rendering includes a process of generating the first resonant sound based on the first resonant information, and a process of generating the second resonant sound based on the second resonant information. The live data distribution method according to claim 17.
19. the second sound source includes a plurality of sound sources; The live data distribution method according to any one of claims 1 to 18.
20. a live data distribution device that distributes, as distribution data, first sound source information relating to a sound of a first sound source generated at a first location in a first venue and position information of the first sound source, and second sound source information relating to a second sound source including environmental sounds generated at a second location in the first venue; a live data reproducing device that renders the distribution data and provides the sound of the first sound source, which has been localized based on the position information of the first sound source, and the sound of the second sound source, to a second venue; A live data distribution system comprising: The live data playback device includes: receiving virtual space information at the second venue, receiving the position of the sound receiving point in the virtual space information; Accepts the specification of any type of audio equipment, performing signal processing on the sound of the first sound source based on a first signal processing model that is a digital filter that reproduces the acoustic characteristics of the acoustic device of the accepted type; generating a sound relating to the reverberation of the space based on the virtual space information, the position information of the first sound source, and the position of the sound receiving point; outputting the sound of the first sound source that has been subjected to signal processing based on the first signal processing model and the sound related to the reverberation of the space; Live data distribution system.
21. Distributing, as distribution data, first sound source information relating to a sound of a first sound source occurring at a first location in a first venue and position information of the first sound source, and second sound source information relating to a second sound source including environmental sounds occurring at a second location in the first venue; causing a live data reproducing device to render the distribution data and provide, to a second venue, the sound of the first sound source that has been localized based on the position information of the first sound source and the sound of the second sound source; A live data distribution device, The live data playback device includes: receiving virtual space information at the second venue, receiving the position of the sound receiving point in the virtual space information; Accepts the specification of any type of audio equipment, performing signal processing on the sound of the first sound source based on a first signal processing model that is a digital filter that reproduces the acoustic characteristics of the acoustic device of the accepted type; generating a sound relating to the reverberation of the space based on the virtual space information, the position information of the first sound source, and the position of the sound receiving point; outputting a sound from a first sound source that has been subjected to signal processing based on the first signal processing model and a sound related to the reverberation of the space; Live data distribution device.
22. receiving, from a live data distribution device that distributes, as distribution data, first sound source information relating to a sound of a first sound source occurring at a first location in a first venue and position information of the first sound source, and second sound source information relating to a second sound source including environmental sound occurring at a second location in the first venue; Rendering the distribution data and providing the sound of the first sound source, which has been localized based on the position information of the first sound source, and the sound of the second sound source to a second venue. A live data playback device, comprising: receiving virtual space information at the second venue, receiving the position of the sound receiving point in the virtual space information; Accepts the specification of any type of audio equipment, performing signal processing on the sound of the first sound source based on a first signal processing model that is a digital filter that reproduces the acoustic characteristics of the acoustic device of the accepted type; generating a sound relating to the reverberation of the space based on the virtual space information, the position information of the first sound source, and the position of the sound receiving point; outputting the sound of the first sound source that has been subjected to signal processing based on the first signal processing model and the sound related to the reverberation of the space; Live data playback device.
Citation Information
Patent Citations
Match watching device, game watching terminal, game watching method, and program therefor
JP2019024157A