Field data transmission method, field data transmission system, transmission device thereof, field data playing device and playing method thereof
By transmitting sound source information and spatial echo information, and combining sound image movement and echo generation processing, the problem of lack of presence when transmitting on-site data is solved, and an immersive spatial listening experience is achieved.
Patent Information
- Application Number
- CN202180009216.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-27
- Filing Date
- 2021-03-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-03-19
AI Technical Summary
Existing technologies cannot effectively provide a sense of presence at the target venue when transmitting live data, and lack an immersive spatial listening experience.
By transmitting encoded information of the sound source and spatial echo, and combining sound image movement, initial reflection generation and rear echo generation processing, the sound field and ambient sound of the venue are reproduced. Signal processing and reproduction are performed at the receiving end using a mixer and playback device.
It achieves the effect of reproducing the sense of presence of the live venue on the receiving end, providing an immersive spatial listening experience, and allowing the audience to perceive the performance environment as if they were at the scene.
Smart Images

Figure CN114945978B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One embodiment of the present application relates to a live data transmission method, a live data transmission system, a live data transmission device, a live data playback device, and a live data playback method. BACKGROUND
[0002] A system that reproduces spatial type audio / content in a listening environment in order to provide a more immersive spatial listening experience is disclosed in Patent Literature 1.
[0003] The system of Patent Literature 1 describes that a pulse response of sound output from a speaker in a listening environment is measured, and a filter processing corresponding to the measured pulse response is performed.
[0004] Patent Literature 1: Japanese Patent Application Laid-Open No. 2015-530043 SUMMARY
[0005] The system of Patent Literature 1 is not a live data transmission system. In the case of transmitting live data, it is desirable to provide a sense of presence of a live venue to a venue of a transmission target as well.
[0006] One embodiment of the present application aims to provide a live data transmission method, a live data transmission system, a live data transmission device, a live data playback device, and a live data playback method that can provide a sense of presence of a live venue to a venue of a transmission target as well in the case of transmitting live data.
[0007] A live data transmission method transmits sound source information related to sound occurring at a first venue and echo information of a space that changes in correspondence with a position of the sound as transmission data, renders the transmission data, and provides sound related to the sound source information and sound related to the echo of the space to a second venue.
[0008] EFFECT OF THE INVENTION
[0009] The live data transmission method can provide a sense of presence of a live venue to a venue of a transmission target as well in the case of transmitting live data. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a block diagram showing a configuration of a live data transmission system 1.
[0011] Figure 2 is a plan view of a first venue 10.
[0012] Figure 3 is a plan view of a second venue 20.
[0013] Figure 4 is a block diagram showing the structure of the mixer 11.
[0014] Figure 5 is a block diagram showing the structure of the transmission device 12.
[0015] Figure 6 is a flowchart showing the action of the transmission device 12.
[0016] Figure 7 is a block diagram showing the structure of the playback device 22.
[0017] Figure 8 is a flowchart showing the action of the playback device 22.
[0018] Figure 9 is a block diagram showing the structure of the live data transmission system 1A to which the modification example 1 is applied.
[0019] Figure 10 is a plan view showing the second venue 20 of the live data transmission system 1A to which the modification example 1 is applied.
[0020] Figure 11 is a block diagram showing the structure of the live data transmission system 1B to which the modification example 2 is applied.
[0021] Figure 12 is a block diagram showing the structure of the AV receiver 32.
[0022] Figure 13 is a block diagram showing the structure of the live data transmission system 1C to which the modification example 3 is applied.
[0023] Figure 14 is a block diagram showing the structure of the terminal 42.
[0024] Figure 15 is a block diagram showing the structure of the live data transmission system 1D to which the modification example 4 is applied.
[0025] Figure 16 is a diagram showing one example of the live image 700 displayed by the playback device at each venue.
[0026] Figure 17 is a block diagram showing an application example of the signal processing by the playback device.
[0027] Figure 18 is a diagram showing the path of the sound arriving at the sound receiving point 75 from the sound source 70 reflected on the wall surface. DETAILED DESCRIPTION
[0028] Figure 1is a block diagram showing the configuration of a live data transmission system 1. The live data transmission system 1 is configured of a plurality of sound equipment and information processing devices provided in a first venue 10 and a second venue 20, respectively.
[0029] Figure 2 is a plan view of the first venue 10, Figure 3 is a plan view of the second venue 20. In this example, the first venue 10 is a live venue where a performer performs a performance. The second venue 20 is a public viewing venue where a remote audience views and listens to the performance by the performer.
[0030] A mixer 11, a transmission device 12, a plurality of microphones 13A to 13F, a plurality of speakers 14A to 14G, a plurality of trackers 15A to 15C, and a camera 16 are provided in the first venue 10. A mixer 21, a playback device 22, a display 23, and a plurality of speakers 24A to 24F are provided in the second venue 20. The transmission device 12 and the playback device 22 are connected via the Internet 5. Furthermore, the number of microphones, the number of speakers, and the number of trackers are not limited to the numbers shown in the present embodiment. In addition, the arrangement of the microphones and the speakers is not limited to the example shown in the present embodiment.
[0031] The mixer 11 is connected to the transmission device 12, the plurality of microphones 13A to 13F, the plurality of speakers 14A to 14G, and the plurality of trackers 15A to 15C. The mixer 11, the plurality of microphones 13A to 13F, and the plurality of speakers 14A to 14G are connected via a network cable or an audio cable. The plurality of trackers 15A to 15C are connected to the mixer 11 via wireless communication. The mixer 11 and the transmission device 12 are connected via a network cable. In addition, the transmission device 12 is connected to the camera 16 via a video cable. The camera 16 captures a live video including a performer.
[0032] The plurality of speakers 14A to 14G are arranged along the wall surface of the first venue 10. The first venue 10 in this example is rectangular in plan view. A stage is arranged in front of the first venue 10. On the stage, a performer performs a performance such as singing or playing a musical instrument. The speaker 14A is arranged on the left side of the stage, the speaker 14B is arranged at the center of the stage, and the speaker 14C is arranged on the right side of the stage. The speaker 14D is arranged on the left side of the front and rear center of the first venue 10, and the speaker 14E is arranged on the right side of the front and rear center of the first venue 10. The speaker 14F is arranged on the left side of the rear of the first venue 10, and the speaker 14G is arranged on the right side of the rear of the first venue 10.
[0033] The microphone 13A is disposed on the left side of the stage, the microphone 13B is disposed on the center of the stage, and the microphone 13C is disposed on the right side of the stage. The microphone 13D is disposed on the left side of the front and rear center of the first venue 10, the microphone 13E is disposed on the rear center of the first venue 10, and the microphone 13F is disposed on the right side of the front and rear center of the first venue 10.
[0034] The mixer 11 receives sound signals from the microphones 13A to 13F. In addition, the mixer 11 outputs sound signals to the speakers 14A to 14G. In the present embodiment, the speakers and the microphones are shown as one example of the sound devices connected to the mixer 11, but actually, a plurality of sound devices are connected to the mixer 11. The mixer 11 receives sound signals from a plurality of sound devices such as the microphones, performs signal processing such as mixing, and outputs sound signals to a plurality of sound devices such as the speakers.
[0035] The microphones 13A to 13F acquire the singing or playing sounds of the respective performers as sounds occurring in the first venue 10. Alternatively, the microphones 13A to 13F acquire the environmental sounds of the first venue 10. In the example of the present embodiment, the microphones 13A to 13C acquire the sounds of the performers, and the microphones 13D to 13F acquire the environmental sounds. The environmental sounds include sounds such as the cheers, applause, shouts, cheers, chorus, or noise of the audience. However, the sounds of the performers can be input by a line. The line input refers to inputting a sound signal from an audio cable or the like connected to a sound source, rather than picking up a sound output from a sound source such as a musical instrument by a microphone. The sounds of the performers are preferably acquired as sounds with a high SN ratio, and do not include other sounds. Figure 2
[0036] The speakers 14A to 14G output the sounds of the performers to the first venue 10. In addition, the speakers 14A to 14G can also output initial reflected sounds or rear reverberation sounds for controlling the sound field of the first venue 10.
[0037] The mixer 21 of the second venue 20 is connected to the playback device 22 and a plurality of speakers 24A to 24F. These sound devices are connected via a network cable or an audio cable. In addition, the playback device 22 is connected to the display 23 via a video cable.
[0038] A plurality of speakers 24A to 24F are arranged along the wall surface of the second venue 20. The second venue 20 of the present example is rectangular in plan view. A display 23 is arranged in front of the second venue 20. The display 23 displays a live image captured in the first venue 10. The speaker 24A is arranged to the left of the display 23, and the speaker 24B is arranged to the right of the display 23. The speaker 24C is arranged to the left of the front and rear center of the second venue 20, and the speaker 24D is arranged to the right of the front and rear center of the second venue 20. The speaker 24E is arranged to the left of the rear of the second venue 20, and the speaker 24F is arranged to the right of the rear of the second venue 20.
[0039] The mixer 21 outputs a sound signal to the speakers 24A to 24F. The mixer 21 receives a sound signal from the playback device 22, performs signal processing such as mixing, and outputs a sound signal to a plurality of sound devices such as speakers.
[0040] The speakers 24A to 24F output the sound of the performer to the second venue 20. In addition, the speakers 24A to 24F output the initial reflected sound or the back echo sound for reproducing the sound field of the first venue 10. In addition, the speakers 24A to 24F output the environmental sound such as the applause of the audience of the first venue 10 to the second venue 20.
[0041] Figure 4 is a block diagram showing the structure of the mixer 11. Further, the mixer 21 has the same structure and functions as the mixer 11, and therefore, the structure of the mixer 11 is shown as a representative case in Figure 4 The mixer 11 has a display 101, a user I / F 102, an audio I / O (Input / Output) 103, a signal processing section (DSP) 104, a network I / F 105, a CPU 106, a flash memory 107, and a RAM 108.
[0042] The CPU 106 is a control section that controls the operation of the mixer 11. The CPU 106 reads a prescribed program stored in the flash memory 107 as a storage medium to the RAM 108 and executes, thereby performing various operations.
[0043] Further, the program read by the CPU 106 does not need to be stored in the flash memory 107 in the present device. For example, the program can be stored in a storage medium of an external device such as a server. In this case, the CPU 106 can read the program to the RAM 108 from the server each time and execute.
[0044] The signal processing section 104 is configured by a DSP for performing various signal processing. The signal processing section 104 performs signal processing such as mixing processing and filter processing on a sound signal input from a microphone or the like via the audio I / O 103 or the network I / F 105. The signal processing section 104 outputs the sound signal after the signal processing to a speaker or the like via the audio I / O 103 or the network I / F 105.
[0045] In addition, the signal processing section 104 can also perform panning processing, initial reflection sound generation processing, and rear reverberation sound generation processing. The panning processing is processing of controlling the volume of a sound signal distributed to the plurality of speakers 14A to 14G in such a manner that the sound image is positioned at the position of a performer. In order to perform the panning processing, the CPU 106 acquires position information of the performer via the trackers 15A to 15C. The position information is information indicating a 2-dimensional or 3-dimensional coordinate with a certain position of the first venue 10 as a reference. The trackers 15A to 15C are, for example, tags that transmit and receive electric waves of Bluetooth (registered trademark) or the like. The trackers 15A to 15C are attached to the performer or the musical instrument. At least three beacons are provided in advance in the first venue 10. Each beacon measures the distance to the trackers 15A to 15C based on the time difference from when the electric wave is transmitted to when the electric wave is received. The CPU 106 is able to uniquely determine the position of the trackers 15A to 15C by measuring the distance from each of the at least three beacons to the tag by acquiring the position information of the beacons in advance.
[0046] The CPU 106 acquires the position information of each performer, that is, the position information of the sound occurring in the first venue 10, via the trackers 15A to 15C in this manner. The CPU 106 determines the volume of each sound signal output to the speakers 14A to 14G in such a manner that the sound image is positioned at the position of the performer based on the acquired position information and the positions of the speakers 14A to 14G. The signal processing section 104 controls the volume of each sound signal output to the speakers 14A to 14G in accordance with the control of the CPU 106. For example, the signal processing section 104 increases the volume of the sound signal output to the speaker close to the position of the performer and decreases the volume of the sound signal output to the speaker away from the position of the performer. Thus, the signal processing section 104 is able to position the sound image of the performance sound or the singing sound of the performer at a prescribed position.
[0047] The initial reflected sound generation processing and the late reverberation sound generation processing are processing of convolving an impulse response to the performer's sound by a FIR filter. The signal processing section 104, for example, convolves an impulse response acquired in a prescribed hall (a hall other than the first hall 10) to the performer's sound. Thereby, the signal processing section 104 controls the sound field of the first hall 10. Alternatively, the signal processing section 104 can further feed back the sound acquired by a microphone disposed near the ceiling or wall surface of the first hall 10 to the speakers 14A to 14G, thereby controlling the sound field of the first hall 10.
[0048] The signal processing section 104 outputs the performer's sound and the performer's position information to the transmission device 12. The transmission device 12 acquires the performer's sound and the performer's position information from the mixer 11.
[0049] In addition, the transmission device 12 acquires an image signal from the camera 16. The camera 16 photographs each performer or the entire first hall 10, and outputs an image signal relating to the live image to the transmission device 12.
[0050] Further, the transmission device 12 acquires echo information of the space of the first hall 10. The echo information of the space is information for generating an indirect sound. The indirect sound is a sound of a sound source reaching a listener by being reflected in a hall, and includes at least an initial reflected sound and a late reverberation sound. The echo information of the space includes, for example, information indicating the size, shape, and material of the wall surface of the space of the first hall 10, and an impulse response relating to the late reverberation sound. The information indicating the size, shape, and material of the wall surface of the space is information for generating the initial reflected sound. The information for generating the initial reflected sound can also be an impulse response. The impulse response is, for example, predetermined in the first hall 10. In addition, the echo information of the space can be information that changes in accordance with the position of the performer. The information that changes in accordance with the position of the performer is, for example, an impulse response predetermined for each position of the performer in the first hall 10. The transmission device 12, for example, acquires a first impulse response in a case where the performer's sound is generated at the front of the stage, a second impulse response in a case where the performer's sound is generated at the left side of the stage, and a third impulse response in a case where the performer's sound is generated at the right side of the stage of the first hall 10. However, the number of impulse responses is not limited to three. In addition, the impulse response need not be actually measured in the first hall 10, and can be calculated by simulation, for example, in accordance with the size, shape, and material of the wall surface of the space of the first hall 10.
[0051] Further, the initial reflected sound is a reflected sound whose arrival direction of sound is fixed, and the late reverberation sound is a reflected sound whose arrival direction of sound is not fixed. The late reverberation sound is less changed than the initial reflected sound due to a change in the position of the performer's voice. Therefore, the echo information of the space can be a form constituted of an impulse response of the initial reflected sound which changes in accordance with the position of the performer, and an impulse response of the late reverberation sound which is constant regardless of the position of the performer.
[0052] In addition, the signal processing section 104 can also acquire environmental information related to the environmental sound and output to the transmission device 12. The environmental sound is sound acquired by the microphones 13D to 13F as described above, and includes background noise, cheers of the audience, applause, shouts, cheers, chorus, or noise. However, the environmental sound can also be acquired by the microphones 13A to 13C on the stage. The signal processing section 104 outputs the sound signal related to the environmental sound as the environmental information to the transmission device 12. Further, the environmental information can also include position information of the environmental sound. The individual "cheers" and the like of the audience in the environmental sound, the shouts of the personal name of the performer, or the exclamation words such as "good" and the like are sounds which are not drowned out by the audience and can be recognized as the sound of the individual audience. The signal processing section 104 can also acquire the position information of these individual sounds. The position information of the environmental sound can be found, for example, from the sound acquired by the microphones 13D to 13F. The signal processing section 104 finds the correlation of the sound signals of the microphones 13D to 13F in a case where the above-described individual sounds are recognized by processing such as voice recognition, and finds the difference in timing at which the individual sounds are picked up by the microphones 13D to 13F, respectively. The signal processing section 104 can uniquely find the position within the first venue 10 where the sound occurs based on the difference in timing at which the sound is picked up by the microphones 13D to 13F. In addition, the position information of the environmental sound can also be considered as the position information of each microphone 13D to 13F.
[0053] The transmission device 12 encodes and transmits the sound source information related to the sound occurring in the first venue 10 and the echo information of the space as transmission data. The sound source information includes at least the voice of the performer, and can include the position information of the voice of the performer. In addition, the transmission device 12 can transmit the environmental information related to the environmental sound included in the transmission data. The transmission device 12 can also transmit the image signal related to the image of the performer included in the transmission data.
[0054] Alternatively, the transmission device 12 can also transmit at least the sound source information related to the voice of the performer and the position information of the performer, and the environmental information related to the environmental sound as the transmission data.
[0055] Figure 5 is a block diagram showing the structure of the transmission device 12. Figure 6is a flowchart showing the operation of the transmission device 12.
[0056] The transmission device 12 is constituted by an information processing device such as a general personal computer. The transmission device 12 has a display 201, a user I / F 202, a CPU 203, a RAM 204, a network I / F 205, a flash memory 206, and a general communication I / F 207.
[0057] The CPU 203 reads a program stored in the flash memory 206 as a storage medium to the RAM 204, thereby implementing a prescribed function. Further, the program read by the CPU 203 need not be stored in the flash memory 206 within the device. For example, the program can be stored in a storage medium of an external device such as a server. In this case, the CPU 203 reads the program to the RAM 204 from the server each time and executes it.
[0058] The CPU 203 acquires the performer's voice and the performer's position information (sound source information) from the mixer 11 via the network I / F 205 (S11). In addition, the CPU 203 acquires the echo information of the space of the first venue 10 (S12). Also, the CPU 203 acquires the environmental information to which the environmental sound is related (S13). In addition, the CPU 203 can acquire the image signal from the camera 16 via the general communication I / F 207.
[0059] The CPU 203 encodes and transmits the data related to the performer's voice and the position information of the voice (sound source information), the data related to the echo information of the space, the data related to the environmental information, and the data related to the image signal as transmission data (S14).
[0060] The playback device 22 receives the transmission data from the transmission device 12 via the Internet 5. The playback device 22 renders the transmission data and provides the voice related to the performer's voice and the echo of the space to the second venue 20. Alternatively, the playback device 22 provides the environmental sound included in the performer's voice and the environmental information to the second venue 20. The playback device 22 can also provide the voice related to the echo of the space corresponding to the environmental information to the second venue 20.
[0061] Figure 7 is a block diagram showing the structure of the playback device 22. Figure 8 is a flowchart showing the operation of the playback device 22.
[0062] The playback device 22 is constituted by an information processing device such as a general personal computer. The playback device 22 has a display 301, a user I / F 302, a CPU 303, a RAM 304, a network I / F 305, a flash memory 306, and an image I / F 307.
[0063] The CPU 303 implements a prescribed function by reading a program stored in the flash memory 306 as a storage medium to the RAM 304. Further, the program read by the CPU 303 need not be stored in the flash memory 306 within the device. For example, the program can be stored in a storage medium of an external device such as a server. In this case, the CPU 303 reads the program from the server to the RAM 304 each time and executes.
[0064] The CPU 303 receives the transmission data from the transmission device 12 via the network I / F 305 (S21). The CPU 303 decodes the transmission data into sound source information, spatial echo information, environmental information, and image signals, and the like (S22), and reproduces the sound source information, spatial echo information, environmental information, and image signals, and the like.
[0065] The CPU 303 causes the mixer 21 to perform the sound image movement processing of the performer's voice as one example of the reproduction of the sound source information (S23). The sound image movement processing is processing that positions the performer's voice at the performer's position as described above. The CPU 303 determines the volume of the sound signal to be allocated to the speakers 24A to 24F in such a manner that the performer's voice is positioned at the position indicated by the position information included in the sound source information. The CPU 303 outputs the sound signal to which the performer's voice is related, and information indicating the output amount of the sound signal to which the performer's voice is related to the speakers 24A to 24F, to the mixer 21, thereby causing the mixer 21 to perform the sound image movement processing.
[0066] Thus, the audience at the second venue 20 can perceive as if the sound is emitted from the performer's position. The audience at the second venue 20 can hear the performer's voice located at the right side of the stage of the first venue 10 from the right side in front at the second venue 20, for example. In addition, the CPU 303 can reproduce the image signal and display the live image on the display 23 via the image I / F 307. Thus, the audience at the second venue 20 can watch the image of the performer displayed on the display 23 while listening to the performer's voice after the sound image movement processing. Thus, the audience at the second venue 20 can obtain the sense of immersion for the live performance because the visual information and the auditory information are consistent.
[0067] Further, the CPU 303 causes the mixer 21 to perform the indirect sound generation process as one example of reproduction of the spatial echo information (S24). The indirect sound generation process includes an initial reflection sound generation process and a late reverberation sound generation process. The initial reflection sound is generated based on the performer's voice included in the sound source information and information indicating the size, shape, wall surface material, and the like of the space of the first venue 10 included in the spatial echo information. The CPU 303 determines the arrival timing of the initial reflection sound based on the size and shape of the space and determines the level of the initial reflection sound based on the wall surface material. More specifically, the CPU 303 calculates the coordinates of the wall surface on which the sound of the sound source is reflected based on the information of the size and shape of the space. Further, the CPU 303 calculates the position of a virtual sound source (virtual sound source) existing as a mirror surface with respect to the position of the sound source based on the position of the sound source, the position of the wall surface, and the position of the pickup point. The CPU 303 calculates the delay amount of the virtual sound source based on the distance from the position of the virtual sound source to the pickup point. In addition, the CPU 303 calculates the level of the virtual sound source based on the information of the wall surface material. The information of the wall surface material corresponds to the energy loss at the time of reflection by the wall surface. Therefore, the CPU 303 calculates the level of the virtual sound source taking into account the energy loss with respect to the sound signal of the sound source. The CPU 303 is able to calculate the delay amount and the level of the sound involved in the spatial echo by repeating such a process. The CPU 303 outputs the calculated delay amount and level to the mixer 21. The mixer 21 convolves these delay amounts and the level tap coefficients corresponding to the levels in the performer's voice. Thus, the mixer 21 reproduces the spatial echo of the first venue 10 in the second venue 20. In addition, in a case where the spatial echo information includes an impulse response of the initial reflection sound, the CPU 303 causes the mixer 11 to perform a process of convolving the impulse response in the performer's voice by a FIR filter. The CPU 303 outputs the spatial echo information (impulse response) included in the transmission data to the mixer 21. The mixer 21 convolves the spatial echo information (impulse response) received from the playback device 22 in the performer's voice. Thus, the mixer 21 reproduces the spatial echo of the first venue 10 in the second venue 20.
[0068] Further, in a case where the spatial echo information changes in correspondence with the position of the performer, the playback device 22 outputs the spatial echo information corresponding to the position of the performer to the mixer 21 on the basis of the position information included in the sound source information. For example, in a case where the performer who is present at the front of the stage of the first venue 10 moves to the left side of the stage, the impulse response with which the sound of the performer is convolved is changed from the first impulse response to the second impulse response. Alternatively, in a case where a virtual sound source is reproduced on the basis of information on the size and shape of the space, the delay amount and the level are recalculated in correspondence with the position of the performer after the movement. Thus, appropriate spatial echo corresponding to the position of the performer is also reproduced in the second venue 20.
[0069] Further, the playback device 22 can also cause the mixer 21 to generate spatial echo sound corresponding to the environmental sound on the basis of the environmental information and the spatial echo information. That is, the sound involved in the spatial echo can include first echo sound corresponding to the sound of the performer (the sound of the first sound source) and second echo sound corresponding to the environmental sound (the sound of the second sound source). Thus, the mixer 21 reproduces the echo of the environmental sound of the first venue 10 in the second venue 20. Further, in a case where the environmental information includes position information, the playback device 22 can also output spatial echo information corresponding to the position of the environmental sound to the mixer 11 on the basis of the position information included in the environmental information. The mixer 21 reproduces the echo sound of the environmental sound on the basis of the position of the environmental sound. For example, in a case where the audience who is present at the left rear of the first venue 10 moves to the right rear, the impulse response with which the cheering sound of the audience is convolved is changed. Alternatively, in a case where a virtual sound source is reproduced on the basis of information on the size and shape of the space, the delay amount and the level are recalculated in correspondence with the position of the audience after the movement. As described above, the spatial echo information includes first echo information which changes in correspondence with the position of the sound of the performer (the first sound source) and second echo information which changes in correspondence with the position of the environmental sound (the second sound source), and the reproduction can include a process of generating the first echo sound on the basis of the first echo information and a process of generating the second echo sound on the basis of the second echo information.
[0070] Further, the rear reverberation sound is reflected sound whose arrival direction of sound is not fixed. The rear reverberation sound has a smaller change due to the change in the position of the sound than the initial reflected sound. Therefore, the playback device 22 can change only the impulse response of the initial reflected sound which changes in correspondence with the position of the performer and fix the impulse response of the rear reverberation sound.
[0071] Further, the playback device 22 can omit the indirect sound generation processing and directly use the echo of the second venue 20. In addition, the indirect sound generation processing can be only the initial reflection sound generation processing. The late reverberation sound can also be directly used by the echo of the second venue 20. Alternatively, the mixer 21 can also enhance the control of the second venue 20 by further feeding back the sound acquired by a microphone (not shown) provided near the ceiling and wall surface of the second venue 20 to the speakers 24A to 24F.
[0072] Further, the CPU 303 of the playback device 22 performs the environmental sound playback processing based on the environmental information (S25). The environmental information includes sound signals of sounds such as background noise, audience's cheers, applause, shouts, cheers, chorus, or commotion. The CPU 303 outputs these sound signals to the mixer 21. The mixer 21 outputs the sound signals received from the playback device 22 to the speakers 24A to 24F.
[0073] The CPU 303 causes the mixer 21 to perform the environmental sound positioning processing by the sound image movement processing in a case where the environmental information includes the position information of the environmental sound. In this case, the CPU 303 determines the volume of the sound signal to be allocated to the speakers 24A to 24F in such a manner that the environmental sound is positioned at the position indicated by the position information included in the environmental information. The CPU 303 causes the mixer 21 to perform the sound image movement processing by outputting the sound signal of the environmental sound and information indicating the output amount of the sound signal involved in the environmental sound to the speakers 24A to 24F to the mixer 21. The same applies in a case where the position information of the environmental sound is the position information of each microphone 13D to 13F. The CPU 303 determines the volume of the sound signal to be allocated to the speakers 24A to 24F in such a manner that the environmental sound is positioned at the position of the microphone. Each microphone 13D to 13F picks up a plurality of environmental sounds (second sound sources) such as background noise, applause, chorus, or cheers such as "Wow", commotion, and the like. The sound of each sound source includes a prescribed delay amount and a level to reach the microphone. That is, background noise, applause, chorus, or cheers such as "Wow", commotion, and the like also reach the microphone as independent sound sources including a prescribed delay amount and a level (information for positioning the sound source). The CPU 303 can easily reproduce the positioning of the independent sound sources by performing the sound image movement processing in such a manner that the sound picked up by the microphone is positioned at the position of the microphone.
[0074] Further, the CPU 303 can also process the sound emitted by a plurality of listeners at the same time, which cannot be recognized as a sound of a single listener, by causing the mixer 21 to perform effect processing such as reverb, thereby perceiving the expansion of the space. For example, background noise, applause, chorus, or cheers such as "Wow", and commotion are sounds that resonate throughout the entire venue. The CPU 303 causes the mixer 21 to perform effect processing that perceives the expansion of the space for these sounds.
[0075] The playback device 22 can also provide the environmental sound based on the above environmental information to the second venue 20. Thereby, the listeners at the second venue 20 can view and hear the live performance with a greater sense of presence as if they were viewing and hearing the live performance at the first venue 10.
[0076] In the above manner, the live data transmission system 1 of the present embodiment transmits the sound source information related to the sound occurring at the first venue 10 and the echo information of the space as transmission data, and provides the sound related to the sound source information and the sound related to the echo of the space to the second venue 20 by reproducing the transmission data. Thereby, the sense of presence of the live venue can be provided to the venue of the transmission target as well.
[0077] In addition, the live data transmission system 1 transmits the first sound source information related to the sound of the first sound source (e.g., the sound of a performer) occurring at a certain first place (e.g., a stage) of the first venue 10 and the position information of the first sound source, and the second sound source information related to the second sound source (e.g., environmental sound) occurring at a second place (e.g., a place where the listeners are located) of the first venue 10 as transmission data, and provides the sound of the first sound source on which positioning processing based on the position information of the first sound source is performed and the sound of the second sound source to the second venue by reproducing the transmission data. Thereby, the sense of presence of the live venue can be provided to the venue of the transmission target as well.
[0078] Next, Figure 9 is a block diagram showing the structure of the live data transmission system 1A to which Modification 1 is applied. Figure 10 is a plan view showing the second venue 20 of the live data transmission system 1A to which Modification 1 is applied. The structure common to the second venue 20 of the live data transmission system 1 shown in FIG. 1 is denoted by the same reference numerals, and the description thereof is omitted. Figure 1 and Figure 3 The structure common to the second venue 20 of the live data transmission system 1 shown in FIG. 1 is denoted by the same reference numerals, and the description thereof is omitted.
[0079] A plurality of microphones 25A to 25C are provided at the second venue 20 of the live data transmission system 1A. The microphone 25A is disposed at the left side of the front and rear center of the stage 80 of the second venue 20, the microphone 25B is disposed at the rear center of the second venue 20, and the microphone 25C is disposed at the right side of the front and rear center of the second venue 20.
[0080] The microphones 25A to 25C acquire the environmental sound of the second venue 20. The mixer 21 outputs the sound signal of the environmental sound as environmental information to the playback device 22. Further, the environmental information can include position information of the environmental sound. The position information of the environmental sound can be calculated from the sound acquired by the microphones 25A to 25C, for example, as described above.
[0081] The playback device 22 transmits the environmental information related to the environmental sound occurring in the second venue 20 as a third sound source to the other venues. For example, the playback device 22 feeds back the environmental sound occurring in the second venue 20 to the first venue 10. Thereby, the performer on the stage of the first venue 10 can hear the sound, applause, cheers, and the like of the audience other than the audience of the first venue 10, and can perform live in an environment filled with a sense of presence. In addition, the audience of the first venue 10 can also hear the sound, applause, cheers, and the like of the audience of the other venues, and can view and listen to the live performance in an environment filled with a sense of presence.
[0082] Further, if the playback device of the other venue reproduces the transmission data to provide the sound of the first venue to the other venue, and provides the environmental sound occurring in the second venue 20 to the other venue, the audience of the other venue can also hear the sound, applause, cheers, and the like of the plurality of audiences, and can view and listen to the live performance in an environment filled with a sense of presence.
[0083] Next, Figure 11 is a block diagram showing the structure of the live data transmission system IB to which Modification 2 is applied. The same structures as those of the live data transmission system 1 are denoted by the same reference numerals, and the description thereof is omitted. Figure 1 The same structures as those of the live data transmission system 1 are denoted by the same reference numerals, and the description thereof is omitted.
[0084] In the live data transmission system IB, the transmission device 12 is connected to an AV receiver 32 of a third venue 20A via the Internet 5. The AV receiver 32 is connected to a display 33, a plurality of speakers 34A to 34F, and a microphone 35. The third venue 20A is, for example, the residence of a certain audience member. The AV receiver 32 is an example of a playback device. The user of the AV receiver 32 is an audience who views and listens to the live performance of the first venue 10 in a remote manner.
[0085] Figure 12 is a block diagram showing the structure of the AV receiver 32. The AV receiver 32 has a display 401, a user I / F 402, an audio I / O (Input / Output) 403, a signal processing section (DSP) 404, a network I / F 405, a CPU 406, a flash memory 407, a RAM 408, and a video I / F 409.
[0086] The CPU 406 is a control unit that controls the operation of the AV receiver 32. The CPU 406 reads a prescribed program stored in the flash memory 407 as a storage medium to the RAM 408 and executes it, thereby performing various operations.
[0087] Further, the program read by the CPU 406 need not be stored in the flash memory 407 in the present apparatus. For example, the program can be stored in a storage medium of an external apparatus such as a server. In this case, the CPU 406 reads the program from the server to the RAM 408 each time and executes it.
[0088] The signal processing unit 404 is constituted by a DSP for performing various signal processing. The signal processing unit 404 performs signal processing on a sound signal input via the audio I / O 403 or the network I / F 405. The signal processing unit 404 outputs the audio signal after signal processing to a sound device such as a speaker via the audio I / O 403 or the network I / F 405.
[0089] The AV receiver 32 performs the same processing as that performed by the mixer 21 and the playback device 22. The CPU 406 receives the transmission data from the transmission device 12 via the network I / F 405. The CPU 406 reproduces the transmission data and provides the sound of the performer and the sound involved in the echo of the space to the 3rd venue 20A. Alternatively, the CPU 406 reproduces the transmission data and provides the environmental sound occurring in the 1st venue 10 to the 3rd venue 20A. Alternatively, the CPU 406 can reproduce the transmission data and display the live image on the display 33 via the video I / F 307.
[0090] The signal processing unit 404 performs the sound image movement processing of the sound of the performer. In addition, the signal processing unit 404 performs the generation processing of the indirect sound. Alternatively, the signal processing unit 404 can perform the sound image movement processing of the environmental sound.
[0091] Thus, the AV receiver 32 can provide the sense of presence of the 1st venue 10 to the 3rd venue 20A as well.
[0092] In addition, the AV receiver 32 acquires the environmental sound (sound of the audience's cheers, applause, or shouts) of the 3rd venue 20A via the microphone 35. The AV receiver 32 transmits the environmental sound of the 3rd venue 20A to other devices. For example, the AV receiver 32 feeds back the environmental sound of the 3rd venue 20A to the 1st venue 10.
[0093] Thus, if the voices from the plurality of audiences are fed back to the first venue 10, the performer on the stage of the first venue 10 can hear the cheers, applause, cheers, and the like of the plurality of audiences other than the audience of the first venue 10, and can perform live in an environment filled with a sense of presence. In addition, the audience of the first venue 10 can also hear the cheers, applause, cheers, and the like of the plurality of audiences at a distance, and can view and listen to the live performance in an environment filled with a sense of presence.
[0094] Alternatively, the AV receiver 32 can also display icon images of "cheers", "applause", "shouts", and "noises" or the like on the display 401, and receive selection operations for these icon images from the audience via the user I / F 402, thereby receiving the reactions of the audience. It can also be that the AV receiver 32 generates sound signals corresponding to each reaction if it receives selection operations of these reactions, and transmits them to other devices as environmental information.
[0095] Alternatively, the AV receiver 32 can also transmit information indicating the kind of environmental sound such as cheers, applause, or shouts of the audience as environmental information. In this case, the receiving side device (for example, the transmission device 12 and the mixer 11) generates corresponding sound signals based on the environmental information, and provides the sound of the audience's cheers, applause, or shouts and the like to the venue. In this way, the environmental information is not a sound signal of the environmental sound, but information indicating a sound that should be generated, and it is possible to perform a process of playing a pre-recorded environmental sound or the like by the transmission device 12 and the mixer 11.
[0096] In addition, the environmental information of the first venue 10 can not be the environmental sound occurring at the first venue 10, but a pre-recorded environmental sound. In this case, the transmission device 12 transmits information indicating a sound that should be generated as environmental information. The playback device 22 or the AV receiver 32 plays the corresponding environmental sound based on the environmental information. In addition, it can also be that the background noise and the noise and the like in the environmental information are recorded sounds, and the other environmental sounds (for example, cheers, applause, shouts, and the like of the audience) are sounds occurring at the first venue 10.
[0097] In addition, the AV receiver 32 can also receive position information of the audience via the user I / F 402. The AV receiver 32 displays an image simulating a plan view or an inclined view or the like of the first venue 10 on the display 401 or the display 33, and receives position information from the audience via the user I / F 402 (for example, refer to Figure 16). The position information is information that specifies an arbitrary position in the first venue 10. The AV receiver 32 transmits the received position information of the audience to the first venue 10. The transfer device 12 and the mixer 11 of the first venue perform processing that positions the environmental sound of the third venue 20A at the specified position based on the environmental sound of the third venue 20A and the position information of the audience received from the AV receiver 32.
[0098] In addition, the AV receiver 32 can also change the content of the sound image movement processing based on the position information received from the user. For example, if the audience specifies a position immediately in front of the stage of the first venue 10, the AV receiver 32 performs sound image movement processing that sets the positioning position of the sound of the performer to the position immediately in front of the audience. Thereby, the audience of the third venue 20A can obtain a sense of presence as if they are located immediately in front of the stage of the first venue 10.
[0099] The sound of the audience of the third venue 20A can not be transmitted to the first venue 10, but can be transmitted to the second venue 20, and can also be transmitted to other venues. For example, the sound of the audience of the third venue 20A can be transmitted only to the friend's house (fourth venue). The audience of the fourth venue can view and listen to the live performance of the first venue 10 while listening to the sound of the audience of the third venue 20A. In addition, a not-illustrated playback device of the fourth venue can also transmit the sound of the audience of the fourth venue to the third venue 20A. In this case, the audience of the third venue 20A can view and listen to the live performance of the first venue 10 while listening to the sound of the audience of the fourth venue. Thereby, the audience of the third venue 20A and the audience of the fourth venue can view and listen to the live performance of the first venue 10 while conversing with each other.
[0100] Figure 13 is a block diagram that shows the structure of the live data transfer system 1C to which Modification 3 is applied. The structures common to those of the live data transfer system 1A shown in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted. Figure 1
[0101] In the live data transfer system 1C, the transfer device 12 is connected to the terminal 42 of the fifth venue 20B via the Internet 5. The terminal 42 is connected to the earphone 43. The fifth venue 20B is, for example, the house of a certain audience. However, in the case where the terminal 42 is mobile, the fifth venue 20B can also be various places such as inside a coffee shop, inside a car, or inside a public transportation device. In this case, all the places can become the fifth venue 20B. The user of the terminal 42 becomes an audience who views and listens to the live performance of the first venue 10 in a remote manner. In this case, the terminal 42 also reproduces the transfer data, and supplies the sound involved in the sound source information and the sound involved in the echo of the space to the second venue (in this example, the fifth venue 20B) via the earphone 43.
[0102] Figure 14 is a block diagram showing the structure of the terminal 42. The terminal 42 is an information processing apparatus such as a personal computer, a smartphone, or a tablet computer. The terminal 42 has a display 501, a user I / F 502, a CPU 503, a RAM 504, a network I / F 505, a flash memory 506, an audio I / O (Input / Output) 507, and a microphone 508.
[0103] The CPU 503 is a control section that controls the operation of the terminal 42. The CPU 503 reads a prescribed program stored in the flash memory 506 as a storage medium to the RAM 504 and executes it, thereby performing various operations.
[0104] Further, the program read by the CPU 503 need not be stored in the flash memory 506 in the apparatus. For example, the program can be stored in a storage medium of an external apparatus such as a server. In this case, the CPU 503 reads the program from the server to the RAM 504 each time and executes it.
[0105] The CPU 503 performs signal processing on a sound signal input via the network I / F 505. The CPU 503 outputs the audio signal after the signal processing to the earphone 43 via the audio I / O 507.
[0106] The CPU 503 receives the transmission data from the transmission apparatus 12 via the network I / F 505. The CPU 503 reproduces the transmission data and provides the sound related to the sound of the performer and the echo of the space to the audience of the 5th venue 20B.
[0107] Specifically, the CPU 503 convolves a sound signal related to the sound of the performer with a head transfer function (hereinafter, referred to as HRTF) to perform a sound image positioning process (binauralization process) in such a manner that the sound of the performer is positioned at the position of the performer. The HRTF corresponds to a transfer function between a prescribed position and the ears of the audience. The HRTF is a transfer function that represents the magnitude, the arrival time, and the frequency characteristics of the sound from a sound source at a certain position to the left and right ears, respectively. The CPU 503 convolves the HRTF with the sound signal of the sound of the performer based on the position of the performer. Thereby, the sound of the performer is positioned at the position corresponding to the position information.
[0108] Further, the CPU 503 performs the generation processing of the indirect sound by convoluting the binaural processing of the HRTF corresponding to the echo information of the space with respect to the sound signal of the sound of the performer. The CPU 503 performs the convolution with respect to the HRTF from the position of the virtual sound source corresponding to each of the initial reflection sounds included in the echo information of the space to the left and right ears, respectively, thereby localizing the initial reflection sounds and the back echo sound. However, the back echo sound is a reflection sound whose arrival direction of the sound is not fixed. Therefore, the CPU 503 can also perform the effect processing such as reverberation without performing the localization processing with respect to the back echo sound. Further, the CPU 503 can also perform the digital filter processing (earphone inverse characteristic processing) which reproduces the inverse characteristic of the acoustic characteristic of the earphone 43 used by the audience.
[0109] Further, the CPU 503 reproduces the environmental information in the transmission data, and provides the environmental sound occurring at the first venue 10 to the audience of the fifth venue 20B. The CPU 503 performs the localization processing based on the HRTF with respect to the environmental sound when the environmental information includes the position information of the environmental sound, and performs the effect processing with respect to the sound whose arrival direction of the sound is not fixed.
[0110] Further, the CPU 503 can also reproduce the image signal in the transmission data, and display the live image on the display 501.
[0111] Thus, the terminal 42 can also provide the sense of presence of the first venue 10 to the audience of the fifth venue 20B.
[0112] Further, the terminal 42 acquires the sound of the audience of the fifth venue 20B via the microphone 508. The terminal 42 transmits the sound of the audience to the other device. For example, the terminal 42 feeds back the sound of the audience to the first venue 10. Alternatively, the terminal 42 can also display the icon images of "cheering", "clapping", "shouting", and "noises" and the like on the display 501, receive the selection operation with respect to these icon images from the audience via the user I / F 502, and receive the reaction. The terminal 42 generates the sound corresponding to the received reaction, and transmits the generated sound as the environmental information to the other device. Alternatively, the terminal 42 can also transmit the information indicating the kind of the environmental sound of the cheering, clapping, or shouting of the audience as the environmental information. In this case, the device on the receiving side (for example, the transmission device 12 and the mixer 11) generates the corresponding sound signal based on the environmental information, and provides the sound of the cheering, clapping, or shouting of the audience to the venue.
[0113] Further, the terminal 42 can also receive the position information of the audience via the user I / F 502. The terminal 42 transmits the received position information of the audience to the first venue 10. The transmission device 12 and the mixer 11 of the first venue perform the process of positioning the sound of the audience at the specified position based on the sound and position information of the audience of the third venue 20A received from the AV receiver 32.
[0114] Further, the terminal 42 can also change the HRTF based on the position information received from the user. For example, if the audience specifies the position immediately in front of the stage of the first venue 10, the terminal 42 sets the position of the sound of the performer to the position immediately in front of the audience, and convolves the HRTF that positions the sound of the performer at the position. Thereby, the audience of the fifth venue 20B can obtain the sense of presence as if he or she is located immediately in front of the stage of the first venue 10.
[0115] The sound of the audience of the fifth venue 20B can be transmitted not to the first venue 10 but to the second venue 20, and can also be transmitted to other venues. As described above, the sound of the audience of the fifth venue 20B can be transmitted only to the friend's house (the fourth venue). Thereby, the audience of the fifth venue 20B and the audience of the fourth venue can watch and listen to the live performance of the first venue 10 while talking to each other.
[0116] Further, in the live data transmission system of the present embodiment, a plurality of users can also specify the same position. For example, a plurality of users can respectively specify the position immediately in front of the stage of the first venue 10. In this case, each audience can obtain the sense of presence as if he or she is located at the position immediately in front of the stage. Thereby, a plurality of audiences can watch and listen to the performance of the performer with the same sense of presence for one position (seat of the venue). In this case, the live operator can provide a service that exceeds the number of accommodatable audiences of the actual space.
[0117] Figure 15 is a block diagram showing the structure of the live data transmission system ID to which Modification 4 is applied. The structures common to those of the live data transmission system 1 shown in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted. Figure 1 The structures common to those of the live data transmission system 1 shown in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted.
[0118] The live data transmission system ID further has a server 50 and a terminal 55. The terminal 55 is provided at a sixth venue 10A. The server 50 is one example of the transmission device, and the hardware structure of the server 50 is the same as that of the transmission device 12. The hardware structure of the terminal 55 is the same as that of the terminal 42 shown in FIG. 1. Figure 14 The hardware structure of the terminal 55 is the same as that of the terminal 42 shown in FIG. 1.
[0119] The 6th venue 10A is a residence or the like of a performer who performs a performance such as playing in a remote manner. The performer at the 6th venue 10A performs a performance such as playing in synchronization with the playing or singing at the 1st venue. The terminal 55 transmits the voice of the performer at the 6th venue 10A to the server 50. In addition, the terminal 55 can also take an image of the performer at the 6th venue 10A by a not-illustrated camera and transmit an image signal to the server 50.
[0120] The server 50 transmits transmission data including the voice of the performer at the 1st venue 10, the voice of the performer at the 6th venue 10A, the echo information of the space at the 1st venue 10, the environmental information at the 1st venue 10, the live image at the 1st venue 10, and the image of the performer at the 6th venue 10A.
[0121] In this case, the playback device 22 reproduces the transmission data and provides the voice of the performer at the 1st venue 10, the voice of the performer at the 6th venue 10A, the echo of the space at the 1st venue 10, the environmental sound at the 1st venue 10, the live image at the 1st venue 10, and the image of the performer at the 6th venue 10A to the 2nd venue 20. For example, the playback device 22 displays the image of the performer at the 6th venue 10A in superposition with the live image at the 1st venue 10.
[0122] The voice of the performer at the 6th venue 10A can not be subjected to the positioning process or can be positioned at a position matching the image displayed on the display. For example, in a case where the performer at the 6th venue 10A is displayed on the right side within the live image, the voice of the performer at the 6th venue 10A is positioned on the right side.
[0123] In addition, the performer at the 6th venue 10A or the transmitter of the transmission data can specify the position of the performer. In this case, the position information of the performer at the 6th venue 10A is included in the transmission data. The playback device 22 positions the voice of the performer at the 6th venue 10A on the basis of the position information of the performer at the 6th venue 10A.
[0124] The image of the performer at the 6th venue 10A is not limited to an image taken by a camera. For example, a 2-dimensional image or a virtual character image (virtual image) composed of 3D modeling can be transmitted as the image of the performer at the 6th venue 10A.
[0125] Further, the transmission data can include audio data in addition to the transmission data. For example, the transmission device can transmit transmission data including the sound of the performer of the first venue 10, audio data, echo information of the space of the first venue 10, environmental information of the first venue 10, live video of the first venue 10, and video data. In this case, the playback device reproduces the transmission data and provides the sound of the performer of the first venue 10, the sound involved in the audio data, the echo of the space of the first venue 10, the environmental sound of the first venue 10, the live video of the first venue 10, and the video involved in the video data to the other venues. The playback device 22 displays the video of the performer corresponding to the video data superimposed on the live video of the first venue 10.
[0126] Further, the transmission device can determine the type of the musical instrument when the sound involved in the audio data is recorded. In this case, the transmission device transmits the transmission data including information indicating the type of the musical instrument determined from the audio data. The playback device generates the video of the corresponding musical instrument based on the information indicating the type of the musical instrument. The playback device can display the video of the musical instrument superimposed on the live video of the first venue 10.
[0127] Further, the transmission data does not necessarily superimpose the video of the performer of the sixth venue 10A on the live video of the first venue 10. For example, the transmission data can transmit the videos of the respective performers of the first venue 10 and the sixth venue 10A and the background video as separate data. In this case, the transmission data includes information indicating the display positions of the respective videos. The playback device reproduces the videos of the respective performers based on the information indicating the display positions.
[0128] Further, the background video is not limited to the video of the venue where the live performance is actually performed, such as the first venue 10. The background video can be the video of a venue different from the venue where the live performance is performed.
[0129] Further, the echo information of the space included in the transmission data does not necessarily correspond to the echo of the space of the first venue 10. For example, the echo information of the space can be virtual space information (information indicating the size, shape, material of the wall surface, and the like of the space of each venue, or an impulse response indicating the transfer function of each venue) for virtually reproducing the echo of the space of the venue corresponding to the background video. The impulse response of each venue can be measured in advance or can be obtained by simulation based on the size, shape, and material of the wall surface of the space of each venue.
[0130] Further, the environmental information can also be changed to content corresponding to the background image. For example, in the case of a background image of a large hall, the environmental information includes sounds of the audience's cheers, applause, cheers, and the like. In addition, an outdoor venue includes background noise that is different from that of an indoor venue. Further, the echo of the environmental sound also changes in correspondence with the echo information of the space described above. Further, the environmental information can also include information indicating the number of audience members, information indicating the degree of crowding (the density of people). The playback device increases or decreases the number of sounds of the audience's cheers, applause, cheers, and the like based on the information indicating the number of audience members. Further, the playback device increases or decreases the volume of the audience's cheers, applause, cheers, and the like based on the information indicating the degree of crowding.
[0131] Alternatively, the environmental information can also be changed in accordance with the performer. For example, in the case of a performer who has many female fans performing live, the sounds of the audience's cheers, shouts, cheers, and the like included in the environmental information are changed to female voices. The environmental information can include sound signals of these audience voices, or can include information indicating the attributes of the audience members, such as the male-to-female ratio or the age ratio. The playback device changes the tone quality of the audience's cheers, applause, cheers, and the like based on the information indicating the attributes.
[0132] Further, the audience at each venue can also specify the background image and the echo information of the space. The audience at each venue specifies the background image and the echo information of the space using the user I / F of the playback device.
[0133] Figure 16 is a diagram indicating one example of a live image 700 displayed by the playback device of each venue. The live image 700 is composed of an image obtained by photographing the 1st venue 10 or another venue, or a virtual image (computer graphics) corresponding to each venue, and the like. The live image 700 is displayed on the display of the playback device. The live image 700 displays images of the background, stage, performer including musical instruments, and audience in the venue, and the like. The images of the background, stage, performer including musical instruments, and audience in the venue can be all actually photographed images, or can be virtual images. Further, it can be that only the background image is an actually photographed image, and the other images are virtual images. Further, the live image 700 displays icon images 751 and 752 for specifying the space. The icon image 751 is an image for specifying the space of a certain venue, i.e., a stage A (Stage A) (for example, the 1st venue 10), and the icon image 752 is an image for specifying the space of another venue, i.e., a stage B (Stage B) (for example, another concert hall, or the like). Further, the live image 700 displays an audience image 753 for specifying the position of the audience.
[0134] The listener using the playback device specifies a desired space by designating either of the icon images 751 or 752 using the user I / F of the playback device. The transmission device transmits the background image and the echo information of the space corresponding to the specified space, included in the transmission data. Alternatively, the transmission device can transmit a plurality of background images and echo information of the spaces included in the transmission data. In this case, the playback device reproduces the background image and the echo information of the space corresponding to the space specified by the listener, included in the received transmission data.
[0135] In the example of FIG. 7, the icon image 751 is designated. The playback device displays the background image corresponding to Stage A of the icon image 751 (e.g., the image of the 1st hall 10), and plays the sound involved in the echo of the space corresponding to the designated Stage A. If the listener designates the icon image 752, the playback device switches to display the background image of the other space, Stage B, corresponding to the icon image 752, and plays the sound involved in the echo of the corresponding other space based on the virtual space information corresponding to Stage B. Figure 16
[0136] Thus, the listener of each playback device can obtain the sense of presence as if the live performance is viewed and heard in the desired space.
[0137] In addition, the listener of each playback device can designate a desired position in the hall by moving the listener image 753 within the live image 700. The playback device performs the localization processing based on the position designated by the user. For example, if the listener moves the listener image 753 to a position immediately in front of the stage, the playback device sets the localization position of the sound of the performer to the position immediately in front of the listener, and performs the localization processing so that the sound of the performer is localized at the position. Thus, the listener of each playback device can obtain the sense of presence as if he or she is immediately in front of the stage.
[0138] In addition, as described above, if the position of the sound source and the position of the listener (the position of the pickup point) change, the sound involved in the echo of the space also changes. The playback device can calculate the initial reflected sound in the case where the space has changed, the case where the position of the sound source has changed, or the case where the position of the pickup point has changed. Therefore, even if the measurement of the impulse response or the like is not performed in the actual space, the playback device can calculate the sound involved in the echo of the space based on the virtual space information. Thus, the playback device can accurately realize the echo generated in the space including the actual space.
[0139] For example, the mixer 11 can function as a transmission device, and the mixer 21 can function as a playback device. In addition, the playback device does not need to be set up in each venue. Figure 15 The server 50 shown may also reproduce the transmission data and transmit the audio signal after signal processing to the terminals at each venue, etc. In this case, the server 50 functions as a playback device.
[0140] The sound source information may also include information indicating the performer's posture (e.g., the performer's left and right orientation). The playback device may also adjust the volume or frequency characteristics based on the performer's posture information. For example, the playback device may use the performer's orientation facing directly opposite as a reference, and the greater the left and right orientation, the lower the volume. In addition, the playback device may also perform a process in which the high-frequency band is attenuated more significantly than the low-frequency band as the left and right orientation increases. Thus, the sound changes in accordance with the performer's posture, so that the audience can view and listen to a more immersive live performance.
[0141] then, Figure 17 This is a block diagram showing an application example of signal processing performed by a playback device. In this example, Figure 13 The terminal 42 and the earphone 43 are shown for reproduction. Figure 13 In the example, the terminal 42) functionally includes a musical instrument model processing unit 551, an amplifier model processing unit 552, a speaker model processing unit 553, a space model processing unit 554, a binaural processing unit 555, and a headphone inverse characteristic processing unit 556.
[0142] The musical instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 perform signal processing to impart the acoustic characteristics of the acoustic device to the sound signal associated with the performance sound. The first digital signal processing model used for this signal processing is, for example, included in the sound source information transmitted by the transmission device 12. The first digital signal processing model is a digital filter that simulates the acoustic characteristics of the musical instrument, the amplifier, and the speaker. The first digital signal processing model is pre-created by the manufacturer of the musical instrument, the amplifier, and the speaker through simulation, etc. The musical instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 perform digital filtering processing that simulates the acoustic characteristics of the musical instrument, the amplifier, and the speaker, respectively. Furthermore, if the musical instrument is an electronic instrument such as a synthesizer, the musical instrument model processing unit 551 inputs note event data (information indicating the timing and pitch of the sound to be produced) instead of the sound signal, generating a sound signal having the acoustic characteristics of an electronic instrument such as a synthesizer.
[0143] Thus, the playback device can reproduce the acoustic characteristics of any musical instrument, etc. Figure 16 In the video, a virtual image (computer graphics) of a live video 700 is displayed. Here, listeners using the playback device can also use the playback device's user interface to change to another virtual instrument image. When a listener changes the instrument displayed in live video 700 to another instrument image, the playback device's instrument model processing unit 551 performs signal processing based on the first digital signal processing model corresponding to the changed instrument. As a result, the playback device outputs a sound that reproduces the acoustic characteristics of the instrument displayed in live video 700.
[0144] Similarly, listeners using the playback device can also use the playback device's user interface to change the type of amplifier and speaker to different types. The amplifier model processing unit 552 and the speaker model processing unit 553 perform digital filtering processing that simulates the acoustic characteristics of the changed amplifier and speaker types. Furthermore, the speaker model processing unit 553 can also simulate the acoustic characteristics of the speaker in each direction. In this case, listeners using the playback device can also use the playback device's user interface to change the speaker's orientation. The speaker model processing unit 553 performs digital filtering processing corresponding to the changed speaker orientation.
[0145] The space model processing unit 554 is a second digital signal processing model that reproduces the acoustic characteristics of the room at the live venue (e.g., the echo in the space). This second digital signal processing model can also be obtained, for example, using test sounds in the actual live venue. Alternatively, the second digital signal processing model can calculate the delay and level of the virtual sound source based on virtual space information (information indicating the size, shape, and wall material of each venue) as described above.
[0146] If the position of the sound source and the listener (the location of the sound pickup point) changes, the sound associated with the spatial echo also changes. The playback device can calculate the delay and level of the virtual sound source even when the space, the sound source, and the sound pickup point change. Therefore, even without measuring impulse responses or other parameters in the actual space, the playback device can determine the sound associated with the spatial echo based on the virtual space information. This allows the playback device to accurately reproduce the echo generated in a space that also includes the actual space.
[0147] Further, the virtual space information can also include information of the position and material of a structure such as a pillar (an obstacle in terms of sound). The playback device reproduces phenomena of reflection, shielding, and diffraction caused by the obstacle in a case where there is an obstacle in the path of the direct sound and the indirect sound from the sound source to the listening point.
[0148] Figure 18 is a schematic view showing the path of the sound from the sound source 70 reflected by the wall surface to reach the listening point 75. Figure 18 The sound source 70 shown can be either a performance sound (1st sound source) or an environmental sound (2nd sound source). The playback device calculates the position of a virtual sound source 70A existing with the wall surface as a mirror surface with respect to the position of the sound source 70 based on the position of the sound source 70, the position of the wall surface, and the position of the listening point 75. Further, the playback device calculates the delay amount of the virtual sound source 70A based on the distance from the virtual sound source 70A to the listening point 75. In addition, the playback device calculates the level of the virtual sound source 70A based on the information of the material of the wall surface. Moreover, the playback device performs a delay process of the sound from the sound source 70 based on the delay amount of the virtual sound source 70A and the level of the virtual sound source 70A. Figure 18 The playback device calculates the frequency characteristics generated by the diffraction through the obstacle 77 in a case where there is an obstacle 77 in the path from the position of the virtual sound source 70A to the listening point 75 as shown in Figure 18 The playback device performs an equalization process of reducing the level of the high frequency band in a case where there is an obstacle 77 in the path from the position of the virtual sound source 70A to the listening point 75 as shown in
[0149] In addition, the playback device can set a new 2nd virtual sound source 77A and a 3rd virtual sound source 77B at positions left and right of the obstacle 77. The 2nd virtual sound source 77A and the 3rd virtual sound source 77B correspond to new sound sources generated by the diffraction. The 2nd virtual sound source 77A and the 3rd virtual sound source 77B are sounds to which the frequency characteristics generated by the diffraction are imparted to the sound from the virtual sound source 70A, respectively. The playback device recalculates the delay amount and the level based on the positions of the 2nd virtual sound source 77A and the 3rd virtual sound source 77B and the position of the listening point 75. Thereby, it is possible to reproduce the diffraction phenomena of the obstacle 77.
[0150] The playback device can also calculate the delay amount and the level of the sound from the virtual sound source 70A reflected by the obstacle 77 and further reflected by the wall surface to reach the listening point 75. In addition, the playback device can eliminate the virtual sound source 70A in a case where it is determined that the virtual sound source 70A is shielded by the obstacle 77. Information of the determination of the shielding can also be included in the virtual space information.
[0151] By performing the above-described processing, the playback device performs first digital signal processing that expresses the acoustic characteristics of the audio equipment and second digital signal processing that expresses the acoustic characteristics of the room, thereby generating sound of the sound source and sound related to the spatial echo.
[0152] The binaural processing unit 555 convolves the sound signal with a head transfer function (hereinafter referred to as HRTF) to perform sound image localization processing for the sound source and various indirect sounds. The headphone inverse characteristic processing unit 556 performs digital filtering processing to reproduce the inverse characteristics of the acoustic characteristics of the headphones used by the listener.
[0153] Through the above-described processing, the user can obtain a sense of presence as if viewing a live performance in a desired space and using desired audio equipment.
[0154] In addition, the playback device does not need to have Figure 17 The instrument model processing unit 551, amplifier model processing unit 552, speaker model processing unit 553, and space model processing unit 554 are all shown. The playback device can perform signal processing using at least one digital signal processing model. In addition, the playback device can perform signal processing on a certain sound signal (for example, the voice of a certain performer) using one digital signal processing model, and can also perform signal processing on multiple sound signals using one digital signal processing model respectively. The playback device can perform signal processing on a certain sound signal (for example, the voice of a certain performer) using multiple digital signal processing models, and can also perform signal processing on multiple sound signals using multiple digital signal processing models. The playback device can also perform signal processing on ambient sound using a digital signal processing model.
[0155] The description of the present embodiment is illustrative in all respects and is not intended to be restrictive. The scope of the present invention is not indicated by the above-described embodiment but by the claims. Furthermore, the scope of the present invention includes all modifications within the meaning and scope equivalent to the claims.
[0156] Description of the label
[0157] 1. 1A, 1B, 1C, 1D…Field data transmission system
[0158] 5…Internet
[0159] 10…Site 1
[0160] 10A…Site 6
[0161] 11…Mixer
[0162] 12…Transmission device
[0163] 13A~13F…Microphone
[0164] 14A to 14G... speakers
[0165] 15A to 15C... trackers
[0166] 16... camera
[0167] 20... 2nd venue
[0168] 20A... 3rd venue
[0169] 20B... 5th venue
[0170] 21... mixer
[0171] 22... player
[0172] 23... display
[0173] 24A to 24F... speakers
[0174] 25A to 25C... microphones
[0175] 32... AV receiver
[0176] 33... display
[0177] 34A... speaker
[0178] 35... microphone
[0179] 42... terminal
[0180] 43... earphone
[0181] 50... server
[0182] 55... terminal
[0183] 101... display
[0184] 102... user I / F
[0185] 103... audio I / O
[0186] 104... signal processing section
[0187] 105... network I / F
[0188] 106... CPU
[0189] 107... flash memory
[0190] 108... RAM
[0191] 201... display
[0192] 202... user I / F
[0193] 203... CPU
[0194] 204... RAM
[0195] 205... Network I / F
[0196] 206... Flash memory
[0197] 207... General-purpose communication I / F
[0198] 301... Display
[0199] 302... User I / F
[0200] 303... CPU
[0201] 304... RAM
[0202] 305... Network I / F
[0203] 306... Flash memory
[0204] 307... Video I / F
[0205] 401... Display
[0206] 402... User I / F
[0207] 403... Audio I / O
[0208] 404... Signal processing section
[0209] 405... Network I / F
[0210] 406... CPU
[0211] 407... Flash memory
[0212] 408... RAM
[0213] 409... Video I / F
[0214] 501... Display
[0215] 503... CPU
[0216] 504... RAM
[0217] 505... Network I / F
[0218] 506... Flash memory
[0219] 507... Audio I / O
[0220] 508... Microphone
[0221] 700... Live video
Claims
1. A live data transmission method wherein, sound source information related to a sound occurring at a first venue and information related to an initial reflected sound varying in correspondence with a position of the sound, i.e., spatial echo information, are transmitted as transmission data, the transmission data is reproduced to provide the sound related to the sound source information and the sound related to the spatial echo to a second venue, in the live data transmission method, the spatial echo information includes information for generating an indirect sound, the reproduction includes a process of generating an indirect sound of the sound of the sound source, the process of generating the indirect sound includes a process of generating an initial reflected sound, the generated initial reflected sound varies in correspondence with the position of the sound, the spatial echo information includes information related to a size, a shape, and an energy loss at a time of reflection of a wall surface of the space, in the reproduction, a delay amount and a level of the sound related to the spatial echo of the space are calculated based on the information related to the size, the shape, and the energy loss at the time of reflection of the wall surface of the space, and the initial reflected sound is generated.
2. The live data transmission method according to claim 1, wherein, the sound source information includes a sound signal of the sound occurring at the first venue and position information of the sound, the reproduction includes a positioning process corresponding to the position of the sound.
3. The live data transmission method according to claim 1 or 2, wherein, the spatial echo information includes an impulse response varying in correspondence with the position of the sound.
4. The live data transmission method according to claim 1 or 2, wherein, environment information related to an ambient sound is included in the transmission data and transmitted, the reproduction includes a process of further providing the ambient sound.
5. The live data transmission method according to claim 1 or 2, wherein, the spatial echo information includes virtual space information for reproducing an echo other than an echo of the first venue, the reproduction is a play of the sound related to the spatial echo based on the virtual space information.
6. The live data transmission method according to claim 5, wherein, an operation of designating a space is received from a user at the second venue, the reproduction is a play of the sound related to the spatial echo based on the virtual space information corresponding to the space received by the operation.
7. The live data transmission method according to claim 5, wherein, a live image is provided to the second venue, the virtual space information corresponds to an echo of a space of the live image.
8. The live data transmission method according to claim 1 or 2, wherein, the transmission data is reproduced to provide the sound related to the sound source information and the sound related to the spatial echo to a third venue, the sound related to the spatial echo is common to the second venue and the third venue.
9. The live data transmission method according to claim 1 or 2, wherein, the sound source information includes a first digital signal processing model exhibiting an acoustic characteristic of an acoustic device, the first digital signal processing model is used in the reproduction. The sound related to the sound source information is provided to the second venue using the first digital signal processing model.
10. The live data transmission method according to claim 1 or 2, wherein The echo information of the space includes a second digital signal processing model that represents the acoustic characteristics of the room, The sound related to the echo of the space is provided to the second venue using the second digital signal processing model.
11. The live data transmission method according to claim 1 or 2, wherein Position information representing the position of a structure provided in the first venue is transmitted, The sound source information includes a sound signal of a sound occurring in the first venue and position information of the sound, The reproduction includes signal processing based on the position information representing the position of the structure and the position information of the sound.
12. A live data transmission system having: a live data transmission device that transmits sound source information related to a sound occurring in a first venue and echo information of a space as transmission data; and a live data playback device that reproduces the transmission data and provides a sound related to the sound source information and a sound related to the echo of the space to a second venue, In the live data transmission system, The echo information of the space includes information for generating a reverberation sound, The reproduction includes processing for generating a reverberation sound of the sound of the sound source, The processing for generating the reverberation sound includes processing for generating an initial reflection sound, The generated initial reflection sound varies in accordance with the position of the sound, The echo information of the space includes information related to the size, shape, and energy loss at the reflection of the wall surface of the space, In the reproduction, the amount of delay and the level of the sound related to the echo of the space are calculated based on the information related to the size, shape, and energy loss at the reflection of the wall surface of the space, and the initial reflection sound is generated.
13. The live data transmission system according to claim 12, wherein The sound source information includes a sound signal of a sound occurring in the first venue and position information of the sound, The reproduction includes positioning processing corresponding to the position of the sound.
14. The live data transmission system according to claim 12, wherein The echo information of the space includes an impulse response that varies in accordance with the position of the sound.
15. The live data transmission system according to any one of claims 12 to 14, wherein The live data transmission device transmits environmental information related to an environmental sound as the transmission data, The reproduction includes processing for further providing the environmental sound.
16. The live data transmission system according to any one of claims 12 to 14, wherein The echo information of the space includes virtual space information for reproducing an echo other than the echo of the first venue, The reproduction is playback of the sound related to the echo of the space based on the virtual space information.
17. The live data transmission system according to claim 16, wherein The live data playback device receives an operation for designating a space from a user of the second venue, The reproduction is playback of a sound involved in an echo of the space based on the virtual space information corresponding to the space received by the operation.
18. The live data transmission system according to claim 16, wherein The live data playback device provides a live image to the second venue, The virtual space information corresponds to an echo of the space of the live image.
19. The live data transmission system according to any one of claims 12 to 14, wherein The live data playback device reproduces the transmission data to provide a sound involved in the sound source information and a sound involved in an echo of the space to a third venue, The sound involved in the echo of the space is common to the second venue and the third venue.
20. The live data transmission system according to any one of claims 12 to 14, wherein The sound source information includes a first digital signal processing model representing an acoustic characteristic of an acoustic device, The live data playback device provides the sound involved in the sound source information to the second venue using the first digital signal processing model.
21. The live data transmission system according to any one of claims 12 to 14, wherein The echo information of the space includes a second digital signal processing model representing an acoustic characteristic of a room, The live data playback device provides the sound involved in the echo of the space to the second venue using the second digital signal processing model.
22. The live data transmission system according to any one of claims 12 to 14, wherein The live data transmission device transmits position information representing a position of a structure provided in the first venue, The sound source information includes a sound signal of a sound occurring in the first venue and position information of the sound, The reproduction includes signal processing based on the position information representing the position of the structure and the position information of the sound.
23. A live data transmission device, wherein The live data transmission device transmits sound source information related to a sound occurring in a first venue and echo information of a space as transmission data, A playback device reproduces the transmission data to provide a sound involved in the sound source information and a sound involved in an echo of the space to a second venue, In the live data transmission device, The echo information of the space includes information for generating a reflected sound, The reproduction includes processing for generating a reflected sound of the sound, The processing for generating the reflected sound includes processing for generating an initial reflected sound, The generated initial reflected sound varies in accordance with a position of the sound, The echo information of the space includes information related to a size, a shape, and an energy loss at a time of reflection of a wall surface of the space, In the reproduction, a delay amount and a level of a sound involved in an echo of the space are calculated based on the information related to the size, the shape, and the energy loss at the time of reflection of the wall surface of the space, and the initial reflected sound is generated.
24. A live data playback apparatus, wherein the live data playback apparatus receives transmission data from a live data transmission apparatus which transmits sound source information relating to a sound occurring at a first venue and spatial echo information as the transmission data, the live data playback apparatus reproduces the transmission data to provide a sound relating to the sound source information and a sound relating to the spatial echo to a second venue, in the live data playback apparatus, the spatial echo information includes information for generating an indirect sound, the reproducing includes a process of generating an indirect sound of the sound of the sound source, the process of generating the indirect sound includes a process of generating an initial reflection sound, the generated initial reflection sound varies in correspondence with a position of the sound, the spatial echo information includes information relating to a size, a shape, and an energy loss at a reflection of a wall surface of the space, in the reproducing, a delay amount and a level of a sound relating to the spatial echo are calculated based on the information relating to the size, the shape, and the energy loss at the reflection of the wall surface of the space, and the initial reflection sound is generated.
25. A live data transmission method, wherein sound source information relating to a sound occurring at a first venue and spatial echo information are transmitted as transmission data, a playback apparatus is caused to reproduce the transmission data to provide a sound relating to the sound source information and a sound relating to the spatial echo to a second venue, in the live data transmission method, the spatial echo information includes information for generating an indirect sound, the reproducing includes a process of generating an indirect sound of the sound of the sound source, the process of generating the indirect sound includes a process of generating an initial reflection sound, the generated initial reflection sound varies in correspondence with a position of the sound, the spatial echo information includes information relating to a size, a shape, and an energy loss at a reflection of a wall surface of the space, in the reproducing, a delay amount and a level of a sound relating to the spatial echo are calculated based on the information relating to the size, the shape, and the energy loss at the reflection of the wall surface of the space, and the initial reflection sound is generated.
26. A live data playback method, wherein transmission data is received from a live data transmission apparatus which transmits sound source information relating to a sound occurring at a first venue and spatial echo information as the transmission data, the transmission data is reproduced to provide a sound relating to the sound source information and a sound relating to the spatial echo to a second venue, in the live data playback method, the spatial echo information includes information for generating an indirect sound, the reproducing includes a process of generating an indirect sound of the sound of the sound source, the process of generating the indirect sound includes a process of generating an initial reflection sound, the generated initial reflection sound varies in correspondence with a position of the sound, the spatial echo information includes information relating to a size, a shape, and an energy loss at a reflection of a wall surface of the space, in the reproducing, a delay amount and a level of a sound relating to the spatial echo are calculated based on the information relating to the size, the shape, and the energy loss at the reflection of the wall surface of the space, and the initial reflection sound is generated. In the reproduction, based on information on the size, shape, and energy loss involved in reflection of the wall surface of the space, the delay amount and the level of the sound involved in the echo of the space are found, and the initial reflected sound is generated.
Citation Information
Patent Citations
Reflected and direct rendering of upmixed content to individually specified drivers.
JP2015530043A
Reproducing device, reproducing method, information processing device, information processing method, and program
CN109983786A
Sound signal processing method and sound field reproduction system
JP2007041164A
Signal processing device, signal processing method, and program
JP2019192975A
Information processing device, information processing method, program, and information processing system
JP2020053791A