Field data transmission method, field data transmission system, field data transmission device, and field data playback device

By transmitting sound source information, spatial echo information and ambient sound information, and performing signal processing on the receiving end, the problem that the existing technology cannot provide a sense of presence in remote venues is solved, and high-quality on-site data transmission and on-site experience are achieved.

CN120148464APending Publication Date: 2025-06-13YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510284440.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-11-27
Filing Date
2021-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art cannot provide a sense of on-site experience to remote venues while transmitting on-site data.

Method used

By transmitting the sound source information, spatial echo information and ambient sound information with the on-site venue, and performing signal processing and reproduction at the receiving end, positioning processing and indirect sound generation are provided to simulate the sound field of the remote venue.

Benefits of technology

It realizes the same on-site experience as the on-site venue while transmitting on-site data in a remote venue, enhancing the audience's immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148464A_ABST
    Figure CN120148464A_ABST
Patent Text Reader

Abstract

A field data transmission method transmits, as transmission data, first sound source information relating to sound of a first sound source generated at a first location of a first venue and position information of the first sound source, and second sound source information relating to a second sound source generated at a second location of the first venue, and reproduces the transmission data. And providing the sound of the first sound source and the sound of the second sound source to a second meeting place, the sound of the first sound source and the sound of the second sound source having been subjected to positioning processing on the basis of the position information of the first sound source.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese national application No. 202180009062.7 (Live Data Transmission Method, Live Data Transmission System, Its Transmission Device, Live Data Playback Device, and Its Playback Method) filed on March 19, 2021, the content of which is hereby incorporated by reference. Technical Field

[0002] One embodiment of the present invention relates to a method for transmitting live data, a live data transmission system, a live data transmission device, a live data playback device, and a live data playback method. Background Art

[0003] Patent Document 1 discloses a game viewing method that enables a user to effectively experience the excitement of a game as if they were in a stadium on a terminal for viewing a sports game.

[0004] The game viewing method of Patent Document 1 sends reaction information representing a user's reaction from each user's terminal. Each user's terminal displays icon information based on the reaction information.

[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2019-024157 Summary of the Invention

[0006] The system of Patent Document 1 only displays icon information and is not a system that provides the sense of presence of the live venue to the destination venue when transmitting live data.

[0007] An object of one embodiment of the present invention is to provide a live data transmission method, a live data transmission system, a live data transmission device, a live data playback device, and a live data playback method that can provide the sense of presence of the live venue to the destination venue even when transmitting live data.

[0008] The live data transmission method transmits, as transmission data, first sound source information related to the sound of a first sound source occurring at a first location in a first venue and the position information of the first sound source, and second sound source information related to a second sound source occurring at a second location in the first venue, reproduces the transmission data, and provides the sound of the first sound source that has undergone positioning processing based on the position information of the first sound source and the sound of the second sound source to a second venue.

[0009] Advantages of the Invention

[0010] The live data transmission method can provide the sense of presence of the live venue to the destination venue even when transmitting live data. Brief Description of the Drawings

[0011] Figure 1 It is a block diagram showing the structure of the live data transmission system 1.

[0012] Figure 2 It is a plan view of the first venue 10.

[0013] Figure 3 It is a plan view of the second venue 20.

[0014] Figure 4 It is a block diagram showing the structure of the mixer 11.

[0015] Figure 5 It is a block diagram showing the structure of the transmission device 12.

[0016] Figure 6 It is a flowchart showing the operation of the transmission device 12.

[0017] Figure 7 It is a block diagram showing the structure of the playback device 22.

[0018] Figure 8 It is a flowchart showing the operation of the playback device 22.

[0019] Figure 9 It is a block diagram showing the structure of the live data transmission system 1A related to Modification Example 1.

[0020] Figure 10 It is a schematic plan view of the second venue 20 of the live data transmission system 1A related to Modification Example 1.

[0021] Figure 11 It is a block diagram showing the structure of the live data transmission system 1B related to Modification Example 2.

[0022] Figure 12 It is a block diagram showing the structure of the AV receiver 32.

[0023] Figure 13 It is a block diagram showing the structure of the live data transmission system 1C related to Modification Example 3.

[0024] Figure 14 It is a block diagram showing the structure of the terminal 42.

[0025] Figure 15 It is a block diagram showing the structure of the live data transmission system 1D related to Modification Example 4.

[0026] Figure 16 It is a diagram showing an example of the live image 700 displayed on the playback device at each venue.

[0027] Figure 17It is a block diagram showing an application example of signal processing performed by a playback device.

[0028] Figure 18 It is a schematic diagram showing the path of sound that is reflected from the sound source 70 on the wall surface and reaches the sound collection point 75. Detailed implementation mode

[0029] Figure 1 It is a block diagram showing the structure of the on-site data transmission system 1. The on-site data transmission system 1 is composed of a plurality of audio devices and information processing devices respectively provided in the first venue 10 and the second venue 20.

[0030] Figure 2 It is a plan view of the first venue 10, Figure 3 It is a plan view of the second venue 20. In this example, the first venue 10 is the on-site venue where the performer performs. The second venue 20 is a public viewing venue where remote audiences watch and listen to the performer's performance.

[0031] In the first venue 10, a mixer 11, a transmission device 12, a plurality of microphones 13A to

[0032] 13F, a plurality of speakers 14A to 14G, a plurality of trackers 15A to 15C, and a camera 16 are provided. In the second venue 20, a mixer 21, a playback device 22, a display 23, and a plurality of speakers 24A to 24F are provided. The transmission device 12 and the playback device 22 are connected via the Internet 5. In addition, the number of microphones, the number of speakers, the number of trackers, etc. are not limited to the numbers shown in this embodiment. Also, the installation methods of the microphones and speakers are not limited to the examples shown in this embodiment.

[0033] The mixer 11 is connected to the transmission device 12, a plurality of microphones 13A to 13F, a plurality of speakers 14A to 14G, and a plurality of trackers 15A to 15C. The mixer 11, a plurality of microphones 13A to 13F, and a plurality of speakers 14A to 14G are connected via network cables or audio cables. The plurality of trackers 15A to 15C are connected to the mixer 11 via wireless communication. The mixer 11 and the transmission device 12 are connected via network cables. In addition, the transmission device 12 is connected to the camera 16 via a video cable. The camera 16 shoots the on-site image including the performer.

[0034] A plurality of speakers 14A to 14G are arranged along the wall surface of the first venue 10. The first venue 10 in this example is rectangular when viewed from above. A stage is arranged in front of the first venue 10. Performers perform shows such as singing or playing on the stage. Speaker 14A is arranged on the left side of the stage, speaker 14B is arranged in the center of the stage, and speaker 14C is arranged on the right side of the stage. Speaker 14D is arranged on the left side of the front-rear center of the first venue 10, and speaker 14E is arranged on the right side of the front-rear center of the first venue 10. Speaker 14F is arranged on the left side of the rear of the first venue 10, and speaker 14G is arranged on the right side of the rear of the first venue 10.

[0035] Microphones 13A are arranged on the left side of the stage, microphone 13B is arranged in the center of the stage, and microphone 13C is arranged on the right side of the stage. Microphone 13D is arranged on the left side of the front-rear center of the first venue 10, microphone 13E is arranged at the rear center of the first venue 10. Microphone 13F is arranged on the right side of the front-rear center of the first venue 10.

[0036] The mixer 11 receives sound signals from microphones 13A to 13F. In addition, the mixer 11 outputs sound signals to speakers 14A to 14G. In the present embodiment, speakers and microphones are shown as an example of audio devices connected to the mixer 11, but actually a plurality of audio devices are connected to the mixer 11. The mixer 11 receives sound signals from a plurality of audio devices such as microphones, performs signal processing such as mixing, and outputs sound signals to a plurality of audio devices such as speakers.

[0037] Microphones 13A to 13F acquire the singing sounds or playing sounds of individual performers as the sounds occurring in the first venue 10. Alternatively, microphones 13A to 13F acquire the ambient sounds of the first venue 10. In Figure 2 the example, microphones 13A to 13C acquire the sounds of the performers, and microphones 13D to 13F acquire the ambient sounds. The ambient sounds include sounds such as the cheering, applause, shouts, cheers, chorus, or noise of the audience. However, the sounds of the performers can be input by line. Line input means that instead of the microphone picking up the sound output from a sound source such as an instrument and inputting it, the sound signal is input from an audio cable or the like connected to the sound source. The sounds of the performers are preferably acquired with a high signal-to-noise ratio and do not include other sounds.

[0038] Speakers 14A to 14G output the sounds of the performers to the first venue 10. In addition, speakers 14A to 14G may also output initial reflection sounds or rear reverberation sounds for controlling the sound field of the first venue 10.

[0039] The mixer 21 in the second venue 20 is connected to the playback device 22 and multiple speakers 24A to 24F. These audio devices are connected via network cables or audio cables. In addition, the playback device 22 is connected to the display 23 via a video cable.

[0040] The multiple speakers 24A to 24F are arranged along the wall surface of the second venue 20. The second venue 20 in this example is rectangular when viewed from above. The display 23 is arranged in front of the second venue 20. The live video captured in the first venue 10 is displayed on the display 23. The speaker 24A is arranged on the left side of the display 23, and the speaker 24B is arranged on the right side of the display 23. The speaker 24C is arranged on the left side of the front-rear center of the second venue 20, and the speaker 24D is arranged on the right side of the front-rear center of the second venue 20. The speaker 24E is arranged on the left side of the rear of the second venue 20, and the speaker 24F is arranged on the right side of the rear of the second venue 20.

[0041] The mixer 21 outputs an audio signal to the speakers 24A to 24F. The mixer 21 receives an audio signal from the playback device 22, performs signal processing such as mixing, and outputs the audio signal to multiple audio devices such as speakers.

[0042] The speakers 24A to 24F output the voices of the performers to the second venue 20. In addition, the speakers 24A to 24F output the initial reflected sound or rear reverberation sound for reproducing the sound field of the first venue 10. In addition, the speakers 24A to 24F output ambient sounds such as the cheers of the audience in the first venue 10 to the second venue 20.

[0043] Figure 4 It is a block diagram showing the structure of the mixer 11. In addition, since the mixer 21 has the same structure and function as the mixer 11, the structure of the mixer 11 is shown as a representative case in Figure 4 The mixer 11 has a display 101, a user I / F 102, an audio I / O (Input / Output) 103, a signal processing unit (DSP) 104, a network I / F 105, a CPU 106, a flash memory 107, and a RAM 108.

[0044] The CPU 106 is a control unit that controls the operation of the mixer 11. The CPU 106 reads a prescribed program stored in the flash memory 107 as a storage medium into the RAM 108 and executes it, thereby performing various operations.

[0045] In addition, the program read by the CPU 106 does not need to be stored in the flash memory 107 within this device. For example, the program can also be stored in a storage medium of an external device such as a server. In this case, the CPU 106 can execute by reading the program from the server into the RAM 108 each time.

[0046] The signal processing unit 104 is composed of a DSP for performing various signal processes. The signal processing unit 104 performs signal processes such as mixing process and filtering process on the sound signal input from a sound device such as a microphone via the audio I / O 103 or the network I / F 105. The signal processing unit 104 outputs the audio signal after the signal process to a sound device such as a speaker via the audio I / O 103 or the network I / F 105.

[0047] In addition, the signal processing unit 104 can also perform panning process, initial reflection sound generation process, and rear reverberation sound generation process. The panning process is a process of controlling the volume of the sound signal distributed to a plurality of speakers 14A to 14G so that the sound image is positioned at the position of the performer. In order to perform the panning process, the CPU 106 acquires the position information of the performer via the trackers 15A to 15C. The position information is information representing two-dimensional or three-dimensional coordinates with a certain position in the first venue 10 as a reference. The trackers 15A to 15C are, for example, tags that transmit and receive radio waves such as Bluetooth (registered trademark). The trackers 15A to 15C are installed on the performer or the musical instrument. At least three beacons are preset in the first venue 10. Each beacon measures the distance to the trackers 15A to 15C based on the time difference from when the radio wave is transmitted to when the radio wave is received. The CPU 106 can uniquely obtain the positions of the trackers 15A to 15C by measuring the distances from at least three beacons to the tags in advance by acquiring the position information of the beacons.

[0048] In this way, the CPU 106 acquires the position information of each performer via the trackers 15A to 15C, that is, the position information of the sound occurring in the first venue 10. The CPU 106 determines the volume of each sound signal output to the speakers 14A to 14G so that the sound image is positioned at the position of the performer based on the acquired position information and the positions of the speakers 14A to 14G. The signal processing unit 104 controls the volume of each sound signal output to the speakers 14A to 14G according to the control of the CPU 106. For example, the signal processing unit 104 increases the volume of the sound signal output to the speaker closer to the position of the performer and decreases the volume of the sound signal output to the speaker farther from the position of the performer. Thereby, the signal processing unit 104 can position the sound image of the performance sound and singing sound of the performer at a specified position.

[0049] The initial reflected sound generation process and the rear reverberation sound generation process are processes of convolving an impulse response with the sound of a performer through an FIR filter. The signal processing unit 104 convolves, for example, an impulse response acquired at a predetermined venue (a venue other than the first venue 10) with the sound of the performer. Thereby, the signal processing unit 104 controls the sound field of the first venue 10. Alternatively, the signal processing unit 104 may further feedback the sound acquired by the microphones provided near the ceiling and the walls of the first venue 10 to the speakers 14A to 14G, thereby controlling the sound field of the first venue 10.

[0050] The signal processing unit 104 outputs the sound of the performer and the position information of the performer to the transmission device 12. The transmission device 12 acquires the sound of the performer and the position information of the performer from the mixer 11.

[0051] In addition, the transmission device 12 acquires an image signal from the camera 16. The camera 16 photographs each performer or the entire first venue 10, etc., and outputs the image signal related to the live image to the transmission device 12.

[0052] Moreover, the transmission device 12 acquires the echo information of the space of the first venue 10. The echo information of the space is information for generating indirect sound. Indirect sound is the sound of the sound source reflected in the venue and reaching the audience, and includes at least the initial reflected sound and the rear reverberation sound. The echo information of the space includes, for example, information indicating the size, shape, and wall material of the space of the first venue 10, and the impulse response related to the rear reverberation sound. The information indicating the size, shape, and wall material of the space is information for generating the initial reflected sound. The information for generating the initial reflected sound may also be an impulse response. The impulse response is, for example, pre-measured in the first venue 10. In addition, the echo information of the space may be information that changes corresponding to the position of the performer. The information that changes corresponding to the position of the performer is, for example, an impulse response pre-measured for each position of the performer in the first venue 10. The transmission device 12 acquires, for example, the first impulse response when the sound of the performer occurs in front of the stage of the first venue 10, the second impulse response when the sound of the performer occurs on the left side of the stage, and the third impulse response when the sound of the performer occurs on the right side of the stage. However, the number of impulse responses is not limited to three. In addition, the impulse response does not need to be actually measured in the first venue 10. For example, it can be obtained by simulation based on the size, shape, and wall material of the space of the first venue 10.

[0053] In addition, the initial reflected sound is a reflected sound with a fixed arrival direction of the sound, and the rear reverberant sound is a reflected sound with an unfixed arrival direction of the sound. Compared with the initial reflected sound, the rear reverberant sound has less change caused by the change in the position of the performer's voice. Therefore, the echo information of the space can be in a form composed of the impulse response of the initial reflected sound that changes corresponding to the position of the performer and the impulse response of the rear reverberant sound that is constant regardless of the position of the performer.

[0054] In addition, the signal processing unit 104 can also acquire the environmental information related to the ambient sound and output it to the transmission device 12. The ambient sound is the sound acquired by the microphones 13D to 13F as described above, and includes background noise, the cheering of the audience, applause, shouts, cheers, chorus, or noise, etc. However, the ambient sound can also be acquired by the microphones 13A to 13C on the stage. The signal processing unit 104 outputs the sound signal related to the ambient sound as environmental information to the transmission device 12. In addition, the environmental information can also include the position information of the ambient sound. Sounds such as the cheering of each audience member in the ambient sound, the shouting of the performer's personal name, or exclamations such as "wonderful" are sounds that can be recognized as the voices of individual audience members without being drowned out by the audience. The signal processing unit 104 can also acquire the position information of these individual sounds. The position information of the ambient sound can be obtained, for example, based on the sound acquired by the microphones 13D to 13F. When the signal processing unit 104 recognizes the above-mentioned individual sounds through processing such as speech recognition, it calculates the correlation of the sound signals of the microphones 13D to 13F and calculates the timing difference of the individual sounds picked up by the microphones 13D to 13F respectively. The signal processing unit 104 can uniquely calculate the position within the first venue 10 where the sound occurred based on the timing difference of the sound picked up by the microphones 13D to 13F. In addition, the position information of the ambient sound can also be regarded as the position information of each of the microphones 13D to 13F.

[0055] The transmission device 12 encodes and transmits the sound source information related to the sound occurring in the first venue 10 and the echo information of the space as transmission data. The sound source information includes at least the performer's voice and may also include the position information of the performer's voice. In addition, the transmission device 12 can include the environmental information related to the ambient sound in the transmission data and transmit it. The transmission device 12 can also include the video signal related to the performer's image in the transmission data and transmit it.

[0056] Alternatively, the transmission device 12 can transmit at least the sound source information related to the performer's voice and the performer's position information and the environmental information related to the ambient sound as transmission data.

[0057] Figure 5 It is a block diagram showing the structure of the transmission device 12. Figure 6It is a flowchart showing the operation of the transmission device 12.

[0058] The transmission device 12 is composed of an information processing device such as a general personal computer. The transmission device 12 has a display 201, a user I / F 202, a CPU 203, a RAM 204, a network I / F 205, a flash memory 206, and a general communication I / F 207.

[0059] The CPU 203 reads the program stored in the flash memory 206 as a storage medium into the RAM 204, thereby realizing a specified function. In addition, the program read by the CPU 203 does not need to be stored in the flash memory 206 within this device. For example, the program can also be stored in a storage medium of an external device such as a server. In this case, the CPU 203 can execute by reading the program from the server into the RAM 204 each time.

[0060] The CPU 203 acquires the voice of the performer and the position information (sound source information) of the performer (S11) from the mixer 11 via the network I / F 205. In addition, the CPU 203 acquires the echo information of the space of the first venue 10 (S12). And the CPU 203 acquires the environmental information related to the ambient sound (S13). In addition, the CPU 203 can also acquire the video signal from the camera 16 via the general communication I / F 207.

[0061] The CPU 203 encodes and transmits the data related to the voice of the performer and the position information of the voice (sound source information), the data related to the echo information of the space, the data related to the environmental information, and the data related to the video signal as transmission data (S14).

[0062] The playback device 22 receives the transmission data from the transmission device 12 via the Internet 5. The playback device 22 performs rendering on the transmission data and provides the voice related to the voice of the performer and the echo in the space to the second venue 20. Or, the playback device 22 provides the voice of the performer and the ambient sound included in the environmental information to the second venue 20. The playback device 22 can also provide the voice related to the echo of the space corresponding to the environmental information to the second venue 20.

[0063] Figure 7 It is a block diagram showing the structure of the playback device 22. Figure 8 It is a flowchart showing the operation of the playback device 22.

[0064] The playback device 22 is composed of an information processing device such as a general personal computer. The playback device 22 has a display 301, a user I / F 302, a CPU 303, a RAM 304, a network I / F 305, a flash memory 306, and a video I / F 307.

[0065] The CPU 303 reads the program stored in the flash memory 306 serving as a storage medium into the RAM 304, thereby implementing a prescribed function. In addition, the program read by the CPU 303 does not necessarily need to be stored in the flash memory 306 within this device. For example, the program can also be stored in a storage medium of an external device such as a server. In this case, the CPU 303 can execute by reading the program from the server into the RAM 304 each time.

[0066] The CPU 303 receives transmission data from the transmission device 12 via the network I / F 305 (S21). The CPU 303 decodes the transmission data into sound source information, spatial echo information, environmental information, and video signals, etc. (S22), and reproduces the sound source information, spatial echo information, environmental information, and video signals, etc.

[0067] As an example of the reproduction of sound source information, the CPU 303 causes the mixer 21 to perform panning processing on the sound of the performer (S23). The panning processing is a process of positioning the sound of the performer at the position of the performer as described above. The CPU 303 determines the volume of the sound signal allocated to the speakers 24A to 24F in such a way that the sound of the performer is positioned at the position indicated by the position information included in the sound source information. The CPU 303 outputs the sound signal related to the sound of the performer and the information indicating the output amount of the sound signal related to the sound of the performer to the speakers 24A to 24F to the mixer 21, thereby causing the mixer 21 to perform panning processing.

[0068] Thereby, the audience in the second venue 20 can perceive as if the sound is emitted from the position of the performer. For example, the audience in the second venue 20 can also hear the sound of the performer on the right side of the stage in the first venue 10 from the front right in the second venue 20. In addition, the CPU 303 can also reproduce the video signal and display the live video on the display 23 via the video I / F 307. Thereby, the audience in the second venue 20 watches the image of the performer displayed on the display 23 while listening to the sound of the performer after panning processing. Thereby, since the visual information and the auditory information of the audience in the second venue 20 are consistent, they can obtain a sense of immersion in the live performance.

[0069] Further, as an example of reproducing the echo information of the space, the CPU 303 causes the mixer 21 to perform the process of generating the indirect sound (S24). The process of generating the indirect sound includes the process of generating the initial reflected sound and the process of generating the rear reverberant sound. The initial reflected sound is generated based on the sound of the performer included in the sound source information and the information indicating the size, shape, wall material, etc. of the space of the first venue 10 included in the echo information of the space. The CPU 303 determines the arrival timing of the initial reflected sound based on the size and shape of the space, and determines the level of the initial reflected sound based on the wall material. More specifically, the CPU 303 obtains the coordinates of the wall where the sound of the sound source is reflected based on the information on the size and shape of the space. Further, the CPU 303 obtains the position of a virtual sound source (virtual sound source) that exists with the wall as a mirror with respect to the position of the sound source based on the position of the sound source, the position of the wall, and the position of the sound collection point. The CPU 303 obtains the delay amount of the virtual sound source based on the distance from the position of the virtual sound source to the sound collection point. In addition, the CPU 303 obtains the level of the virtual sound source based on the information on the wall material. The information on the material corresponds to the energy loss during wall reflection. Therefore, the CPU 303 obtains the level of the virtual sound source considering this energy loss for the sound signal of the sound source. By repeatedly performing such processing, the CPU 303 can calculate the delay amount and level of the sound related to the echo of the space. The CPU 303 outputs the calculated delay amount and level to the mixer 21. The mixer 21 convolves these delay amounts and the level tap coefficients corresponding to the levels with the sound of the performer. Thereby, the mixer 21 reproduces the echo of the space of the first venue 10 in the second venue 20. In addition, when the echo information of the space includes the impulse response of the initial reflected sound, the CPU 303 causes the mixer 11 to perform the process of convolving the impulse response with the sound of the performer through the FIR filter. The CPU 303 outputs the echo information (impulse response) of the space included in the transmission data to the mixer 21. The mixer 21 convolves the echo information (impulse response) of the space received from the playback device 22 with the sound of the performer. Thereby, the mixer 21 reproduces the echo of the space of the first venue 10 in the second venue 20.

[0070] Also, when the echo information in the space changes corresponding to the position of the performer, the playback device 22 outputs the echo information of the space corresponding to the position of the performer to the mixer 21 based on the position information included in the sound source information. For example, when the performer at the front of the stage in the first venue 10 moves to the left side of the stage, the impulse response convolved with the performer's voice changes from the first impulse response to the second impulse response. Or, when reproducing the virtual sound source based on the information on the size and shape of the space, the delay amount and level are recalculated corresponding to the position of the moved performer. Thus, the appropriate echo of the space corresponding to the position of the performer is also reproduced in the second venue 20.

[0071] In addition, the playback device 22 may also cause the mixer 21 to generate the echo sound of the space corresponding to the ambient sound based on the ambient information and the echo information of the space. That is, the sound related to the echo of the space may also include the first echo sound corresponding to the sound of the performer (the sound of the first sound source) and the second echo sound corresponding to the ambient sound (the sound of the second sound source). Thus, the mixer 21 reproduces the echo of the ambient sound in the first venue 10 in the second venue 20. In addition, when the ambient information includes position information, the playback device 22 may also output the echo information of the space corresponding to the position of the ambient sound to the mixer 11 based on the position information included in the ambient information. The mixer 21 reproduces the echo sound of the ambient sound based on the position of the ambient sound. For example, when the audience at the left rear of the first venue 10 moves to the right rear, the impulse response convolved with the cheers of the audience is changed. Or, when reproducing the virtual sound source based on the information on the size and shape of the space, the delay amount and level are recalculated corresponding to the position of the moved audience. As described above, the echo information of the space includes the first echo information that changes corresponding to the position of the sound of the performer (the first sound source) and the second echo information that changes corresponding to the position of the ambient sound (the second sound source), and the reproduction may also include the process of generating the first echo sound based on the first echo information and the process of generating the second echo sound based on the second echo information.

[0072] In addition, the rear reverberation sound is a reflected sound with an unfixed arrival direction of the sound. Compared with the initial reflected sound, the rear reverberation sound changes less due to the change in the position of the sound. Therefore, the playback device 22 can only change the impulse response of the initial reflected sound that changes corresponding to the position of the performer and fix the impulse response of the rear reverberation sound.

[0073] In addition, the playback device 22 may omit the generation process of the indirect sound and directly utilize the echo of the second venue 20. Additionally, the generation process of the indirect sound may be only the initial reflected sound generation process. The rear reverberation sound may also directly utilize the echo of the second venue 20. Alternatively, the mixer 21 may also strengthen the control of the second venue 20 by further feeding back the sound acquired by a microphone (not shown) disposed near the ceiling and walls of the second venue 20 to the speakers 24A to 24F.

[0074] Moreover, the CPU 303 of the playback device 22 performs the playback process of the ambient sound based on the environmental information (S25). The environmental information includes sound signals of sounds such as background noise, the cheering, applause, shouts, cheers, chorus, or noise of the audience. The CPU 303 outputs these sound signals to the mixer 21. The mixer 21 outputs the sound signals received from the playback device 22 to the speakers 24A to 24F.

[0075] When the environmental information includes the position information of the ambient sound, the CPU 303 causes the mixer 21 to perform the localization process of the ambient sound by means of the sound image movement process. In this case, the CPU 303 determines the volume of the sound signals allocated to the speakers 24A to 24F in such a way that the ambient sound is localized at the position indicated by the position information included in the environmental information. The CPU 303 causes the mixer 21 to perform the sound image movement process by outputting the sound signal of the ambient sound and the information indicating the output volume of the sound signal related to the ambient sound to the speakers 24A to 24F. Additionally, the same applies when the position information of the ambient sound is the position information of the microphones 13D to 13F. The CPU 303 determines the volume of the sound signals allocated to the speakers 24A to 24F in such a way that the ambient sound is localized at the position of the microphone. The microphones 13D to 13F pick up multiple ambient sounds (second sound sources) such as background noise, applause, chorus, or cheers such as "wow", and noise. The sound of each sound source reaches the microphone with a prescribed delay amount and level. That is, background noise, applause, chorus, or cheers such as "wow", and noise also reach the microphone as independent sound sources with a prescribed delay amount and level (information for localizing the sound source). The CPU 303 can easily reproduce the localization of the independent sound sources by performing the sound image movement process in such a way that the sound picked up by the microphone is localized at the position of the microphone.

[0076] In addition, the CPU 303 can also process the voices simultaneously emitted by multiple listeners that cannot be recognized as the voices of individual listeners as follows: by causing the mixer 21 to perform effect processing such as reverberation, thereby perceiving the expansion of the space. For example, background noise, applause, chorus, or cheers such as "Wow", or noise is the sound that resounds throughout the entire live venue. The CPU 303 causes the mixer 21 to perform effect processing that perceives the expansion of the space for these sounds.

[0077] The playback device 22 can also provide the ambient sound based on the above environmental information to the second venue 20. As a result, the listeners in the second venue 20 can hear and see a live performance with a stronger sense of presence, just as if they were watching and listening to a live performance at the first venue 10.

[0078] In the above manner, the live data transmission system 1 of the present embodiment transmits the sound source information related to the sound occurring at the first venue 10 and the spatial echo information as transmission data, reproduces the transmission data, and provides the sound related to the sound source information and the sound related to the spatial echo to the second venue 20. As a result, the sense of presence of the live venue can also be provided to the destination venue of the transmission.

[0079] In addition, the live data transmission system 1 transmits, as transmission data, the first sound source information related to the sound of the first sound source (e.g., the voice of a performer) occurring at a certain first location (e.g., the stage) in the first venue 10 and the position information of the first sound source, and the second sound source information related to the second sound source (e.g., ambient sound) occurring at the second location (e.g., the location where the listeners are) in the first venue 10. The transmission data is reproduced, and the sound of the first sound source that has undergone positioning processing based on the position information of the first sound source and the sound of the second sound source are provided to the second venue. As a result, the sense of presence of the live venue can also be provided to the destination venue of the transmission.

[0080] Next, Figure 9 is a block diagram showing the structure of the live data transmission system 1A according to Modification 1. Figure 10 is a schematic plan view of the second venue 20 of the live data transmission system 1A according to Modification 1. The same reference numerals are used for the structures common to Figure 1 and Figure 3 and the description thereof is omitted.

[0081] In the second venue 20 of the live data transmission system 1A, a plurality of microphones 25A to 25C are provided. The microphone 25A is provided on the left side of the front and center, facing the stage 80 of the second venue 20. The microphone 25B is provided at the rear center of the second venue 20. The microphone 25C is provided on the right side of the front and center of the second venue 20.

[0082] Microphones 25A to 25C acquire the ambient sound of the second venue 20. The mixer 21 outputs the sound signal of the ambient sound as ambient information to the playback device 22. In addition, the ambient information may also include the position information of the ambient sound. The position information of the ambient sound can be obtained, for example, from the sound acquired by microphones 25A to 25C as described above.

[0083] The playback device 22 sends the ambient information related to the ambient sound occurring in the second venue 20 to other venues as the third sound source. For example, the playback device 22 feeds back the ambient sound occurring in the second venue 20 to the first venue 10. Thereby, the performers on the stage of the first venue 10 can hear sounds, applause, cheers, etc. other than the audience in the first venue 10, and can perform a live performance in an environment full of a sense of presence. In addition, the audience in the first venue 10 can also hear the sounds, applause, cheers, etc. of the audience in other venues, and can watch and listen to the live performance in an environment full of a sense of presence.

[0084] In addition, further, if the playback devices in other venues reproduce the transmitted data to provide the sound of the first venue to those other venues, and provide the ambient sound occurring in the second venue 20 to those other venues, then the audience in those other venues can also hear the sounds, applause, cheers, etc. of multiple audiences, and can watch and listen to the live performance in an environment full of a sense of presence.

[0085] Next, Figure 11 is a block diagram showing the structure of the live data transmission system 1B related to Modification 2. The same reference numerals are given to the structures Figure 1 common to those in [the previous part], and the description thereof is omitted.

[0086] In the live data transmission system 1B, the transmission device 12 is connected to the AV receiver 32 of the third venue 20A via the Internet 5. The AV receiver 32 is connected to a display 33, multiple speakers 34A to 34F, and a microphone 35. The third venue 20A is, for example, the residence of an individual audience member. The AV receiver 32 is an example of a playback device. The user of the AV receiver 32 is an audience member who watches and listens to the live performance of the first venue 10 remotely.

[0087] Figure 12 is a block diagram showing the structure of the AV receiver 32. The AV receiver 32 has a display 401, a user I / F 402, an audio I / O (Input / Output) 403, a signal processing unit (DSP) 404, a network I / F 405, a CPU 406, a flash memory 407, a RAM 408, and a video I / F 409.

[0088] The CPU 406 is a control unit that controls the operations of the AV receiver 32. The CPU 406 reads a prescribed program stored in the flash memory 407 serving as a storage medium into the RAM 408 and executes it, thereby performing various operations.

[0089] In addition, the program read by the CPU 406 does not necessarily need to be stored in the flash memory 407 within this device. For example, the program can also be stored in a storage medium of an external device such as a server. In this case, the CPU 406 can simply read the program from the server into the RAM 408 each time and execute it.

[0090] The signal processing unit 404 is composed of a DSP for performing various signal processes. The signal processing unit 404 performs signal processing on the sound signals input via the audio I / O 403 or the network I / F 405. The signal processing unit 404 outputs the audio signals after signal processing to a sound device such as a speaker via the audio I / O 403 or the network I / F 405.

[0091] The AV receiver 32 performs the same processes as those performed by the mixer 21 and the playback device 22. The CPU 406 receives transmission data from the transmission device 12 via the network I / F 405. The CPU 406 reproduces the transmission data and provides the sound related to the performer's voice and the spatial echo to the third venue 20A. Or, the CPU 406 reproduces the transmission data and provides the ambient sound occurring at the first venue 10 to the third venue 20A. Or, the CPU 406 can also reproduce the transmission data and display the live video on the display 33 via the video I / F 307.

[0092] The signal processing unit 404 performs panning processing on the performer's voice. Additionally, the signal processing unit 404 performs generation processing of indirect sound. Or, the signal processing unit 404 can also perform panning processing on the ambient sound.

[0093] Thereby, the AV receiver 32 can also provide the sense of presence of the first venue 10 to the third venue 20A.

[0094] In addition, the AV receiver 32 acquires the ambient sound (sounds such as the audience's cheering, applause, or shouts) of the third venue 20A via the microphone 35. The AV receiver 32 sends the ambient sound of the third venue 20A to other devices. For example, the AV receiver 32 feeds back the ambient sound of the third venue 20A to the first venue 10.

[0095] Thus, if the voices of multiple audiences are fed back to the first venue 10, the performers on the stage of the first venue 10 can hear the cheering, applause, and cheers of multiple audiences other than the audiences in the first venue 10, and can perform live in an environment full of a sense of presence. In addition, the audiences in the first venue 10 can also hear the cheering, applause, and cheers of multiple remote audiences, and can watch and listen to the live performance in an environment full of a sense of presence.

[0096] Alternatively, the AV receiver 32 can also display icon images such as "cheering", "applause", "shouting", and "noise" on the display 401, and receive a selection operation for these icon images from the audience via the user I / F 402, thereby receiving the reaction of the audience. It can also be that if the AV receiver 32 receives a selection operation of these reactions, it generates a sound signal corresponding to each reaction and sends it to other devices as environmental information.

[0097] Alternatively, the AV receiver 32 can also send information indicating the type of environmental sound such as the cheering, applause, or shouting of the audience as environmental information. In this case, the receiving-side device (e.g., the transmission device 12 and the mixer 11) generates a corresponding sound signal based on the environmental information and provides the sounds such as the cheering, applause, or shouting of the audience into the venue. In this way, the environmental information is not the sound signal of the environmental sound, but the information indicating the sound to be generated, and the process of playing the pre-recorded environmental sound, etc. by the transmission device 12 and the mixer 11 can be performed.

[0098] In addition, the environmental information of the first venue 10 can also be pre-recorded environmental sound instead of the environmental sound occurring in the first venue 10. In this case, the transmission device 12 transmits information indicating the sound to be generated as environmental information. The playback device 22 or the AV receiver 32 plays the corresponding environmental sound based on the environmental information. In addition, it can also be that the background noise and noise, etc. in the environmental information are pre-recorded sounds, and other environmental sounds (e.g., the cheering, applause, shouting of the audience, etc.) are the sounds occurring in the first venue 10.

[0099] In addition, the AV receiver 32 can also receive the position information of the audience via the user I / F 402. The AV receiver 32 displays an image imitating the floor plan or perspective view, etc. of the first venue 10 on the display 401 or the display 33, and receives the position information from the audience via the user I / F 402 (e.g., refer to Figure 16)). The location information is information that designates an arbitrary position within the first venue 10. The AV receiver 32 transmits the received location information of the audience to the first venue 10. The transmission device 12 and the mixer 11 in the first venue perform processing to localize the ambient sound of the third venue 20A at the designated position based on the ambient sound of the third venue 20A and the location information of the audience received from the AV receiver 32.

[0100] In addition, the AV receiver 32 can also change the content of the sound image movement processing based on the location information received from the user. For example, if the audience designates a position immediately in front of the stage in the first venue 10, the AV receiver 32 sets the localization position of the performer's voice at a position immediately in front of the audience and performs sound image movement processing. Thus, the audience in the third venue 20A can obtain a sense of presence as if they were located immediately in front of the stage in the first venue 10.

[0101] The voices of the audience in the third venue 20A can be sent not to the first venue 10 but to the second venue 20, or can also be sent to other venues. For example, the voices of the audience in the third venue 20A can be sent only to a friend's house (the fourth venue). The audience in the fourth venue can watch and listen to the live performance in the first venue 10 while listening to the voices of the audience in the third venue 20A. In addition, a playback device (not shown) in the fourth venue can also send the voices of the audience in the fourth venue to the third venue 20A. In this case, the audience in the third venue 20A can watch and listen to the live performance in the first venue 10 while listening to the voices of the audience in the fourth venue. Thus, the audience in the third venue 20A and the audience in the fourth venue can watch and listen to the live performance in the first venue 10 while talking to each other.

[0102] Figure 13 is a block diagram showing the configuration of the live data transmission system 1C according to Modification Example 3. The same reference numerals are assigned to the structures common to Figure 1 and the description thereof is omitted.

[0103] In the live data transmission system 1C, the transmission device 12 is connected to the terminal 42 in the fifth venue 20B via the Internet 5. The terminal 42 is connected to the headphones 43. The fifth venue 20B is, for example, the residence of an individual audience member. However, when the terminal 42 is mobile, the fifth venue 20B can also be various places such as inside a coffee shop, inside a car, or inside public transportation equipment. In this case, all places can become the fifth venue 20B. The terminal 42 is an example of a playback device. The user of the terminal 42 becomes an audience member who watches and listens to the live performance in the first venue 10 remotely. In this case, the terminal 42 also reproduces the transmitted data and provides the sound related to the sound source information and the sound related to the echo in the space to the second venue (in this example, the fifth venue 20B) via the headphones 43.

[0104] Figure 14 It is a block diagram showing the structure of the terminal 42. The terminal 42 is an information processing device such as a personal computer, a smartphone, or a tablet computer. The terminal 42 has a display 501, a user I / F 502, a CPU 503, a RAM 504, a network I / F 505, a flash memory 506, an audio I / O (Input / Output) 507, and a microphone 508.

[0105] The CPU 503 is a control unit that controls the operation of the terminal 42. The CPU 503 reads a prescribed program stored in the flash memory 506 serving as a storage medium into the RAM 504 and executes it, thereby performing various operations.

[0106] In addition, the program read by the CPU 503 does not need to be stored in the flash memory 506 within this device. For example, the program can also be stored in a storage medium of an external device such as a server. In this case, the CPU 503 can simply read the program into the RAM 504 from the server each time and execute it.

[0107] The CPU 503 performs signal processing on the sound signal input via the network I / F 505. The CPU 503 outputs the audio signal after signal processing to the earphone 43 via the audio I / O 507.

[0108] The CPU 503 receives transmission data from the transmission device 12 via the network I / F 505. The CPU 503 reproduces the transmission data and provides the sound related to the performer's voice and the echo in the space to the audience in the fifth venue 20B.

[0109] Specifically, the CPU 503 convolves the sound signal related to the performer's voice with a head-related transfer function (hereinafter referred to as HRTF.) so as to perform sound image localization processing (binauralization processing) in such a way that the voice of the performer is localized at the position of the performer. The HRTF corresponds to the transfer function between a prescribed position and the listener's ears. The HRTF is a transfer function that shows the magnitude, arrival time, and frequency characteristics, etc. of the sound reaching the left and right ears respectively from a sound source at a certain position. The CPU 503 convolves the sound signal of the performer's voice with the HRTF based on the position of the performer. Thereby, the voice of the performer is localized at the position corresponding to the position information.

[0110] In addition, the CPU 503 performs binaural processing by convolving the sound signal of the performer's voice with the HRTF corresponding to the echo information of the space to generate indirect sound. The CPU 503 convolves the HRTFs that reach the left and right ears respectively from the positions of the virtual sound sources corresponding to the respective initial reflected sounds included in the echo information of the space, thereby localizing the initial reflected sounds and the late reverberation sounds. However, the late reverberation sound is a reflected sound with an unfixed arrival direction of the sound. Therefore, the CPU 503 may also perform effect processing such as reverberation instead of localizing the late reverberation sound. In addition, the CPU 503 may perform digital filtering processing (headphone inverse characteristic processing) to reproduce the inverse characteristics of the acoustic characteristics of the headphones 43 used by the listener.

[0111] In addition, the CPU 503 reproduces the environmental information in the transmitted data and provides the environmental sound occurring at the first venue 10 to the listeners at the fifth venue 20B. When the environmental information includes the position information of the environmental sound, the CPU 503 performs localization processing based on the HRTF and performs effect processing on the sound with an unfixed arrival direction of the sound.

[0112] In addition, the CPU 503 may also reproduce the video signal in the transmitted data and display the live video on the display 501.

[0113] Thus, the terminal 42 can also provide the sense of presence of the first venue 10 to the listeners at the fifth venue 20B.

[0114] In addition, the terminal 42 acquires the voices of the listeners at the fifth venue 20B via the microphone 508. The terminal 42 sends the voices of the listeners to other devices. For example, the terminal 42 feeds back the voices of the listeners to the first venue 10. Alternatively, the terminal 42 may display icon images such as "cheering", "applauding", "shouting", and "noisy" on the display 501, receive a selection operation for these icon images from the listeners via the user I / F 502, and receive a reaction. The terminal 42 generates a sound corresponding to the received reaction and sends the generated sound as environmental information to other devices. Alternatively, the terminal 42 may also send information indicating the type of environmental sound such as the cheering, applauding, or shouting of the listeners as environmental information. In this case, the receiving device (for example, the transmission device 12 and the mixer 11) generates a corresponding sound signal based on the environmental information and provides the sounds such as the cheering, applauding, or shouting of the listeners to the venue.

[0115] In addition, the terminal 42 can also receive the position information of the audience via the user I / F 502. The terminal 42 sends the received position information of the audience to the first venue 10. The transmission device 12 and the mixer 11 in the first venue perform processing to localize the voice of the audience at the specified position based on the voice and position information of the audience in the third venue 20A received from the AV receiver 32.

[0116] In addition, the terminal 42 can also change the HRTF based on the position information received from the user. For example, if the audience designates a position immediately in front of the stage in the first venue 10, the terminal 42 sets the localization position of the voice of the performer at the position immediately in front of the audience, and convolves the HRTF that localizes the voice of the performer at that position. Thus, the audience in the fifth venue 20B can obtain a sense of presence as if they were located immediately in front of the stage in the first venue 10.

[0117] The voice of the audience in the fifth venue 20B can be sent not to the first venue 10 but to the second venue 20, or can also be sent to other venues. Similarly to the above, the voice of the audience in the fifth venue 20B can be sent only to the friend's house (the fourth venue). Thus, the audience in the fifth venue 20B and the audience in the fourth venue can watch and listen to the live performance in the first venue 10 while talking to each other.

[0118] In addition, in the live data transmission system of the present embodiment, multiple users can also specify the same position. For example, multiple users can each specify a position immediately in front of the stage in the first venue 10. In this case, each audience can obtain a sense of presence as if they were in the position immediately in front of the stage. Thus, for one position (a seat in the venue), multiple audiences can watch and listen to the performance of the performer with the same sense of presence. In this case, the live operator can provide a service with a number of audiences exceeding the number that can be accommodated in the actual space.

[0119] Figure 15 It is a block diagram showing the structure of the live data transmission system 1D related to Modification Example 4. The same reference numerals are assigned to the common structures, and the description is omitted. Figure 1 Common structures are labeled with the same reference numerals, and the description is omitted.

[0120] The live data transmission system 1D also includes a server 50 and terminals 55. The terminals 55 are provided in the sixth venue 10A. The server 50 is an example of a transmission device, and the hardware structure of the server 50 is the same as that of the transmission device 12. The hardware structure of the terminal 55 is the same as that of Figure 14 the terminal 42 shown.

[0121] The 6th venue 10A is the residence etc. of the performers who perform music etc. remotely. The performers at the 6th venue 10A perform music or singing etc. in synchronization with the music or singing at the 1st venue. The terminal 55 sends the voices of the performers at the 6th venue 10A to the server 50. In addition, the terminal 55 may also photograph the performers at the 6th venue 10A with a camera (not shown) and send the video signal to the server 50.

[0122] The server 50 transmits the transmission data including the voices of the performers at the 1st venue 10, the voices of the performers at the 6th venue 10A, the echo information of the space at the 1st venue 10, the environmental information of the 1st venue 10, the live video of the 1st venue 10, and the video of the performers at the 6th venue 10A.

[0123] In this case, the playback device 22 reproduces the transmission data and provides the voices of the performers at the 1st venue 10, the voices of the performers at the 6th venue 10A, the echo of the space at the 1st venue 10, the environmental sound of the 1st venue 10, the live video of the 1st venue 10, and the video of the performers at the 6th venue 10A to the 2nd venue 20. For example, the playback device 22 superimposes and displays the video of the performers at the 6th venue 10A on the live video of the 1st venue 10.

[0124] The voices of the performers at the 6th venue 10A may not be subject to localization processing, or may be localized at a position matching the video displayed on the display. For example, when the performers at the 6th venue 10A are displayed on the right side within the live video, the voices of the performers at the 6th venue 10A are localized on the right side.

[0125] In addition, the performers at the 6th venue 10A, or the transmitter of the transmission data may also specify the positions of the performers. In this case, the transmission data includes the position information of the performers at the 6th venue 10A. The playback device 22 localizes the voices of the performers at the 6th venue 10A based on the position information of the performers at the 6th venue 10A.

[0126] The video of the performers at the 6th venue 10A is not limited to the video captured by a camera. For example, a 2D image or a virtual character image (virtual video) composed of 3D modeling may be transmitted as the video of the performers at the 6th venue 10A.

[0127] In addition, the transmitted data may also include audio recording data. Additionally, the transmitted data may also include video data. For example, the transmitting device may transmit transmitted data that includes the voices of the performers at the first venue 10, audio recording data, echo information of the space at the first venue 10, environmental information of the first venue 10, live images of the first venue 10, and video data. In this case, the playback device reproduces the transmitted data and provides the voices of the performers at the first venue 10, the voices related to the audio recording data, the echo of the space at the first venue 10, the environmental sounds of the first venue 10, the live images of the first venue 10, and the images related to the video data to other venues. The playback device 22 superimposes and displays the images of the performers corresponding to the video data on the live images of the first venue 10.

[0128] In addition, when recording the sound related to the audio recording data, the transmitting device may also determine the category of the musical instrument. In this case, the transmitting device includes information indicating the category of the musical instrument determined to be the audio recording data in the transmitted data and transmits it. The playback device generates an image of the corresponding musical instrument based on the information indicating the category of the musical instrument. The playback device may also superimpose and display the image of the musical instrument on the live images of the first venue 10.

[0129] In addition, there is no need to superimpose the images of the performers at the sixth venue 10A on the live images of the first venue 10 in the transmitted data. For example, the transmitted data may also transmit the images of the respective performers at the first venue 10 and the sixth venue 10A and the background images as separate data. In this case, the transmitted data includes information indicating the display positions of the respective images. The playback device reproduces the images of the respective performers based on the information indicating the display positions.

[0130] In addition, the background image is not limited to the images of the venues where the live performances are actually held, such as the first venue 10. The background image may also be the image of a venue different from the venue where the live performance is held.

[0131] Moreover, the echo information of the space included in the transmitted data does not need to correspond to the echo of the space at the first venue 10. For example, the echo information of the space may also be virtual space information (information indicating the size, shape, wall material, etc. of the space of each venue, or the impulse response indicating the transfer function of each venue) for virtually reproducing the echo of the space corresponding to the background image. The impulse response of each venue can be measured in advance or obtained by simulation based on the size, shape, and wall material, etc. of the space of each venue.

[0132] Moreover, the environmental information can also be changed to content corresponding to the background image. For example, in the case of the background image of a large convention hall, the environmental information includes sounds such as the cheers, applause, and shouts of the audience. In addition, an outdoor venue includes background noise different from that of an indoor venue. Also, the echo of the ambient sound changes accordingly to the echo information of the above space. In addition, the environmental information can also include information indicating the number of the audience and information indicating the degree of crowding (density of people). The playback device increases or decreases the number of sounds such as the cheers, applause, and shouts of the audience based on the information indicating the number of the audience. In addition, the playback device increases or decreases the volume of the cheers, applause, and shouts of the audience based on the information indicating the degree of crowding.

[0133] Alternatively, the environmental information can also be changed according to the performer. For example, in the case of a live performance by a performer with many female audiences or listeners, the sounds such as the cheers, shouts, and applause of the audience included in the environmental information are changed to female voices. The environmental information can include the sound signals of these audience voices, or can include information indicating the attributes of the audience such as the male-female ratio or age ratio. The playback device changes the sound quality of the cheers, applause, and shouts of the audience based on the information indicating the attribute.

[0134] In addition, the audience in each venue can also specify the background image and the echo information of the space. The audience in each venue uses the user I / F of the playback device to specify the background image and the echo information of the space.

[0135] Figure 16 FIG. is an example showing the live image 700 displayed by the playback device in each venue. The live image 700 is composed of an image obtained by photographing the first venue 10 or other venues, or a virtual image (computer graphics) corresponding to each venue, etc. The live image 700 is displayed on the display of the playback device. The live image 700 shows an image of the background, stage, performer including musical instruments, and audience in the venue. The images of the background, stage, performer including musical instruments, and audience in the venue can be all actually photographed images, or can be virtual images. In addition, it can also be that only the background image is an actually photographed image and the other images are virtual images. In addition, icon images 751 and 752 for specifying the space are displayed in the live image 700. The icon image 751 is an image for specifying the space of a certain venue, i.e., Stage A (for example, the first venue 10), and the icon image 752 is an image for specifying the space of another venue, i.e., Stage B (for example, other concert halls, etc.). And, an audience image 753 for specifying the position of the audience is displayed in the live image 700.

[0136] Listeners of the playback device specify a desired space by designating either the icon image 751 or the icon image 752 using the user I / F of the playback device. The transmission device includes the background image corresponding to the specified space and the echo information of the space in the transmission data and transmits it. Alternatively, the transmission device may also include multiple background images and the echo information of the space in the transmission data and transmit it. In this case, the playback device reproduces the background image corresponding to the space specified by the listener in the received transmission data and the echo information of the space.

[0137] In Figure 16 the example, the icon image 751 is designated. The playback device displays the background image corresponding to Stage A of the icon image 751 (for example, the image of the first venue 10) and plays the sound related to the echo of the space corresponding to the designated Stage A. If the listener designates the icon image 752, the playback device switches to display the background image of another space, namely Stage B, corresponding to the icon image 752 and plays the sound related to the echo of the corresponding other space based on the virtual space information corresponding to Stage B.

[0138] Thus, the listeners of each playback device can obtain a sense of presence as if they were watching and listening to a live performance in the desired space.

[0139] In addition, the listeners of each playback device can specify a desired position in the venue by moving the listener image 753 within the live image 700. The playback device performs positioning processing based on the position designated by the user. For example, if the listener moves the listener image 753 to a position right in front of the stage, the playback device sets the positioning position of the performer's voice to a position right in front of the listener and performs positioning processing to make the performer's voice located at this position. Thus, the listeners of each playback device can obtain a sense of presence as if they were right in front of the stage.

[0140] In addition, as described above, if the position of the sound source and the position of the listener (the position of the sound collection point) change, the sound related to the echo of the space also changes. The playback device can also calculate the initial reflected sound in the case where the space has changed, the position of the sound source has changed, or the position of the sound collection point has changed. Therefore, even without measuring impulse responses and the like in the actual space, the playback device can calculate the sound related to the echo of the space based on the virtual space information. Thus, the playback device can accurately reproduce the echo generated in a space including the actual space.

[0141] For example, the mixer 11 can function as a transmission device, and the mixer 21 can also function as a playback device. Additionally, the playback device does not need to be provided in each venue. For example, Figure 15 the server 50 shown can also reproduce the transmitted data and transmit the sound signal after signal processing to terminals in each venue, etc. In this case, the server 50 functions as a playback device.

[0142] The sound source information may also include information indicating the posture of the performer (e.g., the left - right orientation of the performer). The playback device may also perform adjustment processing of volume or frequency characteristics based on the posture information of the performer. For example, the playback device performs processing such that, taking the case where the performer is facing directly forward as a reference, the greater the left - right orientation, the lower the volume. Additionally, the playback device may also perform processing such that the greater the left - right orientation, the greater the attenuation of the high - frequency band compared to the low - frequency band. Thus, the sound changes according to the posture of the performer, so that the audience can watch and listen to a more immersive live performance.

[0143] Next, Figure 17 is a block diagram showing an application example of the signal processing performed by the playback device. In this example, reproduction is performed using Figure 13 the terminal 42 and the headphones 43 shown. The playback device (in the Figure 13 example of, the terminal 42) functionally has an instrument model processing unit 551, an amplifier model processing unit 552, a speaker model processing unit 553, a spatial model processing unit 554, a binauralization processing unit 555, and a headphone inverse characteristic processing unit 556.

[0144] The instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 perform signal processing for imparting the acoustic characteristics of the audio equipment to the sound signal related to the performance sound. The first digital signal processing model for performing this signal processing is included in the sound source information transmitted by the transmission device 12, for example. The first digital signal processing models are digital filters that respectively simulate the acoustic characteristics of the instrument, the acoustic characteristics of the amplifier, and the acoustic characteristics of the speaker. The first digital signal processing models are pre - fabricated by the manufacturer of the instrument, the manufacturer of the amplifier, and the manufacturer of the speaker through simulation, etc. The instrument model processing unit 551, the amplifier model processing unit 552, and the speaker model processing unit 553 respectively perform digital filtering processing that simulates the acoustic characteristics of the instrument, the acoustic characteristics of the amplifier, and the acoustic characteristics of the speaker. In addition, in the case where the instrument is an electronic instrument such as a synthesizer, the instrument model processing unit 551 inputs note event data (information indicating the pronunciation timing, pitch, etc. of the sound to be pronounced) instead of the sound signal and generates a sound signal having the acoustic characteristics of an electronic instrument such as a synthesizer.

[0145] Accordingly, the playback device can reproduce the acoustic characteristics of any musical instrument or the like. For example, in Figure 16 , a live image 700 of a virtual image (computer graphics) is displayed. Here, the listener using the playback device can also change to an image of another virtual musical instrument using the user I / F of the playback device. When the listener changes the musical instrument displayed in the live image 700 to an image of another musical instrument, the musical instrument model processing unit 551 of the playback device performs signal processing corresponding to the first digital signal processing model, which corresponds to the changed musical instrument. Accordingly, the playback device outputs a sound that reproduces the acoustic characteristics of the musical instrument displayed in the live image 700.

[0146] Similarly, the listener using the playback device can also change the type of amplifier and the type of speaker to different types using the user I / F of the playback device. The amplifier model processing unit 552 and the speaker model processing unit 553 perform digital filtering processing that simulates the acoustic characteristics of the amplifier and the speaker of the changed types. In addition, the speaker model processing unit 553 can also simulate the acoustic characteristics in each direction of the speaker. In this case, the listener using the playback device can also change the orientation of the speaker using the user I / F of the playback device. The speaker model processing unit 553 performs digital filtering processing corresponding to the changed orientation of the speaker.

[0147] The space model processing unit 554 is a second digital signal processing model that reproduces the acoustic characteristics of the room of the live venue (for example, the echo in the above space). The second digital signal processing model can also be obtained, for example, by using a test tone or the like in an actual live venue. Or, the second digital signal processing model can also calculate the delay amount and level of the virtual sound source according to the virtual space information (information indicating the size, shape, and wall material of the space of each venue) as described above.

[0148] If the position of the sound source and the position of the listener (the position of the sound collection point) change, the sound related to the echo in the space also changes. The playback device can also calculate the delay amount and level of the virtual sound source when the space changes, the position of the sound source changes, and the position of the sound collection point changes. Therefore, even without measuring the impulse response or the like in the actual space, the playback device can obtain the sound related to the echo in the space based on the virtual space information. Thus, the playback device can accurately reproduce the echo generated in the space including the actual space.

[0149] In addition, the virtual space information may also include information on the position and material of structures such as pillars (acoustic obstacles). When performing sound source localization and indirect sound generation processing, the playback device reproduces the phenomena of reflection, occlusion, and diffraction caused by such obstacles when there are obstacles in the paths of the direct sound and indirect sound arriving from the sound source.

[0150] Figure 18 It is a schematic diagram showing the path of the sound that is reflected from the sound source 70 on the wall surface and reaches the sound collection point 75. Figure 18 The shown sound source 70 can be either a performance sound (the first sound source) or an environmental sound (the second sound source). The playback device calculates the position of the virtual sound source 70A that exists with the wall surface as a mirror with respect to the position of the sound source 70 based on the position of the sound source 70, the position of the wall surface, and the position of the sound collection point 75. Moreover, the playback device calculates the delay amount of the virtual sound source 70A based on the distance from the virtual sound source 70A to the sound collection point 75. In addition, the playback device calculates the level of the virtual sound source 70A based on the information on the material of the wall surface. And, when there is an obstacle 77 in the path from the position of the virtual sound source 70A to the sound collection point 75 as shown in Figure 18 the playback device calculates the frequency characteristics generated by the diffraction through the obstacle 77. Diffraction, for example, attenuates the sound in the high-frequency band. Therefore, when there is an obstacle 77 in the path from the position of the virtual sound source 70A to the sound collection point 75 as shown in Figure 18 the playback device performs equalization processing to reduce the level in the high-frequency band. The frequency characteristics generated by diffraction may also be included in the virtual space information.

[0151] In addition, the playback device may also set new second virtual sound source 77A and third virtual sound source 77B at the positions on the left and right of the obstacle 77. The second virtual sound source 77A and the third virtual sound source 77B correspond to the new sound sources generated by diffraction. The second virtual sound source 77A and the third virtual sound source 77B are sounds obtained by imparting the frequency characteristics generated by diffraction to the sound of the virtual sound source 70A respectively. The playback device recalculates the delay amount and the level based on the positions of the second virtual sound source 77A and the third virtual sound source 77B and the position of the sound collection point 75. Thereby, the diffraction phenomenon of the obstacle 77 can be reproduced.

[0152] The playback device may also calculate the delay amount and the level of the sound that the sound of the virtual sound source 70A is reflected by the obstacle 77 and further reflected by the wall surface and reaches the sound collection point 75. In addition, the playback device may eliminate the virtual sound source 70A when it is determined that the virtual sound source 70A is occluded by the obstacle 77. The information for determining whether there is occlusion may also be included in the virtual space information.

[0153] By performing the above processing, the playback device performs the first digital signal processing that represents the acoustic characteristics of the audio device and the second digital signal processing that represents the acoustic characteristics of the room, and generates the sound related to the sound source and the echo in the space.

[0154] Moreover, the binaural processing unit 555 convolves the sound signal with the head-related transfer function (hereinafter referred to as HRTF) to perform sound image localization processing for the sound source and various indirect sounds. The headphone inverse characteristic processing unit 556 performs digital filtering processing to reproduce the inverse characteristic of the acoustic characteristics of the headphones used by the listener.

[0155] Through the above processing, the user can obtain a sense of presence as if watching and listening to a live performance in a desired space and with a desired audio device.

[0156] In addition, the playback device does not need to have Figure 17 all of the instrument model processing unit 551, amplifier model processing unit 552, speaker model processing unit 553, and space model processing unit 554 shown. The playback device only needs to use at least one digital signal processing model to perform signal processing. In addition, the playback device can perform signal processing on a certain sound signal (for example, the sound of a certain performer) using one digital signal processing model, or can perform signal processing on multiple sound signals using one digital signal processing model respectively. The playback device can perform signal processing on a certain sound signal (for example, the sound of a certain performer) using multiple digital signal processing models, or can perform signal processing on multiple sound signals using multiple digital signal processing models. The playback device can also perform signal processing on ambient sound using a digital signal processing model.

[0157] The description of this embodiment is illustrative in all aspects and is not restrictive. The scope of the present invention is represented not by the above embodiment but by the claims. And, the scope of the present invention includes the equivalent meaning of the claims and all changes within the scope.

[0158] Description of reference numerals

[0159] 1, 1A, 1B, 1C, 1D... Live data transmission system

[0160] 5... Internet

[0161] 10... First venue

[0162] 10A... Sixth venue

[0163] 11... Mixer

[0164] 12... Transmission device

[0165] 13A~13F... Microphones

[0166] 14A - 14G... Loudspeaker

[0167] 15A - 15C... Tracker

[0168] 16... Camera

[0169] 20... Second Venue

[0170] 20A... Third Venue

[0171] 20B... Fifth Venue

[0172] 21... Mixer

[0173] 22... Playback Device

[0174] 23... Display

[0175] 24A - 24F... Loudspeaker

[0176] 25A - 25C... Microphone

[0177] 32... AV Receiver

[0178] 33... Display

[0179] 34A... Loudspeaker

[0180] 35... Microphone

[0181] 42... Terminal

[0182] 43... Headphone

[0183] 50... Server

[0184] 55... Terminal

[0185] 101... Display

[0186] 102... User I / F

[0187] 103... Audio I / O

[0188] 104... Signal Processing Unit

[0189] 105... Network I / F

[0190] 106... CPU

[0191] 107... Flash Memory

[0192] 108... RAM

[0193] 201... Display

[0194] 202... User I / F

[0195] 203... CPU

[0196] 204…RAM

[0197] 205…Network I / F

[0198] 206…Flash memory

[0199] 207…General communication I / F

[0200] 301…Display

[0201] 302…User I / F

[0202] 303…CPU

[0203] 304…RAM

[0204] 305…Network I / F

[0205] 306…Flash memory

[0206] 307…Video I / F

[0207] 401…Display

[0208] 402…User I / F

[0209] 403…Audio I / O

[0210] 404…Signal processing unit

[0211] 405…Network I / F

[0212] 406…CPU

[0213] 407…Flash memory

[0214] 408…RAM

[0215] 409…Video I / F

[0216] 501…Display

[0217] 503…CPU

[0218] 504…RAM

[0219] 505…Network I / F

[0220] 506…Flash memory

[0221] 507…Audio I / O

[0222] 508…Microphone

[0223] 700…Live video

Claims

1. A method for transmitting on-site data, wherein, transmit, as transmission data, first sound source information related to the sound of a first sound source occurring at a first location in a first venue and the location information of the first sound source, and second sound source information related to a second sound source including ambient sound occurring at a second location in the first venue, reproduce the transmission data, and provide the sound of the first sound source subjected to positioning processing based on the location information of the first sound source and the sound of the second sound source to a second venue, the transmission data includes an image signal captured in the first venue, display an image based on the image signal.

2. The on-site data transmission method according to claim 1, wherein, the transmission data includes an image signal capturing a performer, display an image based on the image signal.

3. The on-site data transmission method according to claim 2, wherein, the image of the performer includes a virtual image.

4. The on-site data transmission method according to claim 1, wherein, the transmission data includes an image signal captured in the first venue and an image signal capturing a performer, display them superimposed.

5. The on-site data transmission method according to claim 4, wherein, the transmission data includes information indicating the display positions of the respective images, display each image based on the information indicating the display position.

6. The on-site data transmission method according to claim 2, wherein, locate the sound of the performer at the position of the displayed performer.

7. The on-site data transmission method according to claim 1, wherein, the transmission data includes recording data.

8. The on-site data transmission method according to claim 1, wherein, the transmission data includes information indicating the category of the musical instrument, display an image of the musical instrument based on the information indicating the category of the musical instrument.

9. The on-site data transmission method according to claim 8, wherein, receive an instruction to change the image of the musical instrument from the audience.

10. The on-site data transmission method according to claim 1, wherein, the transmission data includes an image signal captured in the first venue and video data, display the image based on the video data superimposed on the image based on the image signal.

11. The on-site data transmission method according to claim 1, wherein, the transmission data includes a background image and virtual space information for virtually reproducing the echo of the space of the venue corresponding to the background image, display an image based on the background image, generate a sound related to the echo of the space based on the virtual space information.

12. The on-site data transmission method according to claim 11, wherein, generate a sound related to the environmental information corresponding to the background image.

13. The on-site data transmission method according to claim 12, wherein, the environmental information changes according to the performer.

14. The on-site data transmission method according to claim 11, wherein, receive a designation of the background image and the virtual space information from the audience.

15. The on-site data transmission method according to claim 14, wherein, display an icon image for specifying the virtual space information, receive the specification of the virtual space information through the icon image.

16. The on-site data transmission method according to claim 1, wherein, send the environmental information related to the ambient sound of the second venue to outside the second venue.

17. The on-site data transmission method according to claim 16, wherein, feedback the environmental information to the first venue, provide the sound related to the environmental information to the users in the first venue.

18. The on-site data transmission method according to claim 17, wherein, the environmental information includes information corresponding to the reactions of the users, provide the sound corresponding to the reaction to the users in the first venue.

19. The on-site data transmission method according to claim 16, wherein, the environmental information includes the sound picked up by a microphone arranged in the second venue.

20. The on-site data transmission method according to claim 16, wherein, the environmental information includes pre-produced sound.

21. The on-site data transmission method according to claim 20, wherein, the pre-produced sound is different for each venue.

22. The on-site data transmission method according to claim 16, wherein, the environmental information includes information related to the attributes of the users corresponding to the second sound source, the reproduction includes a process of providing sound based on the attributes.

23. The on-site data transmission method according to claim 1, wherein, the second sound source information includes the position information of the second sound source, the reproduction includes a process of providing the sound of the second sound source for which localization processing based on the position information of the second sound source has been performed.

24. The on-site data transmission method according to claim 1, wherein, receive the gesture information of the performer corresponding to the sound of the first sound source, perform adjustment processing on the volume or frequency characteristics of the sound of the first sound source based on the gesture information.

25. The on-site data transmission method according to claim 1, wherein, the sound related to the echo of the space includes a first echo sound corresponding to the sound of the first sound source and a second echo sound corresponding to the sound of the second sound source.

26. The on-site data transmission method according to claim 25, wherein, the echo information of the space includes first echo information that changes corresponding to the position of the first sound source and second echo information that changes corresponding to the position of the second sound source, the reproduction includes a process of generating the first echo sound based on the first echo information and a process of generating the second echo sound based on the second echo information.

27. The on-site data transmission method according to claim 1, wherein, the second sound source includes a plurality of sound sources.

28. An on-site data transmission system, comprising: A live data transmission device that transmits, as transmission data, first sound source information related to the sound of a first sound source that occurs at a first location in a first venue and the position information of the first sound source, and second sound source information related to a second sound source that includes ambient sound occurring at a second location in the first venue; and A live data playback device that reproduces the transmission data and provides the sound of the first sound source that has undergone positioning processing based on the position information of the first sound source and the sound of the second sound source to a second venue, The transmission data includes an image signal captured at the first venue, The live data playback device displays an image based on the image signal.

29. A live data transmission device, wherein, It transmits, as transmission data, first sound source information related to the sound of a first sound source that occurs at a first location in a first venue and the position information of the first sound source, and second sound source information related to a second sound source that includes ambient sound occurring at a second location in the first venue, It causes a live data playback device to reproduce the transmission data and provides the sound of the first sound source that has undergone positioning processing based on the position information of the first sound source and the sound of the second sound source to a second venue, The transmission data includes an image signal captured at the first venue, It causes the live data playback device to display an image based on the image signal.

30. A live data playback device, wherein, It receives transmission data from a live data transmission device that transmits, as the transmission data, first sound source information related to the sound of a first sound source that occurs at a first location in a first venue and the position information of the first sound source, and second sound source information related to a second sound source that includes ambient sound occurring at a second location in the first venue, It reproduces the transmission data and provides the sound of the first sound source that has undergone positioning processing based on the position information of the first sound source and the sound of the second sound source to a second venue, The transmission data includes an image signal captured at the first venue and displays an image based on the image signal.

Citation Information

Patent Citations

  • Match watching device, game watching terminal, game watching method, and program therefor

    JP2019024157A