Information processing device, information processing method, and program

The system generates and outputs individual pseudo-sound data based on participant reactions, addressing transmission issues and enhancing realism in live streaming feedback.

JP7800450B2Active Publication Date: 2026-01-16SONY GROUP CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022578099
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-27
Filing Date
2021-12-07
Publication Date
2026-01-16
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Existing live streaming systems lack the ability to provide performers with realistic and immersive feedback from audience reactions, as high-quality sound delivery is hindered by transmission issues and pre-prepared sounds lack individuality.

Method used

An information processing device and method that generates and outputs individual pseudo-sound data reflecting the characteristics of participant reactions, using a control unit to select and output audio data from an audio output device at the venue, and a process to generate and store pseudo-sound data associated with participant IDs.

Benefits of technology

Delivers realistic and immersive feedback to performers by simulating audience reactions in real-time, overcoming transmission challenges and providing personalized audio feedback without high bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007800450000001
    Figure 0007800450000001
  • Figure 0007800450000002
    Figure 0007800450000002
  • Figure 0007800450000003
    Figure 0007800450000003
Patent Text Reader

Abstract

[Problem] To provide an information processing device, an information processing method, and a program capable of providing audio data that reflect the individuality of a participant and that correspond to a response of the participant, in consideration of transmission problems. [Solution] This information processing device is provided with a control unit which selects individual pseudo-sound data corresponding to reaction information indicting an acquired participant reaction from one or more sets of individual pseudo-sound data reflecting a characteristic of sound issued by the participant, and which controls output of the selected individual pseudo-sound data from an audio output device situated in a venue.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] With the recent advancement of communication technology, so-called live streaming is now available, whereby video footage of events such as concerts, seminars, and plays is distributed in real time. However, with such live streaming, there are challenges in terms of the sense of presence and two-way communication between performers and audiences that are found in traditional live events that attract large audiences.

[0003] Regarding the collection of audience reactions at live streaming events, for example, Patent Document 1 below discloses that quantitative information such as the number of taps made by the audience is acquired in real time as reaction data, and feedback is provided to the performers by displaying the acquired quantitative information on a display that the performers are looking at, or by outputting audio that reflects the quantitative information from earphones or the like that the performers are wearing. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-125647 Summary of the Invention [Problem to be solved by the invention]

[0005] To provide performers with more realistic and immersive feedback, it may be possible to deliver the cheers and other sounds of the audience (hereafter referred to as "participants") watching the live stream to the performers in real time. However, delivering high-quality sound to the performers would require a high bit rate, which could cause transmission problems. It may also be possible to use pre-prepared sound effects such as laughter, applause, and cheers, but these pre-prepared sounds are uniform and lack the sense of realism.

[0006] Therefore, the present disclosure proposes an information processing device, an information processing method, and a program that can provide audio data that reflects the individuality of participants and corresponds to their reactions, while taking transmission issues into consideration. [Means for solving the problem]

[0007] According to the present disclosure, an information processing device is proposed that includes a control unit that selects individual pseudo-sound data corresponding to reaction information indicating the reaction of the acquired participant from one or more individual pseudo-sound data that reflect the characteristics of the sound emitted by the participant, and controls the output of the selected individual pseudo-sound data from an audio output device installed at the venue.

[0008] According to the present disclosure, an information processing device is proposed that includes a control unit that performs a process of generating individual pseudo-sound data by reflecting the characteristics of the sounds emitted by the participants in template sound data, and a process of storing the generated individual pseudo-sound data in association with the participants.

[0009] According to the present disclosure, an information processing method is proposed, which includes a processor selecting individual pseudo-sound data corresponding to reaction information indicating the reaction of the acquired participant from one or more individual pseudo-sound data reflecting the characteristics of the sound emitted by the participant, and controlling the output of the selected individual pseudo-sound data from an audio output device installed in the venue.

[0010] According to the present disclosure, a program is proposed that causes a computer to function as a control unit that selects individual pseudo-sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual pseudo-sound data that reflect the characteristics of the sound emitted by the participant, and controls the output of the selected individual pseudo-sound data from an audio output device installed at the venue. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an overview of a live distribution system according to an embodiment of the present disclosure. [Figure 2]3A to 3C are diagrams illustrating generation and storage of individual artificial sound data according to the present embodiment. [Figure 3] FIG. 2 is a block diagram showing an example of the configuration of a pseudo sound generation server according to the present embodiment. [Figure 4] 10A and 10B are diagrams illustrating a process of superimposing extracted participant characteristics on template sound data according to the present embodiment. [Figure 5] 10 is a flowchart showing an example of a flow of generating individual pseudo clapping sound data according to the present embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a display screen for instructions to participants in collecting applause sounds according to the present embodiment. [Figure 7] 10 is a flowchart showing an example of a flow of generating individual pseudo cheer data according to the present embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a display screen for instructions to participants in collecting cheers according to the present embodiment. [Figure 9] FIG. 2 is a block diagram showing an example of the configuration of a venue server according to the present embodiment. [Figure 10] 10 is a flowchart showing an example of the flow of an operational process for outputting individual dummy sound data by the venue server according to the present embodiment. [Figure 11] 10A and 10B are diagrams illustrating clapping operations on the part of participants according to the present embodiment. [Figure 12] 10A and 10B are diagrams illustrating an example of parameter adjustment of individual pseudo-clap sound data according to the present embodiment. [Figure 13] 10A and 10B are diagrams illustrating cheering operations on the part of participants according to the present embodiment. [Figure 14] 10A and 10B are diagrams illustrating an example of parameter adjustment of individual pseudo cheer data according to the present embodiment. [Figure 15] 10A and 10B are diagrams illustrating an example of parameter adjustment of individual pseudo cheer data according to the present embodiment. [Figure 16] 10A and 10B are diagrams illustrating an example of a foot controller for operating a cheer according to the present embodiment. [Figure 17]10A and 10B are diagrams illustrating an example of parameter adjustment of individual pseudo cheer data when using the foot controller according to the present embodiment. [Figure 18] FIG. 10 is a block diagram showing an example of the configuration of a venue server according to a modified example of the present embodiment. [Figure 19] FIG. 10 is a diagram illustrating a transfer characteristic HO according to a modified example of the present embodiment. [Figure 20] FIG. 10 is a diagram illustrating a transfer characteristic HI according to a modified example of the present embodiment. [Figure 21] 10 is a flowchart showing an example of the flow of a transfer characteristic adding process according to a modified example of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0013] The explanation will be given in the following order: 1. Overview of a live streaming system according to an embodiment of the present disclosure 2.Generating individual pseudo-sound data 2-1. Configuration example of individual pseudo sound generation server 50 2-2. Flow of generating individual pseudo-clap sound data 2-3. Flow of generating individual pseudo cheer data 2-4.Other 3. Output of individual pseudo sound data 3-1. Example of the venue server 20 configuration 3-2. Example of operation processing 3-3. Output of individual pseudo-clap sound data (3-3-1. How to use the applause function) (3-3-2. Adjusting the parameters of individual pseudo-clap sound data) 3-4. Output of individual pseudo cheer data (3-4-1. Cheers operation) (3-4-2. Parameter adjustment for individual pseudo cheer data) 4. Variations 4-1.Generation of individual reverberation pseudo sound data 4-2. Example of the configuration of the venue server 20a 4-3.Additional processing of transfer characteristics 5. Supplementary Information

[0014] <<1. Overview of a Live Streaming System According to an Embodiment of the Present Disclosure>> FIG. 1 is a diagram illustrating an overview of a live streaming system according to an embodiment of the present disclosure. As shown in FIG. 1, the live streaming system according to this embodiment includes a venue server 20 (information processing device) that performs live streaming and participant terminals 10 (10A to 10C, etc.) used by each participant watching the live streaming. The participant terminals 10 and the venue server 20 are connected for communication via a network 70 to send and receive data. Also, at the live venue, there are disposed a pseudo sound output device 30 (audio output device) that outputs audio data according to the participant's reactions, and a venue sound acquisition device 40 (audio collection device) that collects sounds (music, etc.) from the venue. The venue server 20 is connected for communication with the pseudo sound output device 30 and the venue sound acquisition device 40 to send and receive data.

[0015] The participant terminal 10 is an example of an information processing device used by participants when viewing live video streamed by the venue server 20. Participants can view the live stream using the participant terminal 10 at a location different from the live venue. For example, the participant terminal 10 may be realized by a smartphone, tablet terminal, PC (personal computer), HMD, wearable device, projector, etc. The participant terminal 10 may also be configured by multiple devices.

[0016] The live streaming system according to this embodiment is an information processing system that can deliver video and audio from a real venue (also referred to as a live venue in this specification) where a concert, seminar, speech, play, or the like is held to participants in a different location from the real venue in real time via a network 70, and can also deliver the participants' reactions to the real venue in real time. The venue's audio is acquired by a venue sound acquisition device 40 and output to a venue server 20. An example of the venue sound acquisition device 40 is a sound processing device that aggregates and appropriately processes the venue's audio. More specifically, a mixer 42 (see FIGS. 19 and 20) is used. The mixer 42 is a device that individually adjusts and mixes various sound sources input from microphones that capture the voices of performers and performances, electronic musical instruments, various players (e.g., CD players, record players, digital players), and the like, and outputs the resulting sound.

[0017] Furthermore, the live streaming system according to this embodiment provides the performers at the live venue with real-time feedback from participants watching from locations other than the live venue. This eliminates the lack of realism that is a concern with live streaming, which is the highlight of traditional live events that attract large audiences. In this way, this embodiment can provide a sense of realism to performers performing at the live venue. Furthermore, in this embodiment, individual artificial sound data reflecting the individuality of each participant is prepared in advance in the venue server 20, and the artificial sound output device 30 installed at the venue is controlled to output the data in real time based on the participant's reactions. This allows for feedback that provides a more realistic feeling rather than uniform feedback, and also eliminates transmission problems such as increased bit rate and delays. For example, this system can be implemented at a low bit rate.

[0018] Here, the individual pseudo sound data in this embodiment is individual pseudo sound data that is individually generated to simulate sounds that participants may make, such as applause, cheers, and shouts. Examples of "cheers" include exclamations that are expected to be uttered during a live performance (e.g., "Wow!", "Oh!", "Oh!", "Eh!", "Yay!", etc.). Examples of "cheering" include the names of performers, words for an encore, words of praise, etc. In this embodiment, the explanation will focus particularly on the handling of various sounds in live streaming.

[0019] The generation of individual dummy sound data and the output control of the individual dummy sound data performed in the live distribution system according to this embodiment will be described below in order.

[0020] <<2. Generation of individual pseudo sound data>> In this embodiment, before the start of live distribution, individual dummy sound data for each participant is generated in advance and stored in the venue server 20. Here, the generation of individual dummy sound data will be specifically described with reference to FIGS.

[0021] FIG. 2 is a diagram illustrating the generation and storage of individual dummy sound data according to this embodiment. The individual dummy sound data according to this embodiment is generated, for example, by an individual dummy sound generation server 50. The individual dummy sound generation server 50 is an example of an information processing device that generates individual dummy sound data that reflects the individuality of a participant based on applause sound data (the actual applause of the participants) and cheering sound data (the participants' actual voices) collected from the participant terminals 10. The individuality of a participant refers to the characteristics of the sound produced by the participant. More specifically, the individual dummy sound generation server 50 generates individual dummy sound data (i.e., synthetic voice) of applause and cheering by superimposing characteristics (e.g., the results of frequency analysis) extracted from actual sounds (collected applause sound data and cheering sound data) acquired from the participant terminals 10 onto prepared template applause sound data and template cheering sound data (both of which are audio signals). The individual dummy sound generation server 50 also acquires setting information (referred to herein as "operation method information") on the participant's operation method for instructing the output of the generated individual dummy sound data from the participant terminal 10. Then, the individual dummy sound generation server 50 outputs the generated individual dummy sound data of applause and / or cheers and the operation method information to the venue server 20 in association with the participant's ID, and stores them in the venue server 20.

[0022] The generation of such individual artificial sound data will be described in more detail below.

[0023] <2-1. Configuration example of individual pseudo sound generation server 50> 3 is a block diagram showing an example of the configuration of the individual dummy sound generation server 50 according to this embodiment. As shown in FIG. 3, the individual dummy sound generation server 50 includes a communication unit 510, a control unit 520, and a storage unit 530.

[0024] (Communication unit 510) The communication unit 510 can communicate with other devices wirelessly or via a wired connection to send and receive data. The communication unit 510 can be implemented, for example, by a wired / wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), or a mobile communication network (LTE (Long Term Evolution), 3G (third-generation mobile communication system), 4G (fourth-generation mobile communication system), or 5G (fifth-generation mobile communication system)). For example, the communication unit 510 can send and receive data to and from the participant terminal 10 and the venue server 20 via the network 70.

[0025] (Control unit 520) The control unit 520 functions as an arithmetic processing unit and a control device, and controls the overall operation of the individual artificial sound generation server 50 in accordance with various programs. The control unit 520 is realized by electronic circuits such as a CPU (Central Processing Unit) or a microprocessor. The control unit 520 may also include a ROM (Read Only Memory) that stores programs to be used, calculation parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.

[0026] The control unit 520 according to this embodiment also functions as an actual sound analysis unit 521, an individual artificial sound data generation unit 522, and a storage control unit 523. The actual sound analysis unit 521 analyzes the actually collected sounds of participants' applause and cheers (sounds actually produced by the participants) received from the participant terminal 10 via the communication unit 510. The participant terminal 10 uses a microphone (hereinafter referred to as a "mic") to collect sounds such as the participants' actual applause, cheers, and calls, digitizes the collected sounds, and transmits the digitized signals (audio signals) to the individual artificial sound generation server 50. The actual sound analysis unit 521 may also perform frequency analysis as an example of analysis and extract frequency characteristics as features. The actual sound analysis unit 521 may also extract time characteristics as an example of analysis. A program (algorithm) for extracting features may be stored in the storage unit 530.

[0027] The individual pseudo-sound data generation unit 522 superimposes the analysis result (extracted features, such as frequency characteristics) by the actual sound analysis unit 521 on the sound data of a template prepared in advance (applause sound data or cheering sound data) to generate individual pseudo-sound data of the applause sound or cheering sound for each participant. FIG. 4 is a diagram for explaining the process of superimposing the features of the participants extracted according to this embodiment on the template sound data.

[0028] The example shown in the upper part of FIG. 4 is an example of superimposing features in the frequency domain. For example, as shown in the upper part of FIG. 4, when there are characteristic frequencies f1 and f2 of the template sound data A (template applause sound data or cheering sound data), assume that the features (frequency characteristics) of a certain participant are f1' and f2' which are deviated from them. In this case, the individual pseudo-sound data generation unit 522 performs a process of processing or transforming f1 of the template sound data A to f1' and f2 to f2'. In the example shown in the upper part of FIG. 4, since f1 < f1' and f2 < f2', the generated (individualized) individual pseudo-applause / cheering sound data will sound higher than the template sound data A. Note that not limited to the example shown in the upper part of FIG. 4, any process of processing or transforming the template sound data A to reflect the features of a certain participant is acceptable, such as adding a new characteristic frequency f3, or reflecting not only the characteristic frequencies but also the slope of the frequency or a more global trend.

[0029] The lower part of FIG. 4 shows an example of superimposing features in the time domain. For example, as shown in the lower part of FIG. 4, if template sound data B (template applause sound data or cheer data) has a start point t1 and an end point t2, the features (frequency characteristics) of a certain participant are assumed to be t1' and t2', which are offset from these points (taking time characteristics into consideration). In this case, the individual artificial sound data generation unit 522 processes or changes t1 of template sound data B to t1' and t2 to t2'. In the example shown in the lower part of FIG. 4, |t2-t1|>|t2'-t1'|, so the pitch becomes higher, and the generated (individualized) individual artificial applause / cheer data sounds higher-pitched than template sound data B. Note that the example shown in the lower part of FIG. 4 is not limited to this, and features may also reflect the envelope of waveform information or a more general trend. Note that in actual cases of numerous applause or cheers, the start timing of each individual's applause / cheer does not coincide and is scattered. Therefore, by setting the starting points t1 / t1' to random values ​​associated with the IDs of the participants, it is possible to generate more natural pseudo-sound data of applause / cheers.

[0030] The template sound data is sound data of applause or cheers prepared (recorded) in advance for use as a template. Multiple patterns of template applause sound data and template cheer data may be prepared. Even when the same person applauds or cheers, the characteristics of the sound vary depending on the clapping style and vocalization. For example, even an individual may clap differently during an event depending on the melody of the music being watched in a live broadcast or the person's level of excitement. Therefore, multiple patterns of applause sounds with different hand forms may be generated. In this case, when collecting the sounds of the participants' applause on the participant terminal 10, instructions such as displaying an illustration of the clapping form are added, and the microphone captures the sounds and analyzes the collected sounds repeatedly for the number of patterns.

[0031] The individual artificial sound data to be generated is assumed to be, for example, one clap, one cheer, or one shout.

[0032] Then, the storage control unit 523 controls so that the generated individual dummy sound data is stored in the venue server 20 in association with the participant ID. The storage control unit 523 also controls so that the operation method information acquired from the participant terminal 10 is stored in the venue server 20 together with the participant ID and the generated individual dummy sound data.

[0033] The above describes the function of generating dummy sound data by the individual dummy sound generation server 50. The generated dummy sound data is not limited to applause and cheers, but may also include shouts, foot tapping, etc. Examples of "chants" include the names of performers, specific words associated with performers or songs, words for an encore, words of praise, etc.

[0034] Furthermore, in this embodiment, sound data prepared (recorded) in advance for use as a template is commonly used to generate individual pseudo sound data with the characteristics of each participant superimposed. If the audio picked up by the participant terminal 10 is registered and used, there is a risk that sounds other than applause or human voices (noise) may be included, and there is also a possibility that noise or sound dropouts may occur if the recording environment on the participant's side is not necessarily high quality (such as microphone performance), so it is preferable to use sound data prepared in advance for the template (high quality and noise-reduced). However, this embodiment is not limited to this, and it is also possible to save the participants' human voices in advance and output them at the venue in response to the participants' operations during live streaming.

[0035] (Storage unit 530) The storage unit 530 is realized by a ROM (Read Only Memory) that stores programs and calculation parameters used in the processing of the control unit 520, and a RAM (Random Access Memory) that temporarily stores appropriately changing parameters, etc. For example, the storage unit 530 stores template applause sound data, template cheer data, a feature extraction program, etc.

[0036] The configuration of the individual dummy sound generation server 50 according to this embodiment has been described above. Note that the configuration of the individual dummy sound generation server 50 shown in FIG. 3 is an example, and the present disclosure is not limited to this. For example, the individual dummy sound generation server 50 may be a system made up of multiple devices. Furthermore, the function of the individual dummy sound generation server 50 (generation of individual dummy applause sound data) may be realized by the venue server 20. Furthermore, the function of the individual dummy sound generation server 50 (generation of individual dummy applause sound data) may be realized by the participant terminal 10.

[0037] Next, the flow of generating individual pseudo applause sound data and the flow of generating individual pseudo cheer data according to this embodiment will be specifically described.

[0038] <2-2. Flow of generating individual pseudo-clap sound data> 5 is a flowchart showing an example of the flow of generating individual pseudo-clap sound data according to this embodiment. The process shown in FIG. 5 is performed before the start of a live streaming event.

[0039] 5, first, a participant uses the participant terminal 10 to perform a login process to the service (live distribution service) provided by the system, and the control unit 520 of the individual dummy sound generation server 50 acquires the participant ID (participant identification information) (step S103). The login screen may be provided by the individual dummy sound generation server 50.

[0040] Next, the control unit 520 of the individual artificial sound generation server 50 controls the participant terminal 10 to collect the sounds of applause (actual sounds) of the participants (step S106). Specifically, the individual artificial sound generation server 50 displays instructions for collecting the sounds of applause on the display unit of the participant terminal 10, and collects the sounds of applause using the microphone of the participant terminal 10. The display unit of the participant terminal 10 may be a display device such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The display unit of the participant terminal 10 may also be a projector that projects an image onto a screen or a wall. If the participant terminal 10 is a see-through head-mounted display (HMD) worn on the head of a participant, instructions may be displayed in augmented reality (AR) on a see-through display unit placed in front of the participant's eyes. The participant terminal 10 may also be connected to various display devices for communication and control the display of instructions.

[0041] FIG. 6 shows an example of a display screen for instructions to participants in collecting applause sounds according to this embodiment. As shown in the upper part of FIG. 6, first, the control unit 520 of the individual pseudo sound generation server 50 displays a screen 132 on the display unit 130 of the participant terminal 10, indicating that the participants' applause sounds will be collected via the microphone input of the participant terminal 10. As an example, in order to more accurately extract the characteristics of the applause sounds, the time points at which participants should clap are presented to the participants. Specifically, for example, five marks are sequentially lit on the screen every second, and participants are instructed to clap in as uniform a manner as possible in accordance with the lighting. At this time, the clapping form may also be presented with an illustration. Participants clap their hands five times in accordance with the instructions and the time points displayed on the screen. To improve the detection accuracy of feature extraction, participants are instructed to clap multiple times (for example, five times). The actual sound analysis unit 521 of the individual artificial sound generation server 50 may use the second and subsequent clapping sounds as data to be analyzed, since the participants may not be accustomed to the first clapping sound among multiple clapping sounds in time with the time points and the accuracy may be reduced. Also, the actual sound analysis unit 521 of the individual artificial sound generation server 50 may average the clapping sounds of multiple clapping sounds and use the averaged clapping sounds as data to be analyzed.

[0042] Next, the actual sound analysis unit 521 calculates the frequency characteristics of the collected clapping sounds, centering on the time points, each time the light is turned on (steps S106 and S109). Specifically, for example, the actual sound analysis unit 521 performs a spectrum analysis of the clapping sounds, using the time points as a guide, and extracts the frequency characteristics from the spectrum information.

[0043] Next, the actual sound analysis unit 521 generates individual pseudo-applause sound data that reflects the characteristics (individuality) of the participants by superimposing the frequency characteristics onto the template applause sound data (step S115). The superimposition of the characteristics (frequency characteristics) is as described above with reference to FIG.

[0044] While such analysis and generation are being performed, the control unit 520 may display a screen 133 indicating "analysis in progress" on the display unit 130, as shown in the middle part of FIG.

[0045] Furthermore, since it is expected that the same person will clap with different characteristics, the individual imitation sound generation server 50 may repeat the process shown in steps S106 to S115 multiple times to generate multiple individual pseudo-clap sound data. For example, the individual imitation sound generation server 50 may present instructions or illustrations that vary the hand form when clapping, the strength (strong, weak), the timing (fast, slow), etc., to obtain multiple patterns of clapping sounds (actual sounds) from the participants, analyze each of them, and generate multiple individual pseudo-clap sound data.

[0046] Next, when the analysis of the applause sounds and the generation of the individual pseudo-applause sound data are all completed, the individual pseudo-applause generation server 50 sets the operation method for the individual pseudo-applause sound data to be performed during the live streaming event (step S118). The individual pseudo-applause generation server 50 displays a screen 134 showing an explanation of the operation method, for example, as shown in the lower part of Fig. 6, and prompts the participants to set the operation method.

[0047] As an operation method, for example, if microphone input of the participant terminal 10 is permitted during an event, the timing of the actual applause of the participant can be used as reaction information (applause output command). Also, operation methods that do not use a microphone may include clicking an icon displayed on the screen during the event (clicking with a mouse or tapping with a finger or electronic pen), operating a predetermined key on a keyboard, gestures (detected by a camera), operating a button on a controller, waving a controller (for example, a penlight), etc. Also, arm movements detected by a sensor attached to the arm of the participant may be used.

[0048] Then, storage control unit 523 of individual pseudo-clap generation server 50 transmits the one or more pieces of generated individual pseudo-clap sound data and operation method information indicating the set operation method in association with the participant ID to venue server 20 (step S121). Venue server 20 stores the participant ID, the one or more pieces of individual pseudo-clap sound data, and the operation method information in association with each other in a storage unit.

[0049] <2-3. Flow of generating individual pseudo cheer data> Next, the flow of generating individual pseudo cheer data will be described with reference to FIGS.

[0050] Fig. 7 is a flowchart showing an example of the flow of generating individual pseudo cheer data according to this embodiment. As shown in Fig. 7, first, the control unit 520 of the individual pseudo sound generation server 50 acquires a participant ID (step S143). As described with reference to Fig. 5, the participant ID may be acquired from the login process performed by the participant. In the case where the generation of individual pseudo cheer data is performed subsequent to the generation of individual pseudo applause sound data, the participant ID can be said to be acquired subsequent to the login process shown in step S103.

[0051] Next, the control unit 520 of the individual dummy sound generation server 50 controls the participant terminal 10 to collect the cheers (actual sounds) of the participants (step S146). Specifically, the individual dummy sound generation server 50 displays instructions for collecting the cheers on the display unit of the participant terminal 10, and collects the cheers using the microphone of the participant terminal 10. Here, FIG. 8 shows an example of a display screen for instructions to the participants when collecting cheers according to this embodiment. As shown in the upper part of FIG. 8, first, the control unit 520 of the individual dummy sound generation server 50 displays a screen 135 on the display unit 130 of the participant terminal 10, indicating that the cheers of the participants will be collected by the microphone input of the participant terminal 10. Here, as an example, a screen is displayed instructing the participant to input the cheers within three seconds after the beep sounds. As described above, various exclamations can be used as cheers, but the participant may select the cheer they wish to register and then speak it. For example, the cheer may be generated using the same exclamation as the selected cheer, or a different exclamation from the selected cheer. Voice characteristics are extracted from the actual voice of the participant, and the characteristics are reflected in template pseudo-cheer data for the selected exclamation, thereby generating individual pseudo-cheer data. The cheer patterns prepared may be multiple or single.

[0052] Next, the actual sound analysis unit 521 analyzes the collected cheers and extracts features (steps S149 and S152). Specifically, for example, the actual sound analysis unit 521 performs a spectrum analysis on the collected cheers and extracts a spectral envelope and formants as features (frequency characteristics) from the spectrum information.

[0053] Next, the actual sound analysis unit 521 applies frequency characteristics to the cheer data of a prepared template to generate individual pseudo cheer data (step S155). Because it may be impossible to perfectly reproduce one's own cheers in a location that differs from the atmosphere of a live concert venue, such as at home, the actual sound analysis unit 521 generates individual pseudo cheer data by superimposing the voice characteristics of each participant onto the prepared template cheer data.

[0054] While such analysis and generation are being performed, the control unit 520 may display a screen 136 indicating "analysis in progress" on the display unit 130, as shown in the middle part of FIG.

[0055] Next, the individual imitation sound generation server 50 may play back the generated individual pseudo cheer data for the participants to confirm (step S158). For example, as shown in the lower part of FIG. 8, the individual imitation sound generation server 50 displays a screen 137 on the display unit 130 prompting the participants to confirm the individual pseudo cheer data. If the participants wish to redo the generation of the individual pseudo cheer data, they can select the "Back" button on the screen 137 and record the cheers again. In other words, the above steps S146 to S158 are repeated.

[0056] Furthermore, if there are calls or chants that are frequently used at the event, participants can add optional words (step S161). For example, participants can select words to add from the available words (chants) by following the instructions displayed on screen 137 shown in the lower part of FIG. 8. The live streamer can prepare candidate optional words in advance, such as calls for "encore," the name of the artist, or fixed calls for specific songs.

[0057] Next, when an optional word is to be added (step S161 / Yes), first, the individual pseudo sound generation server 50 registers the optional word (step S164). To register the optional word, for example, the participant uses the participant terminal 10 to select the word to be added in each form displayed on the screen 137 shown in the lower part of Fig. 8 (for example, each form presents selectable words in a pull-down menu).

[0058] Next, the individual pseudo-sound generation server 50 checks whether the input word is a word that should not be uttered ethically by checking it against a specific dictionary such as a corpus (for example, a list of prohibited words) (step S167). If the word is selected from candidates prepared in advance by the performer, this ethical judgment process may be skipped. Participants can also freely add optional words, and in that case, the word may be checked against, for example, a list of prohibited words prepared in advance by the performer. If a word included in the list of prohibited words is input, the individual pseudo-sound generation server 50 notifies the participant that it cannot be registered.

[0059] Next, when a word that can be registered is input, the individual pseudo sound generation server 50 controls to collect the voices of the participants (step S170). The participants follow the instructions to input the word to be added aloud into the microphone of the participant terminal 10.

[0060] Next, the actual sound analysis unit 521 of the individual artificial sound generation server 50 performs a spectrum analysis on the collected call, and extracts the spectral envelope and formants as features (frequency characteristics) from the spectrum information (step S176).

[0061] Next, the individual pseudo sound data generating unit 522 generates individual pseudo call data by voice synthesis using the extracted frequency characteristics (step S179). For the voice synthesis, a template call prepared in advance by the performer may be used. Furthermore, in the case of words arbitrarily input by the participant, the individual pseudo sound data generating unit 522 may generate a template call by voice synthesis based on the input word (text), and then superimpose the frequency characteristics on the generated template call to generate individual pseudo call data.

[0062] Next, the individual pseudo-voice generation server 50 may play back the generated individual pseudo-chance data for the participant to confirm (step S182). If the participant inputs an instruction to re-create the individual pseudo-chance data, the process returns to step S170, where voice recording is performed again. If the participant inputs an instruction to add another optional word, the process returns to step S164, where the optional word addition process is repeated.

[0063] In the process shown in steps S164 to S179, voice is collected and analyzed each time an optional word is registered, but this embodiment is not limited to this. For example, it is also possible to collect multiple sample voice data from participants and combine the collected sample data with the input optional word to generate more general-purpose individual pseudo-chants data. This makes it possible to generate individual pseudo-chants data without collecting voices and analyzing them each time.

[0064] Next, when the generation of all the individual pseudo cheer data etc. is completed, the individual pseudo cheer generation server 50 sets the operation method of the individual pseudo cheer data etc. to be performed during the live streaming event (step S185). The individual pseudo cheer generation server 50 displays a screen showing an explanation of the operation method etc. on the display unit 130 of the participant terminal 10, and prompts the participant to set the operation method.

[0065] As an operation method, for example, clicking an icon displayed on the screen during the event (clicking with a mouse or tapping with a finger or electronic pen, etc.) or operating a predetermined key on a keyboard can be used as reaction information (output command for cheers, etc.). For example, if multiple individual pseudo sound data for cheers, shouts, etc. are registered, corresponding icons are displayed on the display unit 130 during live streaming, and participants can select which cheers or shouts to output by operating the icons. Furthermore, if the operation of applause is performed by microphone input, the operation of cheers can be performed by, for example, a foot controller operated by stepping on the foot, so that applause and cheers can be input simultaneously. The foot controller will be described later with reference to FIG. 14.

[0066] The cheer operation method is not limited to the above-mentioned example, and it can also be performed by operating buttons on a handheld controller operated by hand, gestures (detected by a camera, acceleration sensor, etc.), or the like.

[0067] Then, the storage control unit 523 of the individual imitation sound generation server 50 transmits the generated one or more individual pseudo cheer data and operation method information indicating the set operation method in association with the participant ID to the venue server 20 (step S188). The venue server 20 stores the participant ID, the one or more individual pseudo cheer data, and the operation method information in association with each other in the storage unit.

[0068] <2-4.Other> The generation of individual dummy sound data according to this embodiment has been specifically described above. While the generation of individual dummy sound data by the individual dummy sound generation server 50 has been described as an example in this embodiment, the present disclosure is not limited thereto. For example, the actual sound analysis process performed by the actual sound analysis unit 521 and the individual dummy sound data generation process performed by the individual dummy sound data generation unit 522 may be performed by the participant terminal 10. Alternatively, the participant terminal 10 may perform the actual sound analysis process (feature extraction) and transmit the analysis results (extracted features) and operation method information along with the participant ID to the individual dummy sound generation server 50, which then generates individual dummy sound data based on the analysis results. When the participant terminal 10 analyzes the actual sound and generates individual dummy sound data, the individual dummy sound generation server 50 transmits necessary programs, template voices, and the like to the participant terminal 10 as needed.

[0069] <<3. Output of individual pseudo sound data>> Next, we will explain the output of individual artificial sound data during live streaming. In this system, the venue server 20 outputs individual artificial sound data corresponding to the reactions of the live attendees to the live venue in real time during live streaming. Specifically, the venue server 20 controls the output of individual artificial applause data and individual artificial cheer data from the artificial sound output device 30 (speaker) installed at the live venue. This makes it possible to deliver the real-time reactions of a large number of attendees to the performers performing at the live venue, increasing the sense of realism of the live performance.

[0070] The configuration of the venue server 20 that controls the output of individual pseudo sound data in this embodiment and an example of operation processing will be described below.

[0071] <3-1. Example of the venue server 20 configuration> 9 is a block diagram showing an example of the configuration of venue server 20 according to this embodiment. As shown in FIG. 9, venue server 20 has a communication unit 210, a control unit 220, and a storage unit 230.

[0072] (Communication unit 210) The communication unit 210 can communicate with other devices wirelessly or via a wired connection to transmit and receive data. The communication unit 210 can be implemented, for example, by a wired / wireless LAN (Local Area Network). For example, the communication unit 210 can transmit and receive data to and from the participant terminal 10 via the network 70. The communication unit 210 can also transmit individual dummy sound data to an dummy sound output device 30 installed in the live venue, and receive venue audio signals (sound sources collected from microphones that input the performers' voices and musical instruments) from the venue sound acquisition device 40.

[0073] (control unit 220) The control unit 220 functions as a calculation processing unit and control device, and controls the overall operation of the venue server 20 in accordance with various programs. The control unit 220 is realized by electronic circuits such as a CPU (Central Processing Unit) or microprocessor. The control unit 220 may also include a ROM (Read Only Memory) that stores the programs to be used, calculation parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.

[0074] Furthermore, the control unit 220 according to this embodiment also functions as a pseudo sound generation unit 221, a pseudo sound output control unit 222, and a venue sound transmission control unit 223.

[0075] The dummy sound generation unit 221 has a function of generating a dummy sound to be output (played) from the dummy sound output device 30 arranged in the venue. Specifically, the dummy sound generation unit 221 selects individual dummy sound data in accordance with reaction information indicating the reactions of the participants received from the participant terminal 10 via the communication unit 210, and adjusts parameters of the selected individual dummy sound data based on the reaction information.

[0076] Here, "reaction information" refers to operation information relating to operations (actions) such as clapping and cheering by participants, for example. The operation information may include, for example, the number of operations per unit time, the operation timing, the operation amount (pressing amount), or selection operation information (such as the ID of the selected item). The operation information may also include a spectrum obtained by frequency analysis of the clapping sound input by the participant. The operation information is operation information for each unit time (fixed period of time), and may be continuously transmitted from the participant terminal 10.

[0077] Based on this operation information, the artificial sound generation unit 221 selects individual artificial sound data that is pre-associated with the number of operations per unit time (a certain period of time), operation timing information, etc. The artificial sound generation unit 221 may also acquire spectral information of applause actually made by the participants as operation information and select individual artificial sound data similar to the spectral information. In some cases, the selection of individual artificial sound data may be controlled by the performer in accordance with the melody of the music being played at the live venue and the content of the event. For example, it is possible to set the selection so that individual artificial sound data of light applause is selected for a ballad song, and individual artificial sound data of vigorous applause is selected for the exciting parts of the latter half of the event. The operation of applause and cheers by the participants in this embodiment will be specifically described with reference to FIGS. 11 to 17.

[0078] Next, the dummy sound generation unit 221 adjusts the parameters of the selected individual dummy sound data based on the operation information. For example, the dummy sound generation unit 221 adjusts the volume in proportion to the number of operations, adjusts the output timing according to the operation timing, etc. This makes it possible to provide each participant's real-time reaction as more natural and realistic feedback.

[0079] The artificial sound output control unit 222 controls the output of the individual artificial sound data selected by the artificial sound generation unit 221 and parameter-adjusted from the artificial sound output device 30. An example of the artificial sound output device 30 is a small speaker (individual sound output device) placed at each audience seat in a live venue. For example, if a participant ID is associated with a virtual position of the participant at the live venue (hereinafter referred to as a virtual position) (audience seat ID may also be used), the artificial sound output control unit 222 controls the output of the individual artificial sound data of each participant from the small speaker placed at each participant's virtual position. This allows the performers to hear the applause and cheers of each participant from each audience seat in the live venue, giving them the feeling that an audience is actually present in the audience seats.

[0080] A small speaker may be provided in each of the spectator seats, or one small speaker may be provided for each of several spectator seats. In order to give the performers a more realistic sense of presence, it is desirable to provide a small speaker in each of the spectator seats (at least in the positions in the venue assigned to each viewing participant), but this is not necessarily limited to this.

[0081] The venue sound transmission control unit 223 controls the transmission of venue sound (venue sound signal) output from the venue sound acquisition device 40 to each participant terminal 10. An example of the venue sound acquisition device 40 is a small microphone (individual sound collection device) (hereinafter referred to as a small microphone) placed at each audience seat in a live venue. For example, the venue sound transmission control unit 223 acquires a venue sound signal picked up by a small microphone placed at the participant's virtual position in the live venue, which is associated with the participant ID, and transmits the signal to the participant terminal 10 of the participant. By collecting venue sound with a small microphone placed at an audience seat corresponding to the virtual position, venue sound including reverberation, perspective, and directional sense of the venue space can be obtained. This gives participants a sense of realism as if they were actually listening in an audience seat at the live venue. In other words, sounds from nearby audience seats (or small speakers placed therein) sound close by, and the reactions of each participant and the sounds of the live performance sound accompanied by venue reverberation.

[0082] Furthermore, the venue sound transmission control section 223 may transmit the venue sound signal after making fine adjustments (normalization, etc.) to the venue sound signal. For example, the venue sound transmission control section 223 performs dynamic range adjustment, etc.

[0083] (Storage unit 230) The storage unit 230 is realized by a ROM (Read Only Memory) that stores programs and calculation parameters used in the processing of the control unit 220, and a RAM (Random Access Memory) that temporarily stores appropriately changing parameters, etc. For example, the storage unit 230 stores individual dummy sound data, operation method information, virtual positions in the venue, etc. in association with participant IDs.

[0084] The configuration of venue server 20 according to this embodiment has been described above. Note that the configuration of venue server 20 shown in Figure 9 is just an example, and the present disclosure is not limited to this. For example, venue server 20 may be made up of multiple devices.

[0085] <3-2. Operation processing example> Next, the operational processing for outputting individual dummy sound data according to this embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the flow of the operational processing for outputting individual dummy sound data by the venue server 20 according to this embodiment. The processing shown in Fig. 10 can be performed continuously during live distribution.

[0086] 10, first, the venue server 20 acquires the participant ID, the number of operations, timing information, etc. in real time from the participant terminal 10 (step S203). The number of operations and timing information are examples of operation information.

[0087] Next, the dummy sound generator 221 of the venue server 20 selects one individual dummy sound data from the one or more individual dummy sound data associated with the participant ID, in accordance with the number of operations and timing information (step S206).

[0088] Next, the dummy sound generation unit 221 adjusts parameters of the selected individual dummy sound data as necessary (step S209). For example, the dummy sound generation unit 221 adjusts the volume proportional to the number of operations and adjusts the timing (trigger, timing of clapping sounds, etc.) according to the operation timing. More specific examples of parameter adjustment will be described with reference to Figs. 12 and 15 to 17. In some cases, the performer may adjust the parameters by multiplying them by a proportionality coefficient α that is specified in advance according to the content of the event, the melody, the genre of the music, etc. This makes it possible to give individuality to the applause and cheers, and to output applause and cheers that vary in real time according to the atmosphere of the performance, even for the same person.

[0089] Next, the artificial sound output control unit 222 controls the reproduction of the individual artificial sound data from a small speaker (an example of an artificial sound output device 30) placed at a virtual position associated with the participant ID (step S212). Note that in this embodiment, it is assumed that a small speaker is placed as the artificial sound output device 30 in each spectator seat of the venue, as an example.

[0090] Next, the venue sound transmission control unit 223 acquires a venue sound signal picked up by a small microphone placed at a virtual position associated with the participant ID (step S215). Here, as an example, it is assumed that a small microphone is placed at each spectator seat in the venue as a venue sound acquisition device 40.

[0091] Furthermore, the venue sound transmission control unit 223 performs fine adjustment (normalization, etc.) of the venue sound signal (step S218), and controls the transmission of the venue sound signal to the participant terminal 10 (step S221).

[0092] The flow of the output processing of individual artificial sound data according to this embodiment has been specifically described above. Note that the steps in the flowchart shown in Fig. 10 may be processed in parallel as appropriate, or may be processed in the reverse order. Also, not all steps need to be processed. For example, the processing shown in steps S203 to S212 is processing for outputting audience voices (individual artificial sound data) to the venue, and is continuously and repeatedly processed during live distribution. Furthermore, in parallel with the audience voice output processing, processing for transmitting venue voices (venue sound signals) to participants shown in steps S215 to S221 is also continuously and repeatedly processed during live distribution.

[0093] Next, the output of the individual artificial sound data will be described in more detail with a specific example.

[0094] <3-3. Output of individual pseudo-clap sound data> First, the output process of individual pseudo clap sound data, which is an example of individual pseudo sound data, will be described.

[0095] (3-3-1. How to use the applause function) 11 is a diagram illustrating the operation of clapping on the participant side according to this embodiment. As shown in FIG. 11, the participant terminal 10 has a communication unit 110, a control unit 120, a display unit 130, a speaker 150, and a microphone 160. Although not shown in FIG. 11, the participant terminal 10 also has a storage unit and an operation input unit 140. The participant terminal 10 has a function to output live video and audio distributed by the venue server 20.

[0096] The control unit 120 functions as a calculation processing unit and a control device, and controls the overall operation of the participant terminal 10 in accordance with various programs. The control unit 120 is realized by electronic circuits such as a CPU (Central Processing Unit) or a microprocessor. The control unit 120 may also include a ROM (Read Only Memory) that stores the programs to be used, calculation parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.

[0097] In this embodiment, the control unit 120 controls the display unit 130 to display (or project onto a wall or screen) live footage received from the venue server 20 via the network 70 by the communication unit 110, and controls the output of venue sound signals from the speaker 150.

[0098] The display unit 130 may be a display device such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The display unit 130 of the participant terminal 10 may also be a projector that projects an image onto a screen or a wall. If the participant terminal 10 is a see-through head-mounted display (HMD) worn on the participant's head, live video or the like may be displayed in AR (Augmented Reality) on a see-through display unit placed in front of the participant's eyes. The participant terminal 10 may also connect to various display devices via communication and control the display of live video or the like. In the example shown in FIG. 11 , the display unit 130 displays live video and icon images indicating whether the participant's microphone input is on or off and whether applause or cheers are being displayed. For example, if the microphone input is on, the participant P can perform the applause operation by actually clapping their hands. Control unit 120 analyzes the clapping sounds picked up by microphone 160, and transmits the number of clappings per unit time or the timing of clapping as operation information (clap operation command) together with the participant ID from communication unit 110 to venue server 20. Control unit 120 can transmit operation information etc. to venue server 20 for each unit time.

[0099] Furthermore, when microphone input is OFF, participant P can perform the clapping operation by clicking the applause icon image with a mouse, tapping the screen with a finger, or pressing a corresponding key on a keyboard. In this case, control unit 120 transmits the number of clicks per unit time, the click timing, etc., as operation information (applause operation command) together with the participant ID from communication unit 110 to venue server 20. Note that operation methods are not limited to these, and participant P can also perform the clapping operation by waving a controller (which may be a penlight, etc.) held in his / her hand or by performing a predetermined gesture. These actions can be detected by various sensors (acceleration sensor, gyro sensor, camera, etc.).

[0100] The applause icon image displayed on the display unit 130 may be controlled to blink in synchronization with the timing of the applause operation when the applause operation is accepted. This allows feedback to be given to the participant P that the operation has been accepted.

[0101] (3-3-2. Adjusting the parameters of individual pseudo-clap sound data) The imitation sound generation unit 221 of the venue server 20 selects individual fake applause sound data based on the participant ID and operation information transmitted from the participant terminal 10. For example, the imitation sound generation unit 221 selects individual fake applause sound data associated with the number of operations per unit time (number of claps, number of clicks, number of taps, etc.). The imitation sound generation unit 221 then adjusts the parameters of the selected individual fake applause sound data based on the operation information.

[0102] 12 is a diagram illustrating an example of parameter adjustment of individual fake applause sound data according to this embodiment. As shown in FIG. 12, the fake sound generation unit 221 adjusts the volume (amplitude) of first individual fake applause sound data corresponding to the number of operations in unit time b1 (e.g., five operations), for example, to a volume proportional to the number of operations in unit time b1, and further adjusts the timing of each of the five playbacks of the first individual fake applause sound data to match the timing of each of the five operations. Next, the fake sound generation unit 221 adjusts the volume (amplitude) of second individual fake applause sound data corresponding to the number of operations in unit time b2 (e.g., six operations), to a volume proportional to the number of operations in unit time b2, and further adjusts the timing of each of the six playbacks of the second individual fake applause sound data to match the timing of each of the six operations. By adjusting the parameters (volume, timing) appropriately according to the operation information for each unit time and then playing the data, it is possible to play individual fake applause data that more realistically reproduces the participants' actual applause. Furthermore, in this system, individual pseudo-clap sound data can be automatically selected for each unit time in accordance with the number of operations, etc.

[0103] <3-4. Output of individual pseudo cheer data> Next, the output process of individual pseudo cheer data, which is an example of individual pseudo sound data, will be described. Note that while the output process of individual pseudo cheer data will be mainly described here as a representative example, the output process of individual pseudo cheer data can also be performed in a similar manner.

[0104] (3-4-1. Cheers operation) 13 is a diagram illustrating cheering operations on the participant side according to this embodiment. In the example shown in Fig. 13, a keyboard 141, a mouse 142, and a foot controller 143 are given as examples of the operation input unit 140 that the participant terminal 10 has.

[0105] Display unit 130 displays live video and icon images indicating whether the participant's microphone input is on or off, and applause and cheers. The icon images indicating cheers may be displayed according to the cheer pattern. These icon images may be displayed in different colors, for example. Furthermore, each icon image indicating a cheer may display text indicating the cheer pattern. Furthermore, icon images for performing cheer output operations may be displayed in the same way as the cheer icon images (text indicating the content of the cheer may also be displayed).

[0106] The participant P may select a cheer by clicking an icon image of a cheer or the like with a mouse, tapping the screen with a finger, or pressing a corresponding key on a keyboard. In this case, the control unit 120 transmits information indicating the selected cheer pattern (information indicating the selection operation), the number of clicks per unit time, or the click timing, as operation information (operation command for cheers, etc.) from the communication unit 110 to the venue server 20 along with the participant ID. Note that the information indicating the selection operation may be an ID (cheer ID) associated with the selected pattern of cheers, etc. A cheer ID may be assigned to individual pseudo cheer data generated in advance. The control unit 120 may transmit the cheer ID selected by the participant P to the venue server 20. Furthermore, because cheers are sounds having a certain duration, the control unit 120 may record only the time (start timing) when the cheer icon image is clicked, etc., as operation timing information and output this as a trigger to the venue server 20 to start playing the cheers. Furthermore, the control unit 120 may tally up the number of clicks per unit time as the number of operations and transmit this to the venue server 20.

[0107] (3-4-2. Parameter adjustment for individual pseudo cheer data) The dummy sound generation unit 221 of the venue server 20 selects individual pseudo cheer data based on the participant ID transmitted from the participant terminal 10 and the cheer ID (an example of selection operation information) indicating the cheer pattern selected by the participant. Then, the dummy sound generation unit 221 adjusts the parameters of the selected individual pseudo cheer data based on the operation information.

[0108] FIG. 14 is a diagram illustrating an example of parameter adjustment of individual pseudo cheer data according to this embodiment. As shown in FIG. 14, the pseudo sound generation unit 221 may start playback of the selected individual pseudo cheer data when a trigger is input, and adjust the volume (amplitude) in proportion to the number of operations per unit time while the cheers are being played. In the example shown in FIG. 14, for example, playback of first individual pseudo cheer data corresponding to the selected cheer ID is started when the trigger is input, and the first individual pseudo cheer data is adjusted to a volume (amplitude) proportional to the number of operations per unit time b1 (e.g., 5 times). Next, the first individual pseudo cheer data is adjusted to a volume (amplitude) proportional to the number of operations per unit time b2 (e.g., 6 times). The selection of the cheer pattern remains active until the audio output of the selected pattern ends (until the next trigger is input, or until the end of a predetermined duration if a word or other sound needs to be spoken for a certain period of time).

[0109] If the volume is adjusted in proportion to the number of operations, for example, if the call consists of a certain length of words (words or sentences), the participant P must continue to operate the call until the end (for example, by repeatedly tapping the icon image of the call). If the operation time by the participant P is shorter than the duration of the call, the voice will disappear midway. Therefore, the dummy sound generator 221 of the venue server 20 can prevent the words from disappearing midway by setting the baseline volume to a value greater than 0 after a trigger is sent from the participant terminal 10 until the duration of the individual pseudo call data ends. FIG. 15 is a diagram illustrating an example of parameter adjustment of individual pseudo call data according to this embodiment. As shown in FIG. 15, even if the number of operations included in a unit time becomes 0 during the duration of the call played after the trigger is input, the dummy sound generator 221 can adjust the volume to the minimum, thereby preventing the voice from disappearing midway.

[0110] (About Foot Controller 143) The method of operating the cheers and shouts is not limited to clicking on the icon images described above. For example, if clapping is performed by inputting actual clapping into a microphone, it is difficult to simultaneously click on the cheer icon image or operate the keyboard. Therefore, in this embodiment, the foot controller 143, which is operated with the feet, may be used for the cheer operation.

[0111] 16 is a diagram illustrating an example of a foot controller 143 for operating cheers according to this embodiment. As shown in FIG. 16, the foot controller 143 is provided with a plurality of switches that are operated by pressing them with the feet. The plurality of switches have, for example, different colors or shapes, and each switch corresponds to a different pattern of cheer. In the case of the foot controller 143, the strength of the operation (how hard the switch is pressed) can be used for parameter adjustment by the artificial sound generation unit 221, rather than the number of operations.

[0112] Each switch of the foot controller 143 may be provided with a sensor that detects the degree of depression. The degree of depression may be detected by a pressure sensor, or, as shown in the lower part of FIG. 16, the amount of change in the height of the switch may be detected. The switch portion is formed of, for example, a rubber-like elastic material, and the height of the switch portion changes depending on the strength of depression. Also, as shown in the upper part of FIG. 16, the foot controller 143 may be provided with a display unit (a depression force meter) that indicates the degree of depression of each switch.

[0113] FIG. 17 is a diagram illustrating an example of parameter adjustment for individual pseudo cheer data when using the foot controller 143 according to this embodiment. When using the foot controller 143, the operation amount (the strength with which the switch is pressed, the amount of depression, or the amount of change in the height of the switch) changes continuously. The control unit 120 samples this change and transmits the operation amount to the venue server 20. For example, the control unit 120 may sample at a low frequency to reduce the amount of data. Specifically, for example, as shown in the upper part of FIG. 17, sampling may be performed at a frequency per unit time, and only strength information at two points, the start time and the end time, may be transmitted as operation amount information.

[0114] The artificial sound generation unit 221 of the venue server 20 interpolates between two sampled points for each unit time, and creates a smooth approximation signal as shown by the dotted line in the upper part of Fig. 17. The artificial sound generation unit 221 then adjusts the volume (amplitude) of the individual artificial cheer data in accordance with the created approximation signal, as shown in the lower part of Fig. 17. If a trigger time is included within the unit time, the start time is replaced with the trigger time to generate a volume envelope signal.

[0115] Similarly, in the case of the foot controller 143, if the sound is adjusted according to the amount of operation, a call requiring a certain duration must be continued until the end of the word. Therefore, the artificial sound generation unit 221 may adjust parameters so that playback continues at the minimum volume as long as the duration of the artificial sound call is within the duration, even if the operation amount information (information such as pressing strength) included in the unit time is zero. Furthermore, as shown in the upper part of FIG. 16 , a meter indicating the operation time (call duration) may be installed at a position corresponding to each switch on the foot controller 143. By indicating that the call is being made while the meter is lit, participants can be encouraged to consciously continue operating the foot controller 143 until the call duration ends. Considering the possibility that it may be difficult to see such a meter located at the foot during live streaming, the control unit 120 may display the control parameters of the foot controller 143 on the display unit 130. For example, the control unit 120 may display a meter indicating the operation time (duration of the cheer) next to the cheer icon image, and indicate that the cheer is being emitted while the meter is lit, as shown in Fig. 13. Furthermore, the control unit 120 may change the color intensity of the icon image depending on the strength with which the switch of the foot controller 143 is pressed.

[0116] <<4. Modifications>> Next, a modified example of the live distribution system according to this embodiment will be described with reference to FIGS.

[0117] In the above-described embodiment, as an example of the artificial sound output device 30, small speakers are placed in each audience seat at a live venue, and individual artificial sound data for the corresponding participant is output from each small speaker, thereby giving the performers on stage the feeling that an audience is actually present in the audience seats. However, it is also possible to output individual artificial sound data for each participant from large speakers (another example of an artificial sound output device) installed on or around the stage for the performers, rather than using multiple small speakers. In this case, by adding a sense of perspective, a sense of direction, and the reverberation characteristics of the venue (collectively referred to as transfer characteristics) to the individual artificial sound data for each participant to be output, it is possible to give the performers on stage the feeling that they are hearing sounds from the audience seats in the venue.

[0118] Furthermore, in the above-described embodiment, a small microphone was used at each audience seat as an example of venue sound acquisition device 40, but even if a venue cannot provide a large number of small microphones (for example, for the number of audience seats), the venue server 20 can perform predetermined processing on the venue sound signal output from a venue mixer (another example of venue sound acquisition device 40) to give participants the feeling that they are actually listening from the audience seats, such as feeling the reverberations of the venue space. A mixer is a device that individually adjusts and mixes various sound sources input from audio equipment such as microphones that pick up the voices and performances of performers, electronic musical instruments, and various players (for example, CD players, record players, digital players), and outputs the mixed sound, and is an example of an audio processing device that aggregates the venue's audio and processes it appropriately.

[0119] In addition, the specified processing performed on the venue sound signal output from the mixer is a process that adds characteristics such as perspective, directional sense, and reverberation of the venue space (collectively referred to as transmission characteristics) that correspond to the position of the virtual audience seats in the live venue (hereinafter also referred to as virtual position) associated with the participant.

[0120] Furthermore, individual reverberation artificial sound data including reverberations in the venue space may be prepared in advance, and all individual reverberation artificial sound data selected according to the reactions of each participant may be added together and transmitted to the participant terminal 10 together with the venue sound signal. The individual reverberation artificial sound data may be reverberation artificial applause sounds, reverberation artificial cheers, etc. selected according to the reaction of each participant. This allows participants to hear the reactions of all the spectators in the venue, including themselves, with the feeling that they are actually listening to them from the audience seats.

[0121] The use of reverberation-specific artificial sound data and the process of adding transfer characteristics according to this embodiment will be specifically described below.

[0122] <4-1. Generation of individual pseudo-reverberation sound data> In this modification, first, sound data such as applause recorded at an actual live venue, i.e., reverberation template artificial sound data, is prepared. By recording reverberation template artificial sound data (reverberation template applause sound data and reverberation template cheering data) at the actual live venue in advance, sound data including reverberations from the venue space can be obtained. Next, before the start of live streaming, the prepared reverberation template artificial sound data is combined with the characteristics of the sounds emitted by the participants (for example, frequency characteristics) to generate individual reverberation artificial sound data for each participant.

[0123] The reverberation individual dummy sound data may be generated by the individual dummy sound generation server 50, similar to the individual dummy sound data. The generated reverberation individual dummy sound data may be stored in the storage unit 230 of the venue server 20 in association with the participant ID, similar to the individual dummy sound data. Note that the reverberation individual dummy sound data may be associated with individual dummy sound data of the same pattern. In this case, a dummy sound ID may be assigned to each dummy sound data, and the association may be performed using the dummy sound ID.

[0124] The process of generating reverberant individual artificial sound data is the same as the process of generating the individual artificial sound data described above, except for the nature of the template used for generation. The individual artificial sound generation server 50 can generate the individual artificial sound data and the reverberant individual artificial sound data by combining features extracted from the applause and cheers input by the participants into the microphone with the template artificial sound data and the reverberant template artificial sound data.

[0125] The template artificial sound data is sound data such as applause and cheers recorded in an anechoic environment, and the reverberation template artificial sound data is sound data such as applause and cheers recorded in advance at an actual live venue. The reverberation template artificial sound data used may also be reverberation template artificial sound data corresponding to a virtual position of a participant at the live venue (i.e., sound recorded at the actual location corresponding to the virtual position, such as applause and cheers).

[0126] <4-2. Example of the configuration of the venue server 20a> Figure 18 is a block diagram showing an example of the configuration of venue server 20a according to a modification of this embodiment. As shown in Figure 18, venue server 20a has communication unit 210, control unit 220a, and storage unit 230. Note that the configuration of the same reference numerals as in venue server 20 described with reference to Figure 9 is as described above, and therefore detailed description here will be omitted.

[0127] The control unit 220a according to this modification includes a pseudo sound generating unit 221a, a transfer characteristic H O The addition unit 225, the pseudo sound output control unit 222a, and the transfer characteristic H I It also functions as an adding unit 226, an all-participant echo artificial sound synthesizing unit 227, and a hall sound transmission control unit 223a.

[0128] (pseudo sound generation unit 221a) The dummy sound generation unit 221a selects individual dummy sound data based on the participant ID, operation information, etc. acquired from the participant terminal 10 by the communication unit 210. The dummy sound generation unit 221a also selects reverberation individual dummy sound data. For example, the dummy sound generation unit 221a selects reverberation individual dummy sound data of the same pattern as the selected individual dummy sound data. As described above, the reverberation individual dummy sound data can be generated in advance, similar to the individual dummy sound data, and stored in the storage unit 230.

[0129] Then, the artificial sound generation unit 221a adjusts the parameters of the selected individual artificial sound data and reverberation individual artificial sound data. The details of the parameter adjustment are the same as those in the above-described embodiment, and examples of the parameter adjustment include adjusting the volume in proportion to the number of operations and adjusting the output timing according to the operation timing.

[0130] (Transfer characteristics H O Additional section 225) Transfer characteristic H O The adding unit 225 adds a transfer characteristic H of the reverberation of the venue measured in advance to the individual pseudo sound data output from the pseudo sound generating unit 221a. O Add the transfer characteristic H O is the transmission characteristic from the audience seats to the stage (around the performers). O By adding this, even if it is not possible to place small speakers in each audience seat in the venue and only one large speaker 32 (an example of the pseudo sound output device 30) can be placed on the stage at the performers' feet in front of them, it is possible to give the performers the feeling that an audience is present in the space of the venue.

[0131] FIG. 19 shows the transfer characteristic H OAs shown in FIG. 19, a stage and audience seats are set up in a live venue, and an ID (audience seat ID) is assigned to each seat. In the example shown in FIG. 19, the virtual position (virtual position) of participant A and the virtual position of participant B in the live venue are illustrated. The audience seat corresponding to the virtual position is set as the starting point, and the vicinity of the performer (for example, the area surrounded by the dashed line) is set as the sound receiving point. The transfer characteristic (H O(A) , H O(B) ) is measured. O Measurements may be taken for all spectator seats.

[0132] In addition, the receiving point may be changed as appropriate, such as the position of the performer if the performer does not move on the stage, or at least one large speaker 32 (a comprehensive audio output device that outputs individual pseudo-sound data for all participants to the performers) placed in front of the performers on stage at their feet if the performers move to a certain extent or there are multiple performers.

[0133] Measured transfer characteristic H O is stored in the storage unit 230 of the venue server 20 in association with the spectator seat ID. O The adding unit 225 calculates the corresponding transfer characteristic H based on the audience seat ID (virtual position) associated with the participant ID. O Then, the transfer characteristic H O The adding unit 225 adds the acquired transfer characteristic H O Add.

[0134] (dummy sound output control unit 222a) The artificial sound output control section 222a has a transfer characteristic H O The addition unit 225 generates a transfer characteristic H O The individual pseudo sound data to which the above has been added is added for all participants, and the result is output from the large speaker 32.

[0135] (Transfer characteristics H I Additional section 226) Transfer characteristic H IThe adding unit 226 applies a transfer characteristic H from a performer speaker 60 (audio output device) that is provided in the venue facing the audience seats and outputs the venue sound signal input from the mixer 42 to each audience seat to the venue sound signal output from the mixer 42 (an example of the venue sound acquisition device 40). I In this embodiment, sound sources from various audio equipment such as microphones and musical instruments used by performers at a live venue are mixed by a mixer 42, for example, and output from performer speakers 60 at the live venue toward the audience seats, and also distributed to the participant terminal 10. Here, the venue sound signal transmitted to the participant terminal 10 is given a transfer characteristic H I By adding this, it is possible to recreate the feeling of listening to the sound from the venue from each seat in the audience.

[0136] Transfer characteristic H I can be measured in advance before live distribution starts. I As shown in FIG. 20, a live venue is provided with a stage and audience seats, and each seat is assigned an ID (audience seat ID). In the example shown in FIG. 20, the virtual position (virtual position) of participant A in the live venue and the virtual position of participant B are illustrated. As an example, the performer speakers installed in the venue are assumed to be two speakers (performer speaker 60R and performer speaker 60L) installed on the left and right sides of the stage. Then, for each audience seat (A, B) corresponding to the virtual position, the transfer characteristics (H R I(A) , H L I(A) , H R I(B) , H L I(B) ) is measured. I Measurements may be taken for all spectator seats.

[0137] Measured transfer characteristic H I is associated with the spectator seat ID and stored in the venue server 20. I The adding unit 226 calculates the corresponding transfer characteristic H based on the audience seat ID (virtual position) associated with the participant ID.I Then, the transfer characteristic H I The adding unit 226 applies the acquired transfer characteristic H I This creates a sound that mimics the sound space experienced when listening to a performance at a venue from each seat in the audience.

[0138] (All participants echo pseudo sound synthesis section 227) The all-participant artificial echo sound synthesis unit 227 has the function of adding up all of the individual artificial echo sound data for all participants output from the artificial echo sound generation unit 221a. The venue sound signal output from the mixer 42 contains only the output from the performers' microphones, instruments, players, etc. connected to the mixer 42, and therefore does not include the applause and cheers of the entire audience. Therefore, by adding up all of the individual artificial echo sound data for each participant and transmitting it together with the venue sound signal to the participant terminal 10 by the venue sound transmission control unit 223a, it is possible to deliver to the participants the reactions of all participants that mimic the venue's reverberations, i.e., the applause and cheers that match the venue's sound space. This allows participants to listen to and hear the reactions of all the audience in the venue, including themselves, as if they were actually in the audience seats.

[0139] (venue sound transmission control unit 223a) The venue sound transmission control unit 223a controls the transmission of the venue sound (venue sound signal) output from the mixer 42 and the individual reverberation pseudo sound data for all participants synthesized by the all-participant reverberation pseudo sound synthesis unit 227 to the participant terminal 10.

[0140] The configuration of venue server 20a according to this modified example has been specifically described above. Note that the configuration shown in FIG. 18 is just one example, and the present disclosure is not limited to this. For example, venue server 20a may be made up of multiple devices. Also, venue server 20a may not have all of the configuration shown.

[0141] <4-3. Additional processing of transfer characteristics> FIG. 21 is a flowchart showing an example of the flow of the transfer characteristic adding process according to the modified example of this embodiment.

[0142] 21, first, the venue server 20a acquires the participant ID, the number of operations, timing information, etc. in real time from the participant terminal 10 (step S303). The number of operations and timing information are examples of operation information.

[0143] Next, the dummy sound generator 221 of the venue server 20 selects one individual dummy sound data from the one or more individual dummy sound data associated with the participant ID, in accordance with the number of operations and timing information (step S306).

[0144] Next, the dummy sound generation unit 221 adjusts parameters of the selected individual dummy sound data as necessary (step S309).

[0145] Next, the transfer characteristic H O The adding unit 225 adds a transfer characteristic H corresponding to a virtual position (for example, an audience seat ID) in the live venue associated with the participant ID to the individual pseudo sound data. O is added (step S312).

[0146] Next, the artificial sound output control section 222a outputs the artificial sound with the transfer characteristic H O The added individual pseudo sound data is played back (step S315). O As described above, this is the transfer characteristic from a predetermined audience seat to the stage where the performers are located. The artificial sound output control unit 222a calculates the transfer characteristic H O The large speaker 32 is a large speaker placed facing the performer, for example, at the performer's feet on the stage of a live concert venue, and is controlled to output (play) the individual pseudo sound data to which the respective pseudo sound data are added, with a transfer characteristic H O By outputting individual pseudo-sound data with added sound effects, performers on stage can be given a sense of perspective, direction, and reverberation, as if they were hearing applause and cheers from the audience at a live concert.

[0147] Next, the venue server 20a acquires the venue sound signal from the mixer 42 in the venue (step S318).

[0148] Next, the transfer characteristic H I The adding unit 226 adds a transfer characteristic H corresponding to the virtual position (audience seat ID) associated with the participant ID to the hall sound signal. I (Step S321). I As mentioned above, this is the transmission characteristic from, for example, the performer speaker 60 to a specific audience seat. This makes it possible to generate a venue sound signal that reproduces the reverberation and other effects of the space in a live venue, taking into account the virtual positions of the participants.

[0149] Next, the venue sound transmission control unit 223a finely adjusts (normalizes, etc.) the venue sound signal that simulates the reverberation in the venue (step S324).

[0150] On the other hand, the dummy sound generation unit 221a selects one of the reverb-added individual dummy sound data associated with the participant ID based on the operation information received from the participant terminal 10, and adjusts parameters based on the operation information (step S327). This process may be performed in parallel with the process shown in step S306. The dummy sound generation unit 221a may also select reverb-added individual dummy sound data (dummy sound data of the same pattern) associated with the individual dummy sound data selected in the process shown in step S306. Then, the dummy sound generation unit 221a adjusts the volume of the selected reverb-added individual dummy sound data in proportion to the number of operations and adjusts the timing according to the operation timing, similar to the parameter adjustment shown in step S309.

[0151] Next, the all-participant echo artificial sound synthesis unit 227 synthesizes individual echo-added artificial sound data (parameter-adjusted) for all participants (step S330).

[0152] Then, the venue sound transmission control unit 223a controls the transmission of a venue sound signal that simulates the reverberation in the venue and individual pseudo sound data with reverberation for all participants to the participant terminal 10 (step S333). This allows participants to listen to the venue sound signal that reproduces the reverberation in the space of a live venue taking into account the virtual position of the participant, and the reactions of all the audience in the venue, including the participant himself, with the sensation of actually listening to it from the audience seats.

[0153] The flow of the transfer characteristic addition process according to the modified example of this embodiment has been specifically described above. Note that the steps in the flowchart shown in FIG. 21 may be processed in parallel as appropriate, or may be processed in the reverse order. Also, not all steps need to be processed. For example, the process shown in steps S303 to S315 is a process for outputting audience voices (individual artificial sound data) to the venue, and is continuously and repeatedly processed during live distribution. Furthermore, in parallel with the audience voice output process, the process for preparing to return artificial sound to participants shown in steps S327 to S330 and the process for transmitting venue voices (venue sound signals) to participants shown in steps S318 to S333 may be continuously and repeatedly processed during live distribution.

[0154] <<5. Supplementary Information>> Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the present technology is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical ideas described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.

[0155] For example, the above-described embodiments and modifications may be combined as appropriate. As an example, the venue server 20 outputs individual pseudo sound data corresponding to the reactions of each participant from small speakers provided in each audience seat, while the venue sound signal (with transfer characteristics H I The participant terminal 10 may transmit the individual pseudo sound data of the echo of all participants together with the additional sound data.

[0156] In addition, the venue server 20 generates individual pseudo sound data according to the reaction of each participant, with a transfer characteristic H O Alternatively, the audio may be output from at least one large speaker placed on a stage or the like facing the performers without undergoing the additional processing.

[0157] It is also possible to create one or more computer programs for causing hardware such as a CPU, ROM, and RAM built into the participant terminal 10, venue server 20, or individual dummy sound generation server 50 to perform the functions of the participant terminal 10, venue server 20, or individual dummy sound generation server 50. A computer-readable storage medium storing the one or more computer programs is also provided.

[0158] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0159] The present technology can also be configured as follows. (1) An information processing device comprising: a control unit that selects individual dummy sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual dummy sound data that reflect the characteristics of sounds emitted by the participant; and controls the output of the selected individual dummy sound data from an audio output device installed at the venue. (2) The information processing device described in (1) above, wherein the control unit selects individual pseudo-sound data corresponding to the participant's reaction information obtained in real time at a location different from the venue, and controls the output from the audio output device to the performer at the venue. (3) The information processing device described in (1) or (2), wherein the participant's reaction information includes at least one of information indicating the number of operations performed by the participant, information indicating the timing of operations performed by the participant, information indicating the amount of operation, spectral information obtained by frequency analysis of the sound emitted by the participant, or selection operation information performed by the participant. (4) the one or more individual pseudo sound data are one or more different individual pseudo clap sound data, The information processing device described in (3), wherein the control unit selects corresponding individual pseudo-clap sound data from the one or more different individual pseudo-clap sound data based on at least one of the number of applause, the number of click operations, the number of tap operations, or spectrum by the participants in a certain period of time. (5) The participant reaction information includes information indicating the timing of applause by the participant; The information processing device according to (4), wherein the control unit adjusts an output timing of the selected individual pseudo clapping sound data in accordance with a timing of the clapping. (6) The participant reaction information includes information indicating the number of times the participant applauds, The information processing device according to (4) or (5), wherein the control unit adjusts the volume of the output individual pseudo clapping sound data in accordance with the number of clappings within a certain period of time. (7) The one or more individual pseudo sound data are one or more different individual pseudo cheer data or individual pseudo shout data, The information processing device according to (3), wherein the control unit selects corresponding individual pseudo cheer data or individual pseudo cheer data in accordance with a selection operation by the participant. (8) The control unit triggering the start of a selection operation by the participant to start outputting the selected individual pseudo cheer data or individual pseudo shout data; The information processing device according to (7), wherein the volume of the output individual pseudo cheer data or individual pseudo shout data is changed in real time depending on the number of times the selection operation is performed or the amount of operation of the selection operation. (9) The information processing device according to (7) or (8), wherein the control unit adjusts the volume of the individual pseudo shout data to continue outputting the data at least at a minimum volume until the duration of the individual pseudo shout data ends. (10) The information processing device according to any one of (1) to (9), wherein the control unit controls output of the selected individual pseudo sound data from an individual sound output device arranged at a virtual position of the participant in the venue. (11) The information processing device described in (10) wherein the control unit controls the transmission of a venue sound signal acquired from an individual sound collection device placed at the virtual position of the participant in the venue to a participant terminal used by the participant who is located at a location different from the venue. (12) The information processing device described in any one of (1) to (9), wherein the control unit adds a transmission characteristic from the virtual position of the participant in the venue to the performer in the venue to the selected individual pseudo sound data, and controls the output from an integrated audio output device located around the performer in the venue. (13) The information processing device described in any one of (1) to (9), wherein the control unit controls the transmission of a venue sound signal acquired from a sound processing device that aggregates sound sources from the venue's audio equipment to a participant terminal used by the participant who is located in a location other than the venue. (14) The information processing device described in (13) above, wherein the control unit adds a transmission characteristic from an audio output device that outputs the venue sound signal toward an audience seat in the venue to the venue sound signal acquired from the sound processing device, to the virtual position of the participant in the venue, and then controls the transmission to the participant terminal. (15) The control unit Selecting individual reverberation artificial sound data corresponding to reaction information indicating real-time reactions of the participants from one or more individual reverberation artificial sound data that have been generated in advance by reflecting characteristics of sounds emitted by the participants in reverberation artificial sound data including reverberations from the venue; Synthesize the selected individual reverberation pseudo sound data of all participants; The information processing device according to (13) or (14), wherein the synthesized individual reverberation pseudo sound data for all participants is transmitted to the participant terminal together with the hall sound signal. (16) The information processing device described in any one of (1) to (9), wherein the control unit adds a transmission characteristic from an audio output device that outputs the venue sound signal toward the audience seats in the venue to the venue sound signal obtained from an audio processing device that aggregates sound sources from the venue's audio equipment, and then controls the transmission to the participant terminal. (17) A process of generating individual pseudo-sound data by reflecting the characteristics of the sounds made by the participants in the template sound data; a process of storing the generated individual pseudo sound data in association with the participant; An information processing device comprising a control unit that performs the above. (18) The information processing device described in (17), wherein the control unit generates the individual pseudo sound data by synthesizing one or both of frequency characteristics and time characteristics obtained by analyzing the sounds emitted by the participants with the sound data of the template. (19) The information processing device further includes a communication unit, The communication unit receiving characteristics of sounds emitted by the participants that have been collected and analyzed by the participant terminals used by the participants; The information processing device described in (17) or (18) above, which associates the generated individual pseudo sound data with the identification information of the participant and transmits it to a venue server that controls output from an audio output device installed at the venue. (20) The processor: An information processing method including: selecting individual dummy sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual dummy sound data that reflect the characteristics of sounds emitted by the participant; and controlling the output of the selected individual dummy sound data from an audio output device installed in the venue. (twenty one) Computer, A program that functions as a control unit that selects individual pseudo-sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual pseudo-sound data that reflect the characteristics of the sound emitted by the participant, and controls the output of the selected individual pseudo-sound data from an audio output device installed in the venue. (twenty two) The processor: Generate individual pseudo-sound data by reflecting the characteristics of the sounds made by the participants in the template sound data; storing the generated individual pseudo sound data in association with the participant; An information processing method, including: (twenty three) Computer, A process of generating individual pseudo-sound data by reflecting the characteristics of the sounds made by the participants in the template sound data; a process of storing the generated individual pseudo sound data in association with the participant; A program that functions as a control unit to perform the above. (twenty four) The system comprises a participant terminal used by the participant and a server that controls output from an audio output device installed in the venue, The server a communication unit that receives reaction information indicating reactions of the participants from the participant terminals; a control unit that selects individual dummy sound data corresponding to the reaction information indicating the reaction of the participant from one or more individual dummy sound data that reflect the characteristics of the sound emitted by the participant, and controls the audio output device to output the selected individual dummy sound data. system. [Explanation of symbols]

[0160] 10 Participant terminals 110 Communications Department 120 control section 130 Display section 140 Operation input section 150 speakers 160 Mike 20, 20a Venue Server 210 Communications Department 220, 220a control unit 221, 221a Pseudo sound generator 222, 222a Pseudo sound output control section 223, 223a Venue sound transmission control section 225 Transfer characteristics H O Additional part 226 Transfer characteristics H I Additional part 227 All participants echo pseudo sound synthesis section 230 Storage section 30 Pseudo sound output device 40 Venue sound acquisition device 50 Individual pseudo sound generation server 510 Communications Department 520 control section 521 Actual Sound Analysis Unit 522 Individual pseudo sound data generation unit 523 Storage Control Unit 530 Storage section 60 Speaker 70 Network

Claims

1. An information processing device comprising: a control unit that selects individual dummy sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual dummy sound data that reflect the characteristics of sounds emitted by the participant; and controls the output of the selected individual dummy sound data from an audio output device installed at the venue.

2. 2. The information processing device according to claim 1, wherein the control unit selects individual pseudo sound data corresponding to reaction information of the participants acquired in real time at a location different from the venue, and controls the sound output device to output the pseudo sound data to a performer at the venue.

3. 2. The information processing device according to claim 1, wherein the participant's reaction information includes at least one of information indicating the number of operations performed by the participant, information indicating the timing of operations performed by the participant, information indicating the amount of operations performed, information on a spectrum obtained by frequency analysis of a sound emitted by the participant, or information on a selection operation performed by the participant.

4. the one or more individual pseudo sound data are one or more different individual pseudo clap sound data, 4. The information processing device according to claim 3, wherein the control unit selects corresponding individual pseudo-clap sound data from the one or more different individual pseudo-clap sound data based on at least one of the number of applause, the number of click operations, the number of tap operations, or a spectrum by the participants within a certain period of time.

5. The participant reaction information includes information indicating the timing of applause by the participant; The information processing device according to claim 4 , wherein the control unit adjusts an output timing of the selected individual pseudo-clap sound data in accordance with a timing of the clapping.

6. The participant reaction information includes information indicating the number of times the participant applauds, The information processing device according to claim 4 , wherein the control unit adjusts the volume of the individual pseudo-clap sound data to be output in accordance with the number of claps within a certain period of time.

7. the one or more individual pseudo sound data are one or more different individual pseudo cheer data or individual pseudo shout data, The information processing device according to claim 3 , wherein the control unit selects corresponding individual pseudo cheer data or individual pseudo cheer data in accordance with a selection operation by the participant.

8. The control unit triggering the start of a selection operation by the participant to start outputting the selected individual pseudo cheer data or individual pseudo shout data; The information processing device according to claim 7 , wherein the volume of the output individual pseudo cheer data or individual pseudo cheer data is changed in real time according to the number of times the selection operation is performed or the amount of the selection operation.

9. The information processing device according to claim 7 , wherein the control unit adjusts the volume of the individual pseudo cheer data so that the output continues at least at a minimum volume until the duration of the individual pseudo cheer data ends.

10. The information processing apparatus according to claim 1 , wherein the control unit controls output of the selected individual pseudo sound data from an individual sound output device disposed at a virtual position of the participant in the venue.

11. The information processing device according to claim 10, wherein the control unit controls the transmission of a venue sound signal acquired from an individual sound collection device placed at the virtual position of the participant in the venue to a participant terminal used by the participant who is located at a location different from the venue.

12. 2. The information processing device according to claim 1, wherein the control unit adds a transmission characteristic from the virtual position of the participant in the venue to the performer in the venue to the selected individual pseudo-sound data, and controls the output from an integrated audio output device located near the performer in the venue.

13. The information processing device according to claim 1, wherein the control unit controls the transmission of a venue sound signal acquired from an audio processing device that aggregates sound sources from the venue's audio equipment to a participant terminal used by the participant who is located in a location other than the venue.

14. The information processing device according to claim 13, wherein the control unit adds a transmission characteristic from an audio output device that outputs the venue sound signal toward an audience seat in the venue to a virtual position of the participant in the venue to the venue sound signal acquired from the sound processing device, and then controls the transmission to the participant terminal.

15. The control unit selecting individual reverberation artificial sound data corresponding to reaction information indicating real-time reactions of the participants from one or more individual reverberation artificial sound data that have been generated in advance by reflecting characteristics of sounds emitted by the participants in reverberation artificial sound data including reverberations from the venue; Synthesize the selected individual reverberation pseudo sound data of all participants; 14. The information processing apparatus according to claim 13, wherein the synthesized individual reverberation pseudo sound data for all participants is transmitted to the participant terminal together with the hall sound signal.

16. A process of generating individual pseudo-sound data by reflecting the characteristics of the sounds made by the participants in the template sound data; a process of storing the generated individual pseudo sound data in association with the participant; An information processing device comprising a control unit that performs the above.

17. 17. The information processing device according to claim 16, wherein the control unit generates the individual pseudo sound data by synthesizing one or both of frequency characteristics and time characteristics obtained by analyzing the sounds emitted by the participants with the sound data of the template.

18. The information processing device further includes a communication unit, The communication unit receiving characteristics of sounds emitted by the participants that have been collected and analyzed by the participant terminals used by the participants; 17. The information processing device according to claim 16, wherein the generated individual dummy sound data is associated with identification information of the participant and transmitted to a venue server that controls output of the individual dummy sound data from an audio output device installed at the venue.

19. The processor: An information processing method including: selecting individual pseudo-sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual pseudo-sound data that reflect the characteristics of sounds emitted by the participant; and controlling the output of the selected individual pseudo-sound data from an audio output device installed in the venue.

20. Computer, A program that functions as a control unit that selects individual pseudo-sound data corresponding to acquired reaction information indicating the reaction of a participant from one or more individual pseudo-sound data that reflect the characteristics of the sound emitted by the participant, and controls the output of the selected individual pseudo-sound data from an audio output device installed in the venue.

Citation Information

Patent Citations

  • Remote utilization method

    JP1999025188A

  • Actual applause induction type automatic applause device

    JP2000137492A

  • Environmental sound synthesizer, environmental sound transmission system, environmental sound synthesizing method, environmental sound transmission method, and program

    JP2014063145A

  • Sound signal processing system

    JP2015097318A

  • Information communication program, information communication device, and delivery server

    JP2015125647A