Information processing device, information processing method, and storage medium
By using pre-prepared audio data based on real-time metadata, the system addresses server load and communication issues in live streaming, allowing for efficient sharing of audience reactions, thereby improving the event experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2026-03-04
AI Technical Summary
Existing technologies for sharing viewer reactions during live streaming impose a heavy processing load on servers and require significant communication capacity, leading to potential delays.
An information processing device and method that generates viewer audio data using pre-prepared audio data based on real-time audio metadata acquired from viewer terminals, reducing server load and communication requirements.
This approach effectively reduces server processing and communication load while enabling real-time feedback of audience reactions to performers, enhancing the event experience for both viewers and performers.
Smart Images

Figure 0007823668000001 
Figure 0007823668000002 
Figure 0007823668000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a storage medium. [Background technology]
[0002] Events such as sports and live music concerts can be experienced not only at the venue where the event is held, but also through television broadcasts and internet streaming. Particularly in recent years, with the widespread availability and convenience of the internet, events can be streamed in real time, allowing many viewers to participate from public viewing venues or their homes. One major difference between experiencing an event via internet streaming and experiencing it at a venue is that there is no way to communicate audience reactions, such as cheers and applause, to the performers or other viewers. Audience reactions, such as cheers and applause, can motivate performers and further increase the excitement among the audience, making them an important element of any event.
[0003] Regarding such technology for sharing viewer reactions, for example, Patent Document 1 below discloses a method for collecting the vocalizations of each viewer (remote user) watching from a remote location, transmitting the data to a server, adding multiple pieces of audio data on the server, and distributing the data to each remote user, thereby allowing reactions to be shared between remote users. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-129800 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology of Patent Document 1 places a heavy processing load on the server, and constantly uploading the collected audio data to the server during live streaming requires a large amount of communication capacity, which can cause delays.
[0006] Therefore, the present disclosure proposes an information processing device, an information processing method, and a storage medium that can further reduce the load when generating viewer audio data. [Means for solving the problem]
[0007] According to the present disclosure, an information processing device is proposed that acquires audio metadata indicating information about the viewer's vocalizations in real time from one or more information processing terminals, and includes a control unit that controls the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata.
[0008] According to the present disclosure, an information processing method is proposed, which includes a processor acquiring audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controlling the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata.
[0009] According to the present disclosure, a storage medium is proposed that stores a program that causes a computer to function as a control unit that acquires audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controls the computer to generate viewer audio data for output using pre-prepared audio data based on the acquired audio metadata. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an overview of a voice data generation system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of a server and a viewer terminal according to the present embodiment. [Figure 3]FIG. 10 is a sequence diagram showing an example of the flow of a voice data generation process according to the present embodiment. [Figure 4] 10A and 10B are diagrams illustrating data transmission in the voice data generation process according to the present embodiment. [Figure 5] 10A and 10B are diagrams illustrating control for changing viewer voice data according to the scene of an event according to the present embodiment. [Figure 6] FIG. 10 is a sequence diagram showing an example of the flow of a voice data generation process using labeling information according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0012] The explanation will be given in the following order: 1. Overview of the audio data generation system according to one embodiment of the present disclosure 2.Configuration example 2-1. Example of Server 20 configuration 2-2. Example of configuration of viewer terminal 10 3. Operation processing 4. Specific Examples 4-1.Generating viewer voice data according to the number of people 4-2.Generating viewer voice data according to gender and emotion 4-3.Generating viewer voice data according to its characteristics 4-4.Generating viewer voice data according to the viewing environment 4-5.Generating viewer voice data according to the number of viewers in the same location 4-6.Generating viewer voice data according to virtual seating area 4-7.Generating viewer audio data depending on whether the sound pickup unit is enabled or disabled 4-8.Generating viewer voice data according to the event scene 4-9.Generating viewer voice data according to labeling 4-10. Combining audio data and audio metadata 4-11. Use of audio metadata in archive distribution 5. Supplementary Information
[0013] <<1. Overview of Audio Data Generation System According to One Embodiment of the Present Disclosure>> 1 is a diagram illustrating an overview of a voice data generation system according to an embodiment of the present disclosure. As shown in FIG. 1, the voice data generation system according to this embodiment includes an event venue device 30, a server 20, and a viewer terminal 10.
[0014] Event venue device 30 acquires video and audio from the venue where the event is being held and transmits them to server 20. Event venue device 30 may be made up of multiple devices. The event venue may be a facility with a stage and audience seats (such as an arena or concert venue), or may be a room used for recording (a recording studio).
[0015] Server 20 is an information processing device that controls the distribution of video and audio data of the event venue received from event venue device 30 to viewer terminal 10 in real time.
[0016] The viewer terminals 10 (10a to 10c, etc.) are information processing terminals used by viewers to view the event venue. The viewer terminals 10 can be realized, for example, by smartphones, tablet terminals, PCs (personal computers), HMDs (head mounted displays), projectors, television sets, game consoles, etc. The HMDs may have a non-transparent display unit that covers the entire field of view, or may have a transparent display unit. The viewer terminals 10 are connected to the server 20 for communication, and output the video and audio of the event venue received from the server 20.
[0017] (Identifying issues) As mentioned above, one of the major differences between an event experience via internet streaming and an in-venue event is the lack of a way to communicate audience reactions, such as cheers and applause, to the performers and other viewers. Audience reactions, such as cheers and applause, can motivate performers and further increase the excitement among the audience, making them an important element of an event. One possible solution would be to record the vocalizations of each viewer (remote user), send them to a server, process them, and add multiple audio data before distributing them to each remote user. However, this would impose a heavy processing load on the server. Furthermore, constantly uploading recorded audio data to the server during live streaming would require a large amount of communication bandwidth and could result in delays.
[0018] Therefore, in the information processing system according to the present disclosure, it is possible to further reduce the load when generating audio data for a viewer by using audio metadata.
[0019] Specifically, the viewer terminal 10 outputs video and audio from the event venue, while generating audio metadata indicating information about the audience's vocalizations and transmitting it to the server 20. The server 20 acquires audio metadata from one or more viewer terminals 10 in real time, and generates viewer audio data for output based on the acquired audio metadata using audio data prepared in advance. The viewer audio data can be considered the audio data of the entire audience.
[0020] For example, the server 20 counts the number of cheering viewers based on the audio metadata acquired from each viewer terminal 10, selects viewer audio data corresponding to that number from pre-prepared viewer audio data for each number of viewers, and sets the selected viewer audio data as the viewer audio data to be output. The server 20 then transmits the generated viewer audio data to the event venue device 30 and one or more viewer terminals 10. The event venue device 30 outputs the viewer audio data from speakers or the like installed at the event venue, enabling real-time feedback of viewer reactions to the performers. The viewer terminal 10 can provide the viewer with the reactions of other viewers by outputting the viewer audio data.
[0021] In this embodiment, the use of audio metadata reduces the communication load, and the use of pre-prepared audio data can also reduce the processing load on the server 20.
[0022] The outline of the voice data generation system according to an embodiment of the present disclosure has been described above. Next, the configuration of each device included in the voice data generation system according to this embodiment will be described with reference to the drawings.
[0023] <<2. Configuration Example>> 2 is a block diagram showing an example of the configuration of the server 20 and the viewer terminal 10 included in the audio data generation system according to this embodiment. The server 20 and the viewer terminal 10 are connected for communication via a network and can transmit and receive data. The configuration of each device will be described below.
[0024] <2-1. Example of Server 20 Configuration> As shown in FIG. 2, the server 20 includes a communication unit 210, a control unit 220, and a storage unit 230.
[0025] (Communication unit 210) The communication unit 210 transmits and receives data to and from external devices via wired or wireless communication. The communication unit 210 communicates with the viewer terminal 10 and the event venue device 30 using, for example, wired / wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), or a mobile communication network (LTE (Long Term Evolution), 4G (fourth generation mobile communication system), 5G (fifth generation mobile communication system)), etc.
[0026] (control unit 220) The control unit 220 functions as an arithmetic processing unit and a control device, and controls the overall operation of the server 20 in accordance with various programs. The control unit 220 is realized by electronic circuits such as a CPU (Central Processing Unit) or a microprocessor. The control unit 220 may also include a ROM (Read Only Memory) that stores the programs to be used, arithmetic parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.
[0027] The control unit 220 controls the transmission of the video and audio of the event venue received from the event venue device 30 to the viewer terminal 10. The control unit 220 may, for example, stream the video and audio of the event venue where the event is taking place in real time to one or more viewer terminals 10.
[0028] The control unit 220 according to this embodiment also functions as an audio metadata analysis unit 221 and a viewer audio data generation unit 222.
[0029] The audio metadata analysis unit 221 analyzes the audio metadata continuously transmitted from each viewer terminal 10. Specific examples of information contained in the audio metadata will be described later. The audio metadata analysis unit 221 analyzes the audio metadata acquired from each viewer terminal 10 and performs appropriate processing, such as counting the number of viewers cheering. The audio metadata analysis unit 221 outputs the analysis results to the viewer audio data generation unit 222.
[0030] The viewer voice data generation unit 222 generates viewer voice data for output based on the analysis results by the voice metadata analysis unit 221. At this time, the viewer voice data generation unit 222 generates the viewer voice data using voice data prepared in advance (for example, stored in the storage unit 230). The prepared voice data is, for example, cheers ("Wow," "Ahhh," "Whoa," etc.). Such cheers can be prepared for different numbers of people, for example. That is, cheers from 20 people, 50 people, 100 people, etc. are recorded in advance, and the recorded voice data is stored in the storage unit 230.
[0031] For example, the viewer voice data generation unit 222 generates viewer voice data by selecting viewer voice data corresponding to the number of viewers indicated in the analysis results (the number of viewers cheering) from viewer voice data prepared in advance for each number of viewers. Compared to performing audio processing on collected viewer voice data and synthesizing it, selecting viewer voice data from viewer voice data prepared in advance for each number of viewers can significantly reduce the processing load on the server 20. Note that the generation of viewer voice data described here is just one example. Variations in the method of generating viewer voice data will be described later.
[0032] The control unit 220 controls the transmission of the generated viewer voice data from the communication unit 210 to the viewer terminal 10 and the event venue device 30. Note that the control unit 220 may transmit to the viewer terminal 10 voice data that is a combination of the generated viewer voice data and voice data from the event venue.
[0033] The generation and transmission of viewer voice data described above can be performed continuously by the control unit 220. For example, the control unit 220 may generate and transmit the viewer voice data every 0.5 seconds.
[0034] (Storage unit 230) The storage unit 230 is realized by a ROM (Read Only Memory) that stores programs and calculation parameters used in the processing of the control unit 220, and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate. For example, in this embodiment, the storage unit 230 stores audio data used to generate viewer audio data.
[0035] Although the configuration of the server 20 has been specifically described above, the configuration of the server 20 according to the present disclosure is not limited to the example shown in Fig. 2. For example, the server 20 may be realized by a plurality of devices.
[0036] <2-2. Example of configuration of viewer terminal 10> As shown in FIG. 2, the viewer terminal 10 includes a communication unit 110, a control unit 120, a display unit 130, a sound collection unit 140, an audio output unit 150, and a storage unit 160.
[0037] (Communication unit 110) The communication unit 110 transmits and receives data to and from external devices via wired or wireless communication. The communication unit 110 is connected to and communicates with the server 20 using, for example, a wired / wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), a mobile communication network (LTE (Long Term Evolution), 4G (fourth generation mobile communication system), 5G (fifth generation mobile communication system)), or the like.
[0038] (control unit 120) The control unit 120 functions as a calculation processing unit and a control unit, and controls the overall operation of the viewer terminal 10 in accordance with various programs. The control unit 120 is realized by electronic circuits such as a CPU (Central Processing Unit) or a microprocessor. The control unit 120 may also include a ROM (Read Only Memory) that stores the programs to be used, calculation parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.
[0039] The control unit 120 controls the display of the video of the event venue received from the server 20 on the display unit 130, and controls the playback of the audio of the event venue and viewer audio data received from the server 20 from the audio output unit 150. From the server 20, for example, the video and audio of the event venue where the event is taking place are streamed in real time.
[0040] The control unit 120 according to this embodiment also functions as an audio metadata generation unit 121. The audio metadata generation unit 121 generates audio metadata indicating information about the vocalizations of viewers. For example, the audio metadata generation unit 121 generates the audio metadata based on collected sound data obtained by collecting the vocalizations of viewers by the sound collection unit 140. It is assumed that viewers will cheer while watching the broadcast of the event venue, and such cheers (vocalizations) are collected by the sound collection unit 140. The audio metadata generation unit 121 may also generate the audio metadata based on information set or measured in advance. The information about the vocalizations of viewers includes, for example, whether or not the viewers vocalized, the gender of the viewers who vocalized, and the emotions they expressed when they uttered the cheers (specific types of cheers). Specific details of the audio metadata will be described later. The audio metadata generation unit 121 continuously generates audio metadata and transmits it to the server 20 while the server 20 is live streaming the event venue (for example, streaming video and audio from the event venue). For example, the audio metadata generating unit 121 may generate audio metadata every 0.5 seconds and transmit it to the server 20.
[0041] (Display section 130) Display unit 130 has a function of displaying an image of the event venue according to instructions from control unit 120. For example, display unit 130 may be a display panel such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display.
[0042] (Sound pickup unit 140 and sound output unit 150) The sound collection unit 140 has a function of collecting the voice of the viewer (user) and outputs the collected voice data to the control unit 120.
[0043] The audio output unit 150 has a function of outputting (playing back) audio data in accordance with instructions from the control unit 120. The audio output unit 150 may be configured as, for example, a round speaker provided in the viewer terminal 10, headphones or earphones that communicate with the viewer terminal 10 via wired or wireless communication, or a bone conduction speaker.
[0044] (Storage unit 160) The storage unit 160 is realized by a ROM (Read Only Memory) that stores programs and calculation parameters used in the processing of the control unit 120, and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.
[0045] The configuration of the viewer terminal 10 has been specifically described above, but the configuration of the viewer terminal 10 according to the present disclosure is not limited to the example shown in Fig. 2. For example, at least one of the display unit 130, the sound collection unit 140, and the sound output unit 150 may be separate units.
[0046] <<3. Operation Processing>> Next, the flow of the voice data generation process according to this embodiment will be specifically described with reference to the drawings. Fig. 3 is a sequence diagram showing an example of the flow of the voice data generation process according to this embodiment.
[0047] First, as shown in FIG. 3, the viewer terminal 10 acquires collected sound data (input information) from the sound collection unit 140 (step S103).
[0048] Next, the viewer terminal 10 generates audio metadata based on the input information (sound pickup data) (step S106), and transmits the generated audio metadata to the server 20 (step S109).
[0049] Next, the server 20 acquires audio metadata from one or more viewer terminals 10 (step S112), and analyzes the audio metadata (step S115).
[0050] Next, the server 20 generates viewer voice data based on the analysis result (step S118). The viewer voice data can be said to be voice data of the entire audience.
[0051] The server 20 then transmits the viewer audio data together with the event venue audio data (received from the event venue device 30) to each viewer terminal 10 (step S121). Although only one viewer terminal 10 is shown in the example of FIG. 3, the server 20 transmits the viewer audio data to all viewer terminals 10 (all viewers) that are connected to it via communication. The viewer terminal 10 then plays back the event venue audio data and the audio data of all viewers (step S127).
[0052] Server 20 also transmits the viewer voice data to event venue device 30 (step S124). Event venue device 30 reproduces the voice data of all viewers on speakers or the like installed in the event venue (step S130).
[0053] An example of the flow of the audio data generation process according to this embodiment has been described above. Note that the operational process shown in Fig. 3 is just an example, and some of the processes may be performed in a different order or in parallel, or some of the processes may not be performed at all. For example, the process of transmitting viewer audio data to event venue device 30 may not necessarily be performed.
[0054] Furthermore, the above-described processing can be performed continuously while the server 20 is live streaming the event (real-time streaming of video and audio from the event venue). FIG. 4 is a diagram illustrating data transmission in the audio data generation processing according to this embodiment. As shown in FIG. 4, the server 20 generates viewer audio data (audio data of all viewers) at regular intervals (e.g., every 0.5 seconds) based on the audio metadata received from the viewer terminals 10 up to that point, and transmits the data to each viewer terminal 10 and the event venue device 30. Note that the viewer terminals 10 also include information processing terminals corresponding to the public viewing venue. When the event is being streamed to individuals and spectators at the public viewing venue, it becomes possible for the individual's cheers and the cheers of the public viewing spectators to be shared with each other.
[0055] For example, the audio metadata includes the presence or absence of vocalizations (cheers), and the server 20 selects and transmits viewer audio data for each number of viewers according to the number of viewers who cheered. In this case, for example, in the audio data generation process at a certain timing, 50 viewers have vocalizations, so the cheers of 50 viewers are selected and transmitted, and at the next timing, 100 viewers have vocalizations, so the cheers of 100 viewers are selected and transmitted. This allows viewers and performers to share the fact that the audience is gradually getting more excited (the cheers are increasing) in real time.
[0056] <<4. Specific Examples>> Next, the generation of viewer voice data will be described using a specific example.
[0057] <4-1. Generating viewer voice data according to the number of viewers> For example, the audio metadata includes whether or not there is vocalization, and the viewer audio data generator 222 generates viewer audio data according to the number of viewers.
[0058] The audio metadata generation unit 121 of the viewer terminal 10 analyzes the audio data collected by the audio collection unit 140 (audio recognition), determines whether the viewer has spoken, and generates audio metadata including information indicating the presence or absence of vocalization. An example of the generated data may be a "speaking_flag" that is assigned a "1" if vocalization is present and a "2" if vocalization is absent. The viewer terminal 10 determines the presence or absence of vocalization, for example, every second, and generates and transmits the audio metadata.
[0059] The server 20 prepares audio data for each number of viewers in advance. The audio data is, for example, the sound source of cheers and cheers. The audio metadata analysis unit 221 of the server 20 counts the number of viewers who are speaking from information indicating the presence or absence of vocal sounds contained in the audio metadata transmitted from one or more viewer terminals 10. The viewer audio data generation unit 222 then selects audio data that is closest to the counted number from the audio data prepared in advance for each number of viewers, and sets this as the viewer audio data. The server 20 transmits the viewer audio data generated in this way to each viewer terminal 10 and the event venue device 30.
[0060] This means that even if a live stream is being streamed without an audience, it is possible to experience the same live stream as if other viewers (audiences) were cheering. Also, even if a viewer utters inappropriate words, the system only uses information on whether or not they were uttered, so there is no problem with other viewers hearing the inappropriate words as they are.
[0061] Although the example described here illustrates that audio metadata including information indicating that vocalization is absent is transmitted even when there is no vocalization, the present embodiment is not limited to this. The viewer terminal 10 may transmit audio metadata including information indicating that vocalization is present only when vocalization is present.
[0062] <4-2.Generating viewer voice data according to gender and emotion> For example, the audio metadata includes at least one of the gender of the viewer who made the vocalization and the emotion determined from the vocalization, and the viewer audio data generation unit 222 generates viewer audio data according to the gender and emotion.
[0063] The audio metadata generation unit 121 of the viewer terminal 10 analyzes the audio data collected by the audio collection unit 140 (audio recognition), determines whether the voice is female or male, and generates audio metadata including information indicating the gender. If the viewer's gender has been set in advance, that information may be used. The audio metadata generation unit 121 also analyzes the audio data collected by the audio collection unit 140 (audio recognition), determines the emotion evoked by the vocalization, and generates audio metadata including information indicating the emotion. For example, there are various types of cheers, such as voices of disappointment, joy, excitement, impatience, surprise, and screams, which can be considered to evoke emotions. The audio metadata generation unit 121 may also include information indicating the type of cheer as the emotion information. If the analysis of the audio data determines that the viewer is not making any sound, the audio metadata generation unit 121 may also include information indicating that the viewer is not making any sound.
[0064] In the example of generated data, for example, "emotion_type" may be assigned values such as "no vocalization: 0", "dejected: 1", "happy (excited): 2", and "scream: 3".
[0065] The server 20 prepares in advance voice data by gender (such as a sound source of cheers for women only, a sound source of cheers for men only, etc.) and voice data by emotion (such as a sound source of disappointment, a sound source of joy, a sound source of screams, etc.) This voice data may be voice data for one person, or may be prepared for a certain number of people (for example, 1000 people), or may further be prepared for several voice data for different numbers of people.
[0066] The viewer voice data generation unit 222 of the server 20 generates viewer voice data by selecting corresponding voice data from pre-prepared voice data for each gender or voice data for each emotion based on information indicating gender and emotion included in the voice metadata transmitted from one or more viewer terminals 10. More specifically, the viewer voice data generation unit 222 synthesizes voice for each viewer's voice metadata and combines them to generate one voice data.
[0067] Alternatively, for example, the audio metadata analysis unit 221 counts the number of people for each emotion, and the viewer audio data generation unit 222 generates audio data for 50 people using the audio data for disappointment (or selects audio data for a similar number of people) if 50 people express the emotion of disappointment, and generates audio data for 100 people using the audio data for joy (or selects audio data for a similar number of people) if 100 people express the emotion of joy, and combines these to generate final viewer audio data. Furthermore, the final viewer audio data may be generated by adjusting the volume, etc., of emotion-specific audio data for a certain number of people prepared in advance, depending on the proportion of each emotion. The same can be done for gender.
[0068] In this way, by generating viewer voice data from voice data classified by gender and emotion, it is possible to obtain reactions from viewers that are specific to women or men when a performer calls out at a live music concert. As another example, in a soccer goal scene, joy and disappointment may occur simultaneously when the team you are rooting for scores and the opposing team scores, and in such a case, viewer voice data containing both voices can be generated.
[0069] As explained above, by reflecting the gender and emotions of the viewer in the generation of viewer voice data, the viewer reactions shared among viewers and performers can be made closer to the actual reactions.
[0070] <4-3. Generating viewer voice data according to characteristics> To make the generated viewer voice data even closer to the actual reaction, for example, the characteristics of the viewer's voice (gender, pitch (high, low), thickness (thin, thick), etc.) may be used.
[0071] The viewer terminal 10 analyzes the characteristics of the viewer's voice in advance and generates characteristic information. The audio metadata includes information indicating the viewer's gender and voice characteristics. This information is also called voice generation parameters.
[0072] The viewer voice data generation unit 222 of the server 20 generates viewer voice data by appropriately adjusting pre-prepared default voice data based on information indicating the voice characteristics for each piece of voice metadata transmitted from one or more viewer terminals 10, and then combines these to generate one piece of voice data. This makes it possible to generate cheers that are closer to the real thing, reflecting the characteristics of the viewer's voices, rather than simply using the cheers that were originally prepared.
[0073] (Variation 1) The audio metadata described above may include, in addition to gender and voice characteristics, information on emotions determined from vocal sounds (types of cheers). This allows audio data to be generated for each viewer according to their emotions.
[0074] The audio metadata may also include information about the volume of the viewer's voice, which allows the audio data generated for each viewer to correspond to the volume of the viewer's actual voice.
[0075] (Variation 2) The characteristics of the viewer's voice as described above may be set by the viewer at will. This allows the viewer to cheer (shout) in a voice tone different from their actual voice tone. For example, a man may use a female voice. Also, the viewer may be able to select from voice generation parameters (for example, parameters for generating the voice of a celebrity) prepared in advance by the distribution provider. Furthermore, the voice generation parameters prepared by the distribution provider may be sold separately or included only with tickets to a specific event. This allows the viewer to use them as a revenue item for the distribution provider's events.
[0076] (Variation 3) Furthermore, the variations in voice generation parameters to be handled may be limited. The voice metadata generation unit 121 of the viewer terminal 10 selects the characteristics of the viewer's voice from pre-prepared voice generation parameters and includes the selected characteristics in the voice metadata. This reduces the processing load on the server 20 for generating voice data for each viewer. For example, the voice metadata analysis unit 221 of the server 20 counts the number of viewers who have spoken for each voice generation parameter, and the viewer voice data generation unit 222 generates viewer voice data using the pre-prepared voice generation parameters and voice data for each number of viewers.
[0077] In addition, both the process of selecting from pre-prepared voice generation parameters and the process of using the characteristics of the viewer's voice may be performed. In this case, for example, the function of reflecting the characteristics of the viewer's voice may be sold only to specific viewers as an income item for an event on the distribution provider's side.
[0078] <4-4. Generating viewer audio data according to the viewing environment> The audio metadata may include the presence or absence of vocalizations and information about the volume of the viewer's voice. In this case, the server 20 can generate viewer audio data taking into account information about the volume of the viewer's actual voice. However, depending on the viewing environment, some viewers may not be able to speak loudly, resulting in a moderate volume. If there are many viewers in viewing environments where they cannot speak loudly, the generated viewer audio data will also have a moderate volume. Therefore, the viewer terminal 10 may measure the viewer's maximum volume (the maximum volume of the voice that the viewer can speak) in advance and include it in the audio metadata. The audio metadata generation unit 121 of the viewer terminal 10 transmits, for example, information indicating the presence or absence of vocalizations, as well as information indicating the actual volume of the vocalizations and information indicating the previously measured maximum volume, to the server 20.
[0079] Although the case of actual measurement has been described here, the present embodiment is not limited to this, and for example, the maximum volume value may be set by the viewer himself / herself. Also, the audio metadata generation unit 121 may obtain in advance a value measured at a specific timing when the voices are expected to be the loudest at an event (for example, when the artist appears at a live music concert), and use this value as the maximum volume value.
[0080] When generating audio data for each viewer based on each piece of audio metadata, the viewer audio data generation unit 222 of the server 20 may take the maximum volume into consideration and set the volume to be louder than the volume of the voice actually uttered by the user. Also, the viewer audio data generation unit 222 may set a maximum volume setting value A for audio data that can be generated by the viewer audio data generation unit 222, and adjust the maximum volume value of the audio metadata as needed so that it is equal to the maximum volume setting value A.
[0081] <4-5. Generating viewer voice data according to the number of viewers in the same location> In each of the above-described specific examples, it was assumed that the viewer terminal 10 generates audio metadata for a single viewer. However, there are cases where several people watch together, such as family or friends. In this case, for example, the viewer terminal 10 can use voice recognition or a camera to recognize the viewers and then generate audio metadata for the number of people. Also, a field indicating the number of people can be added to the audio metadata, and the rest can be summarized as information for one person.
[0082] Furthermore, when gender information is included in the audio metadata, information indicating that both men and women are included or information indicating the ratio of men to women may be used.
[0083] When one piece of audio metadata contains information indicating the number of viewers, the audio metadata analysis unit 221 of the server 20 performs processing to count the number of viewers as that number, rather than as 1.
[0084] <4-6. Generation of viewer voice data according to virtual seating area> The audio metadata may include information indicating a virtual seating area (viewing position) at an event venue. In an actual live music event, the viewing position is one element of enjoying the event, and performers may also request reactions linked to the viewing position (for example, a performer calling on the audience in the second floor seats to cheer). The virtual seating area may be set in advance for each viewer, or the viewer may select an area of their choice. The viewer terminal 10 includes, for example, information indicating the presence or absence of vocal sounds and information indicating the virtual seating area in the audio metadata.
[0085] The audio metadata analysis unit 221 of the server 20 counts the number of viewers who spoke for each virtual seating area, and the viewer audio data generation unit 222 selects audio data according to the number of viewers for each virtual seating area and generates viewer audio data. The control unit 220 of the server 20 then associates the generated viewer audio data with information about the virtual seating area and transmits it to the event venue device 30. The event venue device 30 controls multiple speakers installed in the audience seats at the event venue to play back the viewer audio data for the virtual seating area corresponding to the position of each speaker. This allows the performers to understand the cheers of the audience at each position.
[0086] The viewer voice data generating unit 222 may generate voice data for each viewer based on each piece of voice metadata, and then compile this data for each virtual viewing area to generate viewer voice data.
[0087] Furthermore, the server 20 may transmit viewer voice data associated with information about virtual seating areas to each viewer terminal 10. When playing back viewer voice data, each viewer terminal 10 may perform processing to localize the sound source to a position corresponding to the virtual seating area based on the virtual seating area information of each viewer voice data. This allows viewers to experience the same atmosphere as when they are seated in an actual venue.
[0088] <4-7. Generation of viewer audio data depending on whether the sound pickup unit is enabled or disabled> In each of the above-described specific examples, audio metadata is generated based on input information (collected audio data) from the audio collection unit 140. However, some viewer terminals 10 may not be provided with or connected to the audio collection unit 140, and some viewers may have no choice but to watch quietly due to their environment. In such cases, the viewer audio data generated by the server 20 may represent cheers from a smaller number of viewers than the actual number of viewers.
[0089] Therefore, information about the sound collection unit 140 is included in the audio metadata. For example, information about the validity (ON / OFF) of the sound collection unit 140 is included. As a result, the viewer audio data generation unit 222 of the server 20 generates viewer audio data by assuming that the same percentage of speakers is present among viewers whose sound collection unit 140 is OFF (invalid, unavailable), based on the percentage of speakers among viewers whose sound collection unit 140 is ON (valid, available). This makes it possible to generate viewer audio data taking into account the number of users who cannot use the sound collection unit 140 (users who cannot speak). Note that the information considered regarding viewers whose sound collection unit 140 is OFF is not limited to the number of speakers. For example, the percentage of genders, types of cheers, voice volume, voice characteristics, etc. may also be appropriately applied to audio metadata of viewers whose sound collection unit 140 is ON.
[0090] Furthermore, in consideration of the possibility that some viewers may be forced to watch quietly due to the environment, the viewer terminal 10 may analyze the viewers' movements captured by a camera and include the analysis results in the audio metadata. For example, it is conceivable that viewers may express their excitement by quietly clapping (such as clapping without actually hitting their hands) or waving their hands because they are in an environment where they cannot make noise. The viewer terminal 10 may grasp such viewers' movements through image analysis, determine whether there is vocalization, the type of cheer is "joyful," etc., and generate audio metadata.
[0091] <4-8. Generation of viewer voice data according to the event scene> For example, in the case of a live music concert, it may be preferable for the audience to cheer quietly during the middle of a song, so that the audience can concentrate on listening to the music as much as possible. On the other hand, the intervals between songs can be a time for the performers and the audience to check the level of excitement. Similarly, in sporting events, there are some sports where players are expected to be quiet while playing, while other sports require players to clap their hands. Thus, the desired volume of sound varies depending on the event, and there are even cases where clapping is preferred over cheering.
[0092] Therefore, the viewer voice data generating unit 222 of the server 20 may change the volume of the viewer voice data to be generated and the type of viewer voice data (cheers, clapping, etc.) depending on the scene of the event. The type of voice data to be generated for each scene may be controlled in real time by the server, may be changed at a preset time, or may be changed depending on the voice of the performers. FIG. 5 is a diagram illustrating the control of changing viewer voice data depending on the scene of the event. As shown in FIG. 5, for example, the type may be changed to "Type: Clapping" and "Volume: Low" during a performance at the event, and to "Type: Cheers" and "Volume: High" during a talk. In this way, the event can be produced using viewer voice data.
[0093] <4-9. Generating viewer voice data according to labeling> In this embodiment, by setting labeling information of the classification to which the viewer belongs, it becomes possible to provide viewer voice data customized for each viewer. In other words, it is possible to provide, for each viewer, viewer voice data corresponding to labeling information of the same classification as the viewer, with emphasis.
[0094] For example, the audio metadata generation unit 121 of the viewer terminal 10 adds information about the team a viewer supports to the audio metadata as labeling information. The viewer audio data generation unit 222 of the server 20 then generates viewer audio data for each piece of labeling information. The control unit 220 of the server 20 then transmits the viewer audio data with the same labeling information as the viewer to the viewer terminal 10 of that viewer. This allows the viewer to primarily hear the cheers of fans of the same soccer team, providing the experience of watching the match among the supporters of the team they support. Not only for soccer teams, but also for live music performances, for example, adding information about a featured artist to the audio metadata as labeling information can provide the viewer with viewer audio data that emphasizes the cheers for that artist. The viewer audio data generation unit 222 may also generate overall viewer audio data for each piece of labeling information, emphasizing (increasing the volume of) the viewer audio data of the corresponding labeling information.
[0095] The voice data generation process using labeling information will be described below with reference to Fig. 6. Fig. 6 is a sequence diagram showing an example of the flow of the voice data generation process using labeling information according to this embodiment.
[0096] 6, first, the viewer or user 10 sets labeling information for viewing (step S203). The labeling information may be set based on a selection by the viewer.
[0097] Next, the viewer terminal 10 acquires collected sound data (input information) from the sound collection unit 140 (step S206).
[0098] Next, the viewer terminal 10 generates audio metadata based on the input information (sound pickup data), and further includes labeling information in the generated audio metadata (step S209), and transmits the audio metadata to the server 20 (step S212).
[0099] Next, in steps S215 to S221, the same processes as those shown in steps S112 to S118 in Fig. 3 are performed. That is, the control unit 220 of the server 20 generates viewer audio data based on the audio metadata acquired from one or more viewer terminals 10.
[0100] Next, the control unit 220 also generates viewer audio data using only the audio metadata of the same labeling information (step S224).
[0101] The server 20 then transmits viewer voice data based on the same labeling information as the labeling information of each viewer to each viewer terminal 10 together with the event venue voice data (received from the event venue device 30) (step S227). The server 20 may transmit only the viewer voice data based on the same labeling information as the viewer's labeling information to the viewer terminal 10, or may generate all viewer voice data in which the viewer voice data based on the same labeling information is emphasized and transmit this to the viewer terminal 10. The viewer terminal 10 plays back the event venue voice data and the viewer voice data (step S233).
[0102] Server 20 also transmits viewer voice data to event venue device 30 (step S230). This viewer voice data is the voice data generated in step S221. Event venue device 30 plays back the voice data of all viewers on speakers or the like installed in the event venue (step S236).
[0103] The above describes the process of generating audio data using labeling information. Note that the operational process shown in Fig. 6 is an example, and the present embodiment is not limited to this.
[0104] <4-10. Combining audio data and audio metadata> The event distribution of this embodiment is not limited to distribution to individuals, and may also be distributed to venues with thousands to tens of thousands of people, such as public viewing. Audio data from the public viewing venue may be collected at the venue and sent directly to server 20, where it may be combined with audio data from other individual viewers (audio data generated based on audio metadata). It is expected that there will be several public viewing venues, and audio from thousands to tens of thousands of people can be combined into one audio data, so the communication capacity and processing load are not as large as when audio data from thousands to tens of thousands of people is sent and processed individually.
[0105] In this way, in this embodiment, it is possible to use both audio data and audio information metadata.
[0106] Also, only specific individual viewers (for example, viewers who have purchased premium tickets) may be allowed to transmit audio data of their vocalizations to the server 20. By adjusting the number of viewers to whom audio data can be transmitted in advance, taking into account the processing load of the server 20, it is possible to provide various services to viewers without causing significant delays.
[0107] <4-11. Use of audio metadata in archive distribution> Each of the above-mentioned specific examples assumes that an event is distributed in real time, that is, so-called live distribution, but this embodiment is not limited to this, and it is also assumed that such event distribution will be archived and distributed at a later date.
[0108] In this case, the control unit 220 of the server 20 may also store audio metadata acquired from each viewer terminal 10 during live distribution, and during archive distribution, generate and distribute viewer audio data using audio metadata that was not used during live distribution. The audio data may include the various information described above, such as the presence or absence of vocalizations, gender, emotion, voice characteristics, voice volume, maximum volume value, number of people, virtual seating area, effectiveness of the sound pickup unit 140, and labeling information. Of these, during live distribution, viewer audio data may be generated and distributed using at least a portion of this information (for example, only the presence or absence of vocalizations) in consideration of processing load, etc., and during archive distribution, viewer audio data may be generated and distributed using various other information as appropriate.
[0109] <<5. Supplementary Information>> Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the present technology is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical ideas described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0110] For example, the event venue audio data and viewer audio data transmitted from the server 20 to the viewer terminal 10 are generated as separate sound sources, and in live streaming or archive streaming, the viewer can optionally turn off one of the audio data and play it back.
[0111] The above-mentioned specific examples may be combined as appropriate. The audio metadata may include at least one of the above-mentioned information such as the presence or absence of vocalization, gender, emotion, voice quality, voice volume, maximum volume value, number of people, virtual seating area, validity of the sound pickup unit 140, and labeling information.
[0112] Other information included in the audio metadata may also include the duration of the vocalization, such as whether the vocalization was momentary or lasted for a certain period of time.
[0113] It is also possible to create one or more computer programs for causing the hardware, such as the CPU, ROM, and RAM, built into the server 20 and the viewer terminal 10 to perform the functions of the server 20 and the viewer terminal 10. A computer-readable storage medium storing the one or more computer programs is also provided.
[0114] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0115] The present technology can also be configured as follows. (1) An information processing device comprising a control unit that acquires audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controls the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata. (2) The information processing device described in (1), wherein the audio metadata is generated based on the analysis results of audio data collected by an audio collection unit that collects the vocalizations of the viewers when data of an event being held in real time is being distributed. (3) the audio metadata includes information indicating the presence or absence of the vocalization; The information processing device described in (2) above, wherein the control unit counts the number of people speaking based on information indicating the presence or absence of the vocalization sound, selects audio data closest to the counted number from among pre-prepared audio data for each number of people, and generates the viewer audio data. (4) the audio metadata includes information indicating the gender of the viewer who produced the vocalization; The information processing device described in (2) above, wherein the control unit selects audio data corresponding to the gender from among pre-prepared gender-specific audio data based on information indicating the gender of the viewer who made the vocalization, and generates the viewer audio data. (5) the audio metadata includes information indicating an emotion determined by an analysis result of the vocalization; The information processing device according to any one of (2) to (4), wherein the control unit selects, based on the information indicating the emotion, audio data corresponding to the emotion from among pre-prepared emotion-specific audio data, and generates the viewer audio data. (6) the audio metadata includes information indicating the characteristics of the utterance, the information being generated as a result of analyzing the utterance; The information processing device according to any one of (2) to (5), wherein the control unit reflects the property in pre-prepared audio data to generate the viewer audio data. (7) the audio metadata includes information indicating a property arbitrarily set by a viewer as a property of the vocalization; The information processing device according to any one of (2) to (5), wherein the control unit reflects the property in pre-prepared audio data to generate the viewer audio data. (8) the audio metadata includes information indicating a property selected from a variety of properties prepared in advance as the property of the utterance; The information processing device according to any one of (2) to (5), wherein the control unit selects audio data that reflects the property from among audio data prepared in advance, and generates the viewer audio data. (9) the audio metadata further includes information indicating the loudness of the utterance determined by an analysis result of the utterance, The information processing device according to any one of (2) to (8), wherein the control unit generates the viewer voice data by further reflecting the volume of the voice of each viewer. (10) The audio metadata further includes information on the maximum volume value of the listener who emitted the utterance; The information processing device according to (9), wherein the control unit generates the viewer voice data by further reflecting a maximum volume value of each viewer. (11) The audio metadata further includes information on the maximum volume value of the listener who emitted the utterance; The information processing device according to (9), wherein the control unit adjusts the maximum volume value to the same volume as a preset maximum volume setting value, and generates and outputs the viewer voice data. (12) the audio metadata includes information about the number of co-located listeners who made the vocalization; The information processing device according to (2), wherein the control unit selects audio data that is close to the number of people from among audio data prepared in advance for each number of people, and generates the viewer audio data. (13) the audio metadata further includes information indicating a virtual seating area of a viewer from whom the vocalization was made; The information processing device according to any one of (2) to (11), wherein the control unit further generates the viewer voice data for each virtual seat area of each viewer. (14) the audio metadata further includes information indicating whether a sound pickup unit that picks up the vocalization sound is valid; The information processing device described in any one of (2) to (13), wherein the control unit applies the ratio of the number of speakers in each viewer for whom the sound collection unit is enabled to the ratio of the assumed number of speakers in each viewer for whom the sound collection unit is disabled, counts the number of speakers, selects audio data from pre-prepared audio data by number of speakers that is closest to the counted number of speakers, and generates the viewer audio data. (15) The audio metadata further includes labeling information of a category to which the viewer belongs; The information processing device according to any one of (2) to (14), wherein the control unit generates the viewer voice data for each classification and outputs the viewer voice data corresponding to the viewer classification to the viewer's information processing terminal. (16) The information processing device according to any one of (1) to (15), wherein the control unit changes at least one of the type and volume of the generated viewer voice data according to the scene of the event to be distributed to each viewer. (17) The information processing device according to any one of (1) to (16), wherein the control unit outputs the generated viewer voice data to the information processing terminal and an event venue device. (18) The information processing device described in any one of (1) to (17), wherein the control unit synthesizes audio data acquired from the public viewing venue with viewer audio data generated based on the audio metadata and outputs the synthesized data to the information processing terminal and the event venue device. (19) The processor: An information processing method including acquiring audio metadata indicating information about the viewer's vocalizations in real time from one or more information processing terminals, and controlling the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata. (20) Computer, A storage medium storing a program that functions as a control unit that acquires audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controls the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata. [Explanation of symbols]
[0116] 10 Viewer terminals 110 Communications Department 120 control section 121 Audio Metadata Generation Unit 130 Display section 140 Sound pickup unit 150 Audio output section 160 Storage section 20 Management Server 210 Communications Department 220 Control Unit 221 Audio Metadata Analysis Unit 222 Viewer voice data generation unit 230 Storage section
Claims
1. An information processing device comprising a control unit that acquires audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controls the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata.
2. The information processing device according to claim 1 , wherein the audio metadata is generated based on an analysis result of audio data collected by a sound collection unit that collects vocalizations of the viewers while data of an event being held in real time is being distributed.
3. the audio metadata includes information indicating the presence or absence of the vocalization; 3. The information processing device according to claim 2, wherein the control unit counts the number of people speaking based on information indicating the presence or absence of the vocalization sound, selects audio data closest to the counted number from among pre-prepared audio data for each number of people, and generates the viewer audio data.
4. the audio metadata includes information indicating the gender of the viewer who produced the vocalization; 3. The information processing device according to claim 2, wherein the control unit selects voice data corresponding to the gender from among gender-specific voice data prepared in advance based on information indicating the gender of the viewer who made the vocalization, and generates the viewer voice data.
5. the audio metadata includes information indicating an emotion determined by an analysis result of the vocalization; The information processing device according to claim 2 , wherein the control unit selects, based on the information indicating the emotion, voice data corresponding to the emotion from among voice data classified by emotion that is prepared in advance, and generates the viewer voice data.
6. the audio metadata includes information indicating the characteristics of the utterance, the information being generated as a result of analyzing the utterance; The information processing device according to claim 2 , wherein the control unit reflects the property in pre-prepared audio data to generate the viewer audio data.
7. the audio metadata includes information indicating a property arbitrarily set by a viewer as a property of the vocalization; The information processing device according to claim 2 , wherein the control unit reflects the property in pre-prepared audio data to generate the viewer audio data.
8. the audio metadata includes information indicating a property selected from a variety of properties prepared in advance as the property of the utterance; The information processing device according to claim 2 , wherein the control unit selects audio data that reflects the property from among audio data prepared in advance, and generates the viewer audio data.
9. the audio metadata further includes information indicating the loudness of the utterance determined by an analysis result of the utterance, The information processing device according to claim 2 , wherein the control unit generates the viewer voice data by further reflecting the volume of the voice of each viewer.
10. The audio metadata further includes information on the maximum volume value of the listener who emitted the utterance, The information processing device according to claim 9 , wherein the control unit generates the viewer voice data by further reflecting a maximum volume value of each viewer.
11. The audio metadata further includes information on the maximum volume value of the listener who emitted the utterance; The information processing device according to claim 9 , wherein the control unit adjusts the maximum volume value to the same volume as a preset maximum volume setting value, and generates and outputs the viewer voice data.
12. the audio metadata includes information about the number of co-located listeners who made the vocalization; The information processing device according to claim 2 , wherein the control unit selects audio data that is closest to the number of people from among audio data prepared in advance for each number of people, and generates the viewer audio data.
13. the audio metadata further includes information indicating a virtual seating area of a viewer from whom the vocalization was made; The information processing device according to claim 2 , wherein the control unit further generates the viewer voice data for each virtual seat area of each viewer.
14. the audio metadata further includes information indicating whether a sound pickup unit that picks up the vocalization sound is valid; 3. The information processing device of claim 2, wherein the control unit counts the number of speakers after applying the ratio of the number of speakers in each viewer for whom the sound collection unit is enabled to the ratio of the assumed number of speakers in each viewer for whom the sound collection unit is disabled, selects audio data from pre-prepared audio data by number of speakers that is closest to the counted number of speakers, and generates the viewer audio data.
15. The audio metadata further includes labeling information of a category to which the viewer belongs; The information processing device according to claim 2 , wherein the control unit generates the viewer voice data for each classification and outputs the viewer voice data corresponding to the viewer classification to the information processing terminal of the viewer.
16. The information processing device according to claim 1 , wherein the control unit changes at least one of the type and volume of the generated viewer voice data in accordance with a scene of an event distributed to each of the viewers.
17. The information processing device according to claim 1 , wherein the control unit outputs the generated viewer voice data to the information processing terminal and an event venue device.
18. The information processing device according to claim 1 , wherein the control unit combines the audio data acquired from the public viewing venue with viewer audio data generated based on the audio metadata, and outputs the combined data to the information processing terminal and the event venue device.
19. The processor: An information processing method including acquiring audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time, and controlling the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata.
20. Computer, A storage medium storing a program that functions as a control unit that acquires audio metadata indicating information about the viewer's vocalizations from one or more information processing terminals in real time and controls the generation of viewer audio data for output using pre-prepared audio data based on the acquired audio metadata.
Citation Information
Patent Citations
Display apparatus and operating method of same
CN111078902A
Audio and video processing method and device and electronic equipment
CN113301359A
Contents transmission apparatus and contents receiving apparatus
JP2005159592A
Video and audio synthesizing unit, and video viewing system of shared remote experience type
JP2007036685A
Information processor, content processing method and program
JP2010232860A