Content distribution system and program

The content distribution system allows voice-based participation actions, adjusting voice content, timing, tone, and volume to match the video content's situation, enhancing viewer engagement and production quality.

JP2025113355AActive Publication Date: 2025-08-01BANDAI NAMCO ENTERTAINMENT INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025083748
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-01
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

Conventional content distribution systems lack the ability for viewers to participate through voice-based actions, limiting the auditory engagement and interaction with the content.

Method used

A content distribution system that includes a sound processing unit to change the content, output timing, tone color, and volume of viewer voice participation actions based on the situation of the video content, allowing for voice donations and enhancing auditory engagement.

Benefits of technology

Enables viewers to participate through voice donations, creating an auditory effect that matches the content's situation, improving viewer engagement and unity, and enhancing the production quality of live streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113355000001_ABST
    Figure 2025113355000001_ABST
Patent Text Reader

Abstract

To provide a content distribution system or the like capable of realizing distribution of viewing contents in which viewers can participate by audio.SOLUTION: A content delivery system includes: a distribution processing unit that performs distribution processing for viewers of viewer terminals to view viewing contents; a reception unit that receives participation actions by a viewer setting or system setting audio as participation actions of the viewers for the viewing contents; and an audio processing unit that at least one of change processing of an audio content, change processing of audio output timing, change processing of audio tone, and control processing of audio volume in participation action on the basis of a status of the viewing content at a timing of the viewer's participation action.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a content distribution system, a program, and the like.

Background Art

[0002] In recent years, content distribution systems that distribute content via networks such as the Internet have become popular. As a conventional technology of such a content distribution system, for example, a system disclosed in Patent Document 1 is known. Patent Document 1 discloses a GUI (Graphical User Interface) for making a donation, which is an action of a viewer participating in viewing content, to a distributor of the viewing content. Patent Document 2 discloses a music performance game machine in which a player enjoys performance operations according to music. Patent Document 2 discloses a method of changing the input sound of an operation unit according to the progress of a piece of music.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional content distribution systems, donations, which are actions of viewers participating, are donations by comments and stamps, and no proposals have been made regarding actions of viewers participating by voice.

[0005] According to some aspects of the present embodiment, it is possible to provide a content distribution system, a program, and the like that can realize distribution of viewing content that enables viewers to participate by voice.

Means for Solving the Problems

[0006] One aspect of the present disclosure relates to a content distribution system including a distribution processing unit that performs distribution processing for a viewer of a viewer terminal to view video content, a reception unit that receives the participation action of the viewer for the video content as voice of viewer settings or system settings, and a sound processing unit that performs at least one of a change process of the content of the voice in the participation action, a change process of the output timing of the voice, a change process of the tone color of the voice, and a control process of the volume of the voice based on the situation of the video content at the timing of the participation action of the viewer. Another aspect of the present disclosure relates to a program that causes a computer to function as each of the above units, or a computer-readable information storage medium storing the program.

[0007] According to one aspect of the present disclosure, since a participation action by voice for video content is received, an auditory effect or the like for the video content becomes possible. Also, based on the situation of the video content at the timing of the participation action, a change process of the content, output timing, or tone color of the voice is performed, or a control process of the volume is performed. Therefore, it is possible to provide a content distribution system or the like that can realize distribution of video content in which a viewer's participation action by voice with content, output timing, tone color, or volume according to the situation of the video content is possible.

[0008] In another aspect of the present disclosure, the sound processing unit may change the voice to a voice with content, tone color, or volume according to the melody or genre of the music of the video content.

[0009] In this way, it becomes possible to output the voice in the participation action with content, tone color, or volume that matches the melody or genre of the music of the video content.

[0010] In another aspect of the present disclosure, the sound processing unit may determine the set timing set for the video content and the timing of the participation action of the viewer, and perform a change process to bring the output timing of the voice closer to the set timing.

[0011] In this way, the voice in the participation action can be output at a timing corresponding to the setting timing of the viewing content.

[0012] Also, in one aspect of the present disclosure, the sound processing unit determines the timing of the participation action of the viewer and the timing of the participation actions of other viewers, and outputs the voice at a representative timing set based on the timing of the participation action of the viewer and the timing of the participation actions of the other viewers. The voice output timing may be changed.

[0013] In this way, the voice in the participation action of the viewer and the voice in the participation action of other viewers can be aggregated and output at the same representative timing.

[0014] Also, in one aspect of the present disclosure, when the voice in the participation action of the viewer and the voice in the participation action of other viewers are output at the viewer terminal, the sound processing unit performs a volume control process of changing at least one of the volume of the voice and the volume of the voice of the other viewers.

[0015] In this way, it becomes possible to perform volume control processing so that the ratio of the volume of the voice in the participation action of the viewer and the volume of the voice in the participation action of other viewers becomes an appropriate output ratio.

[0016] Also, in one aspect of the present disclosure, the sound processing unit determines the setting timing set in the viewing content and the timing of the participation action of the viewer, and performs volume control processing of the voice according to the degree of coincidence between the setting timing and the timing of the participation action.

[0017] In this way, it becomes possible to realize volume control of the voice according to the degree of coincidence between the setting timing of the viewing content and the timing of the participation action by the voice of the viewer.

[0018] In another aspect of the present disclosure, the sound processing unit may perform volume control processing of the voice according to the number of other viewers who have taken the participation action at a timing corresponding to the timing of the participation action of the viewer.

[0019] In this way, it becomes possible to realize volume control of the voice according to the number of participants in the voice-based participation action for the viewing content.

[0020] In another aspect of the present disclosure, the reception unit may receive, as the voice in the participation action, a voice generated based on a phrase selected or input by the viewer and a tone color of the system setting.

[0021] In this way, the phrase selected or input by the viewer can be output as the voice in the participation action with the tone color set in the system.

[0022] In another aspect of the present disclosure, the reception unit may receive, as the voice in the participation action, a voice generated based on a phrase selected or input by the viewer and the tone color of the input voice of the viewer.

[0023] In this way, the phrase selected or input by the viewer can be output as the voice in the participation action with the tone color based on the input voice of the viewer.

Brief Description of Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

MODE FOR CARRYING OUT THE INVENTION

[0025] Hereinafter, the present embodiment will be described. Note that the present embodiment described below does not unduly limit the contents described in the claims. Also, not all of the configurations described in the present embodiment are essential constituent elements.

[0026] 1. Content distribution system First, with reference to FIGS. 1(A) to 1(F), the hardware device that realizes the content distribution system of the present embodiment will be described.

[0027] In Fig. 1(A), a server system 500 (information processing system) is communicatively connected to terminal devices TM1 to TMn via a network 510. For example, the server system 500 is a host, and the terminal devices TM1 to TMn are clients. Note that the content distribution system and its processing according to the present embodiment may be realized by the server system 500, or may be realized by distributed processing of the server system 500 and the terminal devices TM1 to TMn.

[0028] Also, the content distribution system and processing according to the present embodiment can also be realized by a blockchain method. For example, each process of the content distribution system according to the present embodiment may be executed using a program called a smart contract executable on Ethereum. In this case, the terminal devices TM1 to TMn are connected peer-to-peer. Also, various types of information such as content information communicated between the terminal devices TM1 to TMn are transferred using a blockchain. Note that hereinafter, each of the terminal devices TM1 to TMn is appropriately described as the terminal device TM.

[0029] The server system 500 can be realized, for example, by one or more servers (such as a management server, a content delivery server like a game delivery server or a video delivery server, a billing server, a service providing server, an authentication server, a database server, or a communication server, etc.). This server system 500 provides various services for operating content delivery, and can manage data necessary for the delivery of viewing content and perform the delivery of client programs and various data, etc. As a result, viewers can access the server system 500 through the terminal device TM which is a viewer terminal, and view the viewing content provided by the server system 500. Also, through the processing of the server system 500, a delivery function for delivering viewing content to viewers, a participation behavior function (posting function) that enables viewers' participation behaviors (postings) such as paying money for viewing content, etc. are realized. The distributor is, for example, the performer of the viewing content, etc. Also, through the processing of the server system 500, an online shopping function for virtual electronic media such as billing items, a game providing function that enables playing online games, a user management function for registering users and managing user-specific information, etc. are realized.

[0030] The network 510 (distribution network, communication line) is, for example, a communication path using the Internet, a wireless LAN, etc., and can include, in addition to a dedicated line (dedicated cable) for direct connection and a LAN using Ethernet (registered trademark), etc., communication networks such as a telephone communication network, a cable network, and a wireless LAN. Also, the communication method can be either wired or wireless.

[0031] The terminal device TM (user terminal) is a terminal having, for example, a network connection function (Internet connection function). As these terminal devices TM, for example, portable communication terminals such as smartphones and mobile phones shown in FIG. 1(B), portable game devices shown in FIG. 1(C), home game devices (stationary type) shown in FIG. 1(D), business game devices shown in FIG. 1(E), or information processing devices such as personal computers (PCs) and tablet PCs shown in FIG. 1(F) can be used. Alternatively, as the terminal device TM, wearable devices (HMD, watch-type devices, etc.) worn on parts such as the user's head and arm may be used.

[0032] FIG. 2 shows a configuration example of the content distribution system of the present embodiment. Note that the configuration of the content distribution system is not limited to FIG. 2, and various modifications such as omitting a part of its components (each part) or adding other components are possible.

[0033] The content distribution system includes a processing unit 100, a storage unit 170, and a communication unit 196. This content distribution system can be realized, for example, by the server system 500 in FIG. 1(A), and is communicatively connected to the distributor terminal TMP, which is a terminal device TM, and the viewer terminals TMA to TMAm via the network 510. Hereinafter, each of the viewer terminals TMA1 to TMAn will be generically referred to as the viewer terminal TMA as appropriate.

[0034] The processing unit 100 (processor) performs distribution processing, reception processing, management processing, display processing, or sound processing, etc., based on various information, programs, or operation information stored in the storage unit 170.

[0035] Each process (each function) of the present embodiment performed by each part of the processing unit 100 can be realized by a processor (a processor including hardware). For example, each process of the present embodiment can be realized by a processor that operates based on information such as a program and a memory that stores information such as a program. The processor may be, for example, a processor in which the functions of each part are realized by individual hardware, or the functions of each part may be realized by integrated hardware. For example, the processor includes hardware, and the hardware can include at least one of a circuit that processes digital signals and a circuit that processes analog signals. For example, the processor can also be composed of one or more circuit devices (such as ICs, etc.) mounted on a circuit board and one or more circuit elements (such as resistors, capacitors, etc.). The processor may be, for example, a CPU (Central Processing Unit). However, the processor is not limited to the CPU, and various processors such as a GPU (Graphics Processing Unit) or a DSP (Digital Signal Processor) can be used. Also, the processor may be a hardware circuit by an ASIC. Also, the processor may include an amplifier circuit or a filter circuit that processes analog signals. The memory (storage unit) may be a semiconductor memory such as SRAM or DRAM, or may be a register. Or it may be a magnetic storage device such as a hard disk drive (HDD), or an optical storage device such as an optical disk device. For example, the memory stores instructions readable by a computer, and when the instructions are executed by the processor, the processes (functions) of each part of the processing unit 100 are realized. The instructions here may be an instruction set that constitutes a program, or may be instructions that instruct an operation to the hardware circuit of the processor.

[0036] The processing unit 100 includes a distribution processing unit 102, a reception unit 104, a management unit 118, a display processing unit 120, and a sound processing unit 130. Note that the configuration of the processing unit 100 is not limited to this, and various modifications such as omitting some of these components or adding other components are possible.

[0037] The distribution processing unit 102 performs various distribution processes. Specifically, the distribution processing unit 102 performs a distribution process for the viewer of the viewer terminal TMA to view the viewing content. For example, the distribution processing unit 102 performs a transmission process of information using the communication unit 196. Specifically, the distribution processing unit 102 performs a process of transmitting information to the viewer terminal TMA and the distributor terminal TMP via the communication unit 196 and the network 510. The information transmitted to the viewer terminal TMA is, for example, information such as viewing content information. The information transmitted to the distributor terminal TMP is, for example, feedback information from the viewer terminal TMA. The viewing content information is, for example, information for displaying a display image (video) on the display unit 290 of FIG. 3 described later, or information for outputting output sounds such as voices, music, or sound effects from the sound output unit 292. For example, the distribution processing unit 102 performs streaming distribution of the viewing content. Streaming distribution is video distribution by streaming processing. Alternatively, the distribution processing unit 102 may distribute the viewing content to the viewer terminal TMA using reproduction information for reproducing the viewing content at the viewer terminal TMA. The reproduction information is, for example, input information input by the distributor using the distributor terminal TMP, or information for reproducing the performance of the distributor who is the performer. The distribution processing unit 102 can also perform a reception process of information using the communication unit 196. Specifically, the distribution processing unit 102 performs a process of receiving information from the distributor terminal TMP and the viewer terminal TMA via the communication unit 196 and the network 510. The information received from the distributor terminal TMP is, for example, information on the viewing content distributed by the distributor. The information received from the viewer terminal TMA is, for example, input information input by the viewer using the viewer terminal TMA, such as participation action information of the viewer with respect to the viewing content. The participation action information is also called viewer's contribution information.

[0038] The distribution processing unit 102 may also perform various content processes. For example, the distribution processing unit 102 may perform a process of generating information on the viewing content that progresses based on the operation of the distributor terminal TMP by the distributor. That is, it performs a process of advancing the viewing content and a process of generating information on the viewing content. For example, the distribution processing unit 102 may perform a game process for the user to play a game and generate information on the game content that is the viewing content. The game process is, for example, a process of starting the game when the game start condition is satisfied, a process of advancing the started game, a process of ending the game when the game end condition is satisfied, or a process of calculating a game result such as a game score. Taking a browser game as an example, the distribution processing unit 102 controls the progress of the game for each user by managing various information of the user for each user. The user information is stored in the user information storage unit 174. For example, a web page constituting a website that provides a game service is displayed on the terminal device TM such as the distributor terminal TMP or the viewer terminal TMA in response to a request from the terminal device TM. Specifically, the web page is displayed by the web browser provided in the terminal device TM. When a hyperlink on the displayed web page is selected by the user, new HTML data corresponding to the hyperlink is transmitted to the terminal device TM, and a web page based on the new HTML data is displayed on the terminal device TM. In this way, by sequentially providing the web page to the terminal device TM according to the user's operation, it becomes possible to advance the game based on the user's operation on the terminal device TM. In this case, the display image information generated by the distribution processing unit 102 is, for example, HTML data or the like.

[0039] The reception unit 104 performs various reception processes. For example, the reception unit 104 performs a reception process for information input by the distributor using the distributor terminal TMP, or a reception process for information input by the viewer using the viewer terminal TMA. Details of the reception unit 104 will be described later.

[0040] The management department 118 performs, for example, user authentication processing. For example, it performs authentication processing on a user who logs in using the terminal device TM. This authentication processing is performed based on, for example, a password or account information entered by the user. The management department 118 also performs various billing processes. For example, it performs processes such as billing determination processing, billing data creation processing, and storage processing. The management department 118 also performs various management processes. For example, it performs management processes for various services and management processes for various information. The management department 118 can be realized by, for example, a management server.

[0041] For example, in order for a user to use services provided by the server system 500 in FIG. 1(A) etc., the user performs a predetermined procedure to obtain an account. By entering the password associated with the obtained account and logging in, the user can use various services such as live distribution, playing social games and online games, services on live distribution sites and game sites, online shopping for items etc., message exchange between users, and registration of friend users. The management department 118 also performs management processing etc. of such user account information.

[0042] The display processing unit 120 performs processing for displaying an image on the display unit 290 of the terminal device TM in FIG. 3, which is the distributor terminal TMP or the viewer terminal TMA. For example, display image information (image generation data) such as HTML data is transmitted to the terminal device TM via the communication unit 196 and the network 510, and processing for displaying an image corresponding to the display image information on the display unit 290 of the terminal device TM is performed.

[0043] The sound processing unit 130 performs processing for outputting sound from the sound output unit 292 of the terminal device TM. For example, output sound information (sound generation data) is transmitted to the terminal device TM via the communication unit 196 and the network 510, and processing for outputting sound (voice, music, game sound, effect sound) corresponding to the output sound information from the sound output unit 292 of the terminal device TM is performed.

[0044] The storage unit 170 serves as a working area for components such as the processing unit 100 and the communication unit 196, and its functions can be realized by semiconductor memories, HDDs, SSDs, optical disk devices, etc. The storage unit 170 includes a content information storage unit 172 and a user information storage unit 174. The content information storage unit 172 stores information on the viewing content that is the target of content distribution. The viewing content is entertainment content, such as game content, virtual reality (VR) content, or video distribution content. For example, the viewing content is live content. Examples of live content include game live content (game demonstration content) that broadcasts the gameplay of a distributor, and performance live content that broadcasts the performances of real-world idols, singers, bands, actors, etc. and their corresponding VR characters. The information of the viewing content is various information for the viewer to view the viewing content on the viewer terminal TMA, such as display image information (video information), output sound information, or content sequence information for the progress of the content. The user information storage unit 174 stores various information about the user. For example, the user information storage unit 174 stores the user's personal information (name, gender, date of birth, email address, etc.) as user information. For example, the user's account information (user ID) is also stored as user information. For example, the billing information subject to billing processing is associated with each user's account information.

[0045] The communication unit 196 communicates with external devices, and its functions can be realized by hardware such as a communication ASIC or a communication processor, and communication firmware. For example, the communication unit 196, which is a communication interface, performs various communication processes for communicating with terminal devices TM such as the distributor terminal TMP and the viewer terminal TMA via the network 510.

[0046] Also, the content delivery system in FIG. 2 performs each process of this embodiment based on the program of this embodiment. This program is a program for causing a computer (a device including an operation unit, a processing unit, a storage unit, and an output unit) to function as each unit of this embodiment (a program for causing the computer to execute the processing of each unit). This program is stored, for example, in an information storage medium. That is, the content delivery system of this embodiment performs various processes of this embodiment based on the program (data) stored in the information storage medium. The information storage medium, which is a medium readable by a computer, stores programs, data, etc., and its function can be realized by an optical disk, HDD, semiconductor memory, etc. Note that the program (data) for causing a computer to function as each unit of this embodiment may be distributed from the information storage medium of the server system 500 (host device) via the network 510. The use of such an information storage medium by the server system 500 can also be included within the scope of this embodiment.

[0047] FIG. 3 shows a configuration example of the terminal device TM. The terminal device TM includes a processing unit 200, an operation unit 260, an interface unit 262, a storage unit 270, an information storage medium 280, a display unit 290, an audio output unit 292, and a communication unit 296. Note that the configuration of the terminal device TM is not limited to FIG. 3, and various modifications such as omitting a part of its components (each unit) or adding other components are possible.

[0048] The processing unit 200 (processor) executes terminal-side processing in the content delivery system based on operation information from the operation unit 260, programs, etc. For example, it executes processing for content delivery, game processing, etc. The processing unit 200 can be realized by a processor or the like, similarly to the processing unit 100 in FIG. 2 described above. Note that each process of the content delivery system of this embodiment may be realized by distributed processing between the server system 500 and the terminal device TM.

[0049] The operation unit 260 is for the user to input various information such as operation information. Its functions can be realized by operation buttons, direction keys, analog sticks, levers, various sensors (angular velocity sensors, acceleration sensors, etc.), microphones, or touch panel displays. The interface unit 262 performs interface processing with external devices, for example, performs processing to communicate with external devices according to a predetermined interface standard. Also, the interface unit 262 performs interface processing with portable information storage media such as IC cards (memory cards), USB memories, or magnetic cards where various information about the user is stored. The functions of the interface unit 262 can be realized by hardware such as an ASIC for interface processing or a processor for interface processing, or by firmware for interface processing.

[0050] The storage unit 270 serves as a work area for the processing unit 200, communication unit 296, etc. Its functions can be realized by semiconductor memories, HDDs, SSDs, optical disk devices, etc. The information storage medium 280 (a computer-readable medium) stores programs, data, etc. Its functions can be realized by optical disks, HDDs, semiconductor memories, etc. The processing unit 200 performs various processes of this embodiment based on the programs (data) stored in the information storage medium 280. A program (a program for causing a computer to execute the processes of each part, which enables a computer (a device equipped with an operation unit, a processing unit, a storage unit, and an output unit) to function as each part of this embodiment) can be stored in this information storage medium 280.

[0051] The display unit 290 outputs the image generated according to the present embodiment, and its function can be realized by an LCD, an organic EL display, a CRT, or an HMD, etc. The sound output unit 292 outputs the sound generated according to the present embodiment, and its function can be realized by a speaker or headphones, etc. The communication unit 296 (communication interface) communicates with external devices such as the server system 500 and other terminal devices via the network 510, and its function can be realized by hardware such as a communication ASIC or a communication processor, or communication firmware.

[0052] As shown in FIG. 2, the content distribution system (server system) of the present embodiment includes a distribution processing unit 102, a reception unit 104, and a sound processing unit 130.

[0053] The distribution processing unit 102 performs distribution processing for a viewer of the viewer terminal TMA to view the viewing content. For example, the distribution processing unit 102 performs a transmission process of information on the viewing content to the viewer terminal TMA. The distribution processing unit 102 may also perform a reception process of information on the viewing content from the distributor terminal TMP.

[0054] The reception unit 104 performs a process of receiving the viewer's participation actions with respect to the viewed content. The viewer's participation actions with respect to the viewed content are not only actions where the viewer simply watches the viewed content, but also actions where the viewer actively takes actions and has an impact on the viewed content. Due to the viewer's participation actions, for example, the content such as the image or sound of the viewed content changes, or the content progression changes. This participation action is performed based on, for example, the input information entered by the viewer using the viewer terminal TMA, and the reception unit 104 performs a process of receiving the viewer's input information regarding this participation action. For example, in a content distribution system, a viewer of the viewed content performs participation actions such as posting comments or providing items such as stamps to praise the viewed content or the distributor, or to support the distributor who distributes the viewed content. Such posting of comments and providing of items are sometimes called "tipping" in emulation of donations from the audience to street performers, or are called "gifts" and the like.

[0055] And in this embodiment, the reception unit 104 performs a process of receiving, as the viewer's participation actions with respect to the viewed content, participation actions by voice in viewer settings or system settings. That is, the reception unit 104 receives the viewer's participation actions by voice instead of items such as stamps. For example, the reception unit 104 receives voice tipping by the viewer. The voice in the viewer settings is, for example, voice based on the viewer's own voice. This voice in the viewer settings may be voice obtained by performing processing such as voice synthesis on the voice input by the viewer. The voice in the system settings is voice prepared by the system (content distribution system) (default voice), for example, voice held as voice data in the storage unit 170 by the system. The voice in the system settings may be voice obtained by performing processing such as voice synthesis on the voice prepared in advance by the system. Also, the voice that becomes the viewer's participation action may be voice based on both viewer settings and system settings.

[0056] Then, the sound processing unit 130 performs processing for changing the content, output timing, or timbre of the voice in the participation action, or for controlling the volume. For example, based on the situation of the viewing content at the timing of the viewer's participation action, the sound processing unit 130 performs at least one of processing for changing the content of the voice in the participation action, processing for changing the output timing of the voice, processing for changing the timbre of the voice, and processing for controlling the volume of the voice. The situation of the viewing content at the timing of the participation action is the situation of the viewing content in a given period including the timing of the participation action. For example, the timing at which the reception unit 104 receives the viewer's participation action is the timing of the participation action, and for example, the situation of the viewing content in a certain period before and after the participation action timing is determined. The situation of the viewing content includes the situation of the content of the viewing content, the progress of the viewing content, the situation regarding the viewer's participation action with respect to the viewing content, the situation of the data set for the viewing content, or the situation of the distribution device or viewing device of the viewing content, etc. The situation of the content of the viewing content is, for example, the situation of the content such as the music or video that constitutes the viewing content. The content of the voice is the content or type of the message represented by the voice, etc. For example, when the situation of the viewing content is the first situation, the sound processing unit 130 outputs the voice of the first content (the first message) as the voice in the participation action, and when it is the second situation, it outputs the voice of the second content (the second message). Also, when the situation of the viewing content is the first situation, the sound processing unit 130 outputs the voice in the participation action at the first output timing, and when it is the second situation, it outputs the voice in the participation action at the second output timing. Also, when the situation of the viewing content is the first situation, the sound processing unit 130 performs processing for outputting the voice of the first timbre as the voice in the participation action, and when it is the second situation, it outputs the voice of the second timbre. Also, when the situation of the viewing content is the first situation, the sound processing unit 130 performs processing for outputting the voice in the participation action at the first volume, and when it is the second situation, it outputs the voice at the second volume.

[0057] Also, the sound processing unit 130 changes the voice in the participation behavior into a voice with content, timbre, or volume according to the melody or genre of the music of the video content. For example, when the music of the video content is of the first melody or the first genre, the sound processing unit 130 outputs a voice with content (phrase) corresponding to the first melody or the first genre. When the music is of the second melody or the second genre, the sound processing unit 130 outputs a voice with content (phrase) corresponding to the second melody or the second genre as the voice in the participation behavior. This can be realized, for example, by associating the information of the content (phrase) of the voice in the participation behavior with the melody or genre of the music of the video content and storing it in the storage unit 170. Also, when the music of the video content is of the first melody or the first genre, the sound processing unit 130 outputs a voice with the first timbre as the voice in the participation behavior. When the music is of the second melody or the second genre, the sound processing unit 130 performs a process of outputting a voice with the second timbre. This can be realized, for example, by associating the setting information of the timbre of the voice in the participation behavior with the melody or genre of the music of the video content and storing it in the storage unit 170. Also, when the music of the video content is of the first melody or the first genre, the sound processing unit 130 performs a process of outputting the voice in the participation behavior at the first volume. When the music is of the second melody or the second genre, the sound processing unit 130 performs a process of outputting the voice at the second volume. This can be realized, for example, by associating the setting information of the volume of the voice in the participation behavior with the melody or genre of the music of the video content and storing it in the storage unit 170. For example, when the music of the video content is a music with a quiet melody or a quiet genre, the sound processing unit 130 reduces the volume of the voice in the participation behavior. When the music of the video content is a music with a good tempo or a noisy melody or genre, the sound processing unit 130 increases the volume of the voice in the participation behavior.

[0058] Also, the sound processing unit 130 determines the set timing of the video content and the timing of the viewer's participation behavior, and performs a change process to bring the output timing of the voice closer to the set timing. For example, for video content, a plurality of set timings are set in advance. The information on these set timings is stored, for example, in the storage unit 170 (content information storage unit 172) in association with the video content. Examples of the set timing include the call timing set as the timing when a voice such as a call is output in the video content. Then, the sound processing unit 130 compares each of the plurality of set timings of the video content with the timing of the viewer's participation behavior, and searches for, for example, a set timing close to the timing of the participation behavior from among the plurality of set timings. Then, the sound processing unit 130 performs a change process to correct the output timing of the voice so that the output timing of the voice approaches the searched set timing. Then, the sound processing unit 130 controls so that the voice in the participation behavior is output at the output timing after the change process. For example, when a plurality of viewers perform a participation behavior, a change process is performed to correct the output timing of the voices in the participation behaviors of these plurality of viewers so that the output timings of the voices in the participation behaviors of the plurality of viewers are aggregated to the set timing.

[0059] Also, the sound processing unit 130 determines the timing of the viewer's participation behavior and the timing of the participation behavior of other viewers, and performs a process of changing the output timing of the sound so that the sound is output at a representative timing set based on the timing of the viewer's participation behavior and the timing of the participation behavior of other viewers. For example, the sound processing unit 130 determines the timing of the viewer's participation behavior and the timing of the participation behavior of one or more other viewers, and sets a representative timing from these participation behavior timings. The representative timing may be the timing of the participation behavior of any one of a plurality of viewers including the viewer and other viewers, or may be a timing obtained based on the average value of the participation behavior timings of the plurality of viewers. Then, the sound processing unit 130 controls so that the sound in the viewer's participation behavior and the sound in the participation behavior of other viewers are output from the viewer terminal TMA or the like at this representative timing.

[0060] Also, when the sound in the viewer's participation behavior and the sound in the participation behavior of other viewers are output at the viewer terminal, the sound processing unit 130 performs a volume control process of changing at least one of the volume of the viewer's sound and the volume of the sound of other viewers. For example, the sound processing unit 130 performs a volume control process of increasing the volume of one of the volume of the viewer's sound and the volume of the sound of other viewers and decreasing the volume of the other. For example, at the viewer terminal, the volume of the viewer's sound is increased or the volume of the sound of other viewers is decreased. On the other hand, at the viewer terminal of other viewers, the volume of the sound of other viewers is increased or the volume of the viewer's sound is decreased.

[0061] Also, the audio processing unit 130 determines the set timing of the viewing content and the timing of the viewer's participation behavior, and performs audio volume control processing according to the degree of coincidence between the set timing and the participation behavior timing. For example, the audio processing unit 130 performs a matching degree determination process to determine whether the set timing of the viewing content and the timing of the viewer's participation behavior match within a given range. Whether the set timing and the participation behavior timing match within a given range can be determined, for example, by detecting whether the viewer has performed a participation behavior within a given period (judgment period) including the set timing. And when the set timing and the participation behavior timing match within a given range, the audio processing unit 130 performs volume control processing to increase the volume of the audio in the viewer's participation behavior, for example. Also, when the set timing and the participation behavior timing do not match within a given range, the audio processing unit 130 performs volume control processing to decrease the volume of the audio in the viewer's participation behavior, for example.

[0062] In addition, the audio processing unit 130 performs audio volume control processing according to the number of other viewers who have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior. Whether other viewers have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior can be determined, for example, by detecting whether other viewers have performed a participation behavior within a given period (judgment period) including the timing of the viewer's participation behavior. And the audio processing unit 130 performs volume control processing to increase or decrease the volume of the audio according to the number of other viewers who have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior. For example, volume control processing is performed to increase or decrease the volume of the audio according to the number of viewers participating within a given period. For example, when the number of participants is small, the audio processing unit 130 increases the volume of each viewer's audio aiming for a boosting effect. Also, when the number of participants is large, the audio processing unit 130 decreases the volume of each viewer's audio so that the overall volume does not become too large.

[0063] In addition, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer and the voice color of the system settings. For example, the reception unit 104 performs processing to display, on the viewer terminal, a screen for setting the phrase and voice color of the voice in the participation action. Then, the reception unit 104 receives the phrase selected by the viewer from among the plurality of phrases of the system settings displayed on the screen, or receives the phrase of the viewer settings input by the viewer on the screen. In addition, the reception unit 104 receives the voice color selected by the viewer from among the voice colors of the system settings displayed on the screen. Then, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer in this way and the voice color of the system settings. For example, the voice obtained by playing (speaking) the phrase based on the voice color of the system settings is received as the voice in the participation action.

[0064] In addition, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer and the voice color of the viewer's input voice. For example, the reception unit 104 performs processing to display, on the viewer terminal, a screen for setting the phrase and voice color of the voice in the participation action. Then, the reception unit 104, in the same manner as above, receives the phrase selected by the viewer from among the plurality of phrases of the system settings displayed on the screen, or receives the phrase of the viewer settings input by the viewer on the screen. In addition, the reception unit 104 obtains the voice color of the viewer's input voice by having the viewer input their own voice as the viewer settings voice on the screen. Then, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer in this way and the voice color of the viewer's input voice. For example, the voice obtained by playing (speaking) the phrase based on the voice color of the viewer's input voice is received as the voice in the participation action.

[0065] Note that each process of the content distribution system of FIG. 2 and each process of this embodiment such as distribution processing, reception processing, display processing, and sound processing described above can be realized by the server system 500 in FIG. 1(A), or can be realized by a terminal device, or can be realized by distributed processing between the server system 500 and the terminal device. Various modifications are possible. For example, the server system may process only the information necessary for performing each process of this embodiment, or may perform only the process of transmitting the information to the terminal device, and the terminal device may execute other processes. For example, a program for performing each process of this embodiment may be installed in the terminal device, and the terminal device may execute each process of this embodiment based on the installed program. Further, the content distribution system may be realized by an information processing system other than the server system 500.

[0066] 2. Method of this embodiment Next, the method of this embodiment will be described in detail.

[0067] 2.1 Participatory behavior in audiovisual content by voice In the content distribution system, there is a function to display, in a form visible to other viewers, comments, stamps, etc. with more luxurious visual effects than usual due to charging or the like during distribution. Such comments and stamps are called donations, and viewers can perform participatory behaviors (posts) such as making donations to give reactions such as support to the distributor who is the performer. And until now, participatory behaviors such as donations have been performed by posting comments or inserting items such as stamps. In contrast, in this embodiment, participatory behavior by voice is possible as a participatory behavior in audiovisual content, and voice donations are possible. That is, in this embodiment, an auditory donation function that can produce a sense of unity among users enjoying entertainment, such as a live call, is added to the donation function of the conventional content distribution system. As a result, it becomes possible to enhance the interest of entertainment by the content distribution system.

[0068] For example, FIG. 4 shows an example of a live stream by a VR character CH. In this live stream, in a virtual field VFL in a virtual space, a VR character CH (model object) performs various performances. For example, the VR character CH performs performances such as dancing, waving hands, jumping, and clapping hands. Although FIG. 4 shows an example of a live stream by a VR character CH, the live stream may be a live stream by a real-world idol or the like.

[0069] In such a live stream, viewers perform participation actions by voice VCs1, VC2 such as "Yay!" and "Fuffu". That is, instead of making donations by comments or stamps, viewers make donations (broadly speaking, participation actions) by voice VCs1, VC2 set by the viewer or the system. These voice VCs1, VC2 are output as voice at the viewer terminals of the viewers who made the voice donations, for example. Also, the voice VCs1, VC2 are output as voice at the viewer terminals of other viewers. Also, the voice VCs1, VC2 may be output as voice at the distributor terminal, or may be output as voice at the live stream venue, for example. Making donations by comments or stamps adds a visual production effect to the viewing content, while the voice donations in this embodiment add an auditory production effect to the viewing content. For example, at the call timing of the live stream song, each viewer makes a voice donation corresponding to the call, or at the exciting time such as the refrain of the song, each viewer makes a voice donation corresponding to the cheer, thereby enhancing the sense of unity among the viewers and realizing production effects such as livening up the live stream.

[0070] Figure 5 shows an example of a method for setting voices for viewer participation behavior. When making a voice donation, a screen as shown in Figure 5 is displayed. As shown in A1, phrases that are candidates for system settings, such as "Yeah!", "Huff Huff", "Thank you~", and "Amazing", prepared by the system (content delivery system) are displayed. The viewer selects a desired reading phrase from among these system-setting phrases. Also, on the screen of Figure 5, as shown in A2, voice tones that are candidates for system settings, such as boy voice A, boy voice B, young man voice A, young man voice B, and lady voice prepared by the system, are displayed. The viewer selects a desired voice tone from among these system-setting voice tones. Then, for example, as shown in A3, by setting the charging amount, etc., and performing the operation shown in A4, a voice donation is made. For example, in A1 of Figure 5, the phrase "Huff Huff" is selected, and in A2, the voice tone of young man voice A is selected. Therefore, a voice donation is made by playing (uttering) the phrase "Huff Huff" in the voice tone of young man voice A.

[0071] Figures 6, 7(A), and 7(B) show other examples of the method for setting voices for viewer participation behavior. For example, in Figure 6, instead of selecting from system-setting phrases as in Figure 5, the viewer inputs a desired reading phrase as a viewer-setting phrase by themselves. For example, the viewer inputs the characters of the phrase using an operation unit of the viewer terminal, etc. Also, in Figure 6, similar to Figure 5, the viewer selects a desired voice tone from among the voice tones that are candidates for system settings prepared in advance by the system. By doing so, a voice donation is made by playing the viewer-setting phrase input by the viewer in the system-setting voice tone.

[0072] In FIG. 7(A), similar to FIG. 5, the viewer selects a desired phrase from among the phrases of system setting candidates. Also, the viewer inputs input voice by recording his or her own voice. By this input voice of the viewer, the timbre of the voice of the participation behavior is set. By doing so, the phrase of the system setting is made to be the coin - in of the voice reproduced with the timbre of the viewer setting based on the input voice of the viewer. On the other hand, in FIG. 7(B), the viewer inputs the characters of the phrase using the operation unit etc. of the viewer terminal and inputs the input voice by recording his or her own voice. By doing so, the phrase of the viewer setting is made to be the coin - in of the voice reproduced with the timbre of the viewer setting based on the input voice of the viewer.

[0073] As described above, in this embodiment, the voice generated based on the phrase selected or input by the viewer and the timbre of the system setting is received as the voice in the participation behavior. For example, in FIG. 5, the voice generated based on the phrase of the system setting selected by the viewer and the timbre of the system setting is received as the voice in the participation behavior. Also, in FIG. 6, the voice generated based on the phrase of the viewer setting input by the viewer and the timbre of the system setting is received as the voice in the participation behavior. By doing so, it becomes possible to output the phrase selected or input by the viewer as the voice in the participation behavior with the timbre of the system setting. For example, by using the timbre of the system setting, the timbre of the voice in the participation behavior can be unified with the timbre of the system setting. Thereby, for example, it becomes possible to realize an auditory effect with a voice of a timbre suitable for the content of the video content etc.

[0074] In this embodiment, the voice generated based on the phrase selected or input by the viewer and the timbre of the viewer's input voice is accepted as the voice in the participation action. For example, in FIG. 7(A), the voice generated based on the system setting phrase selected by the viewer and the viewer setting timbre based on the viewer's input voice is accepted as the voice in the participation action. Also, in FIG. 7(B), the voice generated based on the viewer setting phrase input by the viewer and the viewer setting timbre based on the viewer's input voice is accepted as the voice in the participation action. In this way, it becomes possible to output the phrase selected or input by the viewer as the voice in the participation action with the timbre based on the viewer's input voice. For example, by using the timbre based on the viewer's input voice, it becomes possible to make the timbre of the voice in the participation action the timbre based on the viewer's voice. Therefore, the viewer can feel as if they are cheering, etc. with their own voice, and it becomes possible to realize the participation action in the viewing content with a realistic voice that reflects the viewer's voice.

[0075] 2.2 Processing Example of this Embodiment FIG. 8 is a flowchart for explaining the processing of this embodiment. First, the content distribution system determines whether a viewer has performed a voice-based participation action (casting money) as a participation action for the viewed content (step S1). If such a participation action has been performed, the content distribution system accepts the viewer's voice-based participation action (step S2). For example, in the screens shown in FIGS. 5 to 7(B), when the viewer sets a voice phrase, tone color, etc. and performs a voice-based participation action for the viewed content, this participation action is accepted. This voice is the voice of the viewer settings or system settings as described in FIGS. 5 to 7(B). Then, the content distribution system acquires the situation of the viewed content at the timing of the viewer's participation action (step S3). The situation of the viewed content at the timing of the viewer's participation action is the situation of the viewed content during a given period (for example, a certain period before and after the timing) including the timing of the participation action. Also, the content distribution system acquires, as the situation of the viewed content, the situation of the content of the viewed content, the progress situation, the situation of participation actions, the situation of setting data, or the situation of the distribution device or viewing device, etc. And the content distribution system performs at least one of the change processing of the content of the voice in the participation action, the change processing of the output timing of the voice, the change processing of the tone color of the voice, and the control processing of the volume of the voice based on the situation of the viewed content at the timing of the viewer's participation action. For example, the content distribution system changes the content of the content such as the music or video constituting the viewed content according to the situation of the viewed content, changes the output timing of the voice on the viewer terminal or the distributor terminal, changes the tone color of the voice to various tone colors corresponding to the viewed content, or changes the volume of the output of the voice on the viewer terminal or the distributor terminal.

[0076] According to this embodiment, when a viewer performs a voice participation action on the viewing content, this participation action is accepted. By performing such a voice participation action on the viewing content, an auditory effect on the viewing content becomes possible. And in this embodiment, based on the situation of the viewing content at the timing of the participation action, a change process of the voice content, output timing or tone color is performed, or a volume control process is performed. Therefore, it becomes possible to realize the distribution of viewing content that enables the viewer's participation action by voice with content, output timing, tone color or volume according to the situation of the viewing content.

[0077] FIG. 9 is a flowchart for explaining a process of controlling the volume of voice according to the melody or genre of a piece of viewing content. First, the content distribution system acquires the melody or genre of the piece of viewing content (step S11). For example, the content distribution system acquires whether the melody of the piece of viewing content is a quiet melody, a lively melody, a bright melody, a dark melody, a ballad melody, a sad melody, a happy melody, an intense melody, a gentle melody, a warm melody or a beautiful melody, etc. Also, the content distribution system acquires whether the genre of the viewing content is pop, rock, idol, punk, hip-hop, club, hard rock, Latin, classical, easy listening, or new age, etc. Then, the content distribution system changes the voice in the participation action to a voice with content, tone color or volume according to the melody or genre (step S12).

[0078] Thus, in this embodiment, the voice of the participation action is changed to a voice with content, timbre, or volume corresponding to the melody or genre of the music of the video content. By doing so, it becomes possible to output the voice of the participation action with content, timbre, or volume that matches the melody or genre of the music of the video content, and it is possible to prevent a situation where the voice of the participation action is output with content, timbre, or volume that does not match the melody or genre. For example, when the music of the content is a quiet melody, the content (phrase) or timbre of the voice that matches the quiet melody is changed, or the volume is decreased. When the music has a lively melody, the content or timbre of the voice that matches the lively melody is changed, or the volume is increased. Also, when the genre of the music is rock, the content or timbre of the voice that matches rock is changed, or the volume is increased. When the genre of the music is classical, the content or timbre of the voice that matches classical is changed, or the volume is decreased. Note that metadata indicating the melody and genre is set for the music of the video content, and data for setting the content (phrase), timbre, or volume of the voice is associated with this metadata and stored in the storage unit 170. The content distribution system changes the voice in the participation action to a voice with content, timbre, or volume corresponding to the melody or genre based on the metadata associated with the music of the video content and the setting information of the content, timbre, or volume of the voice associated with the metadata. Also, as described with reference to FIGS. 5 to 7(B), when the content distribution system displays candidates for phrases or timbre in the system settings on the screen, it displays phrases or timbre associated with the melody or genre of the music of the video content as candidate phrases or timbre.

[0079] FIG. 10 is a flowchart for explaining a process of controlling the output timing of voice based on the setting timing of video content and the participation behavior timing of viewers. First, the content delivery system identifies a setting timing (call timing) close to the timing of the viewer's participation behavior (coin toss) from among a plurality of setting timings (call timings) set for the video content (step S21). For example, a process of searching for a setting timing close to the timing of the viewer's participation behavior is performed from among the setting timings of the video content. Then, the content delivery system changes the output timing of the voice in the participation behavior so as to approach the identified setting timing (step S22). For example, in FIG. 11, setting timings TS1, TS2, TS3, etc. are set for the video content. In FIG. 11, the direction of t is the direction of the time axis. When the viewer's participation behavior is detected, a setting timing close to the timing of the viewer's participation behavior is searched for from among the setting timings TS1, TS2, TS3, etc. of the video content. In FIG. 11, since the setting timing TS2 is searched for as the setting timing close to the timing of the viewer's participation behavior, the output timing TQ of the voice in the participation behavior is changed so as to approach the setting timing TS2. As a result, a voice such as "Fuffhoo" is output at the timing corresponding to the setting timing TS2. The timing corresponding to the setting timing TS2 is, for example, the same timing as the setting timing TS2 or a timing close to the setting timing TS2.

[0080] Thus, in this embodiment, the setting timing set for the video content and the timing of the viewer's participation behavior are judged, and a change process is performed to make the output timing of the voice approach the setting timing. In this way, the voice in the participation behavior can be output at the timing corresponding to the setting timing of the video content. As a result, for example, it is possible to prevent the voice in the participation behavior from being output at a timing unrelated to the setting timing. Therefore, it is possible to prevent the occurrence of a situation where the production effect and atmosphere of the video content are impaired by the voice in the participation behavior being output at a timing unrelated to the setting timing.

[0081] In this embodiment, the timing of the participating actions of the viewer and the timing of the participating actions of other viewers are determined, and based on the timing of the participating actions of the viewer and the timing of the participating actions of other viewers, at the representative timing set, a process for changing the output timing of the voice is performed so that the voice in the participating action is output. For example, TA1, TA2, TA3, TA4, TA5, TA6, TA7, TA8, TA9 in FIG. 12 are the timings of the participating actions of viewers AD1, AD2, AD3, AD4, AD5, AD6, AD7, AD8, AD9, respectively. In this case, TQ1 is set as the representative timing of the participating action timings TA1, TA2, TA3 of viewers AD1, AD2, AD3. Also, TQ2 is set as the representative timing of the participating action timings TA4, TA5, TA6 of viewers AD4, AD5, AD6, and TQ3 is set as the representative timing of the participating action timings TA7, TA8, TA9 of viewers AD7, AD8, AD9. For example, the representative timing of the group of viewers AD1 to AD3 is set to TQ1, the representative timing of the group of viewers AD4 to AD6 is set to TQ2, and the representative timing of the group of viewers AD7 to AD9 is set to TQ3. Then, the voice in the participating actions of viewers AD1 to AD3 is output at the representative timing TQ1, the voice in the participating actions of viewers AD4 to AD6 is output at the representative timing TQ2, and the voice in the participating actions of viewers AD7 to AD9 is output at the representative timing TQ3. For example, when viewer AD1 is regarded as his own viewer and viewers AD2, AD3 are regarded as other viewers, at the representative timing TQ1 set based on the timing of the participating action of viewer AD1 and the timing of the participating actions of other viewers AD2, AD3, the voice in the participating action of viewer AD1 will be output. In this way, the voice in the participating action of viewer AD1 and the voices in the participating actions of other viewers AD2, AD3 will be aggregated and output at the same representative timing TQ1. Therefore, it becomes possible to prevent a situation where the voices in the participating actions of viewer AD1 and other viewers AD2, AD3 are output at scattered timings, resulting in a lack of unity in the output of the voices in the participating actions.When setting timings TS1, TS2, and TS3 are set for the viewing content as shown in FIG. 11, change processing may be performed to make the representative timings TQ1, TQ2, and TQ3 closer to the setting timings.

[0082] FIG. 13 is a flowchart for explaining processing for changing the volume of the viewer's voice and the voices of other viewers. First, the content distribution system acquires voice data in the participation actions of other viewers (step S31). Then, the content distribution system performs processing for mixing the viewer's voice and the voices of other viewers so that, for example, the volume of the viewer's voice is greater than the volume of the voices of other viewers (step S32). Then, the content distribution system performs processing for outputting the mixed voice from the viewer terminal (step S33).

[0083] Thus, in this embodiment, when the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers are output on the viewer terminal, volume control processing is performed to change at least one of the volume of the viewer's voice and the volume of the voice of other viewers. For example, volume control processing is performed to change the output ratio of the viewer's voice and the voice of other viewers. For example, in step S32 of FIG. 13, the volume of the viewer's voice and the volume of the voice of other viewers are changed so that the volume of the viewer's voice is greater than the volume of the voice of other viewers. On the other hand, on the viewer terminal of other viewers, conversely, the volume of the voice of other viewers and the volume of the viewer's voice are changed so that the volume of the voice of other viewers is greater than the volume of the viewer's voice. In this way, volume control processing is performed so that the ratio of the volume of the voice in the participation behavior of the viewer and the volume of the voice in the participation behavior of other viewers becomes an appropriate output ratio, and the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers can be output to the viewer terminal. Therefore, for example, it is possible to prevent the occurrence of a situation such as the voice of other viewers becoming too loud and the volume balance being disrupted. For example, when a plurality of viewers including the viewer and other viewers view viewing content, in order to improve the production effect, it is desirable to output the voice of other viewers from the viewer terminal of the viewer. However, for example, when the voices of a large number of other viewers are output to the viewer terminal, a situation may occur in which the viewer cannot hear the voice in his or her own participation behavior. In this regard, for example, if volume control processing is performed so that the volume of the viewer's voice is greater than the volume of the voice of other viewers, such a situation can be prevented.

[0084] FIG. 14 is a flowchart for explaining a process of controlling the volume of audio based on the degree of coincidence between the timing of the viewer's participation behavior and the setting timing of the viewed content. First, the content distribution system compares the setting timing (call timing) set for the viewed content with the timing of the viewer's participation behavior (tipping) (step S41). That is, it compares the plurality of setting timings TS1, TS2, TS3, etc. as shown in FIG. 11 with the timing of the viewer's participation behavior. Then, the content distribution system determines whether the setting timing and the timing of the participation behavior match within a given range (step S42). For example, it determines whether the timing of the participation behavior falls within a predetermined determination period including the setting timing. And when the setting timing and the timing of the participation behavior match within a given range, the content distribution system increases the volume of the audio in the viewer's participation behavior (step S43), and when they do not match, it decreases the volume of the audio in the viewer's participation behavior (step S44). For example, it increases or decreases the volume of the audio in the viewer's participation behavior output from the viewer terminal or the like.

[0085] Thus, in this embodiment, the setting timing set for the viewing content and the timing of the viewer's participation behavior are determined, and voice volume control processing is performed according to the degree of coincidence between the setting timing and the participation behavior timing. For example, when it is determined that the degree of coincidence is high, the volume of the voice in the viewer's participation behavior is increased, or when it is determined that the degree of coincidence is low, the volume of the voice in the viewer's participation behavior is decreased. In this way, it becomes possible to realize voice volume control according to the degree of coincidence between the setting timing of the viewing content and the timing of the viewer's participation behavior by voice. For example, when the degree of coincidence between the setting timing of the viewing content and the timing of the viewer's participation behavior by voice is high, if the voice volume is increased, the viewer will perform a participation behavior (donation) by voice at a timing that matches the setting timing so that the volume of their own voice becomes large. Therefore, elements of music games such as rhythm games can be added to the distribution of viewing content, improving the viewer's degree of enthusiasm and immersion, and enabling the realization of a content distribution system of a type that has never existed before. For example, if a plurality of viewers perform participation behaviors by voice in accordance with the setting timing of the viewing content, a sense of unity among the viewers will also be generated, making it possible to further enliven the live distribution of the viewing content.

[0086] FIG. 15 is a flowchart for explaining a process of controlling the volume of voice based on the number of viewers participating. First, the content distribution system acquires the number of other viewers N who have taken a participation action at a timing corresponding to the timing of the participation action of the viewer (step S51). That is, it acquires the number of other viewers N who have participated at the same timing as the timing of the participation action of the viewer. For example, the number of participants N can be obtained by counting the number of other viewers who have taken a participation action within a determination period including the timing of the participation action of the viewer. Then, the content distribution system determines whether the number of other viewers N is less than K (step S52), and if N < K, it increases the volume of the voice (step S53). For example, if K = 10, in the case of a situation where the number of participants N is less than 10, in order to liven up the viewing content, the volume of the voice in the participation actions of the viewer and other viewers is increased. Also, when N ≥ K, the content distribution system determines whether the number of participants N is greater than L (L > K) (step S54), and if N > L, it decreases the volume of the voice (step S55). For example, if L = 100, in the case of a lively situation where the number of participants N is more than 100, in order to prevent the volume of the voice from becoming too large, the volume of the voice in the participation actions of the viewer and other viewers is decreased.

[0087] As described above, in this embodiment, a voice volume control process is performed according to the number of other viewers who have taken a participation action at a timing corresponding to the timing of the participation action of the viewer. By doing so, it becomes possible to realize appropriate voice volume control according to the number of participants in the participation action by voice for the viewing content. For example, when the number of participants is small, by increasing the volume of the voice, the cheering and the like for the viewer content become louder, and even when the number of participants is small, it becomes possible to liven up the live distribution of the viewing content by cheering and the like. On the other hand, when the number of participants is large, by decreasing the volume of the voice, it becomes possible to suppress the situation where the cheering and the like for the viewer content become too loud, and it is possible to prevent a situation where each viewer feels that the voice of the participation action is too loud and annoying.

[0088] Although the present embodiment has been described in detail as above, those skilled in the art will easily understand that many modifications can be made without substantially departing from the novel matters and effects of the present disclosure. Therefore, all such modified examples are intended to be included within the scope of the present disclosure. For example, in the specification or drawings, a term (such as "throwing money") described at least once together with a broader or synonymous different term (such as "participation action") can be replaced with the different term anywhere in the specification or drawings. Also, the distribution process, reception process, content of the voice, output timing, voice color change process, voice volume control process, etc. are not limited to those described in the present embodiment, and methods equivalent to these are also included in the scope of the present disclosure.

Explanation of Reference Numerals

[0089] 100... Processing unit, 102... Distribution processing unit, 104... Reception unit, 118... Management unit, 120... Display processing unit, 130... Sound processing unit, 170... Storage unit, 172... Content information storage unit, 174... User information storage unit, 196... Communication unit, 200... Processing unit, 260... Operation unit, 262... Interface unit, 270... Storage unit, 280... Information storage medium, 290... Display unit, 292... Sound output unit, 296... Communication unit, 500... Server system, 510... Network, CH... Character, TA1~TA9... Participation action timing, TM, TM1~TMn... Terminal device, TMA, TMA1~TMAm... Viewer terminal, TMP... Distributor terminal, TQ... Output timing, TQ1~TQ3... Representative timing, TS1~TS3... Setting timing, VC1, VC2... Voice, VFL... Virtual field

Claims

1. a distribution processing unit that performs distribution processing for the viewing content to be viewed by a viewer of a viewer terminal; a receiving unit that receives a participation action by voice in a viewer setting or a system setting as a participation action of the viewer with respect to the viewing content; a sound processing unit that performs at least one of a process of changing the content of the audio at the time of the participation action of the viewer, a process of changing the output timing of the audio, a process of changing the tone of the audio, and a process of controlling the volume of the audio, based on the status of the viewing content at the time of the participation action of the viewer; A content distribution system comprising:

2. In claim 1, The sound processing unit A content distribution system characterized in that the audio is changed to audio with content, tone, or volume that corresponds to the melody or genre of the music of the audiovisual content.

3. In claim 1 or 2, The sound processing unit A content distribution system characterized by determining the set timing set in the viewing content and the timing of the viewer's participation behavior, and performing a change process to bring the audio output timing closer to the set timing.

4. In claim 1 or 2, The sound processing unit A content distribution system characterized by determining the timing of the viewer's participation behavior and the timing of the participation behavior of other viewers, and performing a process to change the output timing of the audio so that the audio is output at a representative timing set based on the timing of the viewer's participation behavior and the timing of the participation behavior of the other viewers.

5. In any one of claims 1 to 4, The sound processing unit A content distribution system characterized by performing volume control processing to change at least one of the volume of the audio and the volume of the audio of the other viewers when the audio of the viewer's participation action and the audio of the other viewers' participation action are output at the viewer terminal.

6. In any one of claims 1 to 5, The sound processing unit A content distribution system characterized by determining the set timing set in the viewing content and the timing of the viewer's participation behavior, and performing volume control processing of the audio according to the degree of match between the set timing and the timing of the participation behavior.

7. In any one of claims 1 to 6, The sound processing unit A content distribution system characterized by performing volume control processing of the audio in accordance with the number of other viewers who participated in the participation behavior at a timing corresponding to the timing of the participation behavior of the viewer.

8. In any one of claims 1 to 7, The reception unit A content distribution system characterized in that a voice generated based on a phrase selected or input by the viewer and a tone set in the system is accepted as the voice in the participation action.

9. In any one of claims 1 to 7, The reception unit A content distribution system characterized in that a voice generated based on a phrase selected or input by the viewer and the tone of the viewer's input voice is accepted as the voice in the participation action.

10. a distribution processing unit that performs distribution processing for the viewing content to be viewed by a viewer of a viewer terminal; a receiving unit that receives a participation action by voice in a viewer setting or a system setting as a participation action of the viewer with respect to the viewing content; a sound processing unit that performs at least one of a process of changing the content of the audio at the time of the participation action of the viewer, a process of changing the output timing of the audio, a process of changing the tone of the audio, and a process of controlling the volume of the audio, based on a state of the viewing content at the time of the participation action of the viewer, A program that causes a computer to function.

Citation Information

Patent Citations

  • Music-enhancing game machine, system for indicating enhancing control for music-enhancing game, and computer-readable storage medium storing program for game

    JP1999151380A

  • Method, system and non-transitory computer-readable recording medium for audio feedback during live broadcast

    JP2019079510A

  • Moving image distribution system, moving image distribution method, and moving image distribution program for live distribution of moving image including animation of character object generated based on movement of distribution user

    JP2020036134A

  • Server system and play data community system

    JP2020162882A

  • Communication device, communication method, and communication program

    JP2020166821A