Content Delivery System and Program
The content distribution system allows viewers to participate by voice and processes these interactions to enhance engagement and match auditory effects with the video content, addressing the limitations of conventional systems.
Patent Information
- Application Number
- JP2021059648
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-03-31
AI Technical Summary
Conventional content delivery systems do not provide a mechanism for viewers to participate by voice, limiting interactive engagement with video content.
A content distribution system that includes a distribution processing unit, a reception unit, and a sound processing unit, which enables viewers to participate by voice and processes this participation to change voice content, timing, timbre, or volume based on the situation of the video content.
Enables viewers to participate interactively with video content through voice, enhancing engagement and creating auditory effects that match the content, output timing, timbre, or volume according to the video content's situation.
Smart Images

Figure 0007685857000001 
Figure 0007685857000002 
Figure 0007685857000003
Abstract
Description
Technical Field
[0001] The present invention relates to a content delivery system, a program, and the like.
Background Art
[0002] In recent years, content delivery systems that deliver content via networks such as the Internet have become popular. As a conventional technology of such a content delivery system, for example, a system disclosed in Patent Document 1 is known. Patent Document 1 discloses a GUI (Graphical User Interface) for making a donation, which is an action of a viewer participating in video content, to the distributor of the video content. Further, Patent Document 2 discloses a music performance game machine in which a player enjoys performance operations according to music. Patent Document 2 discloses a method of changing the input sound of an operation unit according to the progress of a song.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional content delivery systems, donations, which are actions of viewers participating, are donations by comments or stamps, and no proposal has been made regarding actions of viewers participating by voice.
[0005] According to some aspects of the present embodiment, it is possible to provide a content delivery system, a program, and the like that can realize the delivery of video content that enables viewers to participate by voice.
Means for Solving the Problems
[0006] One aspect of the present disclosure relates to a content distribution system including: a distribution processing unit that performs distribution processing for a viewer of a viewer terminal to view video content; a reception unit that receives, as participation actions of the viewer with respect to the video content, the participation actions by voice of viewer settings or system settings; and a sound processing unit that performs at least one of a change process of the content of the voice in the participation action, a change process of the output timing of the voice, a change process of the timbre of the voice, and a control process of the volume of the voice, based on the situation of the video content at the timing of the participation action of the viewer. Another aspect of the present disclosure relates to a program that causes a computer to function as each of the above units, or a computer-readable information storage medium storing the program.
[0007] According to one aspect of the present disclosure, since participation actions by voice with respect to video content are received, auditory effects and the like with respect to the video content become possible. Also, based on the situation of the video content at the timing of the participation action, a change process of the content, output timing, or timbre of the voice is performed, or a control process of the volume is performed. Therefore, it is possible to provide a content distribution system or the like that can realize distribution of video content in which the viewer's participation action by voice with content, output timing, timbre, or volume according to the situation of the video content is possible.
[0008] In another aspect of the present disclosure, the sound processing unit may change the voice to a voice with content, timbre, or volume according to the melody or genre of the music of the video content.
[0009] In this way, it becomes possible to output the voice in the participation action with content, timbre, or volume that matches the melody or genre of the music of the video content.
[0010] In another aspect of the present disclosure, the sound processing unit may determine the set timing set for the video content and the timing of the participation action of the viewer, and perform a change process to bring the output timing of the voice closer to the set timing.
[0011] In this way, the voice in the participation behavior can be output at a timing corresponding to the setting timing of the viewing content.
[0012] Also, in one aspect of the present disclosure, the sound processing unit may perform a change process of the output timing of the voice so that the voice is output at a representative timing set based on the timing of the participation behavior of the viewer and the timing of the participation behavior of other viewers.
[0013] In this way, the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers can be aggregated and output at the same representative timing.
[0014] Also, in one aspect of the present disclosure, when the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers are output at the viewer terminal, the sound processing unit may perform a volume control process of changing at least one of the volume of the voice and the volume of the voice of other viewers.
[0015] In this way, it becomes possible to perform a volume control process so that the ratio of the volume of the voice in the participation behavior of the viewer and the volume of the voice in the participation behavior of other viewers becomes an appropriate output ratio.
[0016] Also, in one aspect of the present disclosure, the sound processing unit may perform a volume control process of the voice according to the degree of coincidence between the set timing set in the viewing content and the timing of the participation behavior of the viewer.
[0017] In this way, it becomes possible to realize volume control of the voice according to the degree of coincidence between the set timing of the viewing content and the timing of the participation behavior by the voice of the viewer.
[0018] In another aspect of the present disclosure, the sound processing unit may perform volume control processing of the sound according to the number of other viewers who have taken the participation action at a timing corresponding to the timing of the participation action of the viewer.
[0019] In this way, it becomes possible to realize volume control of the sound according to the number of participants in the participation action by sound for the viewing content.
[0020] In another aspect of the present disclosure, the reception unit may receive, as the sound in the participation action, sound generated based on a phrase selected or input by the viewer and a tone color of the system setting.
[0021] In this way, the phrase selected or input by the viewer can be output as the sound in the participation action with the tone color of the system setting.
[0022] In another aspect of the present disclosure, the reception unit may receive, as the sound in the participation action, sound generated based on a phrase selected or input by the viewer and the tone color of the input sound of the viewer.
[0023] In this way, the phrase selected or input by the viewer can be output as the sound in the participation action with the tone color based on the input sound of the viewer.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
BEST MODE FOR CARRYING OUT THE INVENTION
[0025] Hereinafter, the present embodiment will be described. Note that the present embodiment described below does not unduly limit the contents described in the claims. Also, not all of the configurations described in the present embodiment are essential constituent elements.
[0026] 1. Content distribution system First, with reference to FIGS. 1(A) to 1(F), a hardware device for realizing the content distribution system of the present embodiment will be described.
[0027] In FIG. 1(A), a server system 500 (information processing system) is communicatively connected to terminal devices TM1 to TMn via a network 510. For example, the server system 500 is a host, and the terminal devices TM1 to TMn are clients. Note that the content distribution system and its processing according to the present embodiment may be realized by the server system 500, or may be realized by distributed processing of the server system 500 and the terminal devices TM1 to TMn.
[0028] Also, the content distribution system and processing according to the present embodiment can also be realized by a blockchain method. For example, each process of the content distribution system according to the present embodiment may be executed using a program called a smart contract executable on Ethereum. In this case, the terminal devices TM1 to TMn will be connected peer-to-peer. Also, various types of information such as content information communicated between the terminal devices TM1 to TMn will be transferred using a blockchain. Note that hereinafter, each of the terminal devices TM1 to TMn will be appropriately described as the terminal device TM.
[0029] The server system 500 can be realized by, for example, one or more servers (such as a management server, a content distribution server such as a game distribution server or a video distribution server, a billing server, a service providing server, an authentication server, a database server, or a communication server, etc.). This server system 500 provides various services for operating content distribution, and can manage data necessary for distributing viewing content, and distribute a client program and various data, etc. Thereby, a viewer can access the server system 500 by using a terminal device TM which is a viewer terminal, and can view the viewing content provided from the server system 500. Also, by the processing of the server system 500, a distribution function for distributing viewing content to a viewer, a participation action function (posting function) that enables a viewer's participation action (posting) such as paying money for the viewing content, etc. are realized. A distributor is, for example, a performer of viewing content, etc. Also, by the processing of the server system 500, an online shopping function for virtual electronic media such as a billing item, a game providing function that enables playing an online game, a user management function for registering users and managing user-specific information, etc. are realized.
[0030] The network 510 (distribution network, communication line) is, for example, a communication path using the Internet, a wireless LAN, etc., and can include a dedicated line (dedicated cable) for direct connection, a LAN using Ethernet (registered trademark), etc., as well as communication networks such as a telephone communication network, a cable network, and a wireless LAN. Also, the communication method can be either wired or wireless.
[0031] The terminal device TM (user terminal) is a terminal having, for example, a network connection function (Internet connection function). As these terminal devices TM, for example, portable communication terminals such as smartphones and mobile phones shown in FIG. 1(B), portable game devices shown in FIG. 1(C), home game devices (stationary type) shown in FIG. 1(D), business game devices shown in FIG. 1(E), or information processing devices such as personal computers (PCs) and tablet PCs shown in FIG. 1(F) can be used. Alternatively, as the terminal device TM, wearable devices (HMD, watch-type devices, etc.) worn on parts such as the user's head and arm may be used.
[0032] FIG. 2 shows a configuration example of the content distribution system of this embodiment. Note that the configuration of the content distribution system is not limited to FIG. 2, and various modifications such as omitting a part of its components (each part) or adding other components are possible.
[0033] The content distribution system includes a processing unit 100, a storage unit 170, and a communication unit 196. This content distribution system can be realized, for example, by the server system 500 in FIG. 1(A), and is communicatively connected to the distributor terminal TMP and the viewer terminals TMA to TMAm, which are terminal devices TM, via the network 510. Hereinafter, each of the viewer terminals TMA1 to TMAn will be generically referred to as the viewer terminal TMA as appropriate.
[0034] The processing unit 100 (processor) performs distribution processing, reception processing, management processing, display processing, or sound processing, etc., based on various information, programs, or operation information stored in the storage unit 170.
[0035] Each process (each function) of the present embodiment performed by each part of the processing unit 100 can be realized by a processor (a processor including hardware). For example, each process of the present embodiment can be realized by a processor that operates based on information such as a program and a memory that stores information such as a program. The processor may be, for example, a processor in which the functions of each part are realized by individual hardware, or the functions of each part may be realized by integrated hardware. For example, the processor includes hardware, and the hardware can include at least one of a circuit for processing digital signals and a circuit for processing analog signals. For example, the processor can also be composed of one or more circuit devices (such as ICs, etc.) mounted on a circuit board and one or more circuit elements (such as resistors, capacitors, etc.). The processor may be, for example, a CPU (Central Processing Unit). However, the processor is not limited to the CPU, and various processors such as a GPU (Graphics Processing Unit) or a DSP (Digital Signal Processor) can be used. Also, the processor may be a hardware circuit by an ASIC. Also, the processor may include an amplifier circuit or a filter circuit for processing analog signals. The memory (storage unit) may be a semiconductor memory such as SRAM or DRAM, or may be a register. Or it may be a magnetic storage device such as a hard disk drive (HDD), or an optical storage device such as an optical disk device. For example, the memory stores instructions readable by a computer, and when the instructions are executed by the processor, the processes (functions) of each part of the processing unit 100 are realized. The instructions here may be an instruction set constituting a program or instructions for instructing operations to the hardware circuit of the processor.
[0036] The processing unit 100 includes a distribution processing unit 102, a reception unit 104, a management unit 118, a display processing unit 120, and a sound processing unit 130. Note that the configuration of the processing unit 100 is not limited to this, and various modifications such as omitting some of these components or adding other components are possible.
[0037] The distribution processing unit 102 performs various distribution processes. Specifically, the distribution processing unit 102 performs a distribution process for the viewer of the viewer terminal TMA to view the viewing content. For example, the distribution processing unit 102 performs a transmission process of information using the communication unit 196. Specifically, the distribution processing unit 102 performs a process of transmitting information to the viewer terminal TMA and the distributor terminal TMP via the communication unit 196 and the network 510. The information transmitted to the viewer terminal TMA is, for example, information on the viewing content. The information transmitted to the distributor terminal TMP is, for example, feedback information from the viewer terminal TMA. The information on the viewing content is, for example, information for displaying a display image (video) on the display unit 290 of FIG. 3 described later, or information for outputting output sound such as voice, music, or sound effects from the sound output unit 292. For example, the distribution processing unit 102 performs streaming distribution of the viewing content. Streaming distribution is video distribution by streaming processing. Alternatively, the distribution processing unit 102 may distribute the viewing content to the viewer terminal TMA using reproduction information for reproducing the viewing content at the viewer terminal TMA. The reproduction information is, for example, input information input by the distributor using the distributor terminal TMP, or information for reproducing the performance of the distributor who is the performer. The distribution processing unit 102 can also perform a reception process of information using the communication unit 196. Specifically, the distribution processing unit 102 performs a process of receiving information from the distributor terminal TMP and the viewer terminal TMA via the communication unit 196 and the network 510. The information received from the distributor terminal TMP is, for example, information on the viewing content distributed by the distributor. The information received from the viewer terminal TMA is, for example, input information input by the viewer using the viewer terminal TMA, and is, for example, participation action information of the viewer with respect to the viewing content. The participation action information is also called viewer post information.
[0038] The distribution processing unit 102 may also perform various content processes. For example, the distribution processing unit 102 may perform a process of generating information on viewing content that progresses based on the operation of the distributor terminal TMP by the distributor. That is, it performs a process of advancing the viewing content and a process of generating information on the viewing content. For example, the distribution processing unit 102 may perform a game process for the user to play a game and generate information on game content that is viewing content. The game process is, for example, a process of starting the game when the game start condition is satisfied, a process of advancing the started game, a process of ending the game when the game end condition is satisfied, or a process of calculating a game result such as a game score. Taking a browser game as an example, the distribution processing unit 102 controls the progress of the game for each user by managing various information of the user for each user. The user information is stored in the user information storage unit 174. For example, a web page constituting a website that provides a game service is displayed on the terminal device TM such as the distributor terminal TMP or the viewer terminal TMA in response to a request from the terminal device TM. Specifically, the web page is displayed by a web browser provided in the terminal device TM. When a hyperlink on the displayed web page is selected by the user, new HTML data corresponding to the hyperlink is transmitted to the terminal device TM, and a web page based on the new HTML data is displayed on the terminal device TM. In this way, by sequentially providing the web page to the terminal device TM according to the user's operation, it becomes possible to advance the game based on the user's operation on the terminal device TM. In this case, the display image information generated by the distribution processing unit 102 is, for example, HTML data or the like.
[0039] The reception unit 104 performs various reception processes. For example, the reception unit 104 performs a reception process for information input by the distributor using the distributor terminal TMP or a reception process for information input by the viewer using the viewer terminal TMA. Details of the reception unit 104 will be described later.
[0040] The management unit 118 performs, for example, user authentication processing. For example, it performs authentication processing for a user who logs in using the terminal device TM. This authentication processing is performed based on, for example, a password or account information input by the user. The management unit 118 also performs various charging processes. For example, it performs processes such as determination processing for charging, creation processing for charging data, and storage processing. The management unit 118 also performs various management processes. For example, it performs management processes for various services and management processes for various information. The management unit 118 can be realized by, for example, a management server.
[0041] For example, in order for a user to use services provided by the server system 500 shown in FIG. 1(A) etc., the user performs a predetermined procedure to obtain an account. By inputting the password associated with the obtained account and logging in, the user can use various services such as live distribution, playing social games and online games, services on live distribution sites and game sites, online shopping for items etc., message exchange between users, and registration of friend users. The management unit 118 also performs management processing etc. for such user account information.
[0042] The display processing unit 120 performs processing for displaying an image on the display unit 290 of the terminal device TM in FIG. 3, which is the distributor terminal TMP or the viewer terminal TMA. For example, display image information (data for image generation) such as HTML data is transmitted to the terminal device TM via the communication unit 196 and the network 510, and processing for displaying an image corresponding to the display image information on the display unit 290 of the terminal device TM is performed.
[0043] The sound processing unit 130 performs processing for outputting sound from the sound output unit 292 of the terminal device TM. For example, output sound information (data for sound generation) is transmitted to the terminal device TM via the communication unit 196 and the network 510, and processing for outputting a sound (voice, music, game sound, effect sound) corresponding to the output sound information from the sound output unit 292 of the terminal device TM is performed.
[0044] The memory unit 170 serves as a working area for components such as the processing unit 100 and the communication unit 196, and its functions can be realized by semiconductor memories, HDDs, SSDs, optical disk devices, etc. The memory unit 170 includes a content information storage unit 172 and a user information storage unit 174. The content information storage unit 172 stores information on the viewing content that is the target of content distribution. The viewing content is entertainment content, such as game content, virtual reality (VR) content, or video distribution content. For example, the viewing content is live content. Examples of live content include game live content (game demonstration content) that broadcasts the gameplay of the distributor, and performance live content that broadcasts the performances of real-world idols, singers, bands, actors, etc. and their corresponding VR characters. The information of the viewing content is various information for the viewer to view the viewing content on the viewer terminal TMA, such as display image information (video information), output sound information, or content sequence information for the progress of the content. The user information storage unit 174 stores various information about the user. For example, the user information storage unit 174 stores the user's personal information (name, gender, date of birth, email address, etc.) as user information. For example, the user's account information (user ID) is also stored as user information. For example, the billing information subject to the billing process is associated with each account information of each user.
[0045] The communication unit 196 communicates with external devices, and its functions can be realized by hardware such as a communication ASIC or a communication processor, and communication firmware. For example, the communication unit 196, which is a communication interface, performs various communication processes for communicating with terminal devices TM such as the distributor terminal TMP and the viewer terminal TMA via the network 510.
[0046] Also, the content distribution system in FIG. 2 performs each process of this embodiment based on the program of this embodiment. This program is a program for causing a computer (a device including an operation unit, a processing unit, a storage unit, and an output unit) to function as each unit of this embodiment (a program for causing the computer to execute the processing of each unit). This program is stored, for example, in an information storage medium. That is, the content distribution system of this embodiment performs various processes of this embodiment based on the program (data) stored in the information storage medium. An information storage medium, which is a medium readable by a computer, stores programs, data, etc., and its function can be realized by an optical disk, an HDD, a semiconductor memory, etc. Note that the program (data) for causing a computer to function as each unit of this embodiment may be distributed from the information storage medium of the server system 500 (host device) via the network 510. The use of such an information storage medium by the server system 500 can also be included within the scope of this embodiment.
[0047] FIG. 3 shows a configuration example of the terminal device TM. The terminal device TM includes a processing unit 200, an operation unit 260, an interface unit 262, a storage unit 270, an information storage medium 280, a display unit 290, a sound output unit 292, and a communication unit 296. Note that the configuration of the terminal device TM is not limited to FIG. 3, and various modifications such as omitting a part of its components (each unit) or adding other components are possible.
[0048] The processing unit 200 (processor) executes terminal-side processing in the content distribution system based on operation information from the operation unit 260, a program, etc. For example, it executes processing for content distribution, game processing, etc. The processing unit 200 can be realized by a processor or the like, similar to the processing unit 100 in FIG. 2 described above. Note that each process of the content distribution system of this embodiment may be realized by distributed processing between the server system 500 and the terminal device TM.
[0049] The operation unit 260 is for the user to input various information such as operation information, and its functions can be realized by operation buttons, direction keys, analog sticks, levers, various sensors (angular velocity sensors, acceleration sensors, etc.), microphones, or touch panel displays. The interface unit 262 performs interface processing with external devices, for example, performs processing to communicate with external devices according to a predetermined interface standard. Also, the interface unit 262 performs interface processing with portable information storage media such as IC cards (memory cards), USB memories, or magnetic cards in which various information about the user is stored. The functions of the interface unit 262 can be realized by hardware such as an ASIC for interface processing or a processor for interface processing, or by firmware for interface processing.
[0050] The storage unit 270 serves as a work area for the processing unit 200, the communication unit 296, etc., and its functions can be realized by semiconductor memories, HDDs, SSDs, optical disk devices, etc. The information storage medium 280 (a medium readable by a computer) stores programs, data, etc., and its functions can be realized by optical disks, HDDs, semiconductor memories, etc. The processing unit 200 performs various processes of this embodiment based on the programs (data) stored in the information storage medium 280. A program (a program for causing a computer to execute the processes of each part) for causing a computer (a device including an operation unit, a processing unit, a storage unit, and an output unit) to function as each part of this embodiment can be stored in this information storage medium 280.
[0051] The display unit 290 outputs the image generated according to this embodiment, and its function can be realized by an LCD, an organic EL display, a CRT, an HMD, or the like. The sound output unit 292 outputs the sound generated according to this embodiment, and its function can be realized by a speaker, headphones, or the like. The communication unit 296 (communication interface) communicates with external devices such as the server system 500 and other terminal devices via the network 510, and its function can be realized by hardware such as a communication ASIC or a communication processor, or communication firmware.
[0052] As shown in FIG. 2, the content distribution system (server system) of this embodiment includes a distribution processing unit 102, a reception unit 104, and a sound processing unit 130.
[0053] The distribution processing unit 102 performs distribution processing for a viewer of the viewing content to view on the viewer terminal TMA. For example, the distribution processing unit 102 performs a process of transmitting information of the viewing content to the viewer terminal TMA. Further, the distribution processing unit 102 may perform a process of receiving information of the viewing content from the distributor terminal TMP.
[0054] The reception unit 104 performs a process of receiving the viewer's participation actions with respect to the viewing content. The viewer's participation actions with respect to the viewing content are not just actions where the viewer simply watches the viewing content, but actions where the viewer actively takes actions with respect to the viewing content to exert an influence. Due to the viewer's participation actions, for example, the content such as the image or sound of the viewing content changes, or the content progression changes. This participation action is performed based on, for example, the input information entered by the viewer using the viewer terminal TMA, and the reception unit 104 performs a process of receiving the viewer's input information regarding this participation action. For example, in a content distribution system, a viewer of the viewing content performs participation actions such as posting comments or providing items such as stamps to praise the viewing content or the distributor, or to support the distributor who distributes the viewing content. Such posting of comments and providing of items are sometimes called "tipping" after the donations from the audience to street performers, or are called "gifts" and the like.
[0055] And in this embodiment, the reception unit 104 performs a process of receiving, as the viewer's participation actions with respect to the viewing content, participation actions by voice in viewer settings or system settings. That is, the reception unit 104 receives the viewer's participation actions by voice instead of items such as stamps. For example, the reception unit 104 receives voice tipping by the viewer. The voice in the viewer settings is, for example, a voice based on the viewer's own voice. This voice in the viewer settings may be a voice obtained by performing processing such as voice synthesis on the voice input by the viewer. The voice in the system settings is a voice prepared by the system (content distribution system) (default voice), for example, a voice held as voice data in the storage unit 170 by the system. The voice in the system settings may be a voice obtained by performing processing such as voice synthesis on the voice prepared in advance by the system. Also, the voice that becomes the viewer's participation action may be a voice based on both the viewer settings and the system settings.
[0056] Then, the sound processing unit 130 performs processing for changing the content, output timing, or timbre of the voice in the participation action, or controlling the volume. For example, based on the situation of the viewing content at the timing of the viewer's participation action, the sound processing unit 130 performs at least one of the following: processing for changing the content of the voice in the participation action, processing for changing the output timing of the voice, processing for changing the timbre of the voice, and processing for controlling the volume of the voice. The situation of the viewing content at the timing of the participation action is the situation of the viewing content in a given period including the timing of the participation action. For example, the timing when the reception unit 104 receives the viewer's participation action is the timing of the participation action, and for example, the situation of the viewing content in a certain period before and after the participation action timing is determined. The situation of the viewing content includes the situation of the content of the viewing content, the progress of the viewing content, the situation regarding the viewer's participation action with respect to the viewing content, the situation of the data set in the viewing content, or the situation of the distribution device or viewing device of the viewing content, etc. The situation of the content of the viewing content is, for example, the situation of the content such as the music or video that constitutes the viewing content. The content of the voice is the content or type of the message represented by the voice, etc. For example, when the situation of the viewing content is the first situation, the sound processing unit 130 outputs the voice of the first content (first message) as the voice in the participation action, and when it is the second situation, it outputs the voice of the second content (second message). Also, when the situation of the viewing content is the first situation, the sound processing unit 130 outputs the voice in the participation action at the first output timing, and when it is the second situation, it outputs the voice in the participation action at the second output timing. Also, when the situation of the viewing content is the first situation, the sound processing unit 130 performs processing to output the voice of the first timbre as the voice in the participation action, and when it is the second situation, it outputs the voice of the second timbre. Also, when the situation of the viewing content is the first situation, the sound processing unit 130 performs processing to output the voice in the participation action at the first volume, and when it is the second situation, it outputs the voice at the second volume.
[0057] Also, the sound processing unit 130 changes the voice in the participation behavior into a voice with content, timbre, or volume according to the melody or genre of the music of the video content. For example, when the music of the video content is the first melody or the first genre, the sound processing unit 130 outputs the voice of the content (phrase) corresponding to the first melody or the first genre, and when it is the second melody or the second genre, as the voice in the participation behavior, it outputs the voice of the content (phrase) corresponding to the second melody or the second genre. This can be realized, for example, by associating and storing in the storage unit 170 the information on the content (phrase) of the voice in the participation behavior with respect to the melody or genre of the music of the video content. Also, when the music of the video content is the first melody or the first genre, the sound processing unit 130 outputs the voice of the first timbre as the voice in the participation behavior, and when it is the second melody or the second genre, it performs the process of outputting the voice of the second timbre. This can be realized, for example, by associating and storing in the storage unit 170 the setting information on the timbre of the voice in the participation behavior with respect to the melody or genre of the music of the video content. Also, when the music of the video content is the first melody or the first genre, the sound processing unit 130 outputs the voice in the participation behavior at the first volume, and when it is the second melody or the second genre, it performs the process of outputting at the second volume. This can be realized, for example, by associating and storing in the storage unit 170 the setting information on the volume of the voice in the participation behavior with respect to the melody or genre of the music of the video content. For example, when the music of the video content is a music with a quiet melody or a quiet genre, the sound processing unit 130 reduces the volume of the voice in the participation behavior, and when the music of the video content is a music with a good tempo or a noisy melody or genre, it performs the process of increasing the volume of the voice in the participation behavior.
[0058] Also, the audio processing unit 130 determines the set timing of the viewing content and the timing of the viewer's participation behavior, and performs a change process to bring the audio output timing closer to the set timing. For example, for the viewing content, a plurality of set timings are set in advance. The information of these set timings is stored, for example, in the storage unit 170 (content information storage unit 172) in association with the viewing content. Examples of the set timing include a call timing set as the timing when audio such as a call is output in the viewing content. Then, the audio processing unit 130 compares each of the plurality of set timings of the viewing content with the timing of the viewer's participation behavior, and searches for, for example, a set timing close to the timing of the participation behavior from among the plurality of set timings. Then, the audio processing unit 130 performs a change process to correct the audio output timing so that the audio output timing approaches the searched set timing. Then, the audio processing unit 130 controls so that the audio in the participation behavior is output at the output timing after the change process. For example, when a plurality of viewers perform a participation behavior, a change process is performed to correct the audio output timing in the participation behavior of the plurality of viewers so that the audio output timings in the participation behaviors of these plurality of viewers are aggregated to the set timing.
[0059] Also, the audio processing unit 130 determines the timing of the viewer's participation behavior and the timing of the participation behavior of other viewers, and performs a process of changing the output timing of the audio so that the audio is output at a representative timing set based on the timing of the viewer's participation behavior and the timing of the participation behavior of other viewers. For example, the audio processing unit 130 determines the timing of the viewer's participation behavior and the timing of the participation behavior of one or more other viewers, and sets a representative timing from these participation behavior timings. The representative timing may be the timing of the participation behavior of any one of a plurality of viewers including the viewer and other viewers, or may be a timing obtained based on the average value of the participation behavior timings of the plurality of viewers. Then, the audio processing unit 130 controls so that the audio in the viewer's participation behavior and the audio in the participation behavior of other viewers are output from the viewer terminal TMA or the like at this representative timing.
[0060] Also, when the audio in the viewer's participation behavior and the audio in the participation behavior of other viewers are output at the viewer terminal, the audio processing unit 130 performs a volume control process of changing at least one of the volume of the viewer's audio and the volume of the audio of other viewers. For example, the audio processing unit 130 performs a volume control process of increasing the volume of one of the volume of the viewer's audio and the volume of the audio of other viewers and decreasing the volume of the other. For example, at the viewer terminal, the volume of the viewer's audio is increased or the volume of the audio of other viewers is decreased. On the other hand, at the viewer terminal of other viewers, the volume of the audio of other viewers is increased or the volume of the viewer's audio is decreased.
[0061] Also, the sound processing unit 130 determines the set timing of the video content and the timing of the viewer's participation behavior, and performs voice volume control processing according to the degree of coincidence between the set timing and the timing of the participation behavior. For example, the sound processing unit 130 performs a matching degree determination process for determining whether the set timing of the video content and the timing of the viewer's participation behavior match within a given range. Whether the set timing and the timing of the participation behavior match within a given range can be determined, for example, by detecting whether the viewer has performed a participation behavior within a given period (judgment period) including the set timing. And when the set timing and the timing of the participation behavior match within a given range, the sound processing unit 130 performs volume control processing to increase the volume of the voice in the viewer's participation behavior, for example. Also, when the set timing and the timing of the participation behavior do not match within a given range, the sound processing unit 130 performs volume control processing to decrease the volume of the voice in the viewer's participation behavior, for example.
[0062] Also, the sound processing unit 130 performs voice volume control processing according to the number of other viewers who have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior. Whether other viewers have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior can be determined, for example, by detecting whether other viewers have performed a participation behavior within a given period (judgment period) including the timing of the viewer's participation behavior. And the sound processing unit 130 performs volume control processing to increase or decrease the volume of the voice according to the number of other viewers who have performed a participation behavior at the timing corresponding to the timing of the viewer's participation behavior. For example, volume control processing is performed to increase or decrease the volume of the voice according to the number of viewers participating within a given period. For example, when the number of participants is small, the sound processing unit 130 increases the volume of each viewer's voice in order to achieve a boosting effect. Also, when the number of participants is large, the sound processing unit 130 decreases the volume of each viewer's voice so that the overall volume does not become too large.
[0063] In addition, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer and the tone color of the system setting. For example, the reception unit 104 performs processing to display, on the viewer terminal, a screen for setting the phrase and tone color of the voice in the participation action. Then, the reception unit 104 receives the phrase selected by the viewer from among the plurality of phrases of the system setting displayed on the screen, or receives the phrase of the viewer setting input by the viewer on the screen. In addition, the reception unit 104 receives the tone color selected by the viewer from among the tone colors of the system setting displayed on the screen. Then, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer in this way and the tone color of the system setting. For example, the voice obtained by reproducing (uttering) the phrase based on the tone color of the system setting is received as the voice in the participation action.
[0064] In addition, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer and the tone color of the viewer's input voice. For example, the reception unit 104 performs processing to display, on the viewer terminal, a screen for setting the phrase and tone color of the voice in the participation action. Then, the reception unit 104, in the same manner as above, receives the phrase selected by the viewer from among the plurality of phrases of the system setting displayed on the screen, or receives the phrase of the viewer setting input by the viewer on the screen. In addition, the reception unit 104 obtains the tone color of the viewer's input voice by allowing the viewer to input their own voice as the viewer-set voice on the screen. Then, the reception unit 104 receives, as the voice in the participation action, the voice generated based on the phrase selected or input by the viewer in this way and the tone color of the viewer's input voice. For example, the voice obtained by reproducing (uttering) the phrase based on the tone color of the viewer's input voice is received as the voice in the participation action.
[0065] Note that each process of the content delivery system of FIG. 2 and the processes of the present embodiment such as the delivery process, the reception process, the display process, and the sound process described above can be implemented by the server system 500 in FIG. 1(A), by the terminal device, or by various modified implementations such as being implemented by distributed processing between the server system 500 and the terminal device. For example, the server system may process only the information necessary to perform each process of the present embodiment, or may perform only the process of transmitting the information to the terminal device, and the terminal device may execute other processes. For example, a program for performing each process of the present embodiment may be installed in the terminal device, and the terminal device may execute each process of the present embodiment based on the installed program. Further, the content delivery system may be realized by an information processing system other than the server system 500.
[0066] 2. Method of the present embodiment Next, the method of the present embodiment will be described in detail.
[0067] 2.1 Participation behavior in video content by voice In the content delivery system, there is provided a function of displaying, during distribution, in a form visible to other viewers, comments, stamps, etc. with more luxurious visual effects than usual due to charging or the like. Such comments and stamps are called donations, and viewers can perform participation actions (posts) such as making donations to give reactions such as support to the distributor who is the performer. And the participation actions by donations and the like up to now have been performed by posting comments or inserting items such as stamps. In contrast, in the present embodiment, participation by voice is possible as a participation action in video content, and donation by voice is possible. That is, in the present embodiment, an auditory donation function that can produce a sense of unity among users enjoying entertainment, such as a live call, is added to the donation function of the conventional content delivery system. Thereby, it becomes possible to enhance the interest of entertainment by the content delivery system.
[0068] For example, FIG. 4 shows an example of a live stream by the VR character CH. In this live stream, in the virtual field VFL of the virtual space, the VR character CH (model object) performs various performances. For example, the VR character CH performs performances such as dancing, waving hands, jumping, and clapping hands. Although FIG. 4 shows an example of a live stream by the VR character CH, the live stream may be a live stream by a real-world idol or the like.
[0069] In such a live stream, viewers perform participation actions using voice VC1 and VC2 such as "Yeah!" and "Fufu". That is, instead of making donations by comments or stamps, viewers make donations (broadly speaking, participation actions) using voice VC1 and VC2 set by the viewer or the system. These voice VC1 and VC2 are output as voices on the viewer terminals of the viewers who made the voice donations, for example. Also, voice VC1 and VC2 are output as voices on the viewer terminals of other viewers. Also, voice VC1 and VC2 may be output as voices on the distributor terminal, or may be output as voices at the live stream venue, for example. Making donations by comments or stamps adds a visual production effect to the viewing content, while the voice donations in this embodiment add an auditory production effect to the viewing content. For example, at the call timing of the live stream song, each viewer makes a voice donation corresponding to the call, or at the exciting time such as the refrain of the song, each viewer makes a voice donation corresponding to the cheer, thereby enhancing the sense of unity among the viewers and realizing production effects such as livening up the live stream.
[0070] FIG. 5 shows an example of a method for setting voice for viewer participation behavior. When making a voice donation, a screen as shown in FIG. 5 is displayed. As shown in A1, phrases that are candidates for system settings such as “Yeah!”, “Huff Huff”, “Thank you~”, and “Amazing” prepared by the system (content distribution system) are displayed. The viewer selects a desired reading phrase from among these system-setting phrases. Also, in the screen of FIG. 5, as shown in A2, tones that are candidates for system settings such as boy voice A, boy voice B, young man voice A, young man voice B, and lady voice prepared by the system are displayed. The viewer selects a desired tone from among these system-setting tones. Then, for example, as shown in A3, by setting the charging amount, etc., and performing the operation shown in A4, a voice donation is made. For example, in A1 of FIG. 5, the phrase “Huff Huff” is selected, and in A2, the tone of young man voice A is selected. Therefore, a voice donation of the phrase “Huff Huff” played (spoken) in the tone of young man voice A is made.
[0071] FIG. 6, FIG. 7(A), and FIG. 7(B) show other examples of the method for setting voice for viewer participation behavior. For example, in FIG. 6, instead of selecting from system-setting phrases as in FIG. 5, the viewer inputs a desired reading phrase as a viewer-setting phrase by themselves. For example, the viewer inputs the characters of the phrase using the operation unit of the viewer terminal, etc. Also, in FIG. 6, similar to FIG. 5, the viewer selects a desired tone from among the tones that are candidates for system settings prepared in advance by the system. By doing so, a voice donation of the viewer-setting phrase input by the viewer played in the system-setting tone is made.
[0072] In FIG. 7(A), similar to FIG. 5, the viewer selects a desired phrase from among the phrases of the system setting candidates. Also, the viewer inputs input voice by recording his or her own voice. By this input voice of the viewer, the timbre of the voice of the participation action is set. By doing so, the coin insertion of the voice reproduced with the timbre of the viewer setting based on the input voice of the viewer is performed for the phrase of the system setting. On the other hand, in FIG. 7(B), the viewer inputs the characters of the phrase using the operation unit of the viewer terminal or the like and inputs the input voice by recording his or her own voice. By doing so, the coin insertion of the voice reproduced with the timbre of the viewer setting based on the input voice of the viewer is performed for the phrase of the viewer setting.
[0073] As described above, in this embodiment, the voice generated based on the phrase selected or input by the viewer and the timbre of the system setting is received as the voice in the participation action. For example, in FIG. 5, the voice generated based on the phrase of the system setting selected by the viewer and the timbre of the system setting is received as the voice in the participation action. Also, in FIG. 6, the voice generated based on the phrase of the viewer setting input by the viewer and the timbre of the system setting is received as the voice in the participation action. By doing so, it becomes possible to output the phrase selected or input by the viewer as the voice in the participation action with the timbre of the system setting. For example, by using the timbre of the system setting, the timbre of the voice in the participation action can be unified with the timbre of the system setting. As a result, for example, it becomes possible to realize an auditory effect with a voice having a timbre suitable for the content of the video content or the like.
[0074] Also, in this embodiment, the voice generated based on the phrase selected or input by the viewer and the timbre of the viewer's input voice is accepted as the voice in the participation action. For example, in FIG. 7(A), the voice generated based on the system setting phrase selected by the viewer and the viewer setting timbre based on the viewer's input voice is accepted as the voice in the participation action. Also, in FIG. 7(B), the voice generated based on the viewer setting phrase input by the viewer and the viewer setting timbre based on the viewer's input voice is accepted as the voice in the participation action. In this way, it becomes possible to output the phrase selected or input by the viewer as the voice in the participation action with the timbre based on the viewer's input voice. For example, by using the timbre based on the viewer's input voice, it becomes possible to make the timbre of the voice in the participation action the timbre based on the viewer's voice. Therefore, the viewer can feel as if they are cheering, etc. with their own voice, and it becomes possible to realize the participation action in the viewing content with a realistic voice that reflects the viewer's voice.
[0075] 2.2 Processing Example of this Embodiment FIG. 8 is a flowchart for explaining the processing of the present embodiment. First, the content distribution system determines whether a viewer has performed a voice-based participation action (tipping) as a participation action for the viewing content (step S1). If such a participation action has been performed, the content distribution system accepts the voice-based participation action by the viewer (step S2). For example, in the screens shown in FIGS. 5 to 7(B), when the viewer sets voice phrases, timbres, etc. and performs a voice-based participation action for the viewing content, this participation action is accepted. This voice is the voice of the viewer settings or system settings as described in FIGS. 5 to 7(B). Then, the content distribution system acquires the situation of the viewing content at the timing of the viewer's participation action (step S3). The situation of the viewing content at the timing of the viewer's participation action is the situation of the viewing content during a given period (for example, a certain period before and after the timing) including the timing of the participation action. Also, the content distribution system acquires, as the situation of the viewing content, the situation of the content of the viewing content, the progress situation, the participation action situation, the setting data situation, or the situation of the distribution device or viewing device, etc. Then, based on the situation of the viewing content at the timing of the viewer's participation action, the content distribution system performs at least one of the following: a change process of the content of the voice in the participation action, a change process of the output timing of the voice, a change process of the timbre of the voice, and a control process of the volume of the voice. For example, the content distribution system changes the content of the content such as the music or video constituting the viewing content according to the situation of the viewing content, changes the output timing of the voice on the viewer terminal or the distributor terminal, changes the timbre of the voice to various timbres corresponding to the viewing content, or changes the volume of the output of the voice on the viewer terminal or the distributor terminal.
[0076] According to this embodiment, when a viewer performs a voice participation action on the viewing content, this participation action is received. By performing such a voice participation action on the viewing content, an auditory effect on the viewing content becomes possible. And in this embodiment, based on the situation of the viewing content at the timing of the participation action, a change process of the voice content, output timing, or tone color is performed, or a volume control process is performed. Therefore, it becomes possible to realize the distribution of viewing content that enables the viewer's participation action by voice with content, output timing, tone color, or volume according to the situation of the viewing content.
[0077] FIG. 9 is a flowchart for explaining a process of controlling the volume of voice according to the tune or genre of the music of the viewing content. First, the content distribution system acquires the tune or genre of the music of the viewing content (step S11). For example, the content distribution system acquires whether the tune of the music of the viewing content is a quiet tune, a lively tune, a bright tune, a dark tune, a ballad tune, a sad tune, a happy tune, an intense tune, a gentle tune, a warm tune, a beautiful tune, or the like. Also, the content distribution system acquires whether the genre of the viewing content is pop, rock, idol, punk, hip-hop, club, hard rock, Latin, classical, easy listening, or new age, or the like. Then, the content distribution system changes the voice in the participation action to a voice with content, tone color, or volume according to the tune or genre (step S12).
[0078] Thus, in this embodiment, the voice of the participation action is changed to a voice with content, timbre, or volume corresponding to the melody or genre of the music of the video content. By doing so, it becomes possible to output the voice of the participation action with content, timbre, or volume that matches the melody or genre of the music of the video content, and it is possible to prevent a situation where the voice of the participation action is output with content, timbre, or volume that does not match the melody or genre. For example, when the music of the content is a quiet melody, the content (phrase) or timbre of the voice that matches the quiet melody is changed, or the volume is decreased. When the music has a lively melody, the content or timbre of the voice that matches the lively melody is changed, or the volume is increased. Also, when the genre of the music is rock, the content or timbre of the voice that matches rock is changed, or the volume is increased. When the genre of the music is classical, the content or timbre of the voice that matches classical is changed, or the volume is decreased. Note that for the information of the music of the video content, metadata indicating the melody and genre is set, and data for setting the content (phrase), timbre, or volume of the voice is associated with this metadata and stored in the storage unit 170. The content distribution system changes the voice in the participation action to a voice with content, timbre, or volume corresponding to the melody or genre based on the metadata associated with the music of the video content and the setting information of the content, timbre, or volume of the voice associated with the metadata. Also, as described with reference to FIGS. 5 to 7(B), when the content distribution system displays candidates for phrases or timbre in the system settings on the screen, the phrases or timbre associated with the melody or genre of the music of the video content are displayed as candidate phrases or timbre.
[0079] FIG. 10 is a flowchart for explaining a process of controlling the output timing of audio based on the setting timing of video content and the participation action timing of viewers. First, the content distribution system identifies a setting timing (call timing) close to the timing of the viewer's participation action (coin insertion) from among a plurality of setting timings set for the video content (step S21). For example, a process of searching for a setting timing close to the timing of the viewer's participation action is performed from among the setting timings of the video content. Then, the content distribution system changes the output timing of the audio in the participation action so as to approach the identified setting timing (step S22). For example, in FIG. 11, setting timings TS1, TS2, TS3, etc. are set for the video content. In FIG. 11, the direction of t is the direction of the time axis. When the viewer's participation action is detected, a setting timing close to the timing of the viewer's participation action is searched from among the setting timings TS1, TS2, TS3, etc. of the video content. In FIG. 11, since the setting timing TS2 is searched as the setting timing close to the timing of the viewer's participation action, the output timing TQ of the audio in the participation action is changed so as to approach the setting timing TS2. As a result, the audio "Fuffhoo" is output at the timing corresponding to the setting timing TS2. The timing corresponding to the setting timing TS2 is, for example, the same timing as the setting timing TS2 or a timing close to the setting timing TS2.
[0080] Thus, in this embodiment, the setting timing set for the video content and the timing of the viewer's participation action are judged, and a change process is performed to make the output timing of the audio approach the setting timing. In this way, the audio in the participation action can be output at the timing corresponding to the setting timing of the video content. Thereby, for example, it is possible to prevent the audio in the participation action from being output at a timing unrelated to the setting timing. Therefore, it is possible to prevent a situation where the production effect and atmosphere of the video content are impaired due to the audio in the participation action being output at a timing unrelated to the setting timing.
[0081] In this embodiment, the timing of the participation behavior of the viewer and the timing of the participation behavior of other viewers are judged, and based on the timing of the participation behavior of the viewer and the timing of the participation behavior of other viewers, at the representative timing set, the voice in the participation behavior is output, and the voice output timing is changed. For example, TA1, TA2, TA3, TA4, TA5, TA6, TA7, TA8, TA9 in FIG. 12 are the timings of the participation behaviors of viewers AD1, AD2, AD3, AD4, AD5, AD6, AD7, AD8, AD9, respectively. In this case, TQ1 is set as the representative timing of the participation behavior timings TA1, TA2, TA3 of viewers AD1, AD2, AD3. Also, TQ2 is set as the representative timing of the participation behavior timings TA4, TA5, TA6 of viewers AD4, AD5, AD6, and TQ3 is set as the representative timing of the participation behavior timings TA7, TA8, TA9 of viewers AD7, AD8, AD9. For example, the representative timing of the group of viewers AD1 to AD3 is set to TQ1, the representative timing of the group of viewers AD4 to AD6 is set to TQ2, and the representative timing of the group of viewers AD7 to AD9 is set to TQ3. Then, the voice in the participation behavior of viewers AD1 to AD3 is output at the representative timing TQ1, the voice in the participation behavior of viewers AD4 to AD6 is output at the representative timing TQ2, and the voice in the participation behavior of viewers AD7 to AD9 is output at the representative timing TQ3. For example, when viewer AD1 is regarded as his own viewer and viewers AD2, AD3 are regarded as other viewers, at the representative timing TQ1 set based on the timing of the participation behavior of viewer AD1 and the timing of the participation behavior of other viewers AD2, AD3, the voice in the participation behavior of viewer AD1 will be output. In this way, the voice in the participation behavior of viewer AD1 and the voices in the participation behaviors of other viewers AD2, AD3 are aggregated and output at the same representative timing TQ1. Therefore, it is possible to prevent the situation where the voices in the participation behaviors of viewer AD1 and other viewers AD2, AD3 are output at scattered timings, resulting in a lack of unity in the output of the voices in the participation behavior.When setting timings TS1, TS2, and TS3 are set for the viewing content as shown in FIG. 11, change processing may be performed to bring the representative timings TQ1, TQ2, and TQ3 closer to the setting timings.
[0082] FIG. 13 is a flowchart for explaining processing for changing the volume of the viewer's voice and the voices of other viewers. First, the content distribution system acquires voice data in the participation actions of other viewers (step S31). Then, the content distribution system performs processing for mixing the viewer's voice and the voices of other viewers so that, for example, the volume of the viewer's voice is greater than the volume of the voices of other viewers (step S32). Then, the content distribution system performs processing for outputting the mixed voice from the viewer terminal (step S33).
[0083] Thus, in this embodiment, when the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers are output on the viewer terminal, a volume control process is performed to change at least one of the volume of the viewer's voice and the volume of the voice of other viewers. For example, a volume control process is performed to change the output ratio of the viewer's voice and the voice of other viewers. For example, in step S32 of FIG. 13, the volume of the viewer's voice and the volume of the voice of other viewers are changed so that the volume of the viewer's voice is greater than the volume of the voice of other viewers. On the other hand, on the viewer terminals of other viewers, conversely, the volume of the voice of other viewers and the volume of the viewer's voice are changed so that the volume of the voice of other viewers is greater than the volume of the viewer's voice. In this way, the volume control process is performed so that the ratio of the volume of the voice in the participation behavior of the viewer and the volume of the voice in the participation behavior of other viewers becomes an appropriate output ratio, and the voice in the participation behavior of the viewer and the voice in the participation behavior of other viewers can be output to the viewer terminal. Therefore, for example, it is possible to prevent the occurrence of a situation such as the voice of other viewers becoming too loud and the volume balance being disrupted. For example, when a plurality of viewers including the viewer and other viewers view viewing content, in order to improve the production effect, it is desirable to output the voices of other viewers from the viewer terminal of the viewer. However, for example, when the voices of a large number of other viewers are output to the viewer terminal, a situation may occur in which the viewer cannot hear his or her own voice in the participation behavior. In this regard, for example, if a volume control process is performed so that the volume of the viewer's voice is greater than the volume of the voice of other viewers, the occurrence of such a situation can be prevented.
[0084] FIG. 14 is a flowchart for explaining a process of controlling the volume of voice based on the degree of coincidence between the timing of the viewer's participation action and the setting timing of the viewing content. First, the content distribution system compares the setting timing (call timing) set for the viewing content with the timing of the viewer's participation action (casting money) (step S41). That is, it compares the plurality of setting timings TS1, TS2, TS3, etc. as shown in FIG. 11 with the timing of the viewer's participation action. Then, the content distribution system determines whether the setting timing and the timing of the participation action match within a given range (step S42). For example, it determines whether the timing of the participation action falls within a predetermined determination period including the setting timing. And when the setting timing and the timing of the participation action match within a given range, the content distribution system increases the volume of the voice in the viewer's participation action (step S43), and when they do not match, it decreases the volume of the voice in the viewer's participation action (step S44). For example, it increases or decreases the volume of the voice in the viewer's participation action output from the viewer terminal or the like.
[0085] Thus, in this embodiment, the setting timing set for the viewing content and the timing of the viewer's participation behavior are determined, and voice volume control processing is performed according to the degree of coincidence between the setting timing and the participation behavior timing. For example, when it is determined that the degree of coincidence is high, the volume of the voice in the viewer's participation behavior is increased, or when it is determined that the degree of coincidence is low, the volume of the voice in the viewer's participation behavior is decreased. In this way, it becomes possible to realize voice volume control according to the degree of coincidence between the setting timing of the viewing content and the timing of the viewer's participation behavior by voice. For example, when the degree of coincidence between the setting timing of the viewing content and the timing of the viewer's participation behavior by voice is high, if the voice volume is increased, the viewer will perform a participation behavior (donation) by voice at a timing that matches the setting timing so that the volume of their own voice becomes large. Therefore, elements of music games such as rhythm games can be added to the distribution of viewing content, and the degree of enthusiasm and immersion of the viewer can be improved, making it possible to realize a type of content distribution system that has never existed before. For example, if multiple viewers perform participation behaviors by voice in accordance with the setting timing of the viewing content, a sense of unity among the viewers will also be generated, making it possible to further liven up the live distribution of the viewing content.
[0086] FIG. 15 is a flowchart for explaining a process of controlling the volume of sound based on the number of viewers participating. First, the content distribution system acquires the number of other viewers N who have taken a participation action at a timing corresponding to the timing of the participation action of the viewer (step S51). That is, it acquires the number of other viewers N who have participated at a timing similar to the timing of the participation action of the viewer. For example, the number of participants N can be obtained by counting the number of other viewers who have taken a participation action within a determination period including the timing of the participation action of the viewer. Then, the content distribution system determines whether the number of other viewers N is less than K (step S52). If N < K, the volume of the sound is increased (step S53). For example, if K = 10, in a situation where the number of participants N is less than 10 and it is idle, in order to liven up the viewing content, the volume of the sound in the participation actions of the viewer and other viewers is increased. Also, when N ≥ K, the content distribution system determines whether the number of participants N is greater than L (L > K) (step S54). If N > L, the volume of the sound is decreased (step S55). For example, if L = 100, in a lively situation where the number of participants N is more than 100, in order to prevent the volume of the sound from becoming too large, the volume of the sound in the participation actions of the viewer and other viewers is decreased.
[0087] Thus, in this embodiment, a sound volume control process is performed according to the number of other viewers who have taken a participation action at a timing corresponding to the timing of the participation action of the viewer. By doing so, it becomes possible to realize appropriate sound volume control according to the number of participants in the participation action by sound for the viewing content. For example, when the number of participants is small, by increasing the volume of the sound, the cheering etc. for the viewer content becomes larger, and even when the number of participants is small, it becomes possible to liven up the live distribution of the viewing content by cheering etc. On the other hand, when the number of participants is large, by decreasing the volume of the sound, it becomes possible to suppress the situation where the cheering etc. for the viewer content becomes too large, and it is possible to prevent a situation where each viewer feels that the sound of the participation action is too loud and annoying.
[0088] Although the present embodiment has been described in detail as above, those skilled in the art will easily understand that many modifications can be made without substantially departing from the novel matters and effects of the present disclosure. Therefore, all such modifications are intended to be included within the scope of the present disclosure. For example, in the specification or drawings, a term (such as "throwing money") described at least once together with a broader or synonymous different term (such as "participation action") can be replaced with the different term anywhere in the specification or drawings. Also, the distribution process, reception process, content of the voice, output timing, voice color change process, voice volume control process, etc. are not limited to those described in the present embodiment, and methods equivalent thereto are also included in the scope of the present disclosure.
Explanation of Reference Numerals
[0089] 100... Processing unit, 102... Distribution processing unit, 104... Reception unit, 118... Management unit, 120... Display processing unit, 130... Sound processing unit, 170... Storage unit, 172... Content information storage unit, 174... User information storage unit, 196... Communication unit, 200... Processing unit, 260... Operation unit, 262... Interface unit, 270... Storage unit, 280... Information storage medium, 290... Display unit, 292... Sound output unit, 296... Communication unit, 500... Server system, 510... Network, CH... Character, TA1~TA9... Participation action timing, TM, TM1~TMn... Terminal device, TMA, TMA1~TMAm... Viewer terminal, TMP... Distributor terminal, TQ... Output timing, TQ1~TQ3... Representative timing, TS1~TS3... Setting timing, VC1, VC2... Voice, VFL... Virtual field
Claims
1. A distribution processing unit that performs distribution processing for a viewer of a viewer terminal to view video content; A receiving unit that receives, as the participation action of the viewer with respect to the video content, the participation action by voice of viewer settings or system settings; Based on the situation of the video content at the timing of the participation action of the viewer, at least one of a change process of the content of the voice in the participation action, a change process of the output timing of the voice, a change process of the tone color of the voice, and a control process of the volume of the voice is performed by a sound processing unit; A content distribution system characterized by including the above.
2. In Claim 1, The sound processing unit A content distribution system characterized by changing the voice into a voice with content, tone color, or volume according to the melody or genre of the music of the video content.
3. In Claim 1 or 2, The sound processing unit A content distribution system characterized by determining the set timing set in the video content and the timing of the participation action of the viewer, and performing a change process of bringing the output timing of the voice closer to the set timing.
4. In Claim 1 or 2, The sound processing unit A content distribution system characterized by determining the timing of the participation action of the viewer and the timing of the participation action of another viewer, and performing a change process of the output timing of the voice so that the voice is output at a representative timing set based on the timing of the participation action of the viewer and the timing of the participation action of the other viewer.
5. In any one of Claims 1 to 4, The sound processing unit A content distribution system characterized by performing a volume control process of changing at least one of the volume of the voice and the volume of the voice of another viewer when the voice in the participation action of the viewer and the voice in the participation action of another viewer are output at the viewer terminal.
6. In any one of Claims 1 to 5, The sound processing unit A content distribution system characterized by determining the set timing set in the video content and the timing of the participation action of the viewer, and performing a volume control process of the voice according to the degree of coincidence between the set timing and the timing of the participation action.
7. In any one of Claims 1 to 6, The sound processing unit A content delivery system characterized by performing volume control processing of the voice according to the number of other viewers who performed a participation action at a timing corresponding to the timing of the participation action of the viewer.
8. In any one of Claims 1 to 7, The receiving unit, A content delivery system characterized by receiving, as the voice in the participation action, a voice generated based on a phrase selected or input by the viewer and the tone color of the system setting.
9. In any one of Claims 1 to 7, The receiving unit, A content delivery system characterized by receiving, as the voice in the participation action, a voice generated based on a phrase selected or input by the viewer and the tone color of the input voice of the viewer.
10. A distribution processing unit that performs distribution processing for a viewer of a viewer terminal to view viewing content, A receiving unit that receives, as a participation action of the viewer with respect to the viewing content, the participation action by voice of the viewer setting or the system setting, Based on the situation of the viewing content at the timing of the participation action of the viewer, as a sound processing unit that performs at least one of change processing of the content of the voice in the participation action, change processing of the output timing of the voice, change processing of the tone color of the voice, and control processing of the volume of the voice, A program characterized by causing a computer to function.
Citation Information
Patent Citations
Music-enhancing game machine, system for indicating enhancing control for music-enhancing game, and computer-readable storage medium storing program for game
JP1999151380A
Voice control device
JP2013195481A
Information terminal device, distribution management device, system, program, and recording medium
JP2018182546A
Server, method, program, and system
JP2019165311A
Information processing apparatus, moving image distribution method, and moving image distribution program
JP2020017870A