Content delivery system, content delivery method, and program

The virtual podcast system addresses the challenge of integrating content and comments in podcast programs by using user service contracts and text-to-speech synthesis, facilitating efficient distribution and monetization while managing copyright.

JP7831563B2Active Publication Date: 2026-03-17SONY GROUP CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing content delivery systems do not facilitate easy integration of content with related comments, particularly in podcast programs, and creators face challenges in managing copyright and monetization.

Method used

A virtual podcast system that utilizes user service contracts and text-to-speech synthesis to provide content with comments, allowing creators to distribute podcast programs efficiently while handling copyright processing and monetization through VPC-type distribution.

Benefits of technology

Enables easy creation and distribution of podcast programs with integrated comments, simplifies copyright management, and ensures fair compensation for creators by leveraging user rights and text-to-speech technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007831563000001
    Figure 0007831563000001
  • Figure 0007831563000002
    Figure 0007831563000002
  • Figure 0007831563000003
    Figure 0007831563000003
Patent Text Reader

Abstract

To easier provide a content and a comment on it.SOLUTION: The present invention provides a content provision system comprising a control part in which a script formed by at least text information in regard to identification information of a content and an advertisement is stored in a predetermined storage medium so as to enable a browsing by a user; in accordance with the script selected by the user, a reading of a content indicated by content identification information contained in their script is controlled so that the content is executed and provided to the user by utilizing the rights that the user has already acquired via a contract with a specific service; and the text information contained in the script is formed by a voice synthesis, and is provided to the user at least either before or after a provision of the content.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a content providing system, a content providing method, and a program, and particularly relates to a content providing system, a content providing method, and a program that enable content and its comments to be provided more easily.

Background Art

[0002] In recent years, with the diversification of methods for providing content, various services and devices have been provided (for example, see Patent Documents 1 and 2).

[0003] Patent Document 1 discloses a device that automatically selects a stream to be played and outputs it to a television monitor based on sequence information for controlling the playback order of the stream. In this device, the stream and the character string to be output are combined and output according to the sequence information.

[0004] Patent Document 2 discloses a program that functions to sequentially extract elements from hypertext downloaded based on an address retrieved from a program list, generate voice by voice synthesis when there is text, and sequentially repeat performing an output corresponding to the material at the link destination when there is a link.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] By the way, when providing content along with comments about that content, it is desirable to provide the content and comments in a simpler way.

[0007] This technology was developed in light of these circumstances, and aims to make it easier to provide content and comments. [Means for solving the problem]

[0008] One aspect of this technology is a content delivery system in which a script consisting of content identification information and text information related to advertisements is stored on a predetermined storage medium in a manner that is accessible to the user, and the system controls the retrieval of content indicated by the content identification information contained in the script, according to the script selected by the user, using rights already acquired by the user through a contract with a specific service, and provides the content to the user, and the script contains Text information relating to the aforementioned advertisement This content delivery system includes a control unit that synthesizes speech and controls the system to provide the content to the user either before or after its delivery.

[0009] One aspect of this technology is a content delivery method in which at least a script consisting of content identification information and text information related to advertisements is stored on a predetermined storage medium in a manner that is accessible to the user, and the method controls the reading of content indicated by the content identification information contained in the script, according to the script selected by the user, and provides it to the user by executing the execution using rights already acquired by the user through a contract with a specific service, and the script contains Text information relating to the aforementioned advertisement This is a content delivery method that synthesizes speech and controls whether it is provided to the user either before or after the content is delivered.

[0010] One aspect of this technology is a program that controls a computer to read content indicated by the content identification information contained in a script, which is stored on a predetermined storage medium accessible to the user, and to provide the content to the user by using the rights already acquired by the user through a contract with a specific service, according to the script selected by the user, and the script contains Text information relating to the aforementioned advertisement This program synthesizes speech and functions as a control unit that manages whether to provide the content to the user either before or after its delivery.

[0011] In one aspect of this technology, a content delivery system, content delivery method, and program, at least a script consisting of content identification information and text information relating to advertisements is stored on a predetermined storage medium in a manner accessible to the user, and the system is controlled to deliver to the user the content indicated by the content identification information contained in the script, according to the script selected by the user, by using the rights already acquired by the user through a contract with a specific service, and the script contains Text information relating to the aforementioned advertisement The speech is synthesized and controlled to be provided to the user either before or after the content is delivered. [Brief explanation of the drawing]

[0012] [Figure 1] This is a representative diagram illustrating the overview of this technology. [Figure 2] This diagram shows an overview of a content delivery system that applies this technology. [Figure 3] This diagram shows the content playback flow using a content delivery system that applies this technology. [Figure 4] This diagram shows an example of rights management when distributing content that includes music. [Figure 5] This diagram shows an example of rights management when performing VPC-type distribution. [Figure 6]This is a diagram showing an example of the overall configuration of a content providing system to which this technology is applied. [Figure 7] This is a diagram showing an example of a script used in a content providing system. [Figure 8] This is a diagram showing an example of the configuration of an embodiment of a content providing system to which this technology is applied. [Figure 9] This is a diagram showing an example of the configuration of a creator terminal device. [Figure 10] This is a diagram showing an example of the functional configuration of a control unit in a creator terminal device. [Figure 11] This is a diagram showing an example of the configuration of a user terminal device. [Figure 12] This is a diagram showing an example of the functional configuration of a control unit in a user terminal device. [Figure 13] This is a diagram showing an example of the configuration of a distribution server. [Figure 14] This is a diagram showing an example of the functional configuration of a control unit in a distribution server. [Figure 15] This is a sequence diagram showing the processing flow when voice synthesis is used to provide a preamble and a postscript of a piece of music together with the programmed music. [Figure 16] This is a diagram showing a first example of the user interface of a program creation tool. [Figure 17] This is a diagram showing an example of a script of a program generated by a program creation tool. [Figure 18] This is a diagram showing a second example of the user interface of a program creation tool. [Figure 19] This is a diagram showing a second example of the user interface of a program creation tool. [Figure 20] This is a diagram showing a second example of the user interface of a program creation tool. [Figure 21] This is a diagram showing a second example of the user interface of a program creation tool. [Figure 22] This is a diagram showing a second example of the user interface of a program creation tool. [Figure 23]This figure shows a second example of a user interface for a program creation tool. [Figure 24] This sequence diagram shows the processing flow when a song is presented as a program and played as a playlist, provided that the introductory and concluding remarks for the song are available using speech synthesis. [Figure 25] This is a sequence diagram illustrating the processing flow when using live voices to provide introductory and concluding remarks for songs that have been featured on a television program. [Figure 26] This sequence diagram shows the processing flow when a song is presented as a program and then played as a playlist, provided that the intro and outro of the song are delivered using live vocals. [Figure 27] This is a flowchart illustrating the overall process in the first embodiment. [Figure 28] This is a flowchart illustrating the overall process in the first embodiment. [Figure 29] This is a sequence diagram showing the processing flow when providing a script to other music distribution services. [Figure 30] This is a flowchart illustrating the overall process in the second embodiment. [Figure 31] This is a flowchart illustrating the overall process in the second embodiment. [Figure 32] This figure shows another example of the functional configuration of the control unit in a distribution server. [Figure 33] This is a sequence diagram showing the processing flow when performing a text check. [Figure 34] This is a flowchart illustrating the overall process in the third embodiment. [Figure 35] This figure shows another example of a configuration of one embodiment of a content delivery system to which this technology is applied. [Figure 36] This figure shows examples of advertisements inserted into a program. [Figure 37] This is a sequence diagram showing the processing flow when inserting advertisements into a program. [Figure 38]This is a flowchart illustrating the overall process in the fourth embodiment. [Figure 39] This figure shows another example of a configuration of one embodiment of a content delivery system to which this technology is applied. [Figure 40] This is a sequence diagram showing the processing flow when managing song IDs and sharing program information. [Figure 41] This is a flowchart illustrating the overall process in the modified example. [Figure 42] This is a flowchart illustrating the overall process in the modified example. [Modes for carrying out the invention]

[0013] The embodiments of this technology will be described below with reference to the drawings. The description will be given in the following order.

[0014] 1. First Embodiment: Basic Configuration 2. Second Embodiment: Integration Function with Other Services 3. Third Embodiment: Minimum License Function 4. Fourth Embodiment: Advertising Function 5. Variations 6. Computer Configuration

[0015] <Representative diagram>

[0016] Figure 1 is a representative diagram illustrating the overview of this technology.

[0017] This technology makes it easier to provide content and related comments when creating a program from content, by utilizing user service contracts and text-to-speech synthesis to provide the content along with comments about that content.

[0018] In Figure 1, the DJ streams their selected tracks from a music distribution server and also provides commentary on those tracks using a microphone. Meanwhile, the user listens to the DJ's selected tracks and commentary, which are streamed from the music distribution server.

[0019] Here, the DJ is considered a virtual entity created by the creator, but the music selected by the DJ is provided to the user through a music streaming service that the user already subscribes to, and the comments made by the DJ are provided to the user through text-to-speech synthesis, making it easier to deliver content and comments.

[0020] <1. First Embodiment>

[0021] (Overview of the Virtual Podcast System) Figure 2 shows an overview of a content delivery system to which this technology is applied. In the example in Figure 2, a virtual podcast system is illustrated as one embodiment of a content delivery system to which this technology is applied.

[0022] A virtual podcast system is a system that allows creators to produce podcast programs simply by operating their own terminal device, selecting music, and writing text. Podcasts are a method of publishing audio and video data files over the internet, and are a type of internet radio or internet television. Note that the text is not limited to text; it can also be provided as an audio file.

[0023] Podcast programs created by creators are registered on a distribution server. This allows users to listen to the podcast programs by operating their own devices and playing the programs distributed from the distribution server.

[0024] By the way, from the perspective of a podcast creator, they would naturally want to efficiently distribute their podcasts and have them listened to by as many users as possible.

[0025] Furthermore, when distributing music via podcasts, the creator is responsible for handling the copyright of the music, which is a cumbersome process for them. Therefore, they would likely want someone else to handle the copyright processing for them.

[0026] In recent years, video streaming sites have seen creators opening their own video streaming channels and disseminating information through video content on various themes. In return for providing video content to users, creators receive compensation such as advertising revenue based on the number of video views and advertising revenue from producing tie-up videos with advertisers.

[0027] For creators who distribute podcast programs, compensation for their podcasts is an extremely important concern, and they undoubtedly hope to receive fair compensation.

[0028] The above-mentioned aspects of program development, distribution copyright processing, and monetization are unavoidable for creators distributing podcast programs, and there is a need for a system that can easily resolve these issues. A virtual podcast system provides a mechanism that allows creators to create and distribute podcast programs, making those podcast programs available for users to listen to, while also addressing the creators' needs for program development, distribution copyright processing, and monetization.

[0029] Figure 3 shows the playback flow of a podcast program generated by a virtual podcast system.

[0030] Figure 3 shows the Nth and N+1th consecutive tracks in a podcast program.

[0031] Each track consists of a warm-up, a song, and an after-song.

[0032] The introduction is an introduction to the song and consists of text. In this example, the introduction is written as, "This song was written when...it's such a fantastic song!" This text corresponding to the introduction can be converted into speech and read aloud using TTS (Text To Speech).

[0033] A song includes an identification ID to identify the song, as well as information about the song's title and artist. For example, by using a song ID such as "1234567", a user can request streaming of the song identified by that song ID from their subscribed music streaming service.

[0034] The postscript is a commentary written after listening to the song, and consists of text. In this example, the postscript is the text, "It was really great..." This text corresponding to the postscript can be read aloud using TTS (Text-to-Speech).

[0035] When distributing a podcast program, there are two possible scenarios for handling the rights to include music in the program.

[0036] Firstly, there is the case of distributing a program that includes music. In this case, as shown in Figure 4, the creator produces a complete podcast program that includes music and spoken parts (introduction and closing remarks), and the podcast program is distributed. Therefore, copyright processing for the music is incurred by the creator who distributes the program.

[0037] Secondly, there is the case of VPC-type distribution. In this case, as shown in Figure 5, the music is distributed using a music distribution service, and the creator only creates and distributes the spoken parts (introductory and concluding remarks), so the creator does not incur any copyright processing for the music.

[0038] In other words, when a creator distributes a podcast program, the program's structure data, as well as the introduction and conclusion, will be distributed. As a result, the musical portions of the program will be distributed by music distribution services, eliminating the need for copyright processing for the music for the creator.

[0039] In VPC-type distribution, music streamed by music distribution services and spoken content (introductions and closing remarks) delivered by creators are combined on the user's terminal device to create a program. Therefore, the rights clearance for the music portion of the podcast program, which is created on the user's terminal device, is handled by the user.

[0040] In this VPC-type distribution, when distributing a podcast, the creator does not create a complete podcast program, but rather distributes song identification information (song ID). This allows the user's terminal device to play the song streamed by the music distribution service based on that song ID.

[0041] In other words, on the user's terminal device, music is played using the rights that the user has already acquired through their contract with the music distribution service, so no copyright processing is required for the creator. On the other hand, for the user, it is possible to play music within the normal music distribution scope of the music distribution service they have contracted with, so they can play music identified by the music ID specified by the creator without paying any additional fees. In a virtual podcast system, podcast programs are distributed using this VPC (Virtual Podcast) type distribution.

[0042] Furthermore, even if a user has a free user account (free user), they can still use the right to play music on a music streaming service, provided that the service only includes advertisements.

[0043] Figure 6 shows an example of the overall configuration of a virtual podcast system.

[0044] As shown in Figure 6, the functions provided by this virtual podcast system can be broadly divided into creator-side functions provided by the creator terminal device, various distribution service-side functions provided by the distribution server, and user-side functions provided by the user terminal device.

[0045] On the creator terminal device, program creation tools and audio creation tools are executed in response to the creator's (PodCaster's) operations, and a podcast program is generated.

[0046] For example, the program creation tool generates a podcast program based on the song ID of a song selected from a song selection list (catalog) provided by a music distribution service, and the introductory and concluding texts of that song, which have been sound-adjusted during speech synthesis by an audio creation tool, and then registers it with the program distribution service.

[0047] The audio creation tool provides TTS (Text-to-Speech) sound adjustment functions based on audio creation data provided by the audio distribution service. By operating the audio creation tool and utilizing the TTS sound adjustment functions, creators can customize the TTS audio played back by the user to their liking.

[0048] The program distribution service provides a service that delivers podcast programs registered using a program creation tool to user terminal devices.

[0049] The music distribution service corresponds to the music distribution service that a user using a user terminal device has subscribed to. The music distribution service distributes songs identified by the song ID set for the podcast program in response to a request from the user terminal device. The music distribution service also provides the creator terminal device with a song list for selection.

[0050] The audio distribution service provides a service that delivers text-to-speech (TTS) audio, obtained by synthesizing the text of the intro and outro of songs set in a podcast program, to user terminal devices. The audio distribution service also provides data for audio creation to creator terminal devices.

[0051] On the user terminal device, the program renderer is executed in response to user (Listener) operations, and the podcast program is played.

[0052] The program renderer, when playing a desired podcast program from among those published by a program distribution service, renders the music distributed by the music distribution service and the TTS audio distributed by the audio distribution service based on the program's structure data (reproduction data).

[0053] This allows the podcast program to be played (reproduced) and made available for viewing by the user. The program renderer running on the user's terminal device can also be considered a playback player.

[0054] Figure 7 shows an example of a script describing the structure of a podcast program.

[0055] As shown in Figure 7, in the virtual podcast system, a podcast program is composed of multiple sets of song IDs for the songs to be turned into a program, along with the introduction and conclusion of each song. The structure of this podcast program is described by the script shown in Figure 7.

[0056] In Figure 7, the script includes information about the program at the beginning, such as the program title, owner, release date, and the name of the service from which the music is distributed.

[0057] The script contains information about the program followed by information about the tracks. Figure 7 shows an example of the description for the first track out of N tracks.

[0058] Each track contains information about the track number, the warm-up, the song, and the after-song.

[0059] A song contains an identifier (ID) to identify the song, as well as information such as the song's title and artist name. For example, by specifying a song ID of "1234567", it is possible to request the distribution of the song identified by that song ID from a music distribution service called "serviceA".

[0060] The warm-up and after-song sections contain comment information corresponding to comments about the song. For example, the warm-up could read, "This song was written in... it's such a great song!", and the after-song could read, "It really is great...", allowing the text to be converted into speech and read aloud using a text-to-speech (TTS) service.

[0061] Figure 7 shows an example of describing only the first track, i.e., information about the first song. However, for the second and subsequent songs, the description should be the same as for the first song, with a song ID, introduction, and conclusion set for each song.

[0062] In this way, the song ID of a song and a script consisting of an introduction and a conclusion about that song are generated by the creator's terminal device and made available to users when registered with the program distribution service.

[0063] On the other hand, the user terminal device used by the user is controlled to ensure that the song indicated by the song ID is streamed using the rights already acquired by the user through their contract with the music distribution service, in accordance with the script published by the program distribution service, and that the TTS audio for the introductory and concluding remarks is provided.

[0064] In other words, the script only contains the song ID and the introduction and conclusion in text format, and does not contain the actual music or audio data that will be played in the podcast program. However, the user's terminal device reproduces the program created by the creator by playing the music and audio data based on the information indicated by the song ID and introduction and conclusion written in the script.

[0065] Furthermore, by generating a script that adds introductory and concluding remarks to songs (or their song IDs) in an existing playlist, it is possible to turn a playlist into a program. Therefore, users can easily turn a playlist into a program simply by entering the introductory and concluding remarks for the songs.

[0066] (System Configuration) Figure 8 shows the configuration of a virtual podcast system as an example of the configuration of one embodiment of a content delivery system to which this technology is applied.

[0067] In Figure 8, the content provision system 1 consists of a creator terminal device 10, a user terminal device 20, a program distribution server 30A, a music distribution server 30B, and an audio distribution server 30C.

[0068] In the content provision system 1, the creator terminal device 10, the user terminal device 20, the program distribution server 30A, the music distribution server 30B, and the audio distribution server 30C are interconnected via the network 50.

[0069] The creator terminal device 10 is a device such as a smartphone, tablet, or personal computer, and is used by a creator.

[0070] The creator terminal device 10 generates a script for the podcast program in response to the creator's operation and sends (uploads) it to the program distribution server 30A via the network 50.

[0071] The user terminal device 20 is a device such as a smartphone, tablet, music player, game console, or personal computer, and is used by the user.

[0072] The user terminal device 20 accesses the program distribution server 30A via the network 50 in response to user operations and receives (downloads) the script of the podcast program.

[0073] The program distribution server 30A consists of one or more servers that provide program distribution services. The program distribution service is a service that distributes podcast programs and is provided by a program distribution provider.

[0074] The program distribution server 30A receives the program script transmitted (uploaded) from the creator terminal device 10 via the network 50 and registers it on a storage medium so that it can be viewed by users using the user terminal device 20.

[0075] When the program distribution server 30A receives a program playback request transmitted from a user terminal device 20 via the network 50, it reads the script of the program from the storage medium and distributes it to the user terminal device 20 that made the playback request.

[0076] The music distribution server 30B consists of one or more servers that provide music distribution services. Music distribution services are services that distribute music over the internet and are provided by music distribution companies. For example, music distribution services are provided in the form of a flat-rate streaming service with unlimited listening.

[0077] When the music distribution server 30B receives a music distribution request transmitted from a user terminal device 20 via the network 50, it identifies the music corresponding to the received distribution request and distributes the streaming data of that music to the user terminal device 20 that made the distribution request.

[0078] The audio distribution server 30C consists of one or more servers that provide audio distribution services. The audio distribution service is a service that distributes audio such as TTS audio and live voices over the internet, and is provided by an audio distribution service provider.

[0079] When the voice distribution server 30C receives a voice distribution request transmitted from a user terminal device 20 via the network 50, it acquires the voice corresponding to the received distribution request and distributes the voice data to the user terminal device 20 that made the distribution request.

[0080] In the following explanation, unless there is a need to distinguish between the program distribution server 30A, the music distribution server 30B, and the audio distribution server 30C, they will be referred to simply as "distribution server 30." Furthermore, the program distribution provider, the music distribution provider, and the audio distribution provider may be the same provider or different providers.

[0081] Network 50 is comprised of communication networks such as the Internet, intranet, or mobile phone network, and enables interconnection between devices using communication protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol).

[0082] (Configuration of the creator terminal device) Figure 9 shows an example of the configuration of the creator terminal device 10 shown in Figure 8.

[0083] As shown in Figure 9, in the creator terminal device 10, the CPU (Central Processing Unit) 101, ROM (Read Only Memory) 102, and RAM (Random Access Memory) 103 are interconnected by a bus 104.

[0084] The CPU 101 controls the operation of each part of the creator terminal device 10 by executing programs recorded in the ROM 102 and the memory unit 107. Various types of data are stored in the RAM 103 as needed.

[0085] An input / output interface 110 is also connected to bus 104. The input / output interface 110 is connected to an input unit 105, an output unit 106, a storage unit 107, a communication unit 108, and a short-range wireless communication unit 109.

[0086] The input unit 105 supplies various input data to each part, including the CPU 101, via the input / output interface 110. For example, the input unit 105 includes an operation unit 111, a camera unit 112, and a sensor unit 113.

[0087] The control unit 111 is operated by the creator and supplies operation data corresponding to that operation to the CPU 101. The control unit 111 consists of physical buttons, a touch panel, etc.

[0088] The camera unit 112 converts light from an incident subject into photoelectric signals, performs signal processing on the resulting electrical signals, and generates and outputs captured image data. The camera unit 112 consists of an image sensor, a signal processing unit, and the like.

[0089] The sensor unit 113 senses spatial information, temporal information, etc., and outputs sensor data obtained as a result of that sensing.

[0090] The sensor unit 113 includes an accelerometer and a gyroscope. The accelerometer measures acceleration in the three directions of the X, Y, and Z axes. The gyroscope measures angular velocity in the three directions of the X, Y, and Z axes. Alternatively, an inertial measurement unit (IMU) may be provided to measure three-dimensional acceleration and angular velocity using three-directional accelerometers and a three-axis gyroscope.

[0091] Furthermore, the sensor unit 113 can include various sensors such as a sound sensor (microphone) for detecting sounds like the creator's voice, a biosensor for measuring information such as the heart rate, body temperature, or posture of a living organism, a proximity sensor for measuring nearby objects, and a magnetic sensor for measuring the magnitude and direction of a magnetic field.

[0092] The output unit 106 outputs various types of information via the input / output interface 110 in accordance with the control from the CPU 101. For example, the output unit 106 includes a display unit 121 and an audio output unit 122.

[0093] The display unit 121 displays images and other data according to the image data, in accordance with the control from the CPU 101. The display unit 121 consists of a panel section such as a liquid crystal panel or an OLED (Organic Light Emitting Diode) panel and a signal processing unit.

[0094] The sound output unit 122 outputs sound according to sound data in accordance with the control from the CPU 101. The sound output unit 122 consists of a speaker or headphones connected to the output terminal.

[0095] The memory unit 107 records various data and programs according to the control from the CPU 101. The CPU 101 reads various data from the memory unit 107 and processes it, or executes programs.

[0096] The memory unit 107 is configured as an auxiliary storage device such as a semiconductor memory. The memory unit 107 may be configured as internal storage or as external storage such as a memory card.

[0097] The communication unit 108 communicates with other devices via the network 50 according to the control from the CPU 101. The communication unit 108 is configured as a communication module that supports cellular communication (e.g., LTE-Advanced or 5G), wireless communication such as Wi-Fi (Local Area Network), or wired communication.

[0098] The short-range wireless communication unit 109 performs wireless communication using short-range wireless communication standards such as Bluetooth (registered trademark) and NFC (Near Field Communication) to exchange various types of data.

[0099] Note that the configuration of the creator terminal device 10 shown in Figure 9 is just one example; for example, a microphone may be provided as an input unit, or an image processing circuit such as a GPU (Graphics Processing Unit) or a power supply circuit may be provided.

[0100] Figure 10 shows an example of the functional configuration of the control unit 100 in the creator terminal device 10. The functions of the control unit 100 are realized by the execution of programs such as program creation tools and audio creation tools by the CPU 101.

[0101] In Figure 10, the control unit 100 includes an input receiving unit 151, a music information acquisition unit 152, a program generation unit 153, an audio information acquisition unit 154, an audio generation unit 155, and a registration unit 156.

[0102] The input receiving unit 151 receives operation data corresponding to the creator's operations, supplied from the input unit 105, and supplies it to the program generation unit 153.

[0103] The music information acquisition unit 152 acquires music information related to songs supplied from the communication unit 108, which communicates with the music distribution server 30B, and supplies it to the program generation unit 153. The music information includes information such as the song list and song ID received from the music distribution server 30B.

[0104] The program generation unit 153 generates a podcast program script by processing the music information supplied from the music information acquisition unit 152 and the comment information regarding the introduction and conclusion, based on the operation data supplied from the input reception unit 151, and supplies it to the registration unit 156.

[0105] The voice information acquisition unit 154 acquires voice information related to the introductory and concluding speeches supplied from the communication unit 108, which communicates with the voice distribution server 30C, and supplies it to the voice generation unit 155. The voice information includes information such as information related to the speech during speech synthesis and speech creation received from the voice distribution server 30C.

[0106] The audio generation unit 155 processes the audio information supplied from the audio information acquisition unit 154 to generate audio for the creator to set the introduction and closing remarks, and supplies it to the program generation unit 153.

[0107] When generating a podcast program, the program generation unit 153 uses the audio supplied from the audio generation unit 155 to provide the creator with information regarding the settings for the introduction and conclusion (such as the audio used during speech synthesis), thereby generating a script for the program and supplying it to the registration unit 156.

[0108] The registration unit 156 controls the communication unit 108 to register the program script supplied from the program generation unit 153 by uploading it to the program distribution server 30A via the network 50.

[0109] (Configuration of user terminal device) Figure 11 shows an example of the configuration of the user terminal device 20 in Figure 8.

[0110] In Figure 11, the configuration of the user terminal device 20 corresponds to the configuration of the creator terminal device 10 shown in Figure 9. That is, the CPU 201 to the short-range wireless communication unit 209 have the same functions as the CPU 101 to the short-range wireless communication unit 109 described above, so their explanation is omitted here.

[0111] Figure 12 shows an example of the functional configuration of the control unit 200 in the user terminal device 20. The functions of the control unit 200 are realized by the execution of programs such as a program renderer by the CPU 201.

[0112] In Figure 12, the control unit 200 includes a program acquisition unit 251, a music acquisition unit 252, an audio acquisition unit 253, a renderer unit 254, and a presentation control unit 255.

[0113] The program acquisition unit 251 acquires the podcast program script in response to user operations, which is supplied from the communication unit 208 that communicates with the program distribution server 30A, and supplies it to the renderer unit 254.

[0114] The music acquisition unit 252 acquires streaming data of a song corresponding to a song ID, supplied from the communication unit 208 which communicates with the music distribution server 30B, and supplies it to the renderer unit 254.

[0115] The audio acquisition unit 253 acquires audio data corresponding to the introductory and concluding remarks, supplied from the communication unit 208 which communicates with the audio distribution server 30C, and supplies it to the renderer unit 254.

[0116] The renderer unit 254 performs rendering processing on the introductory audio data supplied from the audio acquisition unit 253, the music streaming data supplied from the music acquisition unit 252, and the concluding audio data supplied from the audio acquisition unit 253, based on the program script supplied from the program acquisition unit 251, and supplies the resulting data to the presentation control unit 255.

[0117] The presentation control unit 255 presents a program to the user by supplying data from the renderer unit 254 to the output unit 206.

[0118] For example, the presentation control unit 255 can supply the audio data for the introduction, the streaming data for the music, and the audio data for the conclusion to the sound output unit 222, thereby outputting the sounds of the introduction and conclusion set for the program, along with the sound of the music that has been made into a program, before and after the music.

[0119] (Configuration of the distribution server) Figure 13 shows an example of the configuration of the distribution server 30 in Figure 8. Note that the distribution server 30 corresponds to one of the following servers: program distribution server 30A, music distribution server 30B, and audio distribution server 30C shown in Figure 8.

[0120] In the distribution server 30, the CPU 301, ROM 302, and RAM 303 are interconnected by a bus 304. An input / output interface 310 is also connected to the bus 304. An input / output interface 305, an output unit 306, a storage unit 307, a communication unit 308, and a drive 309 are connected to the input / output interface 310.

[0121] The input section 305 consists of a microphone, keyboard, mouse, etc. The output section 306 consists of a speaker, display, etc.

[0122] The storage unit 307 consists of a hard disk drive (HDD) and semiconductor memory. The communication unit 308 is configured as a communication module that supports wireless communication such as wireless LAN or wired communication such as Ethernet®.

[0123] The drive 309 drives a removable recording medium 311, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0124] Figure 14 shows an example of the functional configuration of the control unit 300 in the distribution server 30. The functions of the control unit 300 are realized by the execution of the programs for each service by the CPU 301.

[0125] In Figure 14, the control unit 300 includes a request reception / response unit 351, a distribution processing unit 352, and a database 353.

[0126] The request reception / response unit 351 receives various requests supplied from the communication unit 308, which communicates with the creator terminal device 10 or the user terminal device 20, and supplies them to the distribution processing unit 352.

[0127] The distribution processing unit 352 performs distribution processing in response to various requests supplied from the request reception / response unit 351.

[0128] The database 353 is recorded in the storage unit 307, which consists of a large-capacity storage device such as an HDD or semiconductor memory.

[0129] For example, the database 353 of the program distribution server 30A stores the scripts of podcast programs, etc. The database 353 of the music distribution server 30B stores songs provided by music distribution services, associated with song IDs. Furthermore, the database 353 of the audio distribution server 30C stores information related to audio during speech synthesis and audio creation, as well as audio data for introductions and conclusions.

[0130] When performing distribution processing, the distribution processing unit 352 processes various data stored in the database 353, generates responses corresponding to various requests, and supplies them to the request reception / response unit 351.

[0131] The request reception / response unit 351 controls the communication unit 308 to transmit responses corresponding to various requests supplied from the distribution processing unit 352 to the requesting creator terminal device 10 or user terminal device 20 via the network 50.

[0132] Next, we will explain the processing flow performed by each device in the content delivery system 1.

[0133] (Example 1) Figure 15 is a sequence diagram showing the processing flow when using speech synthesis to provide an introduction and a postscript to a song that has been made into a program.

[0134] In Figure 15, the program creation tool is executed by the creator terminal device 10, and the program renderer is executed by the user terminal device 20. Also in Figure 15, the program distribution service is provided by the program distribution server 30A, the music distribution service is provided by the music distribution server 30B, and the TTS service is provided by the audio distribution server 30C.

[0135] In the creator terminal device 10, the program creation tool is executed by the control unit 100, and the processes in steps S11 to S13 are performed.

[0136] The program creation tool retrieves the song list sent from the music distribution server 30B and presents it to the creator (S11).

[0137] The program creation tool generates a podcast program script based on the song ID of the song selected by the creator from the song list and the introductory and concluding texts of the song entered by the creator (S12), and registers it with the program distribution server 30A (S13).

[0138] As a result, the program distribution server 30A stores the scripts of podcast programs created by creators in the database 353, making them accessible to users using the user terminal device 20.

[0139] In the user terminal device 20, the program renderer is executed by the control unit 200, and the program renderer works in cooperation with each distribution server 30 to execute the processes in steps S14 to S25.

[0140] In the program renderer, when a user instructs the program to play a podcast program that is publicly available for viewing on the program distribution server 30A, the program renderer receives the script of the program distributed from the program distribution server 30A (S14, S15).

[0141] The program renderer requests the audio distribution server 30C to synthesize the text of the introductory remarks, based on the introductory remarks set at the beginning of the received script (S16).

[0142] In the audio distribution server 30C, in response to a request from the program renderer, the text of the introductory remarks is synthesized into speech (S17), and the result of that speech synthesis is distributed (S18).

[0143] As a result, the program renderer receives the results of speech synthesis distributed from the audio distribution server 30C, and rendering is performed, which plays the TTS audio for the introductory part set for the song that has been made into a program.

[0144] Next, the program renderer requests the music distribution server 30B, which provides the music distribution service contracted by the user, to distribute the music identified by the music ID, based on the music ID set after the introductory part of the received script (S19).

[0145] In response to a request from the program renderer, the music distribution server 30B verifies the rights acquired by the user through their contract with the music distribution service (S20). If it determines that the user has legitimate rights and that playback of the music identified by the music ID is possible, the music is streamed (S21).

[0146] As a result, the program renderer receives streaming data of songs distributed from the music distribution server 30B, and rendering is performed so that songs identified by their song IDs are played back as songs that have been turned into a program.

[0147] Subsequently, once the streaming of the song has finished playing, the program renderer requests the audio distribution server 30C to synthesize the text of the subsequent commentary, based on the commentary set after the song ID in the received script (S22).

[0148] In the audio distribution server 30C, in response to a request from the program renderer, the text of the postscript is synthesized into speech (S23), and the result of the speech synthesis is distributed (S24).

[0149] As a result, the program renderer receives the results of speech synthesis distributed from the audio distribution server 30C, and rendering is performed to play the TTS audio for the postscript portion set in the program's music.

[0150] Furthermore, since the podcast program script includes multiple song IDs for each song, along with the introductory and concluding texts for each song, after the processing in steps S16 to S24 is completed, the process returns to the processing in step S16 (S25), and the processing in steps S16 to S25 is repeated according to the number of song IDs.

[0151] As a result, the program renderer repeatedly plays the intro, song, and concluding remarks in the order specified in the script for each song ID, allowing the user to listen to the podcast program.

[0152] The above explains the processing flow performed by each device when using speech synthesis to provide introductory and concluding remarks for songs that have been made into programs.

[0153] (Example of a program creation tool's UI) Refer to Figures 16 to 23 to describe the details of the program creation tools performed on the creator terminal device 10.

[0154] Figure 16 shows the first example of a user interface (UI) for a program creation tool.

[0155] In Figure 16, the program creation screen 410 is a screen displayed when the program creation tool is executed, and is a UI for creating a podcast program in response to the creator's operations.

[0156] The program creation screen 410 includes an operation area 411, a title setting area 412, an opening talk setting area 413, a pre-set music / introduction / conclusion area 414, and a music / introduction / conclusion setting area 415.

[0157] The operation area 411 is the area for controlling and listening to the music set for the program. The operation area 411 includes buttons for playing or stopping the music, buttons for selecting the previous and next music, and a seek bar that indicates the position of the currently playing music.

[0158] Title setting area 412 is the area for setting the program title.

[0159] The opening talk setting area 413 is the area for setting the opening talk. For example, the audio file for the opening talk is set in the opening talk setting area 413, but it is not necessary to set it if an opening talk is not required.

[0160] The pre-set song / intro / post-song area 414 is the area where the pre-set songs and their intros and post-songs are displayed.

[0161] For example, in the pre-configured song / intro / post-speech area 414-1, the song "Song1" has an intro "Speech File1" and a post-speech "Speech File2" assigned to it. Similarly, in the pre-configured song / intro / post-speech area 414-2, the song "Song2" has an intro "Speech File3" and a post-speech "Speech File4" assigned to it.

[0162] The song / intro / post-song setting area 415 is the area for setting up songs and their intros and post-songs.

[0163] For example, the music / introduction / conclusion setting area 415 includes an "Add Music from Song Catalog" button for selecting a desired music file from a music catalog, an "Add Speech File before" button for setting a desired intro, and an "Add Speech File after" button for setting a desired conclusion.

[0164] The introductory and concluding remarks will be set as text files corresponding to the creator's input, but they may also be set as audio files corresponding to, for example, the creator's voice input.

[0165] The creator operates this program creation screen 410, which, for example, creates the program script shown in Figure 17.

[0166] In Figure 17, the podcast program is structured so that the opening talk ("Opening Talk File") is followed by the introduction to the first song ("Speech File 1"), the first song ("Song 1"), and the conclusion to the first song ("Speech File 2"), and then the introduction to the second song ("Speech File 3"), the second song ("Song 2"), and the conclusion to the second song ("Speech File 4") are played in that order.

[0167] Incidentally, this program creation tool may also be provided as a function of an application (hereinafter also referred to as a music distribution app) distributed by a music distribution service contracted by the creator. Figures 18 to 23 show examples of program creation tools provided as a function of a music distribution app executed by the creator terminal device 10.

[0168] In Figure 18, the music playback screen 510 is a screen provided as a function of the music distribution application and is a UI for playing music distributed by a music distribution service via a network 50 such as the internet. The music playback screen 510 has a music playback area 511 and a music operation area 512.

[0169] The playback target song area 511 is a region for displaying the title, artist name, and album art of the song being played.

[0170] The music control area 512 is an area for controlling the music. The music control area 512 includes buttons for playing or stopping the music, buttons for selecting the previous and next tracks, a seek bar that indicates the position of the currently playing track, and so on.

[0171] When a predetermined operation is performed by the creator on the music playback screen 510, the playlist selection screen 520 shown in Figure 19 is displayed.

[0172] In Figure 19, the playlist selection screen 520 is a screen provided as a function of the music distribution app and is a UI for selecting a desired playlist. The playlist selection screen 520 has a playlist list area 521.

[0173] Playlist list area 521 is a region for displaying and selecting playlists provided by music streaming services, or playlists created by the creator themselves or other users (public playlists).

[0174] If the "70 SOUL" playlist, enclosed in frame F1 in the diagram, is selected from the playlists displayed in this playlist list area 521, the playlist editing screen 530 shown in Figure 20 will be displayed.

[0175] In Figure 20, the playlist editing screen 530 is a screen provided as a function of the music distribution app and is a UI for editing the selected playlist. The playlist editing screen 530 has a song list area 531 and an add song button 532.

[0176] The song list area 531 is a region that displays a list of songs registered in the selected playlist and allows the user to select a song. The song add button 532 is a button used to add a new song to the selected playlist.

[0177] If a desired song is selected from the songs displayed in the song list area 531, enclosed in frame F2 in the diagram, and the edit button 533 for the selected song is operated, the song / intro / post-song editing screen 540A in Figure 21 or the song / intro / post-song editing screen 540B in Figure 22 will be displayed.

[0178] In Figure 21, the song / intro / conclusion editing screen 540A is a UI for editing the intro and conclusion of the selected song. The song / intro / conclusion editing screen 540A has an intro description area 541A and a conclusion description area 542A.

[0179] The introductory text area 541A is a region for writing introductory text for the selected song.

[0180] For example, if the creator terminal device 10 is a device such as a smartphone with a touch panel, the creator can input the desired introduction as text by tapping the software keyboard displayed on the display unit 121 overlaid with the touch panel. Alternatively, if the creator terminal device 10 can utilize a cloud-based speech recognition API (Application Programming Interface) service, the creator's voice input of the desired introduction may be converted into text using the said speech recognition service.

[0181] Alternatively, if the creator terminal device 10 is a personal computer or other device with a keyboard, the creator can use the keyboard to input the desired introductory comments.

[0182] The postscript area 542A is a region for writing postscript text for the selected song. The desired postscript text is entered into the postscript area 542A in response to operations such as software keyboard input or voice input by the creator.

[0183] Furthermore, in Figure 22, the song / intro / conclusion editing screen 540B is a UI for editing the intro and conclusion of the selected song. The song / intro / conclusion editing screen 540B has an intro description area 541B and a conclusion description area 542B.

[0184] The introductory text area 541B is a region for writing introductory text for the selected song. The desired introductory text is entered into the introductory text area 541B in response to operations such as software keyboard input or voice input by the creator.

[0185] The postscript area 542B is a region for writing postscript text for the selected song. The desired postscript text is entered into the postscript area 542B in response to operations such as software keyboard input or voice input by the creator.

[0186] In this way, creators can turn a playlist into a program by adding introductory and concluding text to each song in the playlist. Note that the song / introductory / concluding text editing screens 540A and 540B described above are just examples of UIs for setting introductory and concluding text for songs, and introductory and concluding text can also be set using other UIs.

[0187] Returning to the explanation of Figure 20, when a predetermined operation is performed by the creator on the playlist editing screen 530, the playlist settings screen 550 shown in Figure 23 is displayed.

[0188] In Figure 23, the playlist settings screen 550 is a screen provided as a function of the music distribution app, and is a UI for making various settings related to the programmed playlist. The playlist settings screen 550 has a settings area 551.

[0189] Settings area 551 includes items for renaming a program-based playlist, making program-based playlists public, and deleting program-based playlists. If the "Public" option, enclosed in box F3 in the diagram, is tapped from the items displayed in this settings area, the program-based playlist will be made public to other users.

[0190] As a result, the script of the program-formatted playlist, that is, the program script, is stored in the database 353 of the program distribution server 30A and made available for viewing by users using the user terminal device 20.

[0191] In this way, a program creation function can be added as a feature of the music streaming app provided by a music streaming service.

[0192] For example, a creator, as a premium user of a music streaming service, can turn a playlist into a program, set an introduction and a closing for each song, and publish the script of the program-like playlist. In this case, the song ID set in the script generated by the music streaming app will be the song ID managed by the music streaming service to which the creator has a contract.

[0193] (Second example) Figure 24 is a sequence diagram showing the processing flow when a program featuring a song is played as a playlist, provided that the song's introduction and conclusion can be provided using speech synthesis.

[0194] In Figure 24, the program creation tool and program renderer are executed on the creator's and user's respective terminal devices, and the program distribution service, music distribution service, and TTS service are provided by their respective distribution servers, as in the first example described with reference to Figure 15.

[0195] In steps S31 to S33 of Figure 24, similar to steps S11 to S13 of Figure 15, the program script is generated by the program creation tool and registered with the program distribution server 30A.

[0196] In the program renderer, when a user instructs the program to play a program that is publicly available for viewing on the program distribution server 30A, the program renderer receives the script for that program distributed from the program distribution server 30A (S34, S35). Here, for example, let's assume that the user has instructed the program to be played as a playlist.

[0197] The program renderer requests the music distribution server 30B, which provides the music distribution service contracted by the user, to distribute the music identified by the music ID set in the received script (S36).

[0198] In response to a request from the program renderer, the music distribution server 30B verifies the rights acquired by the user through their contract with the music distribution service (S37). If it determines that the user has legitimate rights and that playback of the music identified by the music ID is possible, the music is streamed (S38).

[0199] As a result, the program renderer receives streaming data of songs distributed from the music distribution server 30B, and rendering is performed, allowing the song identified by the song ID to be played.

[0200] Furthermore, since the program script lists the song IDs of multiple songs in playback order, after the processing in steps S36 to S38 is completed, the process returns to step S36 (S39), and the processing in steps S36 to S39 is repeated according to the number of song IDs.

[0201] As a result, the program renderer repeatedly plays the songs in the order of the song IDs set in the script, allowing the user to listen to the songs played as a playlist.

[0202] Thus, when a user requests playlist playback, the system reads the song indicated by the song ID included in the script, without reading the introductory and concluding sections of the script, according to the script selected by the user, and provides it to the user.

[0203] The above explains the processing flow performed by each device when a song is presented as a playlist, provided that the introductory and concluding remarks for the song are available using speech synthesis.

[0204] (Third example) Figure 25 is a sequence diagram showing the processing flow when using live voices to provide introductory and concluding remarks for songs that have been made into programs.

[0205] In Figure 25, the program creation tool and program renderer are executed on the creator's and user's respective terminal devices, and the program distribution service and music distribution service are provided by the respective distribution servers, as in the first example described with reference to Figure 15. On the other hand, in the third example in Figure 25, unlike the first example shown in Figure 15, the audio distribution server 30C provides a live voice distribution service instead of a TTS service.

[0206] In steps S51 to S53 of Figure 25, similar to steps S11 to S13 of Figure 15, the program creation tool generates a script for the podcast program, which is then registered with the program distribution server 30A.

[0207] However, in the sequence diagram of Figure 25, the introduction and conclusion of the song are not included as text in the script, but are read aloud by the creator in their own voice, so the audio data (file) is registered with the audio distribution server 30C (S54).

[0208] In other words, the script registered with the program distribution server 30A includes a song ID that identifies the song, along with link information to the audio data of the introductory and concluding remarks.

[0209] In steps S55 and S56 of Figure 25, similar to steps S14 and S15 of Figure 15, the program renderer receives the script of a podcast program distributed from the program distribution server 30A when it is instructed to play a podcast program that is publicly available for viewing on the program distribution server 30A.

[0210] The program renderer accesses the audio distribution server 30C based on the link information for the introductory remarks set at the beginning of the received script and requests the live voice distribution of the introductory remarks (S57).

[0211] In the audio distribution server 30C, in response to a request from the program renderer, the audio data of the introductory speech registered by the creator is processed (S58), and that audio data of the introductory speech is distributed (S59).

[0212] As a result, the program renderer receives the audio data of the introductory speech streamed from the audio distribution server 30C, and after rendering, the introductory speech portion set for the program's music is played back.

[0213] In steps S60 to S62 of Figure 25, similar to steps S19 to S21 of Figure 15, the program renderer plays the music streamed from the music distribution server 30B based on the music ID set after the introductory text of the received script.

[0214] Subsequently, when the streaming of the song finishes playing, the program renderer accesses the audio distribution server 30C based on the link information for the postscript, which is set after the song ID in the received script, and requests the live voice distribution of the said postscript (S63).

[0215] In the audio distribution server 30C, in response to a request from the program renderer, the audio data of the original voice of the narrator registered by the creator is processed (S64), and the audio data of the original voice is distributed (S65).

[0216] As a result, the program renderer receives the audio data of the post-commentary spoken by the original voice from the audio distribution server 30C, and when rendering is performed, the original voice spoken by the original voice from the post-commentary portion set in the program's music is played back.

[0217] Furthermore, since the podcast program script includes the song IDs of multiple songs along with link information for the introductory and concluding remarks of those songs, after the processing in steps S57 to S65 is completed, the process returns to the processing in step S57 (S66), and the processing in steps S57 to S66 is repeated.

[0218] As a result, the program renderer repeatedly plays the intro, song, and concluding remarks in the order specified in the script for each song ID, allowing the user to listen to the podcast program.

[0219] The above explains the processing flow performed by each device when using live voices to provide introductory and concluding remarks for songs that have been made into programs.

[0220] (Fourth example) Figure 26 is a sequence diagram showing the processing flow when a song that has been made into a program is played as a playlist, in which the introductory and concluding remarks of the song can be provided using live voices.

[0221] In Figure 26, the program creation tool and program renderer are executed on the creator's and user's respective terminal devices, and the program distribution service, music distribution service, and live voice distribution service are provided by their respective distribution servers, as in the third example described with reference to Figure 25.

[0222] In steps S71 to S74 of Figure 26, similar to steps S51 to S54 of Figure 25, the program script is registered with the program distribution server 30A by the program creation tool, and the audio data (files) of the introductory and concluding remarks, read aloud by the creator, are registered with the audio distribution server 30C.

[0223] In the program renderer, when a user instructs the program to play a program that is publicly available for viewing on the program distribution server 30A, the program renderer receives the script for that program distributed from the program distribution server 30A (S75, S76). Here, for example, let's assume that the user has instructed the program to be played as a playlist.

[0224] The program renderer requests the music distribution server 30B, which provides the music distribution service contracted by the user, to distribute the music identified by the music ID set in the received script (S77).

[0225] In response to a request from the program renderer, the music distribution server 30B verifies the rights acquired by the user through their contract with the music distribution service (S78). If it determines that the user has legitimate rights and that playback of the music identified by the music ID is possible, the music is streamed (S79).

[0226] As a result, the program renderer receives streaming data of songs distributed from the music distribution server 30B, and rendering is performed, allowing the song identified by the song ID to be played.

[0227] Furthermore, since the program script lists the song IDs of multiple songs in playback order, after the processing in steps S77 to S79 ​​is completed, the process returns to step S77 (S80), and the processing in steps S77 to S80 is repeated.

[0228] As a result, the program renderer repeatedly plays the songs in the order of the song IDs set in the script, allowing the user to listen to the songs played as a playlist.

[0229] Thus, when a user requests playlist playback, the system reads the song indicated by the song ID included in the script, without reading the introductory and concluding sections of the script, according to the script selected by the user, and provides it to the user.

[0230] The above explains the processing flow performed by each device when a song is presented as a playlist, provided that the introductory and concluding remarks are delivered using live voices.

[0231] (Overall picture of the process) Figures 27 and 28 are flowcharts illustrating the overall process in the first embodiment.

[0232] The processes shown in Figures 27 and 28 are realized through the collaborative operation of the creator terminal device 10 (its control unit 100), the user terminal device 20 (its control unit 200), and the distribution server 30 (its control unit 300) in a content provision system to which this technology is applied.

[0233] In other words, this process is performed by at least one of the control units 100, 200, and 300.

[0234] In the content provision system 1, as shown in Figure 27, when input is received from a creator using the creator terminal device 10 (Yes in S211), a script consisting of content identification information and comment information corresponding to the introduction and conclusion of the content is generated based on that input (S212), and this script is stored in a predetermined storage medium so that it can be viewed by a user using the user terminal device 20 (S213).

[0235] Here, "content" includes songs distributed on music streaming services, and the identifying information for that content includes a song ID that can identify the song. Furthermore, the introduction and conclusion are examples of comments about the content, and it is sufficient if at least one of them is set. For example, when creating a program from songs, an introduction or conclusion can be inserted for each song, or a narration or similar is inserted as an introduction, followed by the playback of three songs consecutively, or four songs are played consecutively, followed by comments or impressions of the songs as a conclusion.

[0236] Comment information includes text representing comments, or link information to the creator's own voice. While comment information is described as corresponding to comments related to content, it is not necessarily limited to comments related to content; it may also correspond to comments unrelated to content. The specified storage medium can be, for example, the storage unit 307 (and its database 353) of the program distribution server 30A.

[0237] Furthermore, in the content provision system 1, as shown in Figure 28, when a user using the user terminal device 20 requests playback of a script stored on a predetermined storage medium (Yes in S231), the system first reads the introduction according to the comment information contained in the script and controls it to provide it to the user (S232).

[0238] Here, the specified storage medium can be, for example, the storage unit 307 (and its database 353) of the program distribution server 30A. The introductory text read out according to the script's comment information includes audio such as TTS audio or live voice.

[0239] Next, in the content provision system 1, the reading of content indicated by the content identification information contained in the script is controlled to be performed using rights already acquired by the user through a contract with a specific service, and to be provided to that user (S233).

[0240] Here, content identification information includes song IDs that can identify songs. Also, for example, a specific service is a music distribution service, and the rights that a user has already acquired include the rights of a paid premium user and the rights of a free user.

[0241] Furthermore, in the content provision system 1, after the content is provided, the system is controlled to read the postscript according to the comment information contained in the script and provide it to the user (S234).

[0242] Here, the postscript, which is read according to the comment information in the script, includes audio such as TTS audio or a live voice. The comments indicated by the comment information (text such as the preface and postscript) may be converted (translated) into a foreign language and provided to the user. For example, if the system knows the user's profile and their native language, the comments may be converted into the native language and provided to the user as text, or the text may be synthesized into speech using TTS and the resulting speech synthesis may be provided to the user.

[0243] As described above, when creating a program from content, we utilize contracts with users for our services and text-to-speech synthesis to provide commentary along with the content, making it easier to deliver content and its related comments.

[0244] <2. Second Embodiment>

[0245] The service provider that distributes songs adapted into a program does not have to be the same service provider that distributes the podcast program; it may be a different service provider. For example, a script for a program created on a specific music distribution app provided by one music distribution service can be provided to other music distribution services.

[0246] Specifically, the scripts of the created programs will be stored in a way that makes them accessible not only to the music streaming service that provides the specific music streaming app, but also to other music streaming services. In addition to identifying information such as song IDs that can be recognized by the specific music streaming app, metadata that allows for song searches will also be stored. For example, this metadata could include the song title, lyricist, composer, and artist name.

[0247] Furthermore, other music distribution services can, in response to playback requests from user terminal devices 20 used by users who have contracted with those other music distribution services, use information such as song titles contained in the metadata to identify songs from their own managed databases and stream the identified songs.

[0248] Furthermore, copyright issues related to the playback of such music are considered resolved through agreements between the user and other music streaming services. In addition, a portion of the revenue generated from the playback of such music distributed from other music streaming services may be returned to the music streaming service that provided the specific music streaming app, or a portion of that revenue may be returned to the creators.

[0249] The configuration of the content provision system 1 in the second embodiment is the same as that of the first embodiment, so its description will be omitted. The following describes the flow of processing performed by each device of the content provision system 1.

[0250] (Processing flow when the distribution provider is different) Figure 29 is a sequence diagram illustrating the processing flow when providing a script created using one music distribution service to another music distribution service.

[0251] In the example shown in Figure 29, the program creation tool and program renderer are run on the creator's and user's respective terminal devices, and the program distribution service, music distribution service, and live voice distribution service are provided by their respective distribution servers, similar to the third example described with reference to Figure 25.

[0252] Furthermore, in the example shown in Figure 29, a music distribution server 30B-1 is provided by Company A, which provides music distribution service A, and a music distribution server 30B-2 is provided by Company B, while the program distribution server 30A is provided by Company A. In addition, the program renderer, which is a music distribution application provided by Company B, is running on the user terminal device 20.

[0253] In steps S91 to S94 of Figure 29, similar to steps S51 to S54 of Figure 25, the program script is registered with Company A's program distribution server 30A by a program creation tool executed by the creator terminal device 10, and the audio data (files) of the introductory and concluding remarks, read aloud in the creator's own voice, are registered with the audio distribution server 30C.

[0254] On the user terminal device 20, the program renderer of company B is executed, and when the user instructs the playback of a program that is publicly available for viewing on company A's program distribution server 30A, the script for that program distributed from company A's program distribution server 30A is received (S95, S96).

[0255] This script includes metadata for searching content, in addition to the song ID, for songs that have been featured on television programs.

[0256] For example, metadata includes the song title, lyricist, composer, artist name, etc. More specifically, the information specified in items such as "title" and "artist" in the script shown in FIG. 7 can be used as metadata.

[0257] Next, the program renderer of Company B requests the music distribution server 30B-2 of Company B that provides the music distribution service B subscribed by the user to distribute the music specified by the metadata based on the received script (S97).

[0258] In the music distribution server 30B-2, in response to the request from the program renderer of Company B, when the rights acquired by the user's contract with the music distribution service B are confirmed and it is determined that the user has legitimate rights, the music is specified using the metadata (S98).

[0259] In the music distribution server 30B-2, when a search for music using the metadata is performed on the music distributed by the music distribution service B and the music can be specified by the metadata, the music is streamed to the user terminal device 20 (S99).

[0260] In this case, before providing the music, the program renderer of Company B accesses the voice distribution server 30C based on the link information of the preface set in the script, and the live voice of the preface part set in the programmed music is played (S101 to S102). Subsequently, the program renderer of Company B plays the music streamed from the music distribution server 30B-2 of Company B.

[0261] After that, when the playback of the streamed music ends, the program renderer of Company B accesses the voice distribution server 30C based on the link information of the epilogue set in the script, and the live voice of the epilogue part set in the programmed music is played (S103 to S105).

[0262] Furthermore, as steps S97 through S106 are repeated, Company B's program renderer repeatedly plays the introduction, song, and concluding segment in the order specified in the script for each song ID, making the podcast program available for the user to listen to.

[0263] On the other hand, if, during the processing in step S98, the music distribution server 30B-2 of company B is unable to identify the music based on the metadata included in the request from company B's program renderer, then, for example, the following processing is performed.

[0264] Firstly, Company B's music distribution server 30B-2 responds to Company B's program renderer to that effect, causing the playback of the unidentified song and its intro and outro to be skipped. As a result, on the user terminal device 20, Company B's program renderer skips the song in question and starts playing the next song (or the intro to the next song).

[0265] Secondly, Company B's music distribution server 30B-2 ensures that sample songs corresponding to metadata are distributed from Company A's music distribution server 30B-1 to the user terminal device 20. As a result, the user terminal device 20 plays sample songs distributed from music distribution service A, which the user does not subscribe to, via Company B's program renderer.

[0266] In the example in Figure 29, live voices are used when providing the introductory and concluding remarks for a song that has been made into a program; however, the introductory and concluding remarks for a song may also be provided using speech synthesis.

[0267] (Overall picture of the process) Figures 30 and 31 are flowcharts illustrating the overall process in the second embodiment.

[0268] The processes shown in Figures 30 and 31 are realized through the collaborative operation of the creator terminal device 10 (its control unit 100), the user terminal device 20 (its control unit 200), and the distribution server 30 (its control unit 300) in a content provision system to which this technology is applied.

[0269] In the content provision system 1, as shown in Figure 30, when input is received from a creator using the creator terminal device 10 (Yes in S311), a script is generated based on that input, consisting of content identification information recognizable by the first service, metadata for searching for content, and comment information corresponding to the introduction and conclusion of the content (S312). This script is then stored in a predetermined storage medium so that it can be viewed by a user using the user terminal device 20 (S313).

[0270] Here, "content" includes songs distributed through music streaming services, and the identifying information for that content includes a song ID that can identify the song. Furthermore, the introduction and conclusion are examples of comments regarding the content, and it is sufficient if at least one of them is provided.

[0271] Furthermore, the first service is, for example, music distribution service A provided by company A. The metadata also includes song titles, etc. The predetermined storage medium can be, for example, the database 353 of the program distribution server 30A.

[0272] Furthermore, in the content provision system 1, as shown in Figure 31, when a user who has contracted with the second service requests playback of a script recorded on a predetermined storage medium (S331, "Yes"), the system performs a process to identify the corresponding content from among the content managed by the second service based on metadata for searching for content (S332), and determines whether or not the corresponding content has been identified (S333).

[0273] Here, the second service is, for example, music distribution service B provided by company B. The content includes songs distributed through the music distribution service, and the metadata includes song titles, etc.

[0274] If the determination process in step S333 determines that the corresponding content has been identified, the system controls the retrieval of the content and its provision to the user by utilizing the rights that the user has already acquired through their contract with the second service (S334).

[0275] In other words, the rights that the user has already acquired include the rights of a paid premium user and a free user of music distribution service B provided by company B, and the user terminal device 20 plays music distributed from music distribution service B.

[0276] On the other hand, if the determination process in step S333 determines that the corresponding content could not be identified, the system is controlled to either skip reading the content and the corresponding comment, or to read sample data corresponding to the content managed by the first service and provide it to the user (S335).

[0277] In other words, on the user terminal device 20, either the song to be played and its introduction and conclusion are skipped, or a sample song distributed from music distribution service A provided by company A is played.

[0278] <3. Third Embodiment>

[0279] For introductory and concluding remarks set for songs that have been turned into programs, the system can automatically identify words and contexts that are inappropriate for introductory or concluding remarks, such as defamatory language. If such words or contexts are found in the text, a warning can be displayed during the program creation stage, or registration can be denied during the program registration stage.

[0280] Also, after a program is registered, if there is an instruction from the copyright holder of a music piece or the like not to permit continuous playback of a specific preface or epilogue and the music piece, for example, only the playback of the preface or epilogue may be prohibited (playback of the music piece is possible), the playback in the order of the preface, the music piece, and the epilogue may be prohibited, or the playback of the music piece itself may be prohibited.

[0281] In this way, when registering and providing a podcast program, a minimum usage permission function can be used.

[0282] (Other configurations of the distribution server) FIG. 32 shows another example of the functional configuration of the control unit 300 in the distribution server 30.

[0283] In FIG. 32, similar to FIG. 14, the control unit 300 has a request reception / response unit 351, a distribution processing unit 352, and a database 353, but a text check unit 161 is further provided.

[0284] In the program distribution server 30A, when the distribution processing unit 352 is requested to register a script of a program from a program creation tool, it supplies the texts of the preface and epilogue included in the script of the program to the text check unit 161.

[0285] The text check unit 161 performs a text check process on the texts of the preface and epilogue supplied from the distribution processing unit 352, and supplies the result of the text check to the distribution processing unit 352.

[0286] In this text check process, for example, natural language processing including morphological analysis and syntactic analysis is performed, the texts of the preface and epilogue are analyzed, and it is checked whether words such as slanderous terms and context (meaning) are included.

[0287] When the result of the text check supplied from the text check unit 161 indicates that the texts of the preface and epilogue are appropriate, the distribution processing unit 352 stores the script of the program requested to be registered from the program creation tool in the database 353.

[0288] The text checking unit 161 can, for example, be installed inside the program distribution server 30A, but it may also be installed as an external server and, in response to a request from the program distribution server 30A, check the text and respond with the result of the text checking.

[0289] (Process flow when performing text checking) Figure 33 is a sequence diagram showing the processing flow when performing text checking.

[0290] In the example shown in Figure 33, the program creation tool and program renderer are executed on the creator's and user's respective terminal devices, and the program distribution service, music distribution service, and TTS service are provided by each distribution server, similar to the first example described with reference to Figure 15. However, the program distribution server 30A also performs processing related to text checking.

[0291] On the creator terminal device 10, the program creation tool is executed and works in cooperation with each distribution server 30 to perform the processes in steps S111 to S124.

[0292] The program creation tool retrieves the song list sent from the music distribution server 30B and presents it to the creator (S111).

[0293] The program creation tool generates a podcast program script based on the song ID of the song selected by the creator from the song list and the introductory and concluding texts of the song entered by the creator (S112), and sends a registration request to the program distribution server 30A (S113).

[0294] At this time, the program distribution server 30A requests a text check of the introductory and concluding remarks by sending the texts of the introductory and concluding remarks, which have been requested to be registered by the program creation tool, to the text check unit 361 (S114).

[0295] In response to a request for text checking, the text checking unit 361 checks the text of the introduction and conclusion (S115) and notifies the user of the results of the text checking (S116).

[0296] Based on the response from the text checking unit 361, the program distribution server 30A notifies the program creation tool of the results of the text checking (S117). As a result, the program creation tool displays the results of the text checking (S118).

[0297] For example, if the text checking unit 361 determines that the text is not suitable as an introduction and a concluding paragraph, it will notify the user that it is not permitted.

[0298] In issuing this disapproval notice, the disapproved portion of the checked text may also be included. This disapproval notice, including the disapproved portion, will be sent to the program creation tool, and the creator will be shown that the introductory and concluding remarks for the song entered by the creator have been disapproved, along with the text of the disapproved portion.

[0299] Creators can revise the introductory and concluding text for songs to be featured in a program based on notifications provided by the program creation tool.

[0300] The program creation tool regenerates the podcast program based on the introductory and concluding texts of the song, which have been modified by the creator (S119), and sends a registration request to the program distribution server 30A again (S120).

[0301] At this time, the program distribution server 30A requests a text check of the introductory and concluding remarks by sending the texts of the introductory and concluding remarks, which have been requested to be registered by the program creation tool, to the text check unit 361 (S121).

[0302] In response to a request for text checking, the text checking unit 361 checks the text of the introduction and conclusion (S122) and notifies the user of the results of the text checking (S123).

[0303] For example, if the text checking unit 361 determines that the text is suitable as an introduction and a concluding paragraph, it will notify the user that permission has been granted.

[0304] On the program distribution server 30A, based on permission notifications from the text checking unit 361, the program script, including the revised introductory and concluding texts for the songs that have been requested to be re-registered by the program creation tool, is stored in the database 353 and made available for viewing by users using the user terminal device 20 (S124).

[0305] (Overall picture of the process) Figure 34 is a flowchart illustrating the overall process in the third embodiment.

[0306] The process shown in Figure 34 is realized through the collaborative operation of the creator terminal device 10 (specifically its control unit 100) and the distribution server 30 (specifically its control unit 300) in a content provision system to which this technology is applied.

[0307] In the content provision system 1, as shown in Figure 34, when a script is generated in response to input from a creator using the creator terminal device 10 (Yes in S411), the content of the comments indicated by the comment information corresponding to the introduction and conclusion of the content included in the script is analyzed (S412), and based on the analysis results, it is determined whether the content of the comments is appropriate as comments related to the content (S413).

[0308] Here, "content" includes songs distributed through music streaming services. Furthermore, the introduction and conclusion are examples of comments regarding the content; at least one of them must be provided.

[0309] If the determination process in step S413 determines that the content of the comment is appropriate as a comment related to the content, the script containing the comment information is stored in a predetermined storage medium so that it can be viewed by the user using the user terminal device 20 (S414).

[0310] In other words, if the introductory and concluding remarks for a song to be featured in a program do not contain defamatory or libelous language or other offensive words or context, they are deemed appropriate for introductory and concluding remarks, and the program script is registered in the database 353 of the program distribution server 30A.

[0311] On the other hand, if the judgment process in step S413 determines that the content of the comment is not appropriate as a comment related to the content, this fact is notified to the creator terminal device 10 used by the creator.

[0312] In other words, if the introductory or concluding remarks of a song to be featured on a program contain defamatory or libelous language or other offensive words or context, these will be automatically identified, and a warning will be displayed during the program creation stage using the program creation tool, or registration will be denied during the program registration stage.

[0313] <4. Fourth Embodiment>

[0314] Podcast programs may include advertisements. For example, the content of at least one of the introductory or concluding remarks set in a programmed song may be analyzed, and relevant advertisements may be inserted before the introductory remarks or after the concluding remarks, depending on the results of that analysis.

[0315] Advertisements inserted into the program can, for example, be written as text in the script, and their speech can be synthesized based on the same phonemes as the text for the introduction and conclusion.

[0316] In other words, the user terminal device 20, in accordance with the script of a program selected by the user, will synthesize the introductory and concluding remarks included in the script using a specific voice and provide them to the user, and will also synthesize the text of the advertisement using the same specific voice and provide it to the user.

[0317] (Other system configurations) Figure 35 shows another example of a configuration of one embodiment of a content delivery system to which this technology is applied.

[0318] In Figure 35, similar to Figure 8, the content provision system 1 includes a creator terminal device 10, a user terminal device 20, a program distribution server 30A, a music distribution server 30B, and an audio distribution server 30C, but an advertising distribution server 30D is also provided.

[0319] The ad delivery server 30D consists of one or more servers that provide ad delivery services. Ad delivery services are services that deliver advertisements over the internet and are provided, for example, by ad delivery companies.

[0320] For example, the ad delivery server 30D, in response to a request from the program delivery server 30A, identifies an ad managed in the ad management or ad database 353 and delivers the identified ad (ad text).

[0321] Furthermore, the ad delivery server 30D has the same configuration as the delivery server 30 and the functional configuration of the control unit 300 shown in Figures 13 and 14.

[0322] (Example of an advertisement) Figure 36 shows an example of an advertisement inserted into a podcast script.

[0323] As mentioned above, podcast scripts typically include the song to be featured, along with a warm-up and after-song for that song. However, in Figure 36, an advertisement is inserted after the after-song.

[0324] In the example in Figure 36, the introductory text includes the sentence, "...I still want to listen to this in the summer...", and the concluding text includes the sentence, "...This is a song that really makes hot summers fun!", and these texts are analyzed.

[0325] In the program's script, based on the analysis of these texts, an advertisement for beer is inserted as an ad related to the keyword "summer," consisting of the text, "On a hot summer day, nothing beats beer. Refreshing dry beer from Company X!" Along with the text of the beer advertisement, the URL (Uniform Resource Locator) of the webpage related to Company X's beer is also included.

[0326] In this case, when the introductory and concluding texts are synthesized into speech and read aloud as text-to-speech audio, natural language processing can be used to analyze the text, allowing for a more accurate analysis compared to analyzing the live voice during live streaming.

[0327] Therefore, by analyzing the introductory and concluding texts, it becomes possible to present more relevant advertisements. For example, if the introductory and concluding narrations include phrases like, "I'd like to drink wine while listening to this song," then advertisements for wine can be presented.

[0328] Furthermore, when the text of the inserted advertisement is read aloud using speech synthesis, it is acceptable to read it as is, but it is also acceptable to read it in a style that matches the DJ's tone, or in a dialect. Alternatively, the tone of the advertisement's TTS audio may be matched to the tone of the introductory or concluding TTS audio.

[0329] In the example shown in Figure 36, an advertisement is inserted after the concluding section. However, advertisements can be inserted at any location, and in particular, due to their relationship with the music, it is preferable to insert advertisements before the introductory section or after the concluding section.

[0330] (Processing flow when inserting advertisements into a program) Figure 37 is a sequence diagram showing the processing flow when inserting advertisements into a program.

[0331] In the example shown in Figure 37, the program renderer runs on the user terminal device 20, and the program distribution service, music distribution service, and TTS service are provided by their respective distribution servers, similar to the first example described with reference to Figure 15. However, an advertising distribution server 30D is provided that offers advertising distribution services and an advertising management database.

[0332] On the program distribution server 30A, when a user instructs the program renderer to play a publicly available podcast program, the script of that program is analyzed (S131, S132).

[0333] The program distribution server 30A sends a request to the advertising distribution server 30D based on the analysis results of the program's script, thereby obtaining a word-advertisement list from the advertising management database 353 (S133, S134).

[0334] For example, during script analysis, the text of the introductory and concluding remarks set in the script is analyzed, the words and meanings contained in the text are extracted, and the placement of advertisements in the program is determined. In addition, since the words and meanings are associated with advertisement IDs in the advertisement management database 353, advertisement IDs corresponding to the introductory and concluding remarks are obtained as a word-advertisement list.

[0335] Furthermore, the program distribution server 30A obtains the advertisement (ad text) by sending a request including the advertisement ID to the advertisement distribution server 30D (S135, S136).

[0336] In other words, since the ad distribution server 30D associates ad IDs with ad text and manages them in the ad database 353, the program distribution server 30A can obtain the ad text identified by the ad ID corresponding to the introductory and concluding segments and insert it at the determined insertion location (S137). The script of the program into which the ad has been inserted is sent to the program renderer (S138).

[0337] In steps S139 to S147, similar to steps S16 to S24 in Figure 15, the program renderer plays the program in the following order: the intro, the music identified by the music ID, and the post-program segment, making the program viewable.

[0338] Furthermore, since the program script includes advertisement text after the postscript, the program renderer requests the audio distribution server 30C, which provides TTS services, to synthesize the text of the advertisement (S148).

[0339] In the audio distribution server 30C, in response to a request from the program renderer, the text of the advertisement is synthesized into speech (S149), and the result of the speech synthesis is distributed (S150).

[0340] As a result, the program renderer receives the results of speech synthesis delivered from the audio distribution server 30C, and rendering is performed to play the TTS audio for the advertisements inserted into the program.

[0341] Furthermore, since the podcast program script includes multiple song IDs along with the introductory and concluding texts for each song, after the processing in steps S139 to S150 is completed, the process returns to step S139 (S151), and steps S139 to S151 are repeated according to the number of song IDs.

[0342] As a result, the program renderer repeatedly plays the intro, song, and post-show in the order set for each song ID in the program's script, allowing users to listen to the podcast program.

[0343] Furthermore, if advertisements related to the introductory or concluding remarks are inserted before or after the introductory or concluding remarks, the TTS audio of the advertisements will also be played, allowing users to link the introductory or concluding remarks, the music, and the advertisements. Note that advertisements are not limited to audio output; they can also be presented as a GUI (Graphical User Interface), for example, by displaying them in a designated area on the screen of a music streaming app.

[0344] Furthermore, for example, when distributing audio data of the original voices of comments such as introductory or concluding remarks, a text-to-speech (TTS) system may be created by using that original voice audio data as training material to create specific music and then performing speech synthesis using that specific music. Using this TTS system, the audio distribution server 30C may synthesize the text of the advertisement into speech and distribute it to the user terminal device 20. This allows the advertisement to be delivered in a voice that is close to the original voice that the user prefers, which may lead to an improvement in the conversion rate of the advertisement.

[0345] Furthermore, in the sequence diagram shown in Figure 37, for the sake of simplicity, it is explained that both the program distribution service and the program analysis processes are executed by the program distribution server 30A, but these processes may be executed by separate servers. Similarly, in the explanation that both the advertising distribution service and the advertising management DB processes are executed by the advertising distribution server 30D, these processes may be executed by separate servers.

[0346] (Overall picture of the process) Figure 38 is a flowchart illustrating the overall process in the fourth embodiment.

[0347] The process shown in Figure 38 is realized through the collaborative operation of the user terminal device 20 (specifically its control unit 200) and the distribution server 30 (specifically its control unit 300) in a content provision system to which this technology is applied.

[0348] In the content provision system 1, as shown in Figure 38, when a user using the user terminal device 20 requests playback of a script stored on a predetermined storage medium (Yes in S511), the content of the comments indicated by the comment information contained in that script is analyzed (S512).

[0349] Here, the comments include at least one of the introductory and concluding remarks of the song as content. For example, the text of these introductory and concluding remarks is analyzed, and the words and meanings contained in the text are extracted to analyze the content of the introductory and concluding remarks.

[0350] Then, in the content provision system 1, based on the analysis results of the comment content, advertising data (ad text) is obtained from the advertising distribution server 30D (S513), and the system is controlled to provide the advertising data to the user either before or after the comment being provided (S514).

[0351] <5. Variation>

[0352] For content identification information such as song IDs written in the script of a podcast program, original IDs owned by each record company may be used.

[0353] (Other system configurations) Figure 39 shows an example of another configuration of one embodiment of a content delivery system to which this technology is applied.

[0354] In Figure 39, similar to Figure 8, the content provision system 1 includes a creator terminal device 10, a user terminal device 20, a program distribution server 30A, a music distribution server 30B, and an audio distribution server 30C, but an ID management server 31 is also provided.

[0355] The ID management server 31 manages song IDs used for each music distribution service by linking them to a management database. The ID management server 31 provides information about the managed song IDs in response to requests from devices such as user terminal devices 20.

[0356] The ID management server 31 has the same configuration as the distribution server 30 shown in Figure 13.

[0357] (Processing flow when sharing program information) Figure 40 is a sequence diagram showing the processing flow when managing song IDs and sharing program information.

[0358] In the example shown in Figure 40, the program creation tool and program renderer are run on the creator's and user's respective terminal devices, and the program distribution service and music distribution service are provided by their respective distribution servers, similar to the first example described with reference to Figure 15.

[0359] Furthermore, in the example shown in Figure 40, there is a music distribution server 30B-1 provided by Company A, which provides music distribution service A, and a music distribution server 30B-2 provided by Company B, which provides music distribution service B, while the program distribution server 30A is provided by Company A.

[0360] Furthermore, an ID management server 31 is provided to manage program IDs. In addition, it is assumed that the program renderer, which is a music distribution application provided by Company B, is running on the user terminal device 20.

[0361] The ID management server 31 receives song IDs used in music distribution service A, transmitted from company A's music distribution server 30B-1, and song IDs used in music distribution service B, transmitted from company B's music distribution server 30B-2 (S171, S172). These song IDs are original IDs owned by company A and company B, respectively.

[0362] The ID management server 31 uses the song master ID to link and manage song IDs used in music distribution service A and song IDs used in music distribution service B.

[0363] In steps S173 to S175, similar to steps S11 to S13 in Figure 15, a program creation tool executed by the creator terminal device 10 generates a program script that includes songs assigned to Company A's song ID, based on a song list transmitted from Company A's music distribution server 30B-1, and registers it with Company A's program distribution server 30A.

[0364] At this time, the user terminal device 20 runs the program renderer of company B, and when a user who has a contract with music distribution service B instructs the playback of a podcast program that is publicly available for viewing on the program distribution server 30A, the script of the program distributed from the program distribution server 30A is received (S176, S177).

[0365] The program renderer of company B sends the list of song IDs of company A set in the received script to the ID management server 31 and requests conversion of the song IDs (S178).

[0366] The ID management server 31, in response to a request from the program renderer, converts the list of song IDs from company A into a list of song IDs from company B and sends it to the program renderer of company B (S179, S180).

[0367] In other words, while the song ID set in the program's script is an original ID of company A, the user has a contract with music distribution service B of company B. Therefore, as things stand, they cannot play the song identified by that song ID using music distribution service B.

[0368] Therefore, the ID management server 31 converts Company A's original song IDs to Company B's original song IDs through ID conversion, thereby enabling playback of songs identified by Company A's song IDs set in the script using Company B's music distribution service B.

[0369] The program renderer requests the music distribution server 30B-2, which provides music distribution service B, to distribute the music identified by the converted music ID (S181).

[0370] In response to a request from the program renderer, the music distribution server 30B-2 of Company B verifies the rights acquired by the user through their contract with music distribution service B (S182). If it is determined that playback of the music identified by the music ID is possible, the music is streamed to Company B's program renderer (S183).

[0371] As a result, B Company's program renderer receives streaming data of songs distributed from B Company's music distribution server 30B-2, performs rendering processing, and plays the songs identified by the song ID.

[0372] For the sake of clarity, we have omitted some details here, but in the podcast program script, the song ID is included in the program along with the introductory and concluding texts. Therefore, the audio for the introductory section introducing the song identified by that song ID is played first, followed by the playback of the song itself, and then the audio for the concluding section is played after the song has finished playing.

[0373] Furthermore, since the podcast program script includes the song IDs of multiple songs along with the introductory and concluding texts for each song, after the processing in steps S181 to S183 is completed, the process returns to step S181 (S184), and the processing in steps S181 to S184 is repeated.

[0374] As a result, B Company's program renderer will repeatedly play the intro, song (a song distributed by B Company's music distribution service B), and post-show in the order specified for each song ID from A Company set in the program's script.

[0375] (Overall picture of the process) Figures 41 and 42 are flowcharts illustrating the overall process in the modified example.

[0376] The processes shown in Figures 41 and 42 are realized through the collaborative operation of the creator terminal device 10 (its control unit 100), the user terminal device 20 (its control unit 200), the distribution server 30 (its control unit 300), and the ID management server 31 (its control unit) in a content provision system to which this technology is applied.

[0377] In the content provision system 1, as shown in Figure 41, when input is received from a creator using the creator terminal device 10 (Yes in S611), a script is generated based on the input from the creator, consisting of first content identification information recognizable by the first service and comment information corresponding to the introduction and conclusion of the content (S612), and this script is stored in a predetermined storage medium so that it can be viewed by the user (S613).

[0378] Here, "content" includes songs distributed through music streaming services, and the identifying information for that content includes a song ID that can identify the song. Furthermore, the introduction and conclusion are examples of comments regarding the content, and it is sufficient if at least one of them is provided.

[0379] The comment information includes text representing the comment, or link information to the original voice recording. The first service is, for example, a music distribution service A provided by Company A. The specified storage medium can be, for example, the database 353 of the program distribution server 30A.

[0380] Furthermore, in the content provision system 1, as shown in Figure 42, when a user who has contracted with the second service requests playback of a script stored on a predetermined storage medium (S631, "Yes"), the second content identification information that the second service can recognize corresponding to the first content identification information is identified (S632), the content corresponding to the second content identification information is identified from the content managed by the second service (S633), and the system is controlled to read the content and provide it to the user using the rights that the user has already acquired through their contract with the second service (S634).

[0381] Here, the second service is, for example, music distribution service B provided by company B, and the rights that the user has already acquired include the rights of a paid premium user and a free user.

[0382] (Other variations) Furthermore, when playing music using the program renderer on the user terminal device 20, it is possible to add characters and images to create a video jockey (VJ) experience.

[0383] <6. Computer Configuration>

[0384] Each step in the flowchart described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on the computer of each device.

[0385] Computer programs can be provided by recording them on removable recording media, such as package media. Furthermore, programs can be provided via wired or wireless transmission media, such as local area networks, the internet, or digital satellite broadcasting.

[0386] In a computer, programs can be installed into the storage unit via an input / output interface by inserting a removable storage medium into a drive. Alternatively, programs can be received by the communication unit via a wired or wireless transmission medium and installed into the storage unit. Furthermore, programs can be pre-installed in ROM or other storage units.

[0387] In this specification, the processes performed by a computer according to a program do not necessarily have to be performed chronologically in the order described in the flowchart. That is, the processes performed by a computer according to a program include processes that are executed in parallel or individually (e.g., parallel processing or object-based processing).

[0388] Furthermore, the program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Moreover, the program may be transferred to a remote computer and executed there.

[0389] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.

[0390] It should be noted that the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the gist of this technology. For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0391] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices. In addition, if a single step includes multiple processes, those processes can be performed by a single device, or they can be divided and performed by multiple devices.

[0392] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0393] Furthermore, this technology can be configured as follows:

[0394] (1) A script, consisting of content identification information and comment information generated by the creator, is stored on a designated storage medium in a manner that allows the user to view it. The system includes a control unit that controls the reading of content indicated by content identifier information contained in a script selected by the user, using rights already acquired by the user through a contract with a specific service, and provides the content to the user, and also controls the reading of comments according to comment information contained in the script, either before or after the provision of the content, and provides the comments to the user. Content delivery system. (2) The control unit controls the system to synthesize the text, which is the comment information contained in the script, into speech and provide it to the user, according to the script selected by the user. The content provision system described in (1) above. (3) The control unit controls the system to access the link information to the audio data, which is the comment information contained in the script, according to the script selected by the user, and to read the audio data and provide it to the user. The content provision system described in (1) above. (4) When the control unit receives a request from the user to play a playlist, it controls the system to read the content indicated by the content identification information contained in the script and provide it to the user, without reading the comments corresponding to the comment information contained in the script, according to the script selected by the user. A content provision system as described in any of (1) to (3) above. (5) The control unit, A script consisting of content identification information generated by the creator, a first comment corresponding to an introductory explanation of the content, and a second comment corresponding to a post-playback explanation of the content, is stored on a designated storage medium in a manner accessible to the user. According to the script selected by the user, the preface is read and provided to the user according to the first comment information contained in that script. The content indicated by the content identification information following the first comment information is read and provided to the user. The following comment information is read and provided to the user according to the second comment information that follows the content identification information. Control it so that A content provision system as described in any of (1) to (4) above. (6) The control unit, A script, consisting of content identification information and comment information recognizable by a specific service, generated by the creator, is stored on a designated storage medium in a manner that allows users to view it. Controls the system to read and provide to the user the content indicated by the content identifier information contained in the script, according to the script selected by the user, using the rights already acquired by the user through their contract with the specific service, and also controls the system to read and provide to the user comments according to the comment information contained in the script, either before or after the provision of the content. The content provision system described in (1) above. (7) The control unit, A script consisting of content identification information recognizable by the first service, metadata for searching content, and comment information, generated by the creator, is stored on a predetermined storage medium in a manner accessible to the user. When the script is selected by a user who has a contract with a second service different from the first service, the metadata identifies the corresponding content managed by the second service, and controls the system to retrieve and provide the content to the user using the rights already acquired by the user through their contract with the second service. The content provision system described in (6) above. (8) The control unit, when the script is selected by a user who has a contract with the second service, controls the system to skip reading the content and the corresponding comment if it cannot identify the corresponding content from the content managed by the second service using the metadata. The content provision system described in (7) above. (9) The control unit, when the script is selected by a user who has a contract with the second service, controls the control unit to read sample data corresponding to the content managed by the first service and provide it to the user if the metadata does not allow the corresponding content to be identified from the content managed by the second service. The content provision system described in (7) above. (10) The control unit, Before the script generated by the creator is stored on the predetermined storage medium, the comment content indicated by the comment information contained in the script is analyzed, Based on the analysis results, if it is determined that a comment is inappropriate regarding the aforementioned content, the system will be controlled to notify the user accordingly. A content provision system as described in any of (1) through (9) above. (11) If the control unit determines that a comment is inappropriate regarding the content, it controls the unit to notify the user of the inappropriate part of the comment. The content provision system described in (10) above. (12) If the control unit determines that a comment is inappropriate for the content, it controls the script containing the comment information corresponding to that comment so that it is not stored on the predetermined storage medium. The content provision system described in (10) or (11) above. (13) The control unit, The content of the comments indicated by the comment information contained in the script selected by the user is analyzed. Based on the analysis results, the system controls whether the advertising data obtained from the ad delivery server is provided to the user before the content is provided to the user, before the comments are provided, or after the content is provided to the user, or after the comments are provided. A content provision system as described in any of (1) to (12) above. (14) The control unit, According to the script selected by the user, the text containing the comment information in that script is synthesized into speech using specific phonemes and provided to the user. The text used as advertising data is also controlled to be synthesized into speech using the specific phonemes and provided to the user. The content provision system described in (13) above. (15) The control unit analyzes the content of the text, which is comment information, included in the script selected by the user. The content provision system described in (13) or (14) above. (16) Master identification information that identifies specific content and content identification information that each service can use to recognize that specific content are managed in association with each other in a management database. The control unit, A script consisting of first content identification information and comment information, generated by a creator and recognizable by the first service, is stored on a predetermined storage medium in a manner that allows the user to view it. When the script is selected by a user who has a contract with a second service different from the first service, the second content identifier, which corresponds to the first content identifier and is recognizable by the second service, is identified according to the management database. The content corresponding to the second content identification information is identified by the content managed by the second service. The system uses the rights that the user has already acquired through their contract with the second service to retrieve and provide the content to the user. The content provision system described in (1) above. (17) The aforementioned content includes music, The aforementioned comment includes at least one of the introductory and concluding remarks set for the aforementioned song, and the aforementioned specific service includes the music streaming service to which the user has a contract. A content provision system as described in any of (1) through (16) above. (18) The first terminal device used by the aforementioned creator, A second terminal device used by the aforementioned user, A first server having the predetermined storage medium in which the script is stored, A second server that distributes the aforementioned content and A content provision system as described in any of (1) to (17) above, including the above. (19) A script, consisting of content identification information and comment information generated by the creator, is stored on a designated storage medium in a manner that allows the user to view it. The system controls the retrieval of content indicated by content identifiers contained in a script selected by the user, using rights already acquired by the user through a contract with a specific service, and also controls the retrieval of comments according to comment information contained in the script, either before or after the provision of such content, and provides these comments to the user. Content delivery methods. (20) Computers, A script, consisting of content identification information and comment information generated by the creator, is stored on a designated storage medium in a manner that allows the user to view it. The system controls the retrieval of content indicated by content identifiers contained in a script selected by the user, using rights already acquired by the user through a contract with a specific service, and also controls the retrieval of comments according to comment information contained in the script, either before or after the provision of such content, and provides these comments to the user. A storage medium containing a program to enable it to function as a control unit. [Explanation of symbols]

[0395] 1 Content provision system, 10 Creator terminal device, 20 User terminal device, 30 Distribution server, 30A Program distribution server, 30B Music distribution server, 30C Audio distribution server, 30D Ad distribution server, 31 ID management server, 50 Network, 100 Control unit, 101 CPU, 102 ROM, 103 RAM, 104 Bus, 105 Input unit, 106 Output unit, 107 Memory unit, 108 Communication unit, 109 Short-range wireless communication unit, 110 Input / Output I / F, 111 Operation unit, 112 Camera unit, 113 Sensor unit, 121 Display unit, 122 Sound output unit, 151 Input reception unit, 152 Music information acquisition unit, 153 Program generation unit, 154 Audio information acquisition unit, 155 Voice generation unit, 156 Registration unit, 200 Control unit, 201 CPU, 202 ROM, 203 RAM, 204 Bus, 205 Input unit, 206 Output unit, 207 Memory unit, 208 Communication unit, 209 Short-range wireless communication unit, 210 Input / Output I / F, 211 Operation unit, 212 Camera unit, 213 Sensor unit, 221 Display unit, 222 Sound output unit, 251 Program acquisition unit, 252 Music acquisition unit, 253 Audio acquisition unit, 254 Renderer unit, 255 Presentation control unit, 300 Control unit, 301 CPU, 302 ROM, 303 RAM, 304 Bus, 305 Input unit, 306 Output unit, 307 Memory unit, 308 Communication Unit, 309 Drive, 310 Input / Output Interface, 351 Input Reception Unit, 352 Distribution Processing Unit, 353 Database, 361 Document Checking Unit

Claims

1. At a minimum, a script consisting of content identification information and text information related to advertisements is stored on a designated storage medium in a manner that is accessible to the user. In accordance with the script selected by the user, the system controls the retrieval of content indicated by the content identifier information contained in that script, using the rights already acquired by the user through a contract with a specific service, and provides the content to the user. The text information relating to the advertisement contained in the script is synthesized into speech and controlled to be provided to the user either before or after the provision of the content. Equipped with a control unit Content delivery system.

2. The script stored in the predetermined storage medium contains comment information that is different from the text information relating to the advertisement, The control unit, if the comment information is in text format, controls the unit to synthesize the comment information contained in the script into speech and provide it to the user either before or after the provision of the content. The content provision system according to claim 1.

3. The control unit controls the speech synthesis of the text information relating to the advertisement and the speech synthesis of the comment information when it is in text format to be performed based on the same phonemes. The content provision system according to claim 2.

4. The speech synthesis of the text information relating to the aforementioned advertisement, or the speech synthesis of the aforementioned comment information when it is in text format, is performed based on a model that has been pre-trained by machine learning using audio data as training material. The content provision system according to claim 2 or 3.

5. The control unit, Analyzing the content of the comments indicated by the aforementioned comment information, The system controls the speech synthesis process to provide users with text information about advertisements based on the analysis results. The content provision system according to claim 2 or 3.

6. At a minimum, a script consisting of content identification information and text information related to advertisements is stored on a designated storage medium in a manner that is accessible to the user. The system controls the retrieval of content indicated by content identifiers contained in a script selected by the user, using rights already acquired by the user through a contract with a specific service, and provides the content to the user. The text information relating to the advertisement contained in the script is synthesized into speech and controlled to be provided to the user either before or after the provision of the content. Content delivery methods.

7. Computers, At a minimum, a script consisting of content identification information and text information related to advertisements is stored on a designated storage medium in a manner that is accessible to the user. In accordance with the script selected by the user, the system controls the retrieval of content indicated by the content identifier information contained in that script, using the rights already acquired by the user through a contract with a specific service, and provides the content to the user. The text information relating to the advertisement contained in the script is synthesized into speech and controlled to be provided to the user either before or after the provision of the content. A program to enable it to function as a control unit.

Citation Information

Patent Citations

  • Information-providing program, information-providing method and recording medium

    JP2002342206A

  • Music playback device and music information distribution server

    JP2007164078A

  • Comment delivery system, terminal device, comment delivery method, and program

    JP2011166833A

  • Information processing system and client terminal

    JP2015060545A

  • Playback control device, playback control method, and program

    JP2015518171A