Program generation system, program generation program, and program generation method
Patent Information
- Application Number
- JP2025023260
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-08-27
AI Technical Summary
【0022】 番組生成システム、番組生成プログラム、および、番組生成方法は、放送番組の放送者による楽曲データや読上げデータの事前準備を不要にできる。
Smart Images

Figure 2026137272000001_ABST
Abstract
Description
Technical Field
[0005] ,
[0001] The present disclosure relates to a program generation system, a program generation program, and a program generation method.
Background Art
[0002] In facilities such as commercial facilities, offices, and schools, a broadcast program including the voice of a speaker in a broadcast facility and music may be broadcast. Patent Document 1 discloses a broadcast device that allows a speaker to broadcast from a microphone device or play sound source data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In facility broadcasts, music requested by listeners in the broadcast area of the facility is broadcast, or the voice of a personality who reads a message from the listener who requested the music is broadcast. However, in conventional broadcasts, it is necessary for the broadcaster of the broadcast program to prepare in advance a broadcast text that organizes the music data of the music related to the request and the messages of the listeners.
Means for Solving the Problems
[0005] (1) A program generation system for solving the problem is a program generation system for generating a broadcast program including music and spoken audio, comprising: an editing unit for editing the broadcast program; and a playback unit for playing the broadcast program in a broadcast area, wherein the editing unit is connectable to a music generation unit and a voice generation unit, the music generation unit is configured to generate music data based on a prompt including composition information relating to the composition of the music, and the voice generation unit is configured to generate spoken audio data based on a prompt including spoken audio information relating to the spoken audio, the editing unit acquires first composition information as the composition information and first spoken audio as the spoken audio information, acquires generated music data generated from the music generation unit by inputting a music generation prompt including the first composition information to the music generation unit, acquires generated spoken audio data generated from the voice generation unit by inputting a spoken audio generation prompt including the first spoken audio information to the voice generation unit, and generates the broadcast program by combining the generated music data and the generated spoken audio data.
[0006] In this configuration, the editorial department generates broadcast programs by combining generated music data based on the first composition information and generated text-to-speech data based on the first text-to-speech information. Since the generated music data and text-to-speech data acquired on the program generation system are used in the generation of broadcast programs, it eliminates the need for broadcasters to prepare music data and text-to-speech data in advance.
[0007] (2) In the program generation system described in (1) above, the composition information includes at least one of the listener preferences of the broadcast program, the musical style of the song, and the type of voice that sings the song. With this configuration, it is possible to generate song data that reflects at least one of the listener preferences, the musical style of the song, and the type of voice that sings the song.
[0008] (3) In the program generation system described in (2) above, the music generation prompt further includes basic composition information which includes at least one of the broadcast time of the broadcast program, information about the broadcast area, information about the listeners of the broadcast program, and the playback time of the music. With this configuration, music data is generated which reflects at least one of the broadcast time of the broadcast program, information about the broadcast area, information about the listeners, and the playback time of the music.
[0009] (4) In any one of the program generation systems described in (1) to (3) above, the spoken information includes at least one of the listener's name who entered the composition information and the song description. This configuration makes it possible to generate spoken data that includes at least one of the listener's name and the song description.
[0010] (5) In the program generation system described in (4) above, the prompt for generating the reading voice further includes basic reading information which includes at least one of the type of reading voice, the broadcast time of the broadcast program, information about the broadcast area, information about the listener, and the playback time of the reading voice. With this configuration, it is possible to generate reading data which reflects at least one of the voice type of the reading voice, the broadcast time of the broadcast program, information about the broadcast area, information about the listener, and the playback time of the reading voice.
[0011] (6) In any one of the program generation systems described in (1) to (5) above, the generated music data includes first generated music data and second generated music data, the generated reading data includes first generated reading data relating to the first generated music data and second generated reading data relating to the second generated music data, and the editing department edits the broadcast program such that the first generated music data is played after the first generated reading data, the second generated reading data is played after the first generated music data, and the second generated music data is played after the second generated reading data.
[0012] With this configuration, the first generated audio data is played after the first generated audio data, making it easier for listeners to understand the content of the first generated audio data compared to when they are played separately. Similarly, the second generated audio data is played after the second generated audio data, making it easier for listeners to understand the content of the second generated audio data compared to when they are played separately.
[0013] (7) In any one of the program generation systems described in (1) to (6) above, when the editorial department acquires multiple generated music data, it sets a musical style for each of the multiple generated music data that is suitable for the music to be played from the generated music data, and edits the broadcast program so that the generated music data with the same musical style are played consecutively. With this configuration, since music data with the same musical style are played consecutively, listeners can easily feel the atmosphere of the music.
[0014] (8) In any one of the program generation systems described in (1) to (7) above, the editorial department, upon acquiring multiple generated music data, acquires the popularity ranking of each of the multiple generated music data and edits the broadcast program so that the multiple generated music data are played in order of increasing popularity. With this configuration, generated music data that is popular with listeners is given priority in playback, making it easier for popular generated music data to be played. As a result, listeners of the broadcast program have more opportunities to hear popular generated music data.
[0015] (9) In any one of the program generation systems described in (1) to (8) above, the editorial department is further connected to an information source unit that stores area information relating to the broadcast area, and obtains the area information from the information source unit and includes additional read-aloud data containing the area information in the broadcast program. With this configuration, the content of the area information is included in the broadcast program, so that listeners of the broadcast program can grasp the area information.
[0016] (10) In any one of the program generation systems described in (1) to (9) above, the editorial department is further connected to a camera that photographs the broadcast area, and obtains area information relating to the broadcast area from the video captured by the camera, and includes additional read-aloud data containing the area information in the broadcast program. With this configuration, the content of the area information is included in the broadcast program, so that listeners of the broadcast program can grasp the area information.
[0017] (11) In any one of the program generation systems described in (1) to (10) above, the editorial department is further connected to a text generation unit configured to generate the music generation prompt based on a prompt containing the music composition information, and obtains the music generation prompt from the text generation unit by inputting a pre-processing prompt containing the first music composition information to the text generation unit. With this configuration, since the music generation prompt is obtained by inputting a pre-processing prompt to the text generation unit, it is possible to generate a highly accurate music generation prompt that reflects the music composition information.
[0018] (12) A program generation program that solves the problem is a program generation program that generates a broadcast program including music and spoken audio, and it executes the following steps: a first step of causing a computer to acquire first composition information as composition information relating to the composition of the music and first spoken information as spoken audio relating to the spoken audio; a second step of causing a music generation unit configured to generate music data based on a prompt including the composition information to input a music generation prompt including the first composition information, thereby acquiring generated music data generated as music data from the music generation unit; a third step of causing a voice generation unit configured to generate spoken audio based on a prompt including the spoken audio information to input a spoken audio generation prompt including the first spoken audio information, thereby acquiring generated spoken audio data generated as spoken audio from the voice generation unit; and a fourth step of generating the broadcast program by combining the generated music data and the generated spoken audio.
[0019] According to this configuration, since a broadcast program is generated from the newly generated generated music data and generated voiceover data, it is possible to eliminate the need for broadcasters to prepare music data and voiceover data in advance.
[0020] (13) A program generation method for solving problems is a program generation method in which a computer generates a broadcast program including music and voiceover, and the computer includes: a first step of acquiring first composition information as composition information related to the composition of the music and first voiceover information as voiceover information related to the voiceover; a second step of inputting a music generation prompt including the first composition information to a music generation unit configured to generate music data based on a prompt including the composition information, and acquiring generated music data generated as the music data from the music generation unit; a third step of inputting a voiceover generation prompt including the first voiceover information to a voice generation unit configured to generate voiceover data based on a prompt including the voiceover information, and acquiring generated voiceover data generated as the voiceover data from the voice generation unit; and a fourth step of generating the broadcast program by combining the generated music data and the generated voiceover data.
[0021] According to this configuration, since the generated music data and generated voiceover data included in the broadcast program are generated in the process of the program generation method, it is possible to eliminate the need for broadcasters to prepare music data and voiceover data in advance.
Effect of the Invention
[0022] A program generation system, a program generation program, and a program generation method can eliminate the need for broadcasters of broadcast programs to prepare music data and voiceover data in advance.
Brief Description of the Drawings
[0023] [Figure 1] It is a block diagram showing the configuration of a program generation system according to the first embodiment. [Figure 2] It is a schematic diagram showing an example of an input form in the connection terminal of FIG. 1. [Figure 3] It is a schematic diagram showing the flow until a broadcast program is reproduced in the program generation system of FIG. 1. [Figure 4] It is a schematic diagram showing an example of the configuration of a broadcast program generated by the flow of FIG. 3. [Figure 5] It is a schematic diagram showing an example of a prompt for music generation. [Figure 6] It is a schematic diagram showing an example of a prompt for voiceover generation. [Figure 7] It is a schematic diagram showing an example of a prompt for additional voiceover generation. [Figure 8] It is a block diagram showing the configuration of a program generation system according to the second embodiment. [Figure 9] It is a schematic diagram showing the flow until a broadcast program is generated in the program generation system of FIG. 8. [Figure 10] It is a schematic diagram showing an example of a first preprocessing prompt. [Figure 11] It is a schematic diagram showing an example of a prompt for music generation. [Figure 12] It is a schematic diagram showing the flow until a broadcast program is generated in the program generation system according to the second modification example. [Figure 13] It is a block diagram showing the configuration of a program generation system according to the third modification example. [Figure 14] It is a schematic diagram showing the flow until a broadcast program is generated in the program generation system according to the fourth modification example. [Figure 15] It is a schematic diagram showing an example of a second preprocessing prompt.
Mode for Carrying Out the Invention
[0024] <First Embodiment> Referring to Figures 1 to 7, the program generation system 10 according to the first embodiment will be described. The program generation system 10 generates a broadcast program 100. The broadcast program 100 includes music and spoken audio. The spoken audio is audio that reads a text aloud. The spoken audio is audio related to music. For example, the spoken audio is audio that introduces music. The spoken audio may also be audio unrelated to music. The broadcast program 100 is broadcast in broadcast area A1.
[0025] Examples of broadcast area A1 include spaces within facilities such as commercial facilities, office buildings, and schools. However, when broadcasting in broadcast area A1, broadcasters previously needed to prepare sound sources in advance at broadcasting facilities inside or outside the facility. Specifically, broadcasters needed to prepare music data for the songs to be broadcast and broadcast text to be read aloud by speakers. The program generation system 10 eliminates the need for broadcasters to make prior preparations when broadcasting in broadcast area A1.
[0026] <Program Generation System Configuration> As shown in Figure 1, the program generation system 10 comprises a computer 11 and a connection terminal 12. The computer 11 is configured, for example, by a server. The computer 11 may be configured by a combination of multiple computers installed in the same location or in different locations.
[0027] Computer 11 comprises a processing circuit and a memory device. The processing circuit consists of a CPU that executes processing according to a program and its peripheral circuits. The memory device consists of ROM where programs are stored, volatile RAM on which data can be temporarily written, and non-volatile storage on which data can be written. Various programs are executed on computer 11 by the processing circuit executing various programs stored in the memory device.
[0028] The connection terminal 12 is a terminal for listeners of the broadcast program 100 to access the program generation system 10. Listeners post messages and other information to the program generation system 10 using the connection terminal 12. The connection terminal 12 is a computer terminal such as a personal computer, tablet, smartphone, or telephone. The connection terminal 12 has, for example, a display for displaying various information, an input unit for inputting text, a speaker unit for playing audio, and a microphone unit for acquiring audio. The connection terminal 12 is connected to the computer 11 via a network so that it can communicate with it. An example of a network is the internet. The network described later is similar.
[0029] A broadcasting device 13 is located in broadcasting area A1. The broadcasting device 13 may be provided separately from the program generation system 10, or it may be configured as part of the program generation system 10. The broadcasting device 13 includes a playback device for playing audio data, an amplifier device for amplifying the audio signal played by the playback device, and a speaker device for broadcasting the audio signal as an audio message.
[0030] The broadcasting device 13 is connected to the computer 11 via a network, for example. The broadcasting device 13 may also be connected via a wired connection. The broadcast program 100 generated by the program generation system 10 is output to the broadcasting device 13 via the network. The broadcasting device 13 amplifies the broadcast program 100 in the broadcasting area A1.
[0031] The program generation system 10 comprises an editing unit 21 and a playback unit 22. Each of the editing unit 21 and the playback unit 22 is constructed, for example, by computational elements of a computer 11 and a program incorporated into the computer 11. Each of the editing unit 21 and the playback unit 22 may be a program function realized within a single computer 11. Each of the editing unit 21 and the playback unit 22 may be composed of a computer 11 that is independent of each other. Each of the editing unit 21 and the playback unit 22 may be constructed by computational elements of a connection terminal 12 and a program incorporated into the connection terminal 12.
[0032] The editorial department 21 edits the broadcast program 100. The editing by the editorial department 21 involves combining the generated music data 110 and the generated text-to-speech data 120, which will be described later. As shown in Figure 4, the broadcast program 100 is generated, for example, by combining the generated music data 110 and the generated text-to-speech data 120.
[0033] The broadcast program 100 is generated, for example, as an audio file. Possible file formats for the broadcast program 100 include AIFF (Audio Interchange File Format), AAC (Advanced Audio Coding), FLAC (Free Lossless Audio Codec), WAV (Waveform Audio File Format), and MP3 (MPEG-1 Audio Layer-3). In this embodiment, the file format for the broadcast program 100 is MP3. The file formats for the generated music data 110 and the generated speech data 120 are the same as the file format for the broadcast program 100.
[0034] The playback unit 22 shown in Figure 1 plays back the broadcast program 100 in broadcast area A1. Playback of the broadcast program 100 by the playback unit 22 indicates that the data of the broadcast program 100 is transmitted to the broadcasting device 13. The broadcast of the broadcast program 100 is carried out by the broadcasting device 13 in broadcast area A1.
[0035] The editing unit 21 of the program generation system 10 can be connected to the music generation unit 30 and the audio generation unit 40. The editing unit 21 can also be connected to the information source unit 50.
[0036] The music generation unit 30 is configured to generate music data based on prompts containing composition information related to the composition of a piece of music. The music generation unit 30 includes, for example, a music generation engine 31. The music generation engine 31 is, for example, a deep learning-based generation model that generates music data in response to prompt inputs containing composition information.
[0037] The music generation unit 30 may be provided separately from the program generation system 10, or it may be configured as part of the program generation system 10. In this embodiment, the music generation unit 30 is provided on an external server configured separately from the computer 11. The music generation unit 30 can communicate with the computer 11 via a network. The music generation unit 30 may also be a program function implemented within the external server.
[0038] The speech generation unit 40 is configured to generate speech data based on prompts containing speech information relating to the speech being read. The speech generation unit 40 includes, for example, a speech synthesis engine 41. The speech synthesis engine 41 is, for example, a text-to-speech (TTS) model that generates speech data from text data.
[0039] The voice generation unit 40 may be provided separately from the program generation system 10, or it may be configured as part of the program generation system 10. In this embodiment, the voice generation unit 40 is provided on an external server configured separately from the computer 11. The voice generation unit 40 can communicate with the computer 11 via a network. The voice generation unit 40 may also be a program function implemented within the external server.
[0040] The information source unit 50 stores area information 131. Area information 131 is information related to broadcast area A1. Examples of area information 131 include weather information, traffic information, disaster prevention information, etc., in the vicinity of broadcast area A1. The vicinity of broadcast area A1 includes not only broadcast area A1 but also facilities that include broadcast area A1 and the surrounding areas of those facilities. Area information 131 may also include warnings, greetings, announcements, etc., in the vicinity of broadcast area A1. Announcements include information about facilities in the vicinity of broadcast area A1, event announcements, product announcements, lost child notices, event start times, and event delay notices.
[0041] The information source unit 50 is provided, for example, on an external server configured separately from the computer 11. The information source unit 50 can communicate with the computer 11 via a network. The program generation system 10 connects to the information source unit 50 sequentially to acquire area information 131. The editing unit 21 in this embodiment acquires area information 131 from the information source unit 50. The editing unit 21 acquires area information 131 from the information source unit 50 using, for example, a function such as an API (Application Programming Interface).
[0042] <Input to the program generation system> Figure 2 shows an example of input form F1. Input form F1 is an interface for listeners of broadcast program 100 to input composition information and reading information into the program generation system 10. Input form F1 is displayed, for example, on the display of the connected terminal 12. Input form F1 is displayed by software running on the OS (Operating System) of the connected terminal 12. The software is, for example, an application such as a web browser.
[0043] The composition information includes at least one of the listener preferences for the broadcast program 100, the musical style of the song, and the type of voice used to sing the song. Listener preferences are the image of the song that listeners want to hear, such as, "I want to refresh myself so I can beat the lingering summer heat and enjoy working in the afternoon," or "It's been raining a lot lately, so I thought a cheerful song that would lift my spirits would be nice." Musical styles include "cheerful," "dark," "fun," and "calm." Musical styles may also include genres such as "J-POP" and "K-POP." Types of voices used to sing the song include "female vocals," "male vocals," "high voice," "low voice," "whispering," and "husky voice."
[0044] In the input form F1 in Figure 2, the listener's preferences are entered as composition information in the "2. Image of music to play during lunch break" field. The musical style of the song is entered as composition information in the "3. Mood" field. The type of voice that sings the song is entered as composition information in the "4. Image of the singer" field. In Figure 2, the fields for musical style and type of voice are selectable.
[0045] The spoken information includes the name of the listener who entered the composition information, and at least one of the song's introductions. Input form F1 is where a pair of composition information and spoken information are entered. The song's introduction in the spoken information is the introduction to the generated song data 110, which is generated based on the corresponding composition information.
[0046] In input form F1 in Figure 2, the name of the listener who entered the composition information is entered in the "1. Your Name (Anonymous is acceptable)" field. The song description is entered in the "5. A final message" field as composition information.
[0047] <Creation and playback of broadcast programs> Referring to Figure 3, the flow from the time the broadcast program 100 generated in the program generation system 10 is played back will be explained.
[0048] The editorial department 21 obtains the first composition information 111 as composition information and the first reading information 121 as reading information (S11). The acquisition of the first composition information 111 and the first reading information 121 is performed, for example, by the input form F1 of the connected terminal 12. The first composition information 111 and the first reading information 121 are entered into the input form F1 by the first listener. One or more listeners input the composition information and the reading information. The editorial department 21 obtains the second composition information as composition information and the second reading information as reading information from, for example, the second listener. The editorial department 21 creates a prompt 112 for music generation by combining the first composition information 111 with the standard composition information 113 (S12).
[0049] As shown in Figure 5, the music generation prompt 112 includes an instruction to the music generation unit 30 that reads, "Generate a song based on the following information." The music generation prompt 112 includes the first composition information 111. Specifically, the music generation prompt 112 is created when the first composition information 111 is input into the composition template information 113. The composition template information 113 is pre-configured. The composition template information 113 includes items such as "musical preference," "musical style," and "type of voice." In the "musical preference" field, the listener's preference is input as the first composition information 111. In the "musical style" field, the musical style of the song is input as the first composition information 111. In the "type of voice" field, the type of voice that sings the song is input as the first composition information 111.
[0050] The music generation prompt 112 further includes basic composition information 114. This basic composition information 114 is pre-set. It is set according to the broadcast program 100. The same basic composition information 114 is used even if the listener changes.
[0051] The composition basic information 114 includes at least one of the following: the broadcast period of the broadcast program 100, information about the broadcast area A1, information about the listeners of the broadcast program 100, and the playback time of the music. The broadcast period of the broadcast program 100 indicates the time of day when the broadcast program 100 is broadcast, such as "morning," "afternoon," "evening," "lunch break," or "open for business." Information about the broadcast area A1 relates to the purpose of the location where the broadcast program 100 is broadcast, such as "company cafeteria," "cafe," or "grocery store." Information about the listeners relates to the attributes of the people who listen to the broadcast program 100, such as "employees," "children," or "couples." The playback time of the music indicates the length of the generated music data 110 within the broadcast program 100. The playback time of the music may be specified as, for example, 3 minutes, or 10 songs broadcast within 30 minutes.
[0052] As shown in Figure 3, the editorial department 21 obtains generated music data 110 generated as music data from the music generation unit 30 by inputting a music generation prompt 112 to the music generation unit 30 (S13). The editorial department 21 obtains one or more generated music data 110 depending on the number of submissions from listeners, for example. The editorial department 21 may also cause the music generation unit 30 to generate multiple generated music data 110 from a single submission.
[0053] The processes in S14 and S15 are performed in parallel with the processes in S12 and S13. The editorial department 21 creates a prompt 122 for reading aloud by combining the first reading information 121 with the standard reading information 123 (S14).
[0054] As shown in Figure 6, the prompt 122 for generating speech includes an instruction to the speech generation unit 40: "Please generate audio for the song introduction based on the following text." The prompt 122 for generating speech includes the first speech information 121. Specifically, the prompt 122 for generating speech is created when the first speech information 121 is input into the standard speech information 123. The standard speech information 123 is pre-set. The standard speech information 123 is set to the text: "This is a submission from [Contributor]. We have received the message [Submission Text]. Please listen." [Contributor] is the name of the listener who entered the first composition information 111 as the first speech information 121. Also, [Submission Text] is the song introduction text entered as the first speech information 121. The generated speech data 120 is generated when the above text is read aloud by the speech synthesis engine 41 of the speech generation unit 40.
[0055] The prompt 122 for generating speech further includes basic speech information 124. The basic speech information 124 is pre-set. The basic speech information 124 is set according to the broadcast program 100. The basic speech information 124 is the same information used even if the listener changes. The basic speech information 124 sets the voice for reading the standard speech information 123 into which the first speech information 121 is inserted.
[0056] The basic reading information 124 includes at least one of the following: the type of reading voice, the broadcast period of the broadcast program 100, information about broadcast area A1, information about the listener, and the playback time of the reading voice. The type of reading voice is the type of voice that reads the reading voice. The broadcast period of the broadcast program 100, information about broadcast area A1, and information about the listener each correspond to, for example, common items in the basic composition information 114. The playback time of the reading voice indicates the length of the reading voice within the broadcast program 100.
[0057] As shown in Figure 3, the editing unit 21 obtains generated text-to-speech data 120, which is generated as text-to-speech data, by inputting a text-to-speech generation prompt 122 to the speech generation unit 40 (S15). If multiple generated music data 110 are generated, the editing unit 21 obtains the generated text-to-speech data 120 corresponding to each of the multiple generated music data 110.
[0058] As shown in Figure 3, the editorial department 21 obtains area information 131 from the information source unit 50 (S16). The editorial department 21 creates an additional read-aloud generation prompt 132 by combining the area information 131 with the additional read-aloud standard information 133 (S17).
[0059] As shown in Figure 7, the additional text-to-speech generation prompt 132 includes an instruction to the music generation unit 30 that reads, "Please generate voice for announcement based on the following text." The additional text-to-speech generation prompt 132 is created, for example, when area information 131 is input into the additional text-to-speech template information 133. In the example in Figure 7, the additional text-to-speech template information 133 is set to the text, "The current weather is [weather]. The temperature around the public address area is [temperature]." As an example, the additional text-to-speech data 130 is generated when the text entered as area information 131 with "weather" for weather and "temperature" for outside temperature is read aloud by the speech synthesis engine 41 of the speech generation unit 40. The additional text-to-speech generation prompt 132 includes basic text-to-speech information 124 that is similar to the text-to-speech generation prompt 122. The basic reading information 124 for the additional text-to-speech generation prompt 132 may be the same as, or different from, the basic reading information 124 for the text-to-speech generation prompt 122.
[0060] As shown in Figure 3, the editorial unit 21 obtains additional read-aloud data 130, including area information 131, by inputting an additional read-aloud generation prompt 132 to the speech generation unit 40 (S18).
[0061] The editorial department 21 generates a broadcast program 100 by combining the generated music data 110 and the generated text-to-speech data 120 (S19). The broadcast program 100 is generated by setting the broadcast order of the generated music data 110 and the generated text-to-speech data 120. The generated broadcast program 100 is stored, for example, in the storage device of the computer 11.
[0062] As shown in Figure 4, the editorial department 21 obtains, for example, multiple generated song data 110 and multiple generated text-to-speech data 120 depending on the number of submissions from listeners. In one example, generated song data 110 includes first generated song data 110A and second generated song data 110B. Generated text-to-speech data 120 includes first generated text-to-speech data 120A and second generated text-to-speech data 120B. First generated text-to-speech data 120A is text-to-speech data relating to first generated song data 110A. Second generated text-to-speech data 120B is text-to-speech data relating to second generated song data 110B.
[0063] The editorial department 21 edits the broadcast program 100 so that, for example, the generated text-to-speech data 120 corresponding to the generated music data 110 is played before the generated music data 110 is played. Specifically, the editorial department 21 edits the broadcast program 100 so that the first generated text-to-speech data 110A is played after the first generated text-to-speech data 120A, the second generated text-to-speech data 120B is played after the first generated music data 110A, and the second generated music data 110B is played after the second generated text-to-speech data 120B.
[0064] Preferably, the editorial department 21 includes additional read-aloud data 130 in the broadcast program 100. The additional read-aloud data 130 is played, for example, between the first generated music data 110A and the second generated read-aloud data 120B, or after the second generated music data 110B. In the example in Figure 4, the additional read-aloud data 130 is played after the second generated music data 110B.
[0065] The playback unit 22 plays the broadcast program 100 in broadcast area A1 (S21). Specifically, the playback unit 22 causes the broadcasting device 13 to play the broadcast program 100 using a streaming protocol such as RTSP (Real Time Streaming Protocol). The broadcasting device 13 may store the audio data of the broadcast program 100, and the audio data of the broadcast program 100 stored by the broadcasting device 13 may be broadcast. The playback unit 22 may also cause the connected terminal 12 to play the broadcast program 100 for each listener.
[0066] <Operation of this embodiment> The first operation of this embodiment will now be described. The program generation system 10 acquires music data generated based on input composition information, and also acquires spoken-word data generated based on input spoken-word information. In this way, a broadcast program 100 can be generated on the program generation system 10 from the composition information and spoken-word information.
[0067] The second operation of this embodiment will now be described. In the program generation system 10, a broadcast program 100 is generated based on the composition information and reading information entered by the listener, so the listener can feel as if they are participating in the broadcast program 100.
[0068] <Effects of this embodiment> The effects of this embodiment will now be explained. (1-1) The program generation system 10 generates a broadcast program 100 that includes music and spoken audio. The program generation system 10 includes an editing unit 21 that edits the broadcast program 100 and a playback unit 22 that plays the broadcast program 100 in broadcast area A1. The playback unit 22 is connectable to the music generation unit 30 and the audio generation unit 40. The music generation unit 30 is configured to generate music data based on prompts that include composition information relating to the composition of music. The audio generation unit 40 is configured to generate spoken audio data based on prompts that include spoken audio relating to spoken audio. The editing unit 21 acquires first composition information 111 as composition information and first spoken audio information 121 as spoken audio information. The editing unit 21 acquires generated music data 110 that is generated as music data by inputting a music generation prompt 112 that includes the first composition information 111 to the music generation unit 30. The editorial department 21 obtains generated speech data 120, which is generated as speech data, by inputting a speech generation prompt 122 containing the first speech information 121 to the speech generation unit 40. The editorial department 21 generates a broadcast program 100 by combining the generated music data 110 and the generated speech data 120.
[0069] In this configuration, the editorial department 21 generates a broadcast program 100 by combining generated music data 110 based on the first composition information 111 and generated reading data 120 based on the first reading information 121. Since the generated music data 110 and generated reading data 120 acquired on the program generation system 10 are used in the generation of the broadcast program 100, it eliminates the need for the broadcaster to prepare music data and reading data in advance.
[0070] (1-2) The composition information includes at least one of the listener preferences of the broadcast program 100, the musical style of the song, and the type of voice that sings the song. This configuration makes it possible to generate song data that reflects at least one of the listener preferences, the musical style of the song, and the type of voice that sings the song.
[0071] (1-3) The music generation prompt 112 further includes basic composition information 114 which includes at least one of the broadcast period of the broadcast program 100, information about the broadcast area A1, information about the listeners, and the playback time of the music. With this configuration, music data is generated which reflects at least one of the broadcast period of the broadcast program 100, information about the broadcast area A1, information about the listeners, and the playback time of the music.
[0072] (1-4) The spoken information includes the name of the listener who entered the composition information, and at least one of the song's introduction. This configuration makes it possible to generate spoken data that includes the listener's name and at least one of the song's introduction.
[0073] (1-5) The prompt 122 for generating speech further includes basic speech information 124, which includes at least one of the voice type of the speech to be spoken, the broadcast time of the broadcast program 100, information about the broadcast area A1, information about the listener, and the playback time of the speech to be spoken.
[0074] This configuration makes it possible to generate speech data that reflects at least one of the following: the type of voice used for the speech, the broadcast period of the broadcast program 100, information about the broadcast area A1, information about the listener, and the playback time of the speech.
[0075] (1-6) The generated music data 110 includes first generated music data 110A and second generated music data 110B. The generated text-to-speech data 120 includes first generated text-to-speech data 120A relating to first generated music data 110A and second generated text-to-speech data 120B relating to second generated music data 110B. The editorial department 21 edits the broadcast program 100 so that first generated music data 110A is played after first generated text-to-speech data 120A, second generated text-to-speech data 120B is played after first generated music data 110A, and second generated music data 110B is played after second generated text-to-speech data 120B.
[0076] With this configuration, the first generated audio data 110A is played after the first generated audio data 120A, making it easier for the listener to understand the content of the first generated audio data 120A compared to when they are played separately. Similarly, the second generated audio data 110B is played after the second generated audio data 120B, making it easier for the listener to understand the content of the second generated audio data 120B compared to when they are played separately.
[0077] (1-7) The program generation system 10 is connected to an information source unit 50 that stores area information 131 relating to broadcast area A1. The editing unit 21 obtains the area information 131 from the information source unit 50. The editing unit 21 includes additional read-aloud data 130 containing the area information 131 in the broadcast program 100. With this configuration, the content of the area information 131 is included in the broadcast program 100, so that listeners of the broadcast program 100 can grasp the area information 131.
[0078] <Second Embodiment> The program generation system 10 according to the second embodiment will be described with reference to Figures 8 to 11. In this embodiment, components that are substantially unchanged from those of the first embodiment are given the same reference numerals as those of the first embodiment, and their descriptions are omitted.
[0079] As shown in Figure 8, the editing unit 21 of the program generation system 10 in this embodiment can be further connected to a text generation unit 60. The text generation unit 60 has, for example, a large language model (LLM) 61. The large language model 61 is a machine learning model that has been trained on a vast amount of linguistic information. The text generation unit 60 may be provided separately from the program generation system 10, or it may be configured as part of the program generation system 10.
[0080] In this embodiment, the text generation unit 60 is located on an external server configured separately from the computer 11. The text generation unit 60 can communicate with the computer 11 via a network. The text generation unit 60 may also be a program function implemented within the external server.
[0081] In the program generation system 10 of this embodiment, the music generation prompt 112 input to the music generation unit 30 is generated by the text generation unit 60. The text generation unit 60 is configured to generate the music generation prompt 112 based on a prompt that includes composition information. The editing unit 21 obtains the music generation prompt 112 by inputting a first pre-processing prompt 140 to the text generation unit 60. The first pre-processing prompt 140 includes first composition information 111.
[0082] Referring to Figure 9, the flow of how a broadcast program 100 is generated in the program generation system 10 of this embodiment will be described. In the flow of Figure 9, steps S16 to S18 in Figure 3 may also be performed. Note that steps S11 to S15 and S19 in Figure 9 are the same as in the flow of Figure 3, so the common explanation will be omitted. The editing unit 21 of this embodiment creates a first pre-processing prompt 140 based on the first composition information 111 (S31). Note that the first pre-processing prompt 140 corresponds to the pre-processing prompt of the "means for solving the problem".
[0083] As shown in Figure 10, the first pre-processing prompt 140 is created by combining the first instruction information 141 with the first composition information 111 and the standard composition information 113 in the first embodiment. In the example in Figure 10, the first instruction information 141 is an instruction to the text generation unit 60 that says, "Please briefly answer in English the Title, Style, and Description of a song that fits the following content." Depending on the model of the music generation engine 31, a favorable result may be obtained by inputting the instruction in English. The first instruction information 141 may also include instructions that specify detailed composition elements such as the instruments and beats to be played in the generated music data 110.
[0084] In S12 of Figure 9, the editorial department 21 creates a music generation prompt 112 by inputting a first pre-processing prompt 140 to the text generation unit 60. As shown in Figure 11, in this embodiment, the music generation prompt 112 associates the first composition information 111, which has been processed by the text generation unit 60 based on the first instruction information 141, with the standard composition information 113. In the example in Figure 11, the instructions for the music generation unit 30 and the basic composition information 114 are displayed in English. The instructions for the music generation unit 30 and the basic composition information 114 may be included in the first pre-processing prompt 140 and also generated by the text generation unit 60 together with the first composition information 111.
[0085] <Effects of this embodiment> The effects of this embodiment will now be explained. (2-1) The editorial department 21 obtains a music generation prompt 112 from the text generation unit 60 by inputting a first preprocessing prompt 140 containing the first composition information 111 to the text generation unit 60. With this configuration, since the music generation prompt 112 is obtained by inputting the first preprocessing prompt 140 to the text generation unit 60, it is possible to generate a highly accurate music generation prompt 112 that reflects the first composition information 111.
[0086] <Variation> The embodiments described above are illustrative of possible forms of the program generation system 10 and are not intended to limit its form. The program generation system 10 may take forms different from those illustrated in the embodiments described above. Examples include forms in which parts of the configuration of each embodiment are replaced, modified, or omitted, or forms in which new configurations are added to the embodiments. Modifications of each embodiment are shown below.
[0087] <First variation> The music played from the generated music data 110 may have a mood assigned to it, such as "bright" or "dark." When the editorial department 21 obtains multiple generated music data 110, it assigns a mood to each of the multiple generated music data 110 that is suitable for the music played from the generated music data 110. The mood is set, for example, based on the mood of the music generation prompt 112. The editorial department 21 associates a label related to the mood with the generated music data 110 generated by the music generation unit 30, for example.
[0088] The editorial department 21 edits the broadcast program 100 so that generated music data 110 with a common musical style are played consecutively. For example, if generated music data 110 with multiple "cheerful" musical styles and generated music data 110 with multiple "dark" musical styles are obtained, playback of the generated music data 110 with "dark" musical styles will begin only after all of the generated music data 110 with "cheerful" musical styles have been played.
[0089] According to the configuration of this modified version, generated music data 110 with a common musical style are played continuously, making it easier for the listener to grasp the atmosphere of the music.
[0090] <Second variation> The editorial department 21, upon acquiring multiple generated music data 110, obtains the popularity ranking of each of the multiple generated music data 110. The editorial department 21 may also edit the broadcast program 100 so that the multiple generated music data 110 are played in order of popularity. The playback unit 22, for example, transmits the generated music data 110 to the connected terminal 12 using a streaming protocol or the like before the broadcast program 100 is broadcast, allowing listeners to listen to each of the multiple generated music data 110. Listeners who have listened to the generated music data 110 vote for their favorite generated music data 110 via the network. The editorial department 21 obtains the popularity ranking of each of the multiple generated music data 110 based on the number of votes from listeners.
[0091] For listeners to hear the generated music data 110, for example, a web page on the internet may contain URLs to each of the multiple generated music data 110. The listener can then listen to the multiple generated music data 110 on this web page.
[0092] Referring to Figure 12, the flow of how the broadcast program 100 is generated in the program generation system 10 of this modified example will be explained. In the flow of Figure 12, steps S16 to S18 in Figure 3 may also be performed. Note that steps S11 to S15 and S19 in Figure 12 are the same as in the flow of Figure 3, so the common explanation will be omitted.
[0093] In this modified example, the editorial department 21, when it obtains multiple generated music data 110 in S13, obtains the popularity ranking of each of the multiple generated music data 110 (S41). In S19, it generates a broadcast program 100 so that the multiple generated music data 110 are played in order of popularity ranking.
[0094] According to this modified configuration, generated music data 110 that are popular with listeners are given priority for playback, making it easier for listeners to hear popular generated music data 110. Therefore, listeners have more opportunities to hear popular generated music data 110.
[0095] <Third variation> As shown in Figure 13, the program generation system 10 may also be connectable to a camera 14. The camera 14 photographs the broadcast area A1.
[0096] The editorial department 21 acquires area information 131 from the video captured by the camera 14. Examples of area information 131 in this modified example include information about the environment in broadcast area A1, such as weather, brightness, scenery, and the degree of crowding. One method for acquiring area information 131 from video is to input the video into a model similar to the large-scale language model 61 in the second embodiment, thereby obtaining text related to the area information 131 as output. Alternatively, area information 131 may be acquired from video by image detection.
[0097] According to the configuration of this modified example, the broadcast program 100 includes the content of area information 131, so listeners of the broadcast program 100 can grasp the area information 131. Since the area information 131 is obtained from the video footage of broadcast area A1 captured by the camera 14, information about the area surrounding broadcast area A1 can be obtained in real time.
[0098] <Fourth variation> As shown in Figure 14, in the second embodiment, the text generation unit 60 may generate the text-to-speech prompt 122. In this modified example, the editing unit 21 creates, for example, a second pre-processing prompt 150 (S51). The text generation unit 60 receives the second pre-processing prompt 150 to generate the text-to-speech prompt 122 (S14).
[0099] As shown in Figure 15, the second preprocessing prompt 150 includes the first composition information 111. The second preprocessing prompt 150 is created by combining the first composition information 111 and the standard composition information 113 in the first embodiment with the second instruction information 151 which states, "Generate an introductory text for the music shown in the following image." In the example in Figure 15, the text generation unit 60 generates the introductory text for the music from the first read-aloud information 121 of the read-aloud generation prompt 122.
[0100] <Other variations> The generated music data 110 output by the audio generation unit 40 may include not only music data but also lyrics, music video, and other video data. In this modified example, the file format of the generated music data 110 is, for example, MP4 (MPEG-4).
[0101] In the broadcast program 100, the playback order of the generated music data 110, the generated text-to-speech data 120, and the additional text-to-speech data 130 may be changed as appropriate. For example, a learning model similar to the large-scale language model 61 in the second embodiment may be made to propose the playback order. The learning model can be made to propose the playback order within the broadcast program 100 by specifying the playback order so that the musical style of the generated music data 110 is not biased, or by prioritizing the playback of the additional text-to-speech data 130, which is of high urgency.
[0102] The broadcast program 100 may play multiple generated music data 110 consecutively, followed by the playback of generated text-to-speech data 120 corresponding to each of the multiple generated music data 110. For example, generated music data 110 with common tastes, musical styles, etc., may be played consecutively.
[0103] The composition basic information 114 may be omitted from the music generation prompt 112. In this modified example, the music generation prompt 112 includes the first composition information 111 and the standard composition information 113.
[0104] The basic reading information 124 may be omitted from the reading generation prompt 122. In this modified example, the reading generation prompt 122 includes the first reading information 121 and the standard reading information 123.
[0105] The editorial department 21 may edit the broadcast program 100 so that generated music data 110 with shared tastes are consecutive. For example, the editorial department 21 determines that the tastes of the generated music data 110 are shared if the keywords included in the listener's tastes among the first composition information 111 are the same.
[0106] • Additional read-aloud data 130 may include date and time information regarding the broadcast date and time of broadcast program 100. Examples of date and time information include holidays, birthdays, the birth flower for the broadcast month, the meaning of that birth flower, and sunrise and sunset times.
[0107] • In generating the generated music data 110, the playback time of the generated music data 110 may be determined according to the number of submissions from listeners. For example, if there is one submission, a longer playback time may be specified than when there are two or more submissions. For example, if there are many submissions, the playback time per song may be shortened to create a medley of songs.
[0108] The editorial department 21 may set the playback order of the multiple generated music data 110 as appropriate when it obtains multiple generated music data 110. For example, the editorial department 21 may edit the broadcast program 100 so that the music is played in order of the listeners' submission times, from earliest to latest. For example, the editorial department 21 may edit the broadcast program 100 so that each of the multiple generated music data 110 is played in a random order.
[0109] In each embodiment, the broadcast program 100 is generated as an audio file, but the broadcast program 100 may also be generated by streaming playback of the generated music data 110 and the generated text-to-speech data 120 in playback order using a streaming protocol or the like. For example, the audio file of the generated music data 110 is generated on a first server. The audio file of the generated text-to-speech data 120 is generated on a second server. In this modified example, the playback unit 22 causes the broadcasting device 13 to stream playback of the generated music data 110 from the first server and the generated text-to-speech data 120 from the second server in the playback order within the broadcast program 100. The first server may be, for example, the server on which the music generation unit 30 is configured, the computer 11 acting as a server, or any other server. The second server may be, for example, the server on which the audio generation unit 40 is configured, the computer 11 acting as a server, or any other server. In this modified example, determining the playback order within the broadcast program 100 corresponds to generating the broadcast program 100. Downloading the generated music data 110 for streaming playback corresponds to obtaining the generated music data 110 from the music generation unit 30. Similarly, downloading the generated text-to-speech data 120 for streaming playback corresponds to obtaining the generated text-to-speech data 120 from the audio generation unit 40.
[0110] A program generation program conforming to the program generation system 10 may be configured. The program generation program is a program that generates a broadcast program 100 including music and spoken audio. The program generation program causes the computer 11 to execute a first step, a second step, a third step, and a fourth step. In the first step, the program generation program causes the computer 11 to acquire first composition information 111 as composition information relating to the composition of music, and first spoken information 121 as spoken audio relating to spoken audio. In the second step, the program generation program causes the computer 11 to acquire generated music data 110 generated as music data from the music generation unit 30 by causing the music generation unit 30 to input a music generation prompt 112 including the first composition information 111. In the second step, the program generation program further causes the computer 11 to create a music generation prompt 112. In the third step, the program generation program causes the computer 11 to input a prompt 122 for generating speech, which includes the first speech information 121, to the speech generation unit 40, thereby obtaining generated speech data 120 generated as speech data from the speech generation unit 40. In the third step, the program generation program further causes the computer 11 to create a prompt 122 for generating speech. In the fourth step, the program generation program causes the computer 11 to generate a broadcast program 100 by combining the generated music data 110 and the generated speech data 120. Each step of the program generation program may be added or modified according to the configuration of the program generation system 10 in each embodiment or each modified example. With this configuration, since the broadcast program 100 is generated using the newly generated generated music data 110 and generated speech data 120, it is not necessary for the broadcaster to prepare music data and speech data in advance.
[0111] In the program generation system 10 of Figure 3, the flow until a broadcast program 100 is generated may be configured as a program generation method. The program generation method is a method by which a computer 11 generates a broadcast program 100 that includes music and spoken audio. The program generation method includes a first step, a second step, a third step, and a fourth step. In the first step, the computer 11 acquires first composition information 111 as composition information relating to the composition of music, and first spoken information 121 as spoken audio relating to spoken audio. In the second step, the computer 11 acquires generated music data 110 that is generated as music data from the music generation unit 30 by inputting a music generation prompt 112 that includes the first composition information 111 to the music generation unit 30. In the second step, the computer 11 creates a music generation prompt 112. In the third step, the computer 11 inputs a prompt 122 containing the first reading information 121 to the speech generation unit 40, thereby obtaining generated reading data 120 generated as reading data from the speech generation unit 40. In the third step, the computer 11 creates the prompt 122. In the fourth step, the computer 11 generates a broadcast program 100 by combining the generated music data 110 and the generated reading data 120. Each step of the program generation method may be added or modified according to the configuration of the program generation system 10 in each embodiment or each modified example. With this configuration, since the generated music data 110 and the generated reading data 120 included in the broadcast program 100 are generated in the process of the program generation method, it is not necessary for the broadcaster to prepare the music data and reading data in advance. [Explanation of symbols]
[0112] 10...Program generation system, 11...Computer, 12...Connection terminal, 13...Broadcasting equipment, 14...Filming equipment, 21...Editing department, 22...Playback department, 30...Music generation department, 31...Music generation engine, 40...Speech generation department, 41...Speech synthesis engine, 50...Information source department, 60...Text generation department, 61...Large-scale language model.
Claims
1. A program generation system that generates broadcast programs including music and spoken audio, The system comprises an editing unit for editing the aforementioned broadcast program, and a playback unit for playing the aforementioned broadcast program in the broadcast area. The aforementioned editorial unit is connectable to the music generation unit and the audio generation unit. The music generation unit is configured to generate music data based on prompts that include composition information relating to the composition of the music. The voice generation unit is configured to generate speech data based on prompts that include speech information relating to the speech being read aloud. The aforementioned editorial department, The first composition information as composition information and the first reading information as reading information are obtained, By inputting a prompt for music generation including the first composition information to the music generation unit, the generated music data generated by the music generation unit as music data is obtained. By inputting a prompt for generating speech including the first speech information to the speech generation unit, the generated speech data generated as speech data from the speech generation unit is obtained. A program generation system that generates the broadcast program by combining the generated music data and the generated text-to-speech data.
2. The program generation system according to claim 1, wherein the composition information includes at least one of the preferences of listeners of the broadcast program, the musical style of the song, and the type of voice that sings the song.
3. The program generation system according to claim 2, wherein the prompt for generating music further includes basic composition information, which includes at least one of the broadcast time of the broadcast program, information regarding the broadcast area, information regarding the listeners of the broadcast program, and the playback time of the music.
4. The program generation system according to claim 1, wherein the read-aloud information includes at least one of the name of the listener who entered the composition information and an introductory text for the song.
5. The program generation system according to claim 4, wherein the prompt for generating the reading aloud further includes basic reading information, which includes at least one of the type of reading voice, the broadcast time of the broadcast program, information regarding the broadcast area, information regarding the listener, and the playback time of the reading voice.
6. The generated music data includes first generated music data and second generated music data. The generated text-to-speech data includes first generated text-to-speech data relating to the first generated music data and second generated text-to-speech data relating to the second generated music data. The program generation system according to claim 1, wherein the editorial department edits the broadcast program such that the first generated music data is played after the first generated text-to-speech data, the second generated text-to-speech data is played after the first generated music data, and the second generated music data is played after the second generated text-to-speech data.
7. The aforementioned editorial department, When multiple generated music data are obtained, a musical style suitable for the music to be played from the generated music data is set for each of the multiple generated music data. The program generation system according to claim 1, wherein the generated music data having a common musical style are edited so that the broadcast program is continuous.
8. The aforementioned editorial department, When multiple generated song data sets are obtained, the popularity ranking of each of the multiple generated song data sets is obtained. The program generation system according to claim 1, wherein the broadcast program is edited so that multiple generated song data are played in order of popularity.
9. The aforementioned editorial department, Furthermore, it can be connected to an information source unit that stores area information related to the aforementioned broadcast area, The program generation system according to claim 1, which acquires the area information from the information source unit and includes additional read-aloud data containing the area information in the broadcast program.
10. The aforementioned editorial department, Furthermore, it can be connected to a camera that films the aforementioned broadcast area, The program generation system according to claim 1, wherein the camera acquires area information relating to the broadcast area from the video footage captured by the camera, and includes additional read-aloud data containing the area information in the broadcast program.
11. The aforementioned editorial department, Furthermore, it is connectable to a text generation unit configured to generate the music generation prompt based on the prompt containing the music composition information, A program generation system according to any one of claims 1 to 10, wherein the program generation system obtains the music generation prompt from the text generation unit by inputting a preprocessing prompt including the first music composition information to the text generation unit.
12. A program generation program that generates a broadcast program including music and spoken audio, On the computer, A first step of obtaining first composition information as composition information relating to the composition of the aforementioned musical piece, and first reading information as reading information relating to the aforementioned reading voice, A second step involves causing a music generation unit, which is configured to generate music data based on a prompt containing the aforementioned composition information, to receive a music generation prompt containing the first composition information, thereby obtaining the generated music data generated as music data from the music generation unit. A third step involves causing a speech generation unit, configured to generate speech data based on a prompt containing the aforementioned speech information, to receive a speech generation prompt containing the first speech information, thereby obtaining the generated speech data generated as the speech data from the speech generation unit. A program generation program that performs a fourth step of generating the broadcast program by combining the generated music data and the generated speech data.
13. A program generation method in which a computer generates a broadcast program that includes music and spoken audio, The first step is for the computer to acquire first composition information as composition information relating to the composition of the musical piece, and first reading information as reading information relating to the reading voice. The second step involves the computer acquiring generated music data generated as music data from a music generation unit, which is configured to generate music data based on a prompt containing the composition information, by inputting a music generation prompt containing the first composition information into the music generation unit. A third step is to obtain generated reading data generated as reading data from a speech generation unit by inputting a prompt for reading generation including the first reading information to a speech generation unit configured to generate reading data based on a prompt including the reading information, A program generation method comprising a fourth step in which the computer generates the broadcast program by combining the generated music data and the generated speech data.
Citation Information
Patent Citations
Private broadcasting apparatus
JP2007067465A