Time stamp generation method, system, device, and program
The use of a multimodal large-scale language model to generate timestamp data for song lyrics addresses the inefficiencies of existing methods, providing a faster and more accurate way to synchronize lyrics with music playback.
Patent Information
- Application Number
- JP2025075424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing methods for generating timestamp data for song lyrics are labor-intensive and time-consuming.
A method and system that utilizes a multimodal large-scale language model to automatically generate timestamp data by inputting lyrics and music data, allowing for efficient and reliable timestamp assignment without the need for specialized machine learning models.
Enables less labor-intensive and time-consuming generation of reliable lyric timestamp data, ensuring accurate synchronization of lyrics with music playback.
Smart Images

Figure 0007770605000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a timestamp generation method, system, device, and program. [Background technology]
[0002] A technique is known in which lyrics at a playback position are displayed in synchronization with playback of a song by using timestamp data that defines the timing of singing lyrics of the song.
[0003] Patent Document 1 discloses that it is possible to synchronize text with speech contained in an audio signal using a machine learning model. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2024-137877 Summary of the Invention [Problem to be solved by the invention]
[0005] There is a need for a less labor-intensive way to generate timestamp data.
[0006] In one aspect, the present disclosure has been made in view of the above circumstances, and an object of the present disclosure is to generate timestamp data for lyrics in a less time-consuming manner. [Means for solving the problem]
[0007] A timestamp generation method according to one aspect of the present disclosure includes: One or more hardware processors The first step is to get the lyrics of the song, a second step of acquiring music data including audio data of the music; The timestamp generation method includes a third step of inputting the lyrics, an instruction to assign a timestamp to the lyrics, and the music data into a multimodal large-scale language model.
[0008] A timestamp generation system according to one aspect of the present disclosure includes: one or more hardware processors, Get the lyrics of the song acquiring music data including audio data of the music piece; The timestamp generation system inputs the lyrics, an instruction to add a timestamp to the lyrics, and the music data into a multimodal large-scale language model.
[0009] A timestamp generation program according to an embodiment of the present disclosure includes: One or more hardware processors The first step is to get the lyrics of the song, a second step of acquiring music data including audio data of the music; and a third step of inputting the lyrics, an instruction to assign a timestamp to the lyrics, and the music data into a multimodal large-scale language model. [Effects of the Invention]
[0010] In one aspect, the present disclosure provides a less labor-intensive method for generating reliable lyric timestamp data. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating the system configuration of a timestamp generation system according to the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating a hardware configuration related to the timestamp generation system of the present disclosure. [Figure 3]FIG. 3 is a diagram illustrating a functional configuration related to the timestamp generation system of the present disclosure. [Figure 4] FIG. 4 is a sequence diagram illustrating the processing flow of the timestamp generation system of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating the structure of a multimodal large-scale language model of the timestamp generation system of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example of a display screen of the timestamp generation system of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating a functional configuration related to a timestamp generation system according to the first modification. [Figure 8] FIG. 8 is a diagram illustrating a functional configuration related to a timestamp generation system according to the second modification. [Figure 9] FIG. 9 is a sequence diagram illustrating the processing flow of the timestamp generation system according to the second modification. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the drawings, identical parts are designated by the same reference numerals and their description will be omitted. The configurations of the embodiments described below, and the actions and effects brought about by such configurations, are merely examples and are not limited to the following description. In addition, ordinal numbers such as "first" and "second" are used as necessary below, but these ordinal numbers are used for convenience of identification and do not indicate any particular meaning, such as a particular priority, unless otherwise specified. When comparing the magnitude relationship of two numerical values, either of the two criteria of "greater than or equal to" and "greater than" may be used, unless otherwise specified.
[0013] [System Overview] The timestamp generation system according to the present disclosure is a computer system that has the function of generating timestamp data that contains timestamp information indicating the lyrics of a song and the timing at which the lyrics are sung. The timestamp generation system generates the timestamp data using song data and its lyrics. The song data is data that includes the voice of the song lyrics being sung.
[0014] The timestamp generation system provides songs to users by transmitting song data and timestamp data to the user terminal. The song data may be content data provided by a distributor terminal 20, which is a user terminal used by the distributor, that has undergone predetermined processing, or may be the content data itself. The content data is, for example, audio data and video data. Examples of predetermined processing include extracting audio data from video data provided by the distributor terminal 20 and re-encoding audio data provided by the distributor terminal 20. A distributor is someone who intends to provide a song to a player using the timestamp generation system. A player is someone who intends to use a song using the timestamp generation system.
[0015] The timestamp generation system stores music data and generated timestamp data in storage devices such as an origin server and a cache server, and transmits the stored music data and timestamp data to a player terminal 30, which is a user terminal used by the player, in response to a request from the player. The player terminal 30 processes the music data and timestamp data to play the music and display the lyrics of the music on a display device. If there are multiple player terminals 30, it is sufficient that at least one of the multiple player terminals 30 plays the music and displays the lyrics of the music. In this disclosure, the distributor and the player will be collectively referred to as the user, and the distributor terminal 20 and the player terminal 30 will be collectively referred to as the user terminal.
[0016] [System Configuration] FIG. 1 is an overall configuration diagram of a timestamp generation system 1 according to an embodiment of the present disclosure. As shown in FIG. 1, the timestamp generation system 1 includes a server 10, a distributor terminal 20, and a player terminal 30. The server 10, the distributor terminal 20, and the player terminal 30 are connected to a network N such as the Internet and are capable of communicating with each other. The timestamp generation system 1 of this embodiment will be described assuming a known client-server timestamp generation system, but is not limited to this. The number of distributor terminals 20, servers 10, and player terminals 30 illustrated in FIG. 1 is not limited to the number shown. The communication network N may include the Internet or an intranet.
[0017] The server 10 may be configured by one or more computers. When the server 10 is configured by multiple computers, these computers may be connected to each other via a communication network N to logically configure a single server 10.
[0018] The distributor terminal 20 is a computer used by a distributor. In one example, the distributor terminal 20 has a function of accessing the timestamp generation system 1 and transmitting content data. The distributor terminal 20 may be a smartphone, a tablet terminal, a head-mounted display (HMD), smart glasses, a smart watch, a mobile terminal such as a laptop personal computer or a mobile phone, a stationary terminal such as a desktop personal computer, or a recording system having a function of recording and transmitting music.
[0019] The player terminal 30 is a computer used by a player. In one example, the player terminal 30 has a function of accessing the timestamp generation system 1 to receive and play music data, and a function of accessing the timestamp generation system 1 to receive timestamp data and display the lyrics of the music. The player terminal 30 may be a smartphone, a tablet terminal, a head-mounted display (HMD), smart glasses, a smart watch, a mobile terminal such as a laptop personal computer or a mobile phone, or a stationary terminal such as a desktop personal computer.
[0020] A distributor operates the distributor terminal 20 to log in to the timestamp generation system 1, thereby being able to provide music to a player. A player operates the player terminal 30 to log in to the timestamp generation system 1, thereby being able to play music. This disclosure is based on the premise that the player and distributor of the timestamp generation system 1 are already logged in. The player may be omitted from logging in to the timestamp generation system 1. In other words, the timestamp generation system 1 may transmit content data to the player terminal 30 of a player who is not logged in. In this case, a player who is not logged in can also play the content.
[0021] FIG. 2 is a block diagram showing a hardware configuration related to a timestamp generation system 1 according to an embodiment of the present disclosure. As an example, a server computer 100 includes a processor 101, a main memory 102, an auxiliary memory 103, and a communication unit 104. The processor 101 is a computing device that executes an operating system and application programs, and is, for example, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). The main memory 102 temporarily stores programs to be executed, computation results, etc., and is generally capable of reading and writing data faster than the auxiliary memory 103 (described later), and is configured with a volatile storage medium such as a dynamic random access memory (DRAM). The auxiliary memory 103 is generally capable of storing larger amounts of data than the main memory 102, and is configured with a nonvolatile storage medium such as a hard disk or flash memory. The auxiliary storage unit 103 stores a server program P1 and various data for causing the server computer 100 to function as the server 10. The communication unit 104 is a device that executes data communication with other computers via the communication network N.
[0022] In this embodiment, the timestamp generation program is implemented as a server program P1. Each functional element of the server 10 is realized by loading the server program P1 onto the processor 101 or the main memory unit 102 and having the processor 101 execute the program. The server program P1 includes code for realizing each functional element of the server 10. The processor 101 operates the communication unit 104 in accordance with the server program P1, and executes reading and writing of data from and to the main memory unit 102 or the auxiliary memory unit 103.
[0023] As an example, the terminal computer 200 includes, as hardware components, a processor 201, a main memory 202, an auxiliary memory 203, a communication unit 204, an input interface 205, and an output interface 206. The processor 201 is a computing device that executes an operating system and application programs, and is, for example, a CPU, a GPU, an ASIC, or an FPGA. The main memory 202 is a device that temporarily stores programs to be executed, calculation results, etc., and is generally capable of reading and writing data faster than the auxiliary memory 203 described below, and is composed of a volatile storage medium such as a DRAM (Dynamic Random Access Memory). The auxiliary memory 203 is generally capable of storing larger amounts of data than the main memory 202, and is composed of a non-volatile storage medium such as a hard disk or flash memory. The auxiliary memory 203 stores a client program P2 and various data for causing the terminal computer 200 to function as the distributor terminal 20 or the player terminal 30. The communication unit 204 is a device that executes data communication with other computers via the communication network N. The input interface 205 is a device that accepts data based on user operations or actions, and is composed of at least one of a keyboard, operation buttons, a pointing device, a touch panel, a microphone, a sensor, and a camera. The output interface 206 is a device that outputs data processed by the terminal computer 200, and is composed of, for example, a display device such as a display and an electro-acoustic conversion device such as a speaker.
[0024] Each functional element of the distributor terminal 20 or the player terminal 30 is realized by loading the corresponding client program P2 into the main memory unit 202 and having the processor 201 execute the program. The client program P2 includes code for realizing each functional element of the distributor terminal 20 or the player terminal 30. The processor 201 operates the communication unit 204, the input interface 205, or the output interface 206 in accordance with the client program P2, and reads and writes data from and to the main memory unit 202 or the auxiliary memory unit 203.
[0025] At least one of the server program P1 and the client program P2 may be provided by being non-temporarily recorded on a tangible recording medium such as a magnetic tape, a magneto-optical disk, an optical disk, a magnetic disk, or a semiconductor memory. Alternatively, at least one of these programs may be provided as a data signal superimposed on a carrier wave via the communications network N. These programs may be provided separately or together.
[0026] FIG. 3 is a diagram illustrating an example of a functional configuration related to the timestamp generation system 1. The server 10 includes functional elements: a content registration unit 11, a lyrics acquisition unit 12, a music data acquisition unit 13, an input unit 14, a generation unit 15, a determination unit 16, and a transmission unit 17. The content registration unit 11 is a functional element that receives content data transmitted from the distributor terminal 20 and records content based on the data in a storage device. The lyrics acquisition unit 12 is a functional element that receives lyrics of music transmitted from the distributor terminal 20. The lyrics acquisition unit 12 may record lyrics transmitted from the distributor terminal 20 together with the content data in a storage device, or may record lyrics transmitted from the distributor terminal 20 in a storage device after recording the content data in a storage device. The music data acquisition unit 13 is a functional element that acquires music data by performing a predetermined process on content data associated with identification information of a music piece specified by the distributor terminal 20, based on a request from the distributor terminal 20. The music data acquisition unit 13 may acquire music data based on receiving lyrics from the distributor terminal 20. The input unit 14 is a functional element that inputs the lyrics of a song, an instruction to add a timestamp to the lyrics, and the song data of the song to the generation unit 15. The generation unit 15 is a functional element that generates timestamp data based on the lyrics of a song, an instruction to add a timestamp to the lyrics, and the song data of the song. The determination unit 16 is a functional element that determines whether the timestamp data satisfies predetermined conditions, and if the conditions are not satisfied, inputs an instruction to the generation unit 15 to correct the timestamp data so that the conditions are satisfied. The transmission unit 17 is a functional element that transmits the song data or the timestamp data to the player terminal 30 in response to receiving a request to transmit the song data or the timestamp data from the player terminal 30.
[0027] The distributor terminal 20 includes, as functional elements, a content transmitting unit 21 and an information transmitting unit 22. The content transmitting unit 21 is a functional element that transmits content data including music to the server 10. The information transmitting unit 22 is a functional element that transmits lyrics of the music to the server 10. The information transmitting unit 22 may transmit identification information of the music along with the lyrics of the music.
[0028] The player terminal 30 includes a content playback unit 31. The content playback unit 31 is a functional element that transmits a transmission request for music data or timestamp data to the server 10, receives the music data and timestamp data transmitted from the server 10 in response to the transmission request, plays the music, and displays the lyrics of the music on a display device.
[0029] [System Operation] FIG. 4 is a sequence diagram showing a processing flow of the timestamp generation system 1 according to an embodiment of the present disclosure.
[0030] In step S11, the content transmitting unit 21 of the distributor terminal 20 transmits content data to the server 10. In one example, the content data is video data including music such as a music video stored in the auxiliary storage unit 203 of the distributor terminal 20. The content data may also be audio data including music but not video.
[0031] In step S12, the content registration unit 11 of the server 10 records the content data and the identification information. In one example, the content data is stored in a web server such as an origin server, and the storage location and the identification information are recorded in a database. If the received content data is video data, the content data may be converted into audio data and stored on the web server. The content registration unit 11 may set any identification information as the identification information, or the content registration unit 11 may receive and record identification information generated by the distributor terminal 20.
[0032] In step S13, the information transmitting unit 22 of the distributor terminal 20 transmits the lyrics and the identification information to the server 10. In one example, the timestamp generation system 1 provides the distributor terminal 20 with a display screen for the distributor to edit information about his or her own song that the distributor has registered in the timestamp generation system 1. The distributor inputs the lyrics of the song through the display screen, and the lyrics of the song are transmitted to the server 10 in response to the distributor's operation.
[0033] In step S14, the lyrics acquisition unit 12 of the server 10 receives the lyrics transmitted from the distributor terminal 20.
[0034] In step S15, the music data acquisition unit 13 of the server 10 acquires music data based on the identification information transmitted from the distributor terminal 20. The music data includes audio data of the music. As a specific example, if the content data is video data and audio data is used as the music data, the server 10 acquires content data with matching identification information and converts the content data into audio data to acquire the music data. If the content data is video data and video data including audio data is used as the music data, the server 10 acquires content data with matching identification information as music data. In this case, the content data may be video data obtained by re-encoding the content data to convert data such as image quality or sound quality. If the content data is audio data and audio data is used as the music data, the server 10 acquires content data with matching identification information as music data. In this case, the content data may be audio data obtained by re-encoding the content data to convert sound quality, etc.
[0035] In step S16, the input unit 14 of the server 10 inputs the received lyrics, an instruction to add a timestamp to the lyrics, and music data of the song to the generation unit 15. For example, the instruction to add a timestamp to the lyrics may be an instruction to output timestamp data in which timestamp information is added to text that matches the lyrics. The instruction to add a timestamp to the lyrics also includes instructions to add a timestamp to the song of the input data, conditions regarding the format of the timestamp data, and an example of how the timestamp data should be written. The generation unit 15 can use a multimodal large-scale language model, which is a language model trained to estimate and output the next token from an input token sequence and can process non-text data such as audio data and video data together with text and output the text. This timestamp data generation method allows timestamp data to be generated using a multimodal large-scale language model without preparing a machine learning model specialized for timestamp data generation, thereby enabling timestamp data to be generated in a simpler manner. Furthermore, highly reliable timestamp data can be generated by inputting lyrics, an instruction to add a timestamp to the lyrics, and music data of the song to the multimodal large-scale language model. The instructions in this case are also called prompts, and for example, "Please output the following audio data song in SRT format with a timestamp. Please output all the lyrics provided, line by line, without changing any symbols or spaces in the lyrics." #lyrics One day in the forest I met a bear at Flowering Forest Path The text may be something like "I met a bear." The instruction to add a timestamp to the lyrics may include an instruction to specify the format of the timestamp data. It may also include an example of the description format of the timestamp data.
[0036] FIG. 5 is a conceptual diagram showing the structure of the multimodal large-scale language model 151 used as the generation unit 15. As shown in FIG. 5, the multimodal large-scale language model 151 includes a conversion unit 152 and a large-scale language model 153. The multimodal large-scale language model 151 converts non-text data NT, such as audio data or video data, into data processable by the large-scale language model 153 using the conversion unit 152, thereby treating data other than text in the same way as text. In this way, the multimodal large-scale language model 151 processes the text data T and the non-text data NT using the large-scale language model 153 and outputs text. In one example, the conversion unit 152 converts input music data into tokens based on feature quantities, such as a spectrogram, that indicate the characteristics of the audio signal included in the music data.
[0037] In step S17, the generation unit 15 of the server 10 generates timestamp data based on the input lyrics, an instruction to add a timestamp to the lyrics, and music data. That is, the multimodal large-scale language model 151 in Fig. 5 outputs timestamp data (output text OT) to which information indicating the singing timing of each part of the lyrics is added, based on the lyrics (text data T), music data (non-text NT), and an instruction to add a timestamp to the lyrics. The generation unit 15 may be configured to generate the timestamp data via an API provided by a business that provides use of the multimodal large-scale language model, which is different from the operator that provides the timestamp generation system 1, or may be configured to generate the timestamp data via an API of the multimodal large-scale language model provided by the operator that provides the timestamp generation system 1.
[0038] In step S18, the determination unit 16 of the server 10 determines whether the generated timestamp data satisfies predetermined conditions. For example, the predetermined conditions include whether the lyrics match the text corresponding to the lyrics included in the timestamp data, whether a timestamp is assigned to each line of lyrics, whether the timestamps are assigned in ascending order, and whether all of the timestamps are within the playback time of the song. If the predetermined conditions are satisfied, the timestamp data is linked to the song identification information and saved. If multiple predetermined conditions are set, the determination unit 16 may be configured to link the timestamp data to the song identification information and save it only if all of the set conditions are satisfied.
[0039] If the determination unit 16 determines that the predetermined condition is not satisfied, step S19 is executed. In step S19, an instruction to modify the timestamp data so that it satisfies the condition that the determination unit 16 determined not to be satisfied is input to the generation unit 15. The instruction to modify the timestamp data so that it satisfies the condition that the determination unit 16 determined not to be satisfied includes, for example, an instruction to modify the timestamp data so that the lyrics match the text corresponding to the lyrics included in the timestamp data, an instruction to modify the timestamp data so that a timestamp is assigned to each line of lyrics, an instruction to modify the timestamps so that the timestamps are assigned in ascending order, or an instruction to modify the timestamps so that all of the timestamps are included within the playback time of the song. This method of generating timestamp data makes it possible to generate more reliable timestamp data.
[0040] Step S18 may also be executed for the corrected timestamp data generated by the generation unit 15. If it is determined that the corrected timestamp data satisfies the predetermined condition, the timestamp data is linked to the identification information of the song and saved. If it is determined that the corrected timestamp data does not satisfy the predetermined condition, step S19 is executed again, and an instruction is input to the generation unit 15 to correct the timestamp data so that it satisfies the condition that the determination unit 16 determined not to be satisfied.
[0041] In step S20, the content playback unit 31 of the player terminal 30 transmits a content request to the server 10 based on an operation by the player. The content request is a request to have the server 10 transmit data necessary for playing a song. The content request may be a request for data necessary for playing a specific song, or a request for data necessary for playing unspecified songs.
[0042] In step S21, the transmitting unit 17 of the server 10 transmits song data and timestamp data to the player terminal 30, which is the sender of the content request, based on the content request. If the content request is a request for data necessary to play a specific song, the transmitting unit 17 transmits song data and timestamp data for the song specified in the content request. If the content request is a request for data necessary to play an unspecified song, the transmitting unit 17 transmits song data and timestamp data for a song selected by any method. The transmission of song data and timestamp data does not need to be performed simultaneously; for example, after the player terminal 30 receives the song data, the player terminal 30 may again transmit a content request and transmit timestamp data to the player terminal 30 in response to an operation by the player.
[0043] In step S22, the content playback unit 31 of the player terminal 30 plays the song based on the received song data and displays the lyrics of the song on the display device based on the received timestamp data. When the playback time of the song reaches the time included in the timestamp data while the lyrics of the song are being displayed, the player terminal 30 changes the display mode of the lyrics associated with the timestamp of that playback time. Changing the display mode of the lyrics is achieved, for example, by changing at least one of the font, character size, character color, and character thickness. When the playback time of the song passes the time of the timestamp data associated with the lyrics whose display mode has been changed, the display mode of the lyrics may remain changed, or the display mode of the lyrics may be restored to the original display mode.
[0044] [Example of display screen] 6A and 6B are exemplary schematic diagrams of display screens according to an embodiment of the present invention. While Fig. 6 shows a smartphone used as the playback terminal 30, the examples shown in these figures can also be executed by the distributor terminal 20.
[0045] 6, the display screen of the player terminal 30 displays a button 301 for switching between playing and stopping the song, a user interface 302 for displaying and changing the song's playback time, song lyrics 303, and a button 304 for switching between displaying and hiding the song lyrics. While a song is being played, the lyrics of the song are displayed on the display screen. In this example, the user interface 302 indicates that the song is being played at a position 7 seconds after the start of the song. In this example, the timestamp data includes the singing timing for the lyrics "I met a bear" 7 seconds after the start of playback of the song, so the lyrics "I met a bear" are displayed in a different bold font from the lyrics for which the singing timing has not yet arrived.
[0046] Various examples of the present disclosure have been described above in detail. However, the present disclosure is not limited to the above examples. Various modifications of the present disclosure are possible without departing from the spirit of the present disclosure. Some modifications of the present embodiment will be described below.
[0047] [Variation 1] 7 is a diagram showing an example of a functional configuration related to a timestamp generation system 1A according to Modification 1. The timestamp generation system 1A according to Modification 1 differs from the timestamp generation system 1 in that it does not include a determination unit.
[0048] The timestamp generation system 1A according to the first modification may be configured so that the distributor or the operator checks whether the timestamp data generated by the generation unit 15A needs to be corrected. In this case, if the distributor or the operator determines that the timestamp data does not need to be corrected, the distributor or the operator uses the user terminal to send a request to the server 10 to associate the timestamp data with the identification information of the song and save it. Upon receiving the request, the server 10 associates the timestamp data with the identification information of the song and saves it.
[0049] When the distributor or operator determines that the timestamp data needs to be corrected, the distributor or operator uses the user terminal to send an instruction to the server 10 to correct the data to the desired timestamp data. The server 10 inputs the instruction to the generation unit 15A, and generates the timestamp data corrected in accordance with the instruction.
[0050] [Variation 2] 8 is a diagram showing an example of a functional configuration related to a timestamp generation system 1B according to Modification 2. The timestamp generation system 1B according to Modification 2 differs from the timestamp generation system 1 in that it does not include an input unit, that a generation unit 14B generates timestamp data without using a multimodal large-scale language model, and that it includes a correction unit 16B that corrects the generated timestamp data using a large-scale language model so that the generated timestamp data satisfies predetermined conditions.
[0051] [Operation of the System According to Modification 2] FIG. 9 is a sequence diagram showing the processing flow of a timestamp generation system 1B according to the second modification of the present disclosure.
[0052] In step S11B, content transmitter 21B of distributor terminal 20B transmits content data to server 10B. In one example, the content data is video data including music such as a music video stored in auxiliary storage unit 203B of distributor terminal 20B. The content data may also be audio data that does not include video.
[0053] In step S12B, the content registration unit 11B of the server 10B records the content data and the identification information. In one example, the content data is stored in a web server such as an origin server, and the storage location and the identification information are recorded in a database. If the received content data is video data, the content data may be converted into audio data and stored in the web server.
[0054] In step S13B, the information transmitting unit 22B of the distributor terminal 20B transmits the lyrics and the identification information to the server 10B. In one example, the timestamp generation system 1B provides the distributor terminal 20B with a display screen for editing information about the distributor's own music that the distributor has registered in the timestamp generation system 1B. The distributor inputs the lyrics of the music through the display screen, and the lyrics of the music are transmitted to the server 10B in response to the distributor's operation.
[0055] In step S14B, the lyrics acquisition unit 12B of the server 10B receives the lyrics transmitted from the distributor terminal 20B.
[0056] In step S15B, music data acquisition unit 13B of server 10B acquires music data based on the identification information transmitted from distributor terminal 20B. If the content data is video data and audio data is used as the music data, server 10B acquires content data with matching identification information and converts the content data into audio data to acquire music data. If the content data is video data and video data is used as the music data, server 10B acquires content data with matching identification information as music data. In this case, the content data may be video data obtained by re-encoding the content data to convert data such as image quality or sound quality. If the content data is audio data and audio data is used as the music data, server 10B acquires content data with matching identification information as music data. In this case, the content data may be audio data obtained by re-encoding the content data to convert sound quality, etc.
[0057] In step S16B, the generation unit 14B of the server 10B generates timestamp data based on the music data. An existing timestamp assignment tool can be used to generate the timestamp data. As an example, the generation unit 14B extracts lyric text from the music data using acoustic physics analysis technology, and generates timestamp data by associating the pronunciation position with the extracted text.
[0058] In step S17B, the determination unit 15B of the server 10 determines whether the generated timestamp data satisfies predetermined conditions. For example, the predetermined conditions include whether the lyrics match the text corresponding to the lyrics included in the timestamp data, whether a timestamp is assigned to each line of lyrics, whether the timestamps are assigned in ascending order, and whether all of the timestamps are within the playback time of the song. If the predetermined conditions are satisfied, the timestamp data is stored in association with the song identification information. If multiple predetermined conditions are set, the determination unit 15B may store the timestamp data in association with the song identification information only if all of the set conditions are satisfied.
[0059] If the determination unit 15B determines that the predetermined condition is not satisfied, step S17B is executed. In step S18B, an instruction to modify the timestamp data so that it satisfies the condition that the determination unit 15B determined not to be satisfied is input to the modification unit 16B, and the modification unit 16B generates modified timestamp data based on the instruction. The instruction to modify the timestamp data so that it satisfies the condition that the determination unit 15B determined not to be satisfied includes, for example, an instruction to modify the timestamp data so that the lyrics match the text corresponding to the lyrics included in the timestamp data, an instruction to modify the timestamp data so that a timestamp is assigned to each line of lyrics, an instruction to modify the timestamps so that they are assigned in ascending order, or an instruction to modify the timestamps so that all of them are included in the playback time of the song.
[0060] Correction unit 16B can use a multimodal large-scale language model or a language model trained to estimate and output the next token from an input token sequence, and capable of processing text and outputting the text. Correction unit 16B may be configured to generate timestamp data via an API provided by a business that provides the use of a large-scale language model different from the operator that provides timestamp generation system 1, or may be configured to generate timestamp data via an API of a large-scale language model provided by the operator that provides timestamp generation system 1. According to this method of generating timestamp data, the large-scale language model corrects the timestamp data, making it possible to generate highly reliable timestamp data in a less time-consuming manner.
[0061] Step S17B may also be executed for the corrected timestamp data generated by the correction unit 16B. If it is determined that the corrected timestamp data satisfies the predetermined condition, the timestamp data is linked to the identification information of the song and stored. If it is determined that the corrected timestamp data does not satisfy the predetermined condition, step S18B is executed again, and an instruction to correct the timestamp data so that it satisfies the condition determined by the determination unit 15B not to be satisfied is input to the correction unit 16B, and the correction unit 16B generates corrected timestamp data based on the instruction.
[0062] The timestamp generation system 1B according to the second modification may be configured such that the distributor or the operator checks whether the timestamp data generated by the generation unit 14B needs to be corrected. In this case, if the distributor or the operator determines that the timestamp data does not need to be corrected, the distributor or the operator uses the user terminal to send a request to the server 10B to associate the timestamp data with the identification information of the song and save it. Upon receiving the request, the server 10B associates the timestamp data with the identification information of the song and saves it.
[0063] If the distributor or operator determines that the timestamp data needs to be corrected, the distributor or operator uses the user terminal to send an instruction to the server 10B to correct the data to the desired timestamp data. The server 10B inputs the instruction to the correcting unit 16B, thereby generating timestamp data corrected in accordance with the instruction.
[0064] In step S20B, the content playback unit 31B of the player terminal 30B transmits a content request to the server 10B based on the player's operation. The content request is a request to have the server 10B transmit data necessary for playing a song. The content request may be a request for data necessary for playing a specific song, or a request for data necessary for playing unspecified songs.
[0065] In step S21B, the transmitting unit 17B of the server 10B transmits the song data and timestamp data to the player terminal 30B, which is the sender of the content request, based on the content request. If the content request is for data necessary to play a specific song, the transmitting unit 17B transmits the song data and timestamp data of the song specified in the content request. If the content request is for data necessary to play an unspecified song, the transmitting unit 17B transmits the song data and timestamp data of a song selected by any method. The transmission of the song data and the timestamp data does not need to be performed simultaneously; for example, after the player terminal 30B receives the song data, the player terminal 30B may again transmit a content request and send timestamp data to the player terminal 30B in response to an operation by the player.
[0066] In step S22B, the content playback unit 31B of the player terminal 30B plays the song based on the received song data and displays the lyrics of the song on the display device based on the received timestamp data. While displaying the lyrics of the song, when the playback time of the song reaches the time included in the timestamp data, the player terminal 30B changes the display mode of the lyrics associated with the timestamp of the playback time. Changing the display mode of the lyrics is achieved, for example, by changing at least one of the font, character size, character color, and character thickness. When the playback time of the song passes the time of the timestamp data associated with the lyrics whose display mode has been changed, the display mode of the lyrics may be restored to the original display mode.
[0067] Various embodiments of the timestamp generation system of the present disclosure have been described in detail above. The above description of the timestamp generation system has been given on the assumption that the content data includes music. However, the content data is not limited to this. For example, the content data may be video data including speech, such as movies, or audio data including speech, such as podcasts and radio. In this case, timestamp data can be generated by inputting transcribed text of speech, instructions to assign a timestamp to the text, and the audio data of the speech into a multimodal large-scale language model.
[0068] Any part or all of the functional units described herein may be realized by a program. The programs described herein may be distributed by being non-temporarily recorded on a computer-readable recording medium, distributed via a communication line (including wireless communication) such as the Internet, or distributed in a state where they are installed on any terminal. Alternatively, the programs may be programs that run on a web browser (so-called web apps). In this case, a computer may receive a program written in a markup language file (e.g., an HTML file) from a server and execute the program using the web browser.
[0069] Based on the above description, a person skilled in the art may be able to conceive additional effects and various modifications of the present invention, but the aspects of the present invention are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents.
[0070] For example, what is described in this specification as a single device (or component, the same applies hereinafter) (including what is depicted as a single device in the drawings) may be realized by multiple devices. Conversely, what is described in this specification as multiple devices (including what is depicted as multiple devices in the drawings) may be realized by a single device. Alternatively, some or all of the means or functions included in a certain device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may consist of a single device, or may consist of two or more devices (e.g., a server and a user terminal, or multiple user terminals).
[0071] Furthermore, not all of the features described in this specification are essential requirements. In particular, features described in this specification but not included in the claims can be considered optional additional features.
[0072] It should be noted that the applicant is merely aware of the inventions disclosed in the documents listed in the "Prior Art Documents" section of this specification, and the present invention does not necessarily aim to solve the problems of the disclosed inventions. The problem that the present invention aims to solve should be determined by taking into consideration the entire specification. For example, if this specification states that a specific configuration achieves a certain effect, it can also be said that the present invention solves a problem that is the reverse of that effect. However, it is not necessarily intended that such a specific configuration be an essential requirement.
[0073] [Note] The present disclosure discloses the following configurations. [Section 1] One or more hardware processors The first step is to get the lyrics of the song, a second step of acquiring music data including audio data of the music; a third step of inputting the lyrics, an instruction to assign a timestamp to the lyrics, and the music data into a multimodal large-scale language model. [Section 2] Item 1. The timestamp generation method according to Item 1, wherein the instruction to add a timestamp to the lyrics is an instruction to output timestamp data, which is data in which timestamp information is added to text that matches the lyrics. [Section 3] the one or more hardware processors: The timestamp generation method described in item 1 further includes a fourth step of inputting an instruction to the multimodal large-scale language model to modify timestamp data, which is data in which timestamp information is added to text corresponding to the lyrics output from the multimodal large-scale language model. [Section 4] In the fourth step, Item 4. The method of generating timestamps according to item 3, further comprising inputting instructions to the multimodal large-scale language model to modify the timestamp data so that text corresponding to the lyrics matches the lyrics. [Section 5] In the fourth step, Item 3. The timestamp generation method according to Item 3, wherein if a timestamp is not assigned to each line of the lyrics, an instruction to modify the multimodal large-scale language model to assign a timestamp to each line of the lyrics is input. [Section 6] In the fourth step, Item 4. The timestamp generation method according to item 3, wherein if the timestamps are not assigned in ascending order, an instruction to correct the assignment of timestamps in ascending order is input to the multimodal large-scale language model. [Section 7] In the fourth step, Item 3. A timestamp generation method as described in Item 3, wherein if there is a timestamp that is not included in the playback time of the music data, an instruction to modify the multimodal large-scale language model to add a timestamp that is included in the playback time of the music is input. [Section 8] In the second step, acquiring identification information of the song; acquiring video data corresponding to the music piece based on the identification information; The timestamp generating method according to claim 1 , wherein the video data is converted into the music data. [Section 9] one or more hardware processors, Get the lyrics of the song acquiring music data including audio data of the music piece; A timestamp generation system inputs the lyrics, an instruction to add a timestamp to the lyrics, and the music data into a multimodal large-scale language model. [Section 10] the one or more hardware processors: acquire timestamp data, which is data in which a timestamp is added to text corresponding to the lyrics output from the multimodal large-scale language model; Playing the song, displaying the lyrics on a display device while the music is being played; Item 10. The timestamp generation system according to item 9, wherein when the playback time of the music reaches the time included in the timestamp data, the display mode of the text associated with the timestamp of the playback time is changed. [Section 11] One or more hardware processors The first step is to get the lyrics of the song, a second step of acquiring music data including audio data of the music; a third step of inputting the lyrics, an instruction to assign a timestamp to the lyrics, and the music data into a multimodal large-scale language model. [Explanation of symbols]
[0074] 1...timestamp generation system, 10...server, 11...content registration unit, 12...lyrics acquisition unit, 13...music data acquisition unit, 14...input unit, 15...generation unit, 16...determination unit, 17...transmission unit, 20...distributor terminal, 21...content transmission unit, 22...information transmission unit, 30...player terminal, 31...content playback unit
Claims
1. one or more hardware processors, A first step is to obtain the lyrics of the song; a second step of acquiring music data including audio data of the music; a third step of inputting the lyrics, an instruction to timestamp the lyrics, and the music data into a multimodal large-scale language model; A timestamp generation method in which the lyrics are input to the multimodal large-scale language model as text data without being converted into phonemes and short pauses where no phonemes are present.
2. 2. The timestamp generating method according to claim 1, wherein the instruction to add a timestamp to the lyrics is an instruction to output timestamp data, which is data in which timestamp information is added to text that matches the lyrics.
3. the one or more hardware processors:
2. The timestamp generation method according to claim 1, further comprising a fourth step of inputting, into the multimodal large-scale language model, an instruction to modify timestamp data, which is data in which timestamp information is added to text corresponding to the lyrics and which is output from the multimodal large-scale language model.
4. In the fourth step, The method of claim 3 , further comprising inputting instructions to the multimodal large-scale language model to modify the timestamp data so that text corresponding to the lyrics matches the lyrics.
5. In the fourth step, The timestamp generation method according to claim 3, wherein if a timestamp is not assigned to each line of the lyrics, an instruction to modify the multimodal large-scale language model so that a timestamp is assigned to each line of the lyrics is input to the multimodal large-scale language model.
6. In the fourth step, 4. The timestamp generation method according to claim 3, further comprising inputting, into the multimodal large-scale language model, an instruction to correct the timestamps so that they are assigned in ascending order if the timestamps are not assigned in ascending order.
7. In the fourth step, The timestamp generation method according to claim 3, wherein if there is a timestamp that is not included in the playback time of the music data, an instruction to modify the model so that a timestamp that is included in the playback time of the music is added is input to the multimodal large-scale language model.
8. In the second step, acquiring identification information of the song; acquiring video data corresponding to the music piece based on the identification information; The timestamp generating method according to claim 1 , wherein the video data is converted into the music data.
9. one or more hardware processors, Get the lyrics of the song acquiring music data including audio data of the music piece; inputting the lyrics, an instruction to add a timestamp to the lyrics, and the music data into a multimodal large-scale language model; A timestamp generation system in which the lyrics are input to the multimodal large-scale language model as text data without being converted into phonemes and short pauses where no phonemes are present.
10. the one or more hardware processors: acquire timestamp data, which is data in which a timestamp is added to text corresponding to the lyrics output from the multimodal large-scale language model; Playing the song, displaying the lyrics on a display device while the music is being played; The timestamp generation system according to claim 9 , wherein when the playback time of the music reaches the time included in the timestamp data, a display mode of text linked to the timestamp of the playback time is changed.
11. one or more hardware processors, A first step is to obtain the lyrics of the song; a second step of acquiring music data including audio data of the music; a third step of inputting the lyrics, an instruction to timestamp the lyrics, and the music data into a multimodal large-scale language model; A timestamp generation program in which the lyrics are input to the multimodal large-scale language model as text data without being converted into phonemes and short pauses where no phonemes are present.
12. The audio data includes a voice singing lyrics of the song, The timestamp generation method according to claim 1 , wherein information regarding vocal and non-vocal sections in the music data is not input to the multimodal large-scale language model.
Citation Information
Patent Citations
Automatic system and method for temporal alignment of music audio signal with lyric
JP2008134606A
Lyrics output data correction apparatus and program
JP2013029762A
Audio signal processor and method for synchronizing speech and text using machine learning model
JP2024137877A