System

The system addresses the challenge of creating professional-quality music by using generative AI to automatically generate lyrics and melodies that match user-specified themes, enabling easy music creation and enjoyment for enthusiasts.

JP2026025620APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128429
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Music enthusiasts, particularly those in their teens to 30s, lack the proper tools and skills to easily create professional-quality music, leading many to abandon their creative endeavors.

Method used

A system that includes an input means for specifying a theme or atmosphere, a transmission means for data transfer, a generation means using generative AI models to create lyrics and melodies, and a display means for outputting the results, allowing users to easily generate and enjoy high-quality original songs.

Benefits of technology

Enables users to create and enjoy high-quality original songs without specialized knowledge, with the system automatically generating lyrics and melodies that match the specified theme or atmosphere, and providing easy display and playback options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025620000001_ABST
    Figure 2026025620000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: inputting means for a user to designate a theme or mood; transmitting means for transmitting the designated theme or mood to a server; generating means for receiving the designated theme or mood and generating lyrics and a melody using a generative AI model; transmitting means for transmitting the generated lyrics and melody; and displaying means for displaying or reproducing the lyrics and the melody.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Currently, many music enthusiasts are interested in songwriting and composing, but the technical barriers are high and creating professional-quality music is not easy. Music enthusiasts in their teens to 30s, in particular, often lack the proper tools and skills, leading many to give up on creative endeavors. There is a need for a system that can solve this problem and allow anyone to easily create high-quality original songs. [Means for solving the problem]

[0005] The system of the present invention includes an input means for a user to specify a theme or atmosphere, a transmission means for transmitting the theme and atmosphere data to a server, a generation means for generating lyrics and a melody using a generative AI model, a transmission means for transmitting the generated data to a terminal, and a display means for displaying or playing the data. Furthermore, by including a natural language generation model and a music generation model in the generation means, lyrics and a melody that match the specified theme or atmosphere can be automatically generated with high accuracy. Furthermore, by including a playback means for displaying the lyrics as a text file on the display means and playing the melody as a music file, the user can easily check the generated music. This system allows music lovers to create original songs easily and at low cost.

[0006] A "user" is someone who uses the system to specify a theme and atmosphere and create original lyrics and melodies.

[0007] "Input means" is an interface that allows the user to specify a theme or atmosphere to the system.

[0008] The "transmission means" is a function for transmitting the theme and atmosphere data acquired through the input means to the server.

[0009] The "generation method" is a function that automatically generates lyrics and melodies using a generative AI model based on theme and atmosphere data.

[0010] A "generative AI model" is an algorithm or model that uses artificial intelligence to generate lyrics and melodies that fit a specified theme or atmosphere.

[0011] A "natural language generation model" is an AI model that generates natural-looking sentences based on language data.

[0012] A "music generation model" is an AI model that generates the melody and rhythm of a song based on music data.

[0013] "Data" refers to information about lyrics and melody generated by the generating means.

[0014] The "display means" is an interface for allowing the user to see and hear the generated lyrics and melody data.

[0015] The "playback means" is a part of the display means, and is a function for playing back the melody generated as a music file. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0038] System configuration

[0039] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0040] User Input

[0041] The user inputs the theme or mood text using the interface provided at the terminal, for example, the theme "a refreshing and uplifting summer song."

[0042] Data transmission

[0043] The device sends the entered data about the theme and mood to the server, typically using an HTTP request, with the entered text packaged in JSON format.

[0044] Receiving data and starting generation AI

[0045] The server receives and analyzes the data sent from the device. A generative AI model is then activated to generate lyrics and a melody that match the specified theme and atmosphere. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0046] Lyric and melody generation

[0047] The generative AI model automatically generates lyrics and melodies based on user specifications. For example, it generates bright lyrics and a fun melody that matches the "refreshing and uplifting feeling of summer."

[0048] Data processing

[0049] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0050] Sending data

[0051] The server then sends the generated file back to the terminal, usually as an HTTP response containing BASE64-encoded file data.

[0052] Data display and playback

[0053] The device decodes the received file and displays it to the user. The lyrics are displayed as text and the melody is played as a music file. The user can check and enjoy the generated song through the device interface.

[0054] Specific examples

[0055] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0056] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0060] Step 2:

[0061] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0062] Step 3:

[0063] The device sends the data converted to JSON format to the server using an HTTP request, using the appropriate endpoint (URL).

[0064] Step 4:

[0065] The server receives the HTTP request and parses the JSON data, extracting the theme and mood specified by the user.

[0066] Step 5:

[0067] The server then passes the received theme and atmosphere data to the generative AI model, specifically, launching a natural language generation model and a music generation model.

[0068] Step 6:

[0069] The generative AI model (natural language generation model) generates lyrics based on the theme and mood specified by the user, such as "The sky is dyed in the colors of a sunset, and a lonely wind blows."

[0070] Step 7:

[0071] Similarly, generative AI models (music generation models) can generate melodies that fit a given theme or mood, such as a lonely piano melody.

[0072] Step 8:

[0073] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0074] Step 9:

[0075] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0076] Step 10:

[0077] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0078] Step 11:

[0079] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0080] Step 12:

[0081] Through the device interface, the user can read the generated lyrics and press the play button to listen to the melody.

[0082] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme and atmosphere.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] Conventional music generation systems have had difficulty generating original lyrics and melodies that perfectly match the theme and atmosphere specified by the user. Furthermore, they lacked the means to process the generated lyrics and melodies into an appropriate format and to display and play them in a way that is easy for the user to understand. As a result, the user experience was poor and the quality of the generated music was insufficient.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes an analysis means for analyzing data on the theme and atmosphere, a generation means for generating lyrics and melody using a generative AI model, and a processing means for converting the generated lyrics and melody into an appropriate format and saving them. This makes it possible to automatically generate high-quality lyrics and melodies that perfectly match the theme and atmosphere specified by the user and provide them to the user in an appropriate format.

[0088] The "input means" is a means by which the user specifies the theme and atmosphere.

[0089] The "transmission means" is a means for transmitting data on the theme and atmosphere designated through the input means to the server.

[0090] The "generation means" is a means for receiving the theme and atmosphere data and generating lyrics and melody using a generative AI model.

[0091] "Expression means" refers to a means for transmitting the generated lyrics and melody data to a terminal.

[0092] The "display means" is a means for displaying or reproducing the lyrics and melody data.

[0093] The "processing means" is a means for converting the generated lyrics and melody into an appropriate format and saving it.

[0094] "Analysis methods" are methods for analyzing thematic and atmospheric data.

[0095] A "natural language generation model" is an algorithm or program that generates natural language sentences based on input text data.

[0096] A "music generation model" is an algorithm or program that generates music based on specified parameters and input data.

[0097] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0098] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0099] First, the user uses the terminal to input a specific theme or atmosphere. For example, a specific theme such as "a refreshing and uplifting summer song" can be input. An input field is displayed on the terminal, and the user inputs the theme or atmosphere into the specified field. This is called the "input means."

[0100] Next, the device sends the input data to the server. This transmission is usually done via an HTTP request. The input theme and atmosphere data is converted to JSON format and sent to the server. This is called the "transmission method."

[0101] The server receives and analyzes this transmitted data. This analysis method accurately understands the theme and mood data entered. The server then uses generative AI models to generate lyrics and a melody that match the specified theme and mood. These generative models include natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0102] The generated lyrics and melody are then converted into the appropriate format on the server. The lyrics are saved as a text file, and the melody is saved as a music file (e.g., .mp3 or .wav). This process is called the "processing method."

[0103] The server sends the generated file to the terminal. This transmission is also usually done as an HTTP response. The transmitted data is in BASE64 encoded format.

[0104] The terminal decodes the received data and displays it to the user. The lyrics are displayed as text, and the melody is played as a music file. This process is called "display means." The user can check and enjoy the generated music through the terminal's interface.

[0105] Specific examples

[0106] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0107] Prompt Sentence Examples

[0108] "Generate lyrics and a melody for a song that evokes the lonely mood of an autumn evening."

[0109] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0110] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0111] Step 1:

[0112] The user uses the terminal to input text about the theme or atmosphere. Specifically, the user enters "A refreshing and uplifting summer song" into the input field and presses the send button. The input at this point is text data about the theme or atmosphere specified by the user.

[0113] Step 2:

[0114] The device sends the text data of the theme and atmosphere entered to the server. Specifically, the device converts this data into JSON format and sends it to the server via an HTTP POST request. The input is the text data entered by the user, and the output is text data in JSON format.

[0115] Step 3:

[0116] The server receives and analyzes the data sent from the device. Specifically, the server parses the received HTTP request and extracts the "theme" field from the JSON data. The input is JSON format data, and the output is parsed text data (theme or atmosphere).

[0117] Step 4:

[0118] The server launches a generative AI model based on the analyzed theme and atmosphere. Specifically, it passes a prompt to the generative AI model, which then generates lyrics and a melody. The input is the extracted theme and atmosphere text data, and the output is the generated lyrics and melody data.

[0119] Step 5:

[0120] A generative AI model generates lyrics and a melody that fit a specified theme or atmosphere. Specifically, a natural language generation model (e.g., GPT-3) generates lyrics based on the theme, and a music generation model (e.g., MusicVAE) generates the melody. The input is a prompt to the generative AI model, and the output is text data of the lyrics and music data of the melody.

[0121] Step 6:

[0122] The server converts the generated lyrics and melody into the appropriate format and saves them. Specifically, it saves the lyrics as a text file and the melody as an MP3 or WAV file. The input is the generated lyrics and melody data, and the output is a text file and a music file.

[0123] Step 7:

[0124] The server sends the generated lyrics and melody files to the terminal. Specifically, it encodes the files in BASE64 and sends them to the terminal as an HTTP response. The input is a text file and a music file, and the output is the BASE64-encoded file data.

[0125] Step 8:

[0126] The terminal decodes the received file and displays and plays it for the user. Specifically, it decodes the lyric data into text format and displays it on the screen, and passes the melody data to the music player for playback. The input is BASE64-encoded file data, and the output is the lyrics displayed on the screen and the melody played.

[0127] (Application example 1)

[0128] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0129] In today's entertainment and content distribution industries, there is a growing demand for tools that allow individual users to enjoy creative activities. However, conventional systems require specialized knowledge and an advanced system environment to automatically generate high-quality original lyrics and melodies that match a theme or atmosphere. This makes it difficult for ordinary users to easily enjoy creating music, and there is also a lack of means for sharing these creations with other users in real time. The present invention aims to solve these problems and provide a system that allows users to easily create original music and share and enjoy it.

[0130] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0131] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, generation means for receiving the data on the theme and atmosphere and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, display means for displaying or playing the lyrics and melody data, and sharing means for using the generated lyrics and melody to share them with other users in real time. This enables users to easily generate high-quality original music and further share the generated music with other users in real time.

[0132] "Input means" refers to the interface that the user uses to specify the theme and atmosphere.

[0133] The "transmission means" refers to a device or system having a function of transmitting data on the theme or atmosphere designated through the input means to the server.

[0134] "Generator" means a device or system that executes a process to generate lyrics and melody using a generative AI model based on received theme and mood data.

[0135] "Display means" refers to a device or system for displaying or playing back the generated lyrics and melody data on a terminal.

[0136] "Sharing means" refers to a device or system that has the function of sharing the created lyrics and melody with other users in real time.

[0137] A "generative AI model" refers to an artificial intelligence model that generates lyrics and melodies based on a theme or atmosphere specified by the user.

[0138] "Theme" refers to a subject or concept that allows the user to specify the content and atmosphere of a piece of music.

[0139] "Atmosphere" refers to the emotional tone or mood of a song.

[0140] "Song" refers to a musical composition that combines generated lyrics and melody.

[0141] "Server" refers to a central processing unit or system for managing data transmission, reception, and generation processes.

[0142] "Terminal" refers to electronic devices such as computers, smartphones, smart glasses, and head-mounted displays that are operated by users.

[0143] The present invention relates to a system that allows a user to specify a theme or atmosphere, automatically generate high-quality original lyrics and melodies, and share them. Specific embodiments of the system are described below.

[0144] System configuration

[0145] The system consists of the following elements:

[0146] 1. Input means: Provides an interface for users to specify the theme and atmosphere. Specifically, this applies to devices such as smartphones, smart glasses, and head-mounted displays.

[0147] 2. Transmission method: The input theme and atmosphere data is sent to the server using an HTTP request, and the data is packaged in JSON format.

[0148] 3. Generation: The server generates lyrics and a melody based on the received data using a generative AI model (e.g., GPT-3 or MusicVAE). This generation process includes both natural language generation and music generation.

[0149] 4. Display: The generated lyrics and melody are displayed or played on the user's device. The lyrics are provided as text and the melody as a music file.

[0150] 5. Sharing: We provide a function to share the created music with other users in real time, so that users can instantly enjoy the music they have created with other users.

[0151] Program processing

[0152] Server side:

[0153] The server receives the HTTP request, analyzes the theme and mood data, and activates a generative AI model. This uses a natural language generation model (GPT-3) and a music generation model (MusicVAE). These models generate lyrics and a melody that match the specified theme and mood, respectively. The generated data is converted into the appropriate format (lyrics as a text file, melody as a music file), BASE64 encoded, and sent to the device.

[0154] Terminal processing:

[0155] The device decodes the data received from the server and displays or plays it according to the media format. For example, lyrics are displayed as text, and melodies are played using a music player. This allows users to check and share the generated music in real time.

[0156] Hardware and software used

[0157] Hardware:

[0158] Input devices (smartphones, smart glasses, head-mounted displays)

[0159] Server (for data processing and generation)

[0160] Output device (same as above)

[0161] software:

[0162] Natural Language Generation Model (GPT-3)

[0163] Music generation model (MusicVAE)

[0164] Communication library (requests)

[0165] Programming language (Python)

[0166] Specific examples

[0167] As a concrete example, consider a case where a user uses a smartphone to specify a theme such as "a song for having fun with friends while watching fireworks on a summer night." When the user inputs this theme and presses the send button, the device sends the input data to a server. The server receives the theme data and generates lyrics and a melody using a generative AI model. The generated lyrics and melody are then sent back to the device, where they are displayed and played in a format that the user can view. This song can also be shared with other users in real time.

[0168] Example prompt sentence:

[0169] "A song for enjoying summer nights with friends while watching fireworks"

[0170] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0171] Step 1:

[0172] The user inputs the theme and atmosphere. Specifically, the user uses a smartphone, smart glasses, or a head-mounted display to input the theme and atmosphere in text format, such as "songs for enjoying a fun summer night with friends while watching fireworks." This input generates theme and atmosphere data.

[0173] Step 2:

[0174] The device sends the input theme and atmosphere data (text information) to the server. Specifically, it uses an HTTP request to package the theme and atmosphere data in JSON format and send it to the server. The input is text data from the user, and the output is JSON format data sent to the server.

[0175] Step 3:

[0176] The server analyzes the received theme and mood data. Specifically, it analyzes the text data of the theme and mood and converts it into a format suitable for the generative AI model. This analysis process generates appropriate instructions (prompts) based on the content of the text data. The input is theme data in JSON format, and the output is the prompts to be passed to the generative AI model.

[0177] Step 4:

[0178] The server then uses the analyzed data to launch a generative AI model to generate lyrics and a melody. Specifically, it generates lyrics using a natural language generation model (GPT-3) and a music generation model (MusicVAE) to generate a melody. The input is a prompt sentence, and the output is the generated lyrics and melody data.

[0179] Step 5:

[0180] The generated lyrics and melody data are converted into the appropriate format on the server. Specifically, the lyrics are encoded as a text file (.txt) and the melody is encoded as a music file (.mp3 or .wav). The input of this step is the generated lyrics and melody data, and the output is a text file and a music file.

[0181] Step 6:

[0182] The server sends the generated lyrics and melody data to the terminal. Specifically, it encodes this data in BASE64 and sends it back to the terminal as an HTTP response. The input is a text file and a music file, and the output is sent to the terminal as BASE64-encoded data.

[0183] Step 7:

[0184] The device decodes the received files and displays them to the user. Specifically, it decodes the BASE64 encoded data back into text and music files, displays the lyrics in text format, and plays the melody using a music player. The input to this step is the BASE64 encoded data, and the output is the text displayed on the user interface and the audio played.

[0185] Step 8:

[0186] The user checks the generated song and uses the sharing function of the device to share it with other users in real time. Specifically, the user selects the sharing option and sends the generated song to the other user's device. The input of this step is the generated song data, and the output is the shared data sent to other users.

[0187] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0188] The present invention relates to a system that not only allows a user to specify a theme and atmosphere, but also uses an emotion engine to recognize the user's emotions and automatically generates original lyrics and melodies based on those emotions. Specific embodiments of this system are described below.

[0189] System configuration

[0190] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion engine that recognizes user emotions.

[0191] User Input

[0192] The user inputs the text of the theme or mood using the interface provided on the device. For example, the user can input the theme "refreshing and uplifting summer music" and then use the voice input function to express their current mood.

[0193] Data transmission

[0194] The device converts the input data about the theme and mood into JSON format. In the case of voice input, the emotion engine performs another level of analysis, so the voice data is also sent along with the input data, making it easier for the server to analyze.

[0195] Emotion analysis

[0196] The server parses the received theme and mood data, as well as the audio data. The emotion engine analyzes the audio data to identify the user's emotion. The resulting emotional state is passed along with the text input to the generative AI model.

[0197] Starting the Generative AI

[0198] The server then uses the received data to activate a generative AI model, which generates lyrics and a melody based on the specified theme, mood, and analyzed emotions. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0199] Lyric and melody generation

[0200] The generative AI model automatically generates lyrics and melodies based on the theme, atmosphere, and emotion specified by the user. For example, if a user requests a song with a lonely autumn evening mood and the emotion engine identifies the emotion as sad, the generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will be a melancholic piano melody.

[0201] Data processing

[0202] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0203] Sending data

[0204] The server then BASE64-encodes the generated file and sends it to the terminal as an HTTP response, which also typically includes the BASE64-encoded file data and, if necessary, metadata.

[0205] Data display and playback

[0206] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0207] User Verification

[0208] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0209] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme, mood, and emotions. This invention will further enrich users' creative activities as music lovers and help them create high-quality music even without special technical skills.

[0210] The processing flow will be explained below.

[0211] Step 1:

[0212] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0213] Step 2:

[0214] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0215] Step 3:

[0216] The user provides voice input to the emotion engine. For example, the user may say, "I feel a little sad today."

[0217] Step 4:

[0218] The emotion engine analyzes voice input and identifies the user's emotion. Specifically, it uses voice recognition technology to analyze the emotion "sadness."

[0219] Step 5:

[0220] The device compiles the text input data and analyzed emotion data into JSON format and sends it to the server as an HTTP request.

[0221] Step 6:

[0222] The server receives the HTTP request and parses the JSON data, extracting the theme "A song with a lonely autumn evening mood" and the emotion "Sadness."

[0223] Step 7:

[0224] The server passes the extracted data to a generative AI model, which then uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE).

[0225] Step 8:

[0226] The natural language generation model generates lyrics based on the theme "A song with a lonely autumn evening mood" and the emotion "sadness." For example, it generates lyrics like "The sky is dyed in the colors of the sunset, and a lonely wind blows..."

[0227] Step 9:

[0228] Similarly, music generation models can generate melodies that fit a given theme, mood, or emotion, such as a lonely piano melody.

[0229] Step 10:

[0230] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0231] Step 11:

[0232] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0233] Step 12:

[0234] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0235] Step 13:

[0236] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0237] Step 14:

[0238] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0239] Through this series of steps, users can easily create and check original lyrics and melodies that match their specified theme, atmosphere, and emotions at the time.

[0240] Example 2

[0241] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0242] Conventional music generation systems often fail to respond to users' requests for specific themes and moods. Furthermore, there are no systems that can recognize users' emotions and generate music based on them, making it difficult to generate original music that perfectly matches the user's emotions. Furthermore, there is a need for a system that allows users to easily view the generated music.

[0243] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0244] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, emotion analysis means for receiving the theme and atmosphere data and analyzing the user's voice input data to identify emotions, generation means for generating lyrics and melody using a generative AI model based on the analyzed emotional state and theme and atmosphere data, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data. This allows an original piece of music to be automatically generated that reflects not only the theme and atmosphere specified by the user but also the user's emotions, and makes it easy to check the music.

[0245] The "input means" is an interface that allows the user to specify a theme or atmosphere.

[0246] The "transmission means" is a mechanism for transmitting data on the theme and atmosphere designated through the input means to the server.

[0247] "Emotion analysis means" refers to a system or algorithm for analyzing a user's voice input data to identify emotions.

[0248] "Generation means" refers to a process or mechanism for generating lyrics and melody using a generative AI model based on the analyzed emotional state and thematic or atmospheric data.

[0249] The "display means" is a mechanism for displaying or playing back the generated lyrics and melody data on a terminal.

[0250] "Generative AI models" are artificial intelligence models that automatically generate lyrics and melodies based on user input data, including natural language generation models and music generation models.

[0251] A "prompt" is input data or an instruction given to a generative AI model, and is a sentence that influences the generated results.

[0252] This invention relates to a system that not only allows a user to specify a theme and atmosphere, but also recognizes the user's emotions using emotion analysis technology, and automatically generates original lyrics and melodies based on those emotions. A specific embodiment of this system will be described.

[0253] System configuration

[0254] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion analysis engine that recognizes user emotions.

[0255] User Input

[0256] The user inputs text about a theme or mood using the provided interface on the device. For example, the user can input a theme such as "a refreshing and uplifting summer song" and express their current mood by voice using the voice input function. This input method uses hardware such as a touchscreen and microphone, and software such as a web application.

[0257] Data transmission

[0258] The device converts the input theme and mood data into JSON format. In the case of voice input, the voice data is also sent to the server for further analysis by the sentiment analysis engine. Specifically, the device creates a file called input.json and sends it to the server via an HTTP POST request.

[0259] Emotion analysis

[0260] The server analyzes the received theme and mood data, as well as the audio data. An emotion analysis engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Emotion API) analyzes the audio data to identify the user's emotion. The resulting emotional state, along with the text input, is passed to a generative AI model. This analysis is performed using a cloud-based analysis service.

[0261] Starting the Generative AI

[0262] The server then launches a generative AI model based on the received data. Specifically, a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) are used. This generation process is performed using Python scripts and the API of the generative AI model.

[0263] Specific examples

[0264] For example, if a user requests a song with a lonely autumn evening feeling and the emotion analysis engine identifies the emotion as "sadness," the generative AI model will generate the following lyrics and melody:

[0265] Example prompt sentence:

[0266] Theme: Autumn Dusk

[0267] Emotion: Sadness

[0268] Generated lyrics: "The sky is dyed in the colors of sunset, a lonely wind blows..."

[0269] Generated Melody: A melancholic piano melody

[0270] Lyric and melody generation

[0271] The generative AI model automatically generates lyrics and melodies based on the theme, mood, and emotion specified by the user. In this process, the natural language generation model generates the lyrics, and the music generation model creates the melody.

[0272] Data processing and transmission

[0273] The generated lyrics and melody are converted into the appropriate format on the server. The lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav). These files are then BASE64 encoded and sent to the terminal as an HTTP response.

[0274] Data display and playback

[0275] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0276] User Verification

[0277] Users can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing them to easily create and check original music that matches their emotions.

[0278] This allows users to create high-quality music that matches their specified theme, atmosphere, and emotions at the time, even without detailed technical skills. The generated music also perfectly matches the user's emotions, enriching the creative activities of music lovers.

[0279] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0280] Step 1: Accepting User Input

[0281] The user uses the device to input a theme or mood as text, and also uses the voice input function to express their current mood aloud. For example, the user enters the text "A refreshing and uplifting summer song" into the device's interface and speaks "I'm feeling very excited right now" into the microphone.

[0282] Input: Theme and atmosphere text, user voice data

[0283] Output: JSON format input data, audio data file

[0284] Step 2: Sending data

[0285] The device converts the input text data into JSON format and sends it to the server along with the audio data. Specifically, it parses the text data, converts it into a JSON file, saves the audio data, and sends it to the server via an HTTP POST request.

[0286] Input: JSON format input data, audio data file

[0287] Output: HTTP POST request data sent to the server

[0288] Step 3: Receiving and analyzing themes and emotions

[0289] The server analyzes the received JSON data and voice data. In particular, it uses a sentiment analysis engine to analyze the voice data and identify the user's emotion (e.g., "excited"). In this process, the server calls the sentiment analysis API and obtains the analysis results.

[0290] Input: JSON format input data, audio data file

[0291] Output: Sentiment analysis results (text format), theme and mood data

[0292] Step 4: Launching the generative AI model

[0293] The server then launches a generative AI model based on the analysis results. Specifically, it invokes a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) to generate lyrics and a melody using the prompt sentence.

[0294] Input: Sentiment analysis results, theme and mood data

[0295] Output: Generated lyrics data, generated melody data

[0296] Step 5: Processing the generated data

[0297] The generated lyrics and melody data are converted into the appropriate format: lyrics are saved as a text file (.txt), and melodies are saved as music files (.mp3 or .wav).

[0298] Input: Generated lyrics data, generated melody data

[0299] Output: Text file and music file

[0300] Step 6: Sending data

[0301] The server encodes the generated file in BASE64 and sends it to the terminal as an HTTP response. Specifically, the server encodes the file in BASE64 format, adds it to the body of the HTTP response, and sends it to the terminal.

[0302] Input: Text file, music file

[0303] Output: BASE64 encoded data contained in the HTTP response

[0304] Step 7: Receive and decode data

[0305] The device decodes the received HTTP response and converts the BASE64-encoded data back into the original text and music files. Specifically, it decodes the BASE64 data and saves it as a file.

[0306] Input: BASE64 encoded data included in the HTTP response

[0307] Output: Original text file, music file

[0308] Step 8: View and Play Data

[0309] The device displays the lyrics as text and the melody as a playable music file, using HTML5 audio tags and text display functions.

[0310] Input: Original text file, music file

[0311] Output: Lyric text displayed, melody played

[0312] Step 9: Verify the user

[0313] The user can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing the user to check the quality of the generated music.

[0314] Input: Lyric text to be displayed, melody to be played

[0315] Output: User confirmation results (subjective evaluation)

[0316] (Application example 2)

[0317] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0318] Conventional music generation systems simply generate lyrics and melodies based on a theme or atmosphere specified by the user, and are unable to reflect the user's emotions. This makes it difficult to provide music that perfectly matches the user's current emotions. Furthermore, the music generation process is not intuitive for users, making it difficult to customize the system to meet individual needs.

[0319] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0320] In this invention, the server includes input means for the user to specify a theme and atmosphere, voice analysis means for recording and analyzing the user's voice input, transmission means for transmitting the theme and atmosphere specified through the input means and voice analysis means and analyzed emotion data to the server, generation means for receiving the theme, atmosphere, and emotion data and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data, thereby enabling the generation and playback of original lyrics and melodies that match the user's emotions.

[0321] The "input means for the user to specify the theme or atmosphere" is a device or interface function that allows the user to input the theme or atmosphere of the music piece that he or she desires in text form.

[0322] The "voice analysis means for recording and analyzing the user's voice input" is a device or function for recording the voice uttered by the user and analyzing the voice data to extract information such as emotions.

[0323] The "transmission means for transmitting data on theme, atmosphere, and emotion to the server" is a device or function for transmitting the theme and atmosphere input by the user and analyzed emotion data to the server.

[0324] A "generation means for generating lyrics and melodies using a generative AI model" is a device or function that automatically generates lyrics and melodies using an artificial intelligence model based on input data.

[0325] The "transmission means for transmitting lyric and melody data to a terminal" is a device or function for transmitting the generated lyric and melody data to a user's terminal.

[0326] The "display means for displaying or reproducing the lyric and melody data" is a device or function for visually displaying the transmitted lyric and melody data on the user terminal or reproducing it as sound.

[0327] The present invention relates to a system that automatically generates original lyrics and melodies based on a user's emotions, and displays them visually and plays them audibly on a user terminal. Specific embodiments of this system are described below.

[0328] First, the user inputs the theme and mood of the song in text using a device such as a smartphone or smart glasses, and simultaneously inputs voice to express their current emotions.

[0329] Next, the device's built-in voice analysis unit records the user's voice input and analyzes their emotions. The analysis results, along with the theme and mood data entered as text, are converted into JSON format and sent to the server.

[0330] The server generates lyrics and melodies using a generative AI model based on the received theme and mood data and the analyzed emotional data. This generation process uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE). The generated lyrics and melodies are saved in text files (.txt) and music files (.mp3 or .wav), respectively.

[0331] The server then encodes the generated lyrics and melody data into BASE64 format and sends it as an HTTP response to the user's device, where it decodes the received BASE64 data and restores it to the original text and music files.

[0332] Finally, the user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[0333] As a specific example, if a user inputs the theme "A song to cheer me up when I'm a little tired," and expresses in voice input, "I've been working all day today and I'm a little tired. But tomorrow is a day off, so I have to do my best," then based on the analysis results, a bright, uplifting song with a rhythmic melody will be generated.

[0334] An example of a prompt might be:

[0335] "A refreshing and uplifting summer song"

[0336] "I feel great today!"

[0337] This allows users to easily create and enjoy original music that perfectly matches their emotions.

[0338] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0339] Step 1:

[0340] Using a device such as a smartphone or smart glasses, the user inputs the theme and mood of the song in text, and simultaneously inputs voice to express their current emotions.

[0341] Input: Text (e.g., "I'm a little tired, so I need a song to cheer me up"), audio data (e.g., "I've been working all day today and I'm a little tired. But tomorrow is a holiday, so I have to do my best.")

[0342] Output: User-entered text and audio data of the theme and mood

[0343] Step 2:

[0344] A voice analysis means of the terminal records the user's voice input and analyzes the voice data to extract emotion data.

[0345] Input: User's voice data

[0346] Data processing: Analysis of voice data and emotion recognition (e.g., emotion classification such as happiness, sadness, fatigue, etc.)

[0347] Output: Parsed emotion data (e.g., "tired")

[0348] Step 3:

[0349] The device converts the theme and atmosphere data entered by the user and the analyzed emotional data into JSON format and sends it to the server.

[0350] Input: Theme and mood of text input, analyzed sentiment data

[0351] Data processing: Conversion to JSON format

[0352] Output: Send data in JSON format (e.g., {"theme": "Songs to cheer you up when you're tired", "emotion": "Tired"})

[0353] Step 4:

[0354] The server extracts theme, mood, and emotion data from the received JSON data, and then invokes a generative AI model to generate lyrics and a melody based on the specified theme, mood, and emotion.

[0355] Input: JSON format data (theme, mood, emotion data)

[0356] Data processing: Generate lyrics using natural language generation models (e.g., GPT-3) and melodies using music generation models (e.g., MusicVAE).

[0357] Output: Generated lyrics (text format) and melody (music file format)

[0358] Step 5:

[0359] The server encodes the generated lyrics and melody data into BASE64 format and sends it to the user's terminal as an HTTP response.

[0360] Input: Generated lyrics and melody data

[0361] Data processing: Encoding to BASE64 format and creating HTTP responses

[0362] Output: BASE64 encoded lyrics and melody data (sent as HTTP response)

[0363] Step 6:

[0364] The terminal decodes the received BASE64 encoded data and restores it to the original text and music files.

[0365] Input: BASE64 encoded lyrics and melody data

[0366] Data processing: Decoding BASE64 data

[0367] Output: Original text file (lyrics) and music file (melody)

[0368] Step 7:

[0369] The user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[0370] Input: Decoded lyrics text file and music file

[0371] Specific behavior: Display a text file and play a music file

[0372] Output: User sees lyrics visually and hears melody (e.g., lyrics displayed on device screen and melody played through speaker)

[0373] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0374] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0375] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0376] [Second embodiment]

[0377] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0378] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0379] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0380] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0381] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0383] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0384] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0385] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0386] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0387] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0388] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0389] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0390] System configuration

[0391] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0392] User Input

[0393] The user inputs the theme or mood text using the interface provided at the terminal, for example, the theme "a refreshing and uplifting summer song."

[0394] Data transmission

[0395] The device sends the entered data about the theme and mood to the server, typically using an HTTP request, with the entered text packaged in JSON format.

[0396] Receiving data and starting generation AI

[0397] The server receives and analyzes the data sent from the device. A generative AI model is then activated to generate lyrics and a melody that match the specified theme and atmosphere. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0398] Lyric and melody generation

[0399] The generative AI model automatically generates lyrics and melodies based on user specifications. For example, it generates bright lyrics and a fun melody that matches the "refreshing and uplifting feeling of summer."

[0400] Data processing

[0401] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0402] Sending data

[0403] The server then sends the generated file back to the terminal, usually as an HTTP response containing BASE64-encoded file data.

[0404] Data display and playback

[0405] The device decodes the received file and displays it to the user. The lyrics are displayed as text and the melody is played as a music file. The user can check and enjoy the generated song through the device interface.

[0406] Specific examples

[0407] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0408] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0409] The processing flow will be explained below.

[0410] Step 1:

[0411] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0412] Step 2:

[0413] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0414] Step 3:

[0415] The device sends the data converted to JSON format to the server using an HTTP request, using the appropriate endpoint (URL).

[0416] Step 4:

[0417] The server receives the HTTP request and parses the JSON data, extracting the theme and mood specified by the user.

[0418] Step 5:

[0419] The server then passes the received theme and atmosphere data to the generative AI model, specifically, launching a natural language generation model and a music generation model.

[0420] Step 6:

[0421] The generative AI model (natural language generation model) generates lyrics based on the theme and mood specified by the user, such as "The sky is dyed in the colors of a sunset, and a lonely wind blows."

[0422] Step 7:

[0423] Similarly, generative AI models (music generation models) can generate melodies that fit a given theme or mood, such as a lonely piano melody.

[0424] Step 8:

[0425] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0426] Step 9:

[0427] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0428] Step 10:

[0429] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0430] Step 11:

[0431] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0432] Step 12:

[0433] Through the device interface, the user can read the generated lyrics and press the play button to listen to the melody.

[0434] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme and atmosphere.

[0435] Example 1

[0436] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0437] Conventional music generation systems have had difficulty generating original lyrics and melodies that perfectly match the theme and atmosphere specified by the user. Furthermore, they lacked the means to process the generated lyrics and melodies into an appropriate format and to display and play them in a way that is easy for the user to understand. As a result, the user experience was poor and the quality of the generated music was insufficient.

[0438] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0439] In this invention, the server includes an analysis means for analyzing data on the theme and atmosphere, a generation means for generating lyrics and melody using a generative AI model, and a processing means for converting the generated lyrics and melody into an appropriate format and saving them. This makes it possible to automatically generate high-quality lyrics and melodies that perfectly match the theme and atmosphere specified by the user and provide them to the user in an appropriate format.

[0440] The "input means" is a means by which the user specifies the theme and atmosphere.

[0441] The "transmission means" is a means for transmitting data on the theme and atmosphere designated through the input means to the server.

[0442] The "generation means" is a means for receiving the theme and atmosphere data and generating lyrics and melody using a generative AI model.

[0443] "Expression means" refers to a means for transmitting the generated lyrics and melody data to a terminal.

[0444] The "display means" is a means for displaying or reproducing the lyrics and melody data.

[0445] The "processing means" is a means for converting the generated lyrics and melody into an appropriate format and saving it.

[0446] "Analysis methods" are methods for analyzing thematic and atmospheric data.

[0447] A "natural language generation model" is an algorithm or program that generates natural language sentences based on input text data.

[0448] A "music generation model" is an algorithm or program that generates music based on specified parameters and input data.

[0449] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0450] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0451] First, the user uses the terminal to input a specific theme or atmosphere. For example, a specific theme such as "a refreshing and uplifting summer song" can be input. An input field is displayed on the terminal, and the user inputs the theme or atmosphere into the specified field. This is called the "input means."

[0452] Next, the device sends the input data to the server. This transmission is usually done via an HTTP request. The input theme and atmosphere data is converted to JSON format and sent to the server. This is called the "transmission method."

[0453] The server receives and analyzes this transmitted data. This analysis method accurately understands the theme and mood data entered. The server then uses generative AI models to generate lyrics and a melody that match the specified theme and mood. These generative models include natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0454] The generated lyrics and melody are then converted into the appropriate format on the server. The lyrics are saved as a text file, and the melody is saved as a music file (e.g., .mp3 or .wav). This process is called the "processing method."

[0455] The server sends the generated file to the terminal. This transmission is also usually done as an HTTP response. The transmitted data is in BASE64 encoded format.

[0456] The terminal decodes the received data and displays it to the user. The lyrics are displayed as text, and the melody is played as a music file. This process is called "display means." The user can check and enjoy the generated music through the terminal's interface.

[0457] Specific examples

[0458] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0459] Prompt Sentence Examples

[0460] "Generate lyrics and a melody for a song that evokes the lonely mood of an autumn evening."

[0461] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0462] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0463] Step 1:

[0464] The user uses the terminal to input text about the theme or atmosphere. Specifically, the user enters "A refreshing and uplifting summer song" into the input field and presses the send button. The input at this point is text data about the theme or atmosphere specified by the user.

[0465] Step 2:

[0466] The device sends the text data of the theme and atmosphere entered to the server. Specifically, the device converts this data into JSON format and sends it to the server via an HTTP POST request. The input is the text data entered by the user, and the output is text data in JSON format.

[0467] Step 3:

[0468] The server receives and analyzes the data sent from the device. Specifically, the server parses the received HTTP request and extracts the "theme" field from the JSON data. The input is JSON format data, and the output is parsed text data (theme or atmosphere).

[0469] Step 4:

[0470] The server launches a generative AI model based on the analyzed theme and atmosphere. Specifically, it passes a prompt to the generative AI model, which then generates lyrics and a melody. The input is the extracted theme and atmosphere text data, and the output is the generated lyrics and melody data.

[0471] Step 5:

[0472] A generative AI model generates lyrics and a melody that fit a specified theme or atmosphere. Specifically, a natural language generation model (e.g., GPT-3) generates lyrics based on the theme, and a music generation model (e.g., MusicVAE) generates the melody. The input is a prompt to the generative AI model, and the output is text data of the lyrics and music data of the melody.

[0473] Step 6:

[0474] The server converts the generated lyrics and melody into the appropriate format and saves them. Specifically, it saves the lyrics as a text file and the melody as an MP3 or WAV file. The input is the generated lyrics and melody data, and the output is a text file and a music file.

[0475] Step 7:

[0476] The server sends the generated lyrics and melody files to the terminal. Specifically, it encodes the files in BASE64 and sends them to the terminal as an HTTP response. The input is a text file and a music file, and the output is the BASE64-encoded file data.

[0477] Step 8:

[0478] The terminal decodes the received file and displays and plays it for the user. Specifically, it decodes the lyric data into text format and displays it on the screen, and passes the melody data to the music player for playback. The input is BASE64-encoded file data, and the output is the lyrics displayed on the screen and the melody played.

[0479] (Application example 1)

[0480] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0481] In today's entertainment and content distribution industries, there is a growing demand for tools that allow individual users to enjoy creative activities. However, conventional systems require specialized knowledge and an advanced system environment to automatically generate high-quality original lyrics and melodies that match a theme or atmosphere. This makes it difficult for ordinary users to easily enjoy creating music, and there is also a lack of means for sharing these creations with other users in real time. The present invention aims to solve these problems and provide a system that allows users to easily create original music and share and enjoy it.

[0482] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0483] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, generation means for receiving the data on the theme and atmosphere and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, display means for displaying or playing the lyrics and melody data, and sharing means for using the generated lyrics and melody to share them with other users in real time. This enables users to easily generate high-quality original music and further share the generated music with other users in real time.

[0484] "Input means" refers to the interface that the user uses to specify the theme and atmosphere.

[0485] The "transmission means" refers to a device or system having a function of transmitting data on the theme or atmosphere designated through the input means to the server.

[0486] "Generator" means a device or system that executes a process to generate lyrics and melody using a generative AI model based on received theme and mood data.

[0487] "Display means" refers to a device or system for displaying or playing back the generated lyrics and melody data on a terminal.

[0488] "Sharing means" refers to a device or system that has the function of sharing the created lyrics and melody with other users in real time.

[0489] A "generative AI model" refers to an artificial intelligence model that generates lyrics and melodies based on a theme or atmosphere specified by the user.

[0490] "Theme" refers to a subject or concept that allows the user to specify the content and atmosphere of a piece of music.

[0491] "Atmosphere" refers to the emotional tone or mood of a song.

[0492] "Song" refers to a musical composition that combines generated lyrics and melody.

[0493] "Server" refers to a central processing unit or system for managing data transmission, reception, and generation processes.

[0494] "Terminal" refers to electronic devices such as computers, smartphones, smart glasses, and head-mounted displays that are operated by users.

[0495] The present invention relates to a system that allows a user to specify a theme or atmosphere, automatically generate high-quality original lyrics and melodies, and share them. Specific embodiments of the system are described below.

[0496] System configuration

[0497] The system consists of the following elements:

[0498] 1. Input means: Provides an interface for users to specify the theme and atmosphere. Specifically, this applies to devices such as smartphones, smart glasses, and head-mounted displays.

[0499] 2. Transmission method: The input theme and atmosphere data is sent to the server using an HTTP request, and the data is packaged in JSON format.

[0500] 3. Generation: The server generates lyrics and a melody based on the received data using a generative AI model (e.g., GPT-3 or MusicVAE). This generation process includes both natural language generation and music generation.

[0501] 4. Display: The generated lyrics and melody are displayed or played on the user's device. The lyrics are provided as text and the melody as a music file.

[0502] 5. Sharing: We provide a function to share the created music with other users in real time, so that users can instantly enjoy the music they have created with other users.

[0503] Program processing

[0504] Server side:

[0505] The server receives the HTTP request, analyzes the theme and mood data, and activates a generative AI model. This uses a natural language generation model (GPT-3) and a music generation model (MusicVAE). These models generate lyrics and a melody that match the specified theme and mood, respectively. The generated data is converted into the appropriate format (lyrics as a text file, melody as a music file), BASE64 encoded, and sent to the device.

[0506] Terminal processing:

[0507] The device decodes the data received from the server and displays or plays it according to the media format. For example, lyrics are displayed as text, and melodies are played using a music player. This allows users to check and share the generated music in real time.

[0508] Hardware and software used

[0509] Hardware:

[0510] Input devices (smartphones, smart glasses, head-mounted displays)

[0511] Server (for data processing and generation)

[0512] Output device (same as above)

[0513] software:

[0514] Natural Language Generation Model (GPT-3)

[0515] Music generation model (MusicVAE)

[0516] Communication library (requests)

[0517] Programming language (Python)

[0518] Specific examples

[0519] As a concrete example, consider a case where a user uses a smartphone to specify a theme such as "a song for having fun with friends while watching fireworks on a summer night." When the user inputs this theme and presses the send button, the device sends the input data to a server. The server receives the theme data and generates lyrics and a melody using a generative AI model. The generated lyrics and melody are then sent back to the device, where they are displayed and played in a format that the user can view. This song can also be shared with other users in real time.

[0520] Example prompt sentence:

[0521] "A song for enjoying summer nights with friends while watching fireworks"

[0522] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0523] Step 1:

[0524] The user inputs the theme and atmosphere. Specifically, the user uses a smartphone, smart glasses, or a head-mounted display to input the theme and atmosphere in text format, such as "songs for enjoying a fun summer night with friends while watching fireworks." This input generates theme and atmosphere data.

[0525] Step 2:

[0526] The device sends the input theme and atmosphere data (text information) to the server. Specifically, it uses an HTTP request to package the theme and atmosphere data in JSON format and send it to the server. The input is text data from the user, and the output is JSON format data sent to the server.

[0527] Step 3:

[0528] The server analyzes the received theme and mood data. Specifically, it analyzes the text data of the theme and mood and converts it into a format suitable for the generative AI model. This analysis process generates appropriate instructions (prompts) based on the content of the text data. The input is theme data in JSON format, and the output is the prompts to be passed to the generative AI model.

[0529] Step 4:

[0530] The server then uses the analyzed data to launch a generative AI model to generate lyrics and a melody. Specifically, it generates lyrics using a natural language generation model (GPT-3) and a music generation model (MusicVAE) to generate a melody. The input is a prompt sentence, and the output is the generated lyrics and melody data.

[0531] Step 5:

[0532] The generated lyrics and melody data are converted into the appropriate format on the server. Specifically, the lyrics are encoded as a text file (.txt) and the melody is encoded as a music file (.mp3 or .wav). The input of this step is the generated lyrics and melody data, and the output is a text file and a music file.

[0533] Step 6:

[0534] The server sends the generated lyrics and melody data to the terminal. Specifically, it encodes this data in BASE64 and sends it back to the terminal as an HTTP response. The input is a text file and a music file, and the output is sent to the terminal as BASE64-encoded data.

[0535] Step 7:

[0536] The device decodes the received files and displays them to the user. Specifically, it decodes the BASE64 encoded data back into text and music files, displays the lyrics in text format, and plays the melody using a music player. The input to this step is the BASE64 encoded data, and the output is the text displayed on the user interface and the audio played.

[0537] Step 8:

[0538] The user checks the generated song and uses the sharing function of the device to share it with other users in real time. Specifically, the user selects the sharing option and sends the generated song to the other user's device. The input of this step is the generated song data, and the output is the shared data sent to other users.

[0539] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0540] The present invention relates to a system that not only allows a user to specify a theme and atmosphere, but also uses an emotion engine to recognize the user's emotions and automatically generates original lyrics and melodies based on those emotions. Specific embodiments of this system are described below.

[0541] System configuration

[0542] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion engine that recognizes user emotions.

[0543] User Input

[0544] The user inputs the text of the theme or mood using the interface provided on the device. For example, the user can input the theme "refreshing and uplifting summer music" and then use the voice input function to express their current mood.

[0545] Data transmission

[0546] The device converts the input data about the theme and mood into JSON format. In the case of voice input, the emotion engine performs another level of analysis, so the voice data is also sent along with the input data, making it easier for the server to analyze.

[0547] Emotion analysis

[0548] The server parses the received theme and mood data, as well as the audio data. The emotion engine analyzes the audio data to identify the user's emotion. The resulting emotional state is passed along with the text input to the generative AI model.

[0549] Starting the Generative AI

[0550] The server then uses the received data to activate a generative AI model, which generates lyrics and a melody based on the specified theme, mood, and analyzed emotions. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0551] Lyric and melody generation

[0552] The generative AI model automatically generates lyrics and melodies based on the theme, atmosphere, and emotion specified by the user. For example, if a user requests a song with a lonely autumn evening mood and the emotion engine identifies the emotion as sad, the generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will be a melancholic piano melody.

[0553] Data processing

[0554] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0555] Sending data

[0556] The server then BASE64-encodes the generated file and sends it to the terminal as an HTTP response, which also typically includes the BASE64-encoded file data and, if necessary, metadata.

[0557] Data display and playback

[0558] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0559] User Verification

[0560] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0561] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme, mood, and emotions. This invention will further enrich users' creative activities as music lovers and help them create high-quality music even without special technical skills.

[0562] The processing flow will be explained below.

[0563] Step 1:

[0564] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0565] Step 2:

[0566] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0567] Step 3:

[0568] The user provides voice input to the emotion engine. For example, the user may say, "I feel a little sad today."

[0569] Step 4:

[0570] The emotion engine analyzes voice input and identifies the user's emotion. Specifically, it uses voice recognition technology to analyze the emotion "sadness."

[0571] Step 5:

[0572] The device compiles the text input data and analyzed emotion data into JSON format and sends it to the server as an HTTP request.

[0573] Step 6:

[0574] The server receives the HTTP request and parses the JSON data, extracting the theme "A song with a lonely autumn evening mood" and the emotion "Sadness."

[0575] Step 7:

[0576] The server passes the extracted data to a generative AI model, which then uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE).

[0577] Step 8:

[0578] The natural language generation model generates lyrics based on the theme "A song with a lonely autumn evening mood" and the emotion "sadness." For example, it generates lyrics like "The sky is dyed in the colors of the sunset, and a lonely wind blows..."

[0579] Step 9:

[0580] Similarly, music generation models can generate melodies that fit a given theme, mood, or emotion, such as a lonely piano melody.

[0581] Step 10:

[0582] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0583] Step 11:

[0584] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0585] Step 12:

[0586] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0587] Step 13:

[0588] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0589] Step 14:

[0590] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0591] Through this series of steps, users can easily create and check original lyrics and melodies that match their specified theme, atmosphere, and emotions at the time.

[0592] Example 2

[0593] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0594] Conventional music generation systems often fail to respond to users' requests for specific themes and moods. Furthermore, there are no systems that can recognize users' emotions and generate music based on them, making it difficult to generate original music that perfectly matches the user's emotions. Furthermore, there is a need for a system that allows users to easily view the generated music.

[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0596] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, emotion analysis means for receiving the theme and atmosphere data and analyzing the user's voice input data to identify emotions, generation means for generating lyrics and melody using a generative AI model based on the analyzed emotional state and theme and atmosphere data, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data. This allows an original piece of music to be automatically generated that reflects not only the theme and atmosphere specified by the user but also the user's emotions, and makes it easy to check the music.

[0597] The "input means" is an interface that allows the user to specify a theme or atmosphere.

[0598] The "transmission means" is a mechanism for transmitting data on the theme and atmosphere designated through the input means to the server.

[0599] "Emotion analysis means" refers to a system or algorithm for analyzing a user's voice input data to identify emotions.

[0600] "Generation means" refers to a process or mechanism for generating lyrics and melody using a generative AI model based on the analyzed emotional state and thematic or atmospheric data.

[0601] The "display means" is a mechanism for displaying or playing back the generated lyrics and melody data on a terminal.

[0602] "Generative AI models" are artificial intelligence models that automatically generate lyrics and melodies based on user input data, including natural language generation models and music generation models.

[0603] A "prompt" is input data or an instruction given to a generative AI model, and is a sentence that influences the generated results.

[0604] This invention relates to a system that not only allows a user to specify a theme and atmosphere, but also recognizes the user's emotions using emotion analysis technology, and automatically generates original lyrics and melodies based on those emotions. A specific embodiment of this system will be described.

[0605] System configuration

[0606] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion analysis engine that recognizes user emotions.

[0607] User Input

[0608] The user inputs text about a theme or mood using the provided interface on the device. For example, the user can input a theme such as "a refreshing and uplifting summer song" and express their current mood by voice using the voice input function. This input method uses hardware such as a touchscreen and microphone, and software such as a web application.

[0609] Data transmission

[0610] The device converts the input theme and mood data into JSON format. In the case of voice input, the voice data is also sent to the server for further analysis by the sentiment analysis engine. Specifically, the device creates a file called input.json and sends it to the server via an HTTP POST request.

[0611] Emotion analysis

[0612] The server analyzes the received theme and mood data, as well as the audio data. An emotion analysis engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Emotion API) analyzes the audio data to identify the user's emotion. The resulting emotional state, along with the text input, is passed to a generative AI model. This analysis is performed using a cloud-based analysis service.

[0613] Starting the Generative AI

[0614] The server then launches a generative AI model based on the received data. Specifically, a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) are used. This generation process is performed using Python scripts and the API of the generative AI model.

[0615] Specific examples

[0616] For example, if a user requests a song with a lonely autumn evening feeling and the emotion analysis engine identifies the emotion as "sadness," the generative AI model will generate the following lyrics and melody:

[0617] Example prompt sentence:

[0618] Theme: Autumn Dusk

[0619] Emotion: Sadness

[0620] Generated lyrics: "The sky is dyed in the colors of sunset, a lonely wind blows..."

[0621] Generated Melody: A melancholic piano melody

[0622] Lyric and melody generation

[0623] The generative AI model automatically generates lyrics and melodies based on the theme, mood, and emotion specified by the user. In this process, the natural language generation model generates the lyrics, and the music generation model creates the melody.

[0624] Data processing and transmission

[0625] The generated lyrics and melody are converted into the appropriate format on the server. The lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav). These files are then BASE64 encoded and sent to the terminal as an HTTP response.

[0626] Data display and playback

[0627] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0628] User Verification

[0629] Users can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing them to easily create and check original music that matches their emotions.

[0630] This allows users to create high-quality music that matches their specified theme, atmosphere, and emotions at the time, even without detailed technical skills. The generated music also perfectly matches the user's emotions, enriching the creative activities of music lovers.

[0631] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0632] Step 1: Accepting User Input

[0633] The user uses the device to input a theme or mood as text, and also uses the voice input function to express their current mood aloud. For example, the user enters the text "A refreshing and uplifting summer song" into the device's interface and speaks "I'm feeling very excited right now" into the microphone.

[0634] Input: Theme and atmosphere text, user voice data

[0635] Output: JSON format input data, audio data file

[0636] Step 2: Sending data

[0637] The device converts the input text data into JSON format and sends it to the server along with the audio data. Specifically, it parses the text data, converts it into a JSON file, saves the audio data, and sends it to the server via an HTTP POST request.

[0638] Input: JSON format input data, audio data file

[0639] Output: HTTP POST request data sent to the server

[0640] Step 3: Receiving and analyzing themes and emotions

[0641] The server analyzes the received JSON data and voice data. In particular, it uses a sentiment analysis engine to analyze the voice data and identify the user's emotion (e.g., "excited"). In this process, the server calls the sentiment analysis API and obtains the analysis results.

[0642] Input: JSON format input data, audio data file

[0643] Output: Sentiment analysis results (text format), theme and mood data

[0644] Step 4: Launching the generative AI model

[0645] The server then launches a generative AI model based on the analysis results. Specifically, it invokes a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) to generate lyrics and a melody using the prompt sentence.

[0646] Input: Sentiment analysis results, theme and mood data

[0647] Output: Generated lyrics data, generated melody data

[0648] Step 5: Processing the generated data

[0649] The generated lyrics and melody data are converted into the appropriate format: lyrics are saved as a text file (.txt), and melodies are saved as music files (.mp3 or .wav).

[0650] Input: Generated lyrics data, generated melody data

[0651] Output: Text file and music file

[0652] Step 6: Sending data

[0653] The server encodes the generated file in BASE64 and sends it to the terminal as an HTTP response. Specifically, the server encodes the file in BASE64 format, adds it to the body of the HTTP response, and sends it to the terminal.

[0654] Input: Text file, music file

[0655] Output: BASE64 encoded data contained in the HTTP response

[0656] Step 7: Receive and decode data

[0657] The device decodes the received HTTP response and converts the BASE64-encoded data back into the original text and music files. Specifically, it decodes the BASE64 data and saves it as a file.

[0658] Input: BASE64 encoded data included in the HTTP response

[0659] Output: Original text file, music file

[0660] Step 8: View and Play Data

[0661] The device displays the lyrics as text and the melody as a playable music file, using HTML5 audio tags and text display functions.

[0662] Input: Original text file, music file

[0663] Output: Lyric text displayed, melody played

[0664] Step 9: Verify the user

[0665] The user can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing the user to check the quality of the generated music.

[0666] Input: Lyric text to be displayed, melody to be played

[0667] Output: User confirmation results (subjective evaluation)

[0668] (Application example 2)

[0669] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0670] Conventional music generation systems simply generate lyrics and melodies based on a theme or atmosphere specified by the user, and are unable to reflect the user's emotions. This makes it difficult to provide music that perfectly matches the user's current emotions. Furthermore, the music generation process is not intuitive for users, making it difficult to customize the system to meet individual needs.

[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0672] In this invention, the server includes input means for the user to specify a theme and atmosphere, voice analysis means for recording and analyzing the user's voice input, transmission means for transmitting the theme and atmosphere specified through the input means and voice analysis means and analyzed emotion data to the server, generation means for receiving the theme, atmosphere, and emotion data and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data, thereby enabling the generation and playback of original lyrics and melodies that match the user's emotions.

[0673] The "input means for the user to specify the theme or atmosphere" is a device or interface function that allows the user to input the theme or atmosphere of the music piece that he or she desires in text form.

[0674] The "voice analysis means for recording and analyzing the user's voice input" is a device or function for recording the voice uttered by the user and analyzing the voice data to extract information such as emotions.

[0675] The "transmission means for transmitting data on theme, atmosphere, and emotion to the server" is a device or function for transmitting the theme and atmosphere input by the user and analyzed emotion data to the server.

[0676] A "generation means for generating lyrics and melodies using a generative AI model" is a device or function that automatically generates lyrics and melodies using an artificial intelligence model based on input data.

[0677] The "transmission means for transmitting lyric and melody data to a terminal" is a device or function for transmitting the generated lyric and melody data to a user's terminal.

[0678] The "display means for displaying or reproducing the lyric and melody data" is a device or function for visually displaying the transmitted lyric and melody data on the user terminal or reproducing it as sound.

[0679] The present invention relates to a system that automatically generates original lyrics and melodies based on a user's emotions, and displays them visually and plays them audibly on a user terminal. Specific embodiments of this system are described below.

[0680] First, the user inputs the theme and mood of the song in text using a device such as a smartphone or smart glasses, and simultaneously inputs voice to express their current emotions.

[0681] Next, the device's built-in voice analysis unit records the user's voice input and analyzes their emotions. The analysis results, along with the theme and mood data entered as text, are converted into JSON format and sent to the server.

[0682] The server generates lyrics and melodies using a generative AI model based on the received theme and mood data and the analyzed emotional data. This generation process uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE). The generated lyrics and melodies are saved in text files (.txt) and music files (.mp3 or .wav), respectively.

[0683] The server then encodes the generated lyrics and melody data into BASE64 format and sends it as an HTTP response to the user's device, where it decodes the received BASE64 data and restores it to the original text and music files.

[0684] Finally, the user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[0685] As a specific example, if a user inputs the theme "A song to cheer me up when I'm a little tired," and expresses in voice input, "I've been working all day today and I'm a little tired. But tomorrow is a day off, so I have to do my best," then based on the analysis results, a bright, uplifting song with a rhythmic melody will be generated.

[0686] An example of a prompt might be:

[0687] "A refreshing and uplifting summer song"

[0688] "I feel great today!"

[0689] This allows users to easily create and enjoy original music that perfectly matches their emotions.

[0690] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0691] Step 1:

[0692] Using a device such as a smartphone or smart glasses, the user inputs the theme and mood of the song in text, and simultaneously inputs voice to express their current emotions.

[0693] Input: Text (e.g., "I'm a little tired, so I need a song to cheer me up"), audio data (e.g., "I've been working all day today and I'm a little tired. But tomorrow is a holiday, so I have to do my best.")

[0694] Output: User-entered text and audio data of the theme and mood

[0695] Step 2:

[0696] A voice analysis means of the terminal records the user's voice input and analyzes the voice data to extract emotion data.

[0697] Input: User's voice data

[0698] Data processing: Analysis of voice data and emotion recognition (e.g., emotion classification such as happiness, sadness, fatigue, etc.)

[0699] Output: Parsed emotion data (e.g., "tired")

[0700] Step 3:

[0701] The device converts the theme and atmosphere data entered by the user and the analyzed emotional data into JSON format and sends it to the server.

[0702] Input: Theme and mood of text input, analyzed sentiment data

[0703] Data processing: Conversion to JSON format

[0704] Output: Send data in JSON format (e.g., {"theme": "Songs to cheer you up when you're tired", "emotion": "Tired"})

[0705] Step 4:

[0706] The server extracts theme, mood, and emotion data from the received JSON data, and then invokes a generative AI model to generate lyrics and a melody based on the specified theme, mood, and emotion.

[0707] Input: JSON format data (theme, mood, emotion data)

[0708] Data processing: Generate lyrics using natural language generation models (e.g., GPT-3) and melodies using music generation models (e.g., MusicVAE).

[0709] Output: Generated lyrics (text format) and melody (music file format)

[0710] Step 5:

[0711] The server encodes the generated lyrics and melody data into BASE64 format and sends it to the user's terminal as an HTTP response.

[0712] Input: Generated lyrics and melody data

[0713] Data processing: Encoding to BASE64 format and creating HTTP responses

[0714] Output: BASE64 encoded lyrics and melody data (sent as HTTP response)

[0715] Step 6:

[0716] The terminal decodes the received BASE64 encoded data and restores it to the original text and music files.

[0717] Input: BASE64 encoded lyrics and melody data

[0718] Data processing: Decoding BASE64 data

[0719] Output: Original text file (lyrics) and music file (melody)

[0720] Step 7:

[0721] The user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[0722] Input: Decoded lyrics text file and music file

[0723] Specific behavior: Display a text file and play a music file

[0724] Output: User sees lyrics visually and hears melody (e.g., lyrics displayed on device screen and melody played through speaker)

[0725] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0726] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0727] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0728] [Third embodiment]

[0729] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0730] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0731] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0732] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0733] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0734] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0735] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0736] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0737] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0738] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0739] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0740] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0741] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0742] System configuration

[0743] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0744] User Input

[0745] The user inputs the theme or mood text using the interface provided at the terminal, for example, the theme "a refreshing and uplifting summer song."

[0746] Data transmission

[0747] The device sends the entered data about the theme and mood to the server, typically using an HTTP request, with the entered text packaged in JSON format.

[0748] Receiving data and starting generation AI

[0749] The server receives and analyzes the data sent from the device. A generative AI model is then activated to generate lyrics and a melody that match the specified theme and atmosphere. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0750] Lyric and melody generation

[0751] The generative AI model automatically generates lyrics and melodies based on user specifications. For example, it generates bright lyrics and a fun melody that matches the "refreshing and uplifting feeling of summer."

[0752] Data processing

[0753] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0754] Sending data

[0755] The server then sends the generated file back to the terminal, usually as an HTTP response containing BASE64-encoded file data.

[0756] Data display and playback

[0757] The device decodes the received file and displays it to the user. The lyrics are displayed as text and the melody is played as a music file. The user can check and enjoy the generated song through the device interface.

[0758] Specific examples

[0759] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0760] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0761] The processing flow will be explained below.

[0762] Step 1:

[0763] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0764] Step 2:

[0765] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0766] Step 3:

[0767] The device sends the data converted to JSON format to the server using an HTTP request, using the appropriate endpoint (URL).

[0768] Step 4:

[0769] The server receives the HTTP request and parses the JSON data, extracting the theme and mood specified by the user.

[0770] Step 5:

[0771] The server then passes the received theme and atmosphere data to the generative AI model, specifically, launching a natural language generation model and a music generation model.

[0772] Step 6:

[0773] The generative AI model (natural language generation model) generates lyrics based on the theme and mood specified by the user, such as "The sky is dyed in the colors of a sunset, and a lonely wind blows."

[0774] Step 7:

[0775] Similarly, generative AI models (music generation models) can generate melodies that fit a given theme or mood, such as a lonely piano melody.

[0776] Step 8:

[0777] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0778] Step 9:

[0779] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0780] Step 10:

[0781] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0782] Step 11:

[0783] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0784] Step 12:

[0785] Through the device interface, the user can read the generated lyrics and press the play button to listen to the melody.

[0786] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme and atmosphere.

[0787] Example 1

[0788] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0789] Conventional music generation systems have had difficulty generating original lyrics and melodies that perfectly match the theme and atmosphere specified by the user. Furthermore, they lacked the means to process the generated lyrics and melodies into an appropriate format and to display and play them in a way that is easy for the user to understand. As a result, the user experience was poor and the quality of the generated music was insufficient.

[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0791] In this invention, the server includes an analysis means for analyzing data on the theme and atmosphere, a generation means for generating lyrics and melody using a generative AI model, and a processing means for converting the generated lyrics and melody into an appropriate format and saving them. This makes it possible to automatically generate high-quality lyrics and melodies that perfectly match the theme and atmosphere specified by the user and provide them to the user in an appropriate format.

[0792] The "input means" is a means by which the user specifies the theme and atmosphere.

[0793] The "transmission means" is a means for transmitting data on the theme and atmosphere designated through the input means to the server.

[0794] The "generation means" is a means for receiving the theme and atmosphere data and generating lyrics and melody using a generative AI model.

[0795] "Expression means" refers to a means for transmitting the generated lyrics and melody data to a terminal.

[0796] The "display means" is a means for displaying or reproducing the lyrics and melody data.

[0797] The "processing means" is a means for converting the generated lyrics and melody into an appropriate format and saving it.

[0798] "Analysis methods" are methods for analyzing thematic and atmospheric data.

[0799] A "natural language generation model" is an algorithm or program that generates natural language sentences based on input text data.

[0800] A "music generation model" is an algorithm or program that generates music based on specified parameters and input data.

[0801] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[0802] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[0803] First, the user uses the terminal to input a specific theme or atmosphere. For example, a specific theme such as "a refreshing and uplifting summer song" can be input. An input field is displayed on the terminal, and the user inputs the theme or atmosphere into the specified field. This is called the "input means."

[0804] Next, the device sends the input data to the server. This transmission is usually done via an HTTP request. The input theme and atmosphere data is converted to JSON format and sent to the server. This is called the "transmission method."

[0805] The server receives and analyzes this transmitted data. This analysis method accurately understands the theme and mood data entered. The server then uses generative AI models to generate lyrics and a melody that match the specified theme and mood. These generative models include natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0806] The generated lyrics and melody are then converted into the appropriate format on the server. The lyrics are saved as a text file, and the melody is saved as a music file (e.g., .mp3 or .wav). This process is called the "processing method."

[0807] The server sends the generated file to the terminal. This transmission is also usually done as an HTTP response. The transmitted data is in BASE64 encoded format.

[0808] The terminal decodes the received data and displays it to the user. The lyrics are displayed as text, and the melody is played as a music file. This process is called "display means." The user can check and enjoy the generated music through the terminal's interface.

[0809] Specific examples

[0810] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[0811] Prompt Sentence Examples

[0812] "Generate lyrics and a melody for a song that evokes the lonely mood of an autumn evening."

[0813] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[0814] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0815] Step 1:

[0816] The user uses the terminal to input text about the theme or atmosphere. Specifically, the user enters "A refreshing and uplifting summer song" into the input field and presses the send button. The input at this point is text data about the theme or atmosphere specified by the user.

[0817] Step 2:

[0818] The device sends the text data of the theme and atmosphere entered to the server. Specifically, the device converts this data into JSON format and sends it to the server via an HTTP POST request. The input is the text data entered by the user, and the output is text data in JSON format.

[0819] Step 3:

[0820] The server receives and analyzes the data sent from the device. Specifically, the server parses the received HTTP request and extracts the "theme" field from the JSON data. The input is JSON format data, and the output is parsed text data (theme or atmosphere).

[0821] Step 4:

[0822] The server launches a generative AI model based on the analyzed theme and atmosphere. Specifically, it passes a prompt to the generative AI model, which then generates lyrics and a melody. The input is the extracted theme and atmosphere text data, and the output is the generated lyrics and melody data.

[0823] Step 5:

[0824] A generative AI model generates lyrics and a melody that fit a specified theme or atmosphere. Specifically, a natural language generation model (e.g., GPT-3) generates lyrics based on the theme, and a music generation model (e.g., MusicVAE) generates the melody. The input is a prompt to the generative AI model, and the output is text data of the lyrics and music data of the melody.

[0825] Step 6:

[0826] The server converts the generated lyrics and melody into the appropriate format and saves them. Specifically, it saves the lyrics as a text file and the melody as an MP3 or WAV file. The input is the generated lyrics and melody data, and the output is a text file and a music file.

[0827] Step 7:

[0828] The server sends the generated lyrics and melody files to the terminal. Specifically, it encodes the files in BASE64 and sends them to the terminal as an HTTP response. The input is a text file and a music file, and the output is the BASE64-encoded file data.

[0829] Step 8:

[0830] The terminal decodes the received file and displays and plays it for the user. Specifically, it decodes the lyric data into text format and displays it on the screen, and passes the melody data to the music player for playback. The input is BASE64-encoded file data, and the output is the lyrics displayed on the screen and the melody played.

[0831] (Application example 1)

[0832] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0833] In today's entertainment and content distribution industries, there is a growing demand for tools that allow individual users to enjoy creative activities. However, conventional systems require specialized knowledge and an advanced system environment to automatically generate high-quality original lyrics and melodies that match a theme or atmosphere. This makes it difficult for ordinary users to easily enjoy creating music, and there is also a lack of means for sharing these creations with other users in real time. The present invention aims to solve these problems and provide a system that allows users to easily create original music and share and enjoy it.

[0834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0835] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, generation means for receiving the data on the theme and atmosphere and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, display means for displaying or playing the lyrics and melody data, and sharing means for using the generated lyrics and melody to share them with other users in real time. This enables users to easily generate high-quality original music and further share the generated music with other users in real time.

[0836] "Input means" refers to the interface that the user uses to specify the theme and atmosphere.

[0837] The "transmission means" refers to a device or system having a function of transmitting data on the theme or atmosphere designated through the input means to the server.

[0838] "Generator" means a device or system that executes a process to generate lyrics and melody using a generative AI model based on received theme and mood data.

[0839] "Display means" refers to a device or system for displaying or playing back the generated lyrics and melody data on a terminal.

[0840] "Sharing means" refers to a device or system that has the function of sharing the created lyrics and melody with other users in real time.

[0841] A "generative AI model" refers to an artificial intelligence model that generates lyrics and melodies based on a theme or atmosphere specified by the user.

[0842] "Theme" refers to a subject or concept that allows the user to specify the content and atmosphere of a piece of music.

[0843] "Atmosphere" refers to the emotional tone or mood of a song.

[0844] "Song" refers to a musical composition that combines generated lyrics and melody.

[0845] "Server" refers to a central processing unit or system for managing data transmission, reception, and generation processes.

[0846] "Terminal" refers to electronic devices such as computers, smartphones, smart glasses, and head-mounted displays that are operated by users.

[0847] The present invention relates to a system that allows a user to specify a theme or atmosphere, automatically generate high-quality original lyrics and melodies, and share them. Specific embodiments of the system are described below.

[0848] System configuration

[0849] The system consists of the following elements:

[0850] 1. Input means: Provides an interface for users to specify the theme and atmosphere. Specifically, this applies to devices such as smartphones, smart glasses, and head-mounted displays.

[0851] 2. Transmission method: The input theme and atmosphere data is sent to the server using an HTTP request, and the data is packaged in JSON format.

[0852] 3. Generation: The server generates lyrics and a melody based on the received data using a generative AI model (e.g., GPT-3 or MusicVAE). This generation process includes both natural language generation and music generation.

[0853] 4. Display: The generated lyrics and melody are displayed or played on the user's device. The lyrics are provided as text and the melody as a music file.

[0854] 5. Sharing: We provide a function to share the created music with other users in real time, so that users can instantly enjoy the music they have created with other users.

[0855] Program processing

[0856] Server side:

[0857] The server receives the HTTP request, analyzes the theme and mood data, and activates a generative AI model. This uses a natural language generation model (GPT-3) and a music generation model (MusicVAE). These models generate lyrics and a melody that match the specified theme and mood, respectively. The generated data is converted into the appropriate format (lyrics as a text file, melody as a music file), BASE64 encoded, and sent to the device.

[0858] Terminal processing:

[0859] The device decodes the data received from the server and displays or plays it according to the media format. For example, lyrics are displayed as text, and melodies are played using a music player. This allows users to check and share the generated music in real time.

[0860] Hardware and software used

[0861] Hardware:

[0862] Input devices (smartphones, smart glasses, head-mounted displays)

[0863] Server (for data processing and generation)

[0864] Output device (same as above)

[0865] software:

[0866] Natural Language Generation Model (GPT-3)

[0867] Music generation model (MusicVAE)

[0868] Communication library (requests)

[0869] Programming language (Python)

[0870] Specific examples

[0871] As a concrete example, consider a case where a user uses a smartphone to specify a theme such as "a song for having fun with friends while watching fireworks on a summer night." When the user inputs this theme and presses the send button, the device sends the input data to a server. The server receives the theme data and generates lyrics and a melody using a generative AI model. The generated lyrics and melody are then sent back to the device, where they are displayed and played in a format that the user can view. This song can also be shared with other users in real time.

[0872] Example prompt sentence:

[0873] "A song for enjoying summer nights with friends while watching fireworks"

[0874] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0875] Step 1:

[0876] The user inputs the theme and atmosphere. Specifically, the user uses a smartphone, smart glasses, or a head-mounted display to input the theme and atmosphere in text format, such as "songs for enjoying a fun summer night with friends while watching fireworks." This input generates theme and atmosphere data.

[0877] Step 2:

[0878] The device sends the input theme and atmosphere data (text information) to the server. Specifically, it uses an HTTP request to package the theme and atmosphere data in JSON format and send it to the server. The input is text data from the user, and the output is JSON format data sent to the server.

[0879] Step 3:

[0880] The server analyzes the received theme and mood data. Specifically, it analyzes the text data of the theme and mood and converts it into a format suitable for the generative AI model. This analysis process generates appropriate instructions (prompts) based on the content of the text data. The input is theme data in JSON format, and the output is the prompts to be passed to the generative AI model.

[0881] Step 4:

[0882] The server then uses the analyzed data to launch a generative AI model to generate lyrics and a melody. Specifically, it generates lyrics using a natural language generation model (GPT-3) and a music generation model (MusicVAE) to generate a melody. The input is a prompt sentence, and the output is the generated lyrics and melody data.

[0883] Step 5:

[0884] The generated lyrics and melody data are converted into the appropriate format on the server. Specifically, the lyrics are encoded as a text file (.txt) and the melody is encoded as a music file (.mp3 or .wav). The input of this step is the generated lyrics and melody data, and the output is a text file and a music file.

[0885] Step 6:

[0886] The server sends the generated lyrics and melody data to the terminal. Specifically, it encodes this data in BASE64 and sends it back to the terminal as an HTTP response. The input is a text file and a music file, and the output is sent to the terminal as BASE64-encoded data.

[0887] Step 7:

[0888] The device decodes the received files and displays them to the user. Specifically, it decodes the BASE64 encoded data back into text and music files, displays the lyrics in text format, and plays the melody using a music player. The input to this step is the BASE64 encoded data, and the output is the text displayed on the user interface and the audio played.

[0889] Step 8:

[0890] The user checks the generated song and uses the sharing function of the device to share it with other users in real time. Specifically, the user selects the sharing option and sends the generated song to the other user's device. The input of this step is the generated song data, and the output is the shared data sent to other users.

[0891] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0892] The present invention relates to a system that not only allows a user to specify a theme and atmosphere, but also uses an emotion engine to recognize the user's emotions and automatically generates original lyrics and melodies based on those emotions. Specific embodiments of this system are described below.

[0893] System configuration

[0894] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion engine that recognizes user emotions.

[0895] User Input

[0896] The user inputs the text of the theme or mood using the interface provided on the device. For example, the user can input the theme "refreshing and uplifting summer music" and then use the voice input function to express their current mood.

[0897] Data transmission

[0898] The device converts the input data about the theme and mood into JSON format. In the case of voice input, the emotion engine performs another level of analysis, so the voice data is also sent along with the input data, making it easier for the server to analyze.

[0899] Emotion analysis

[0900] The server parses the received theme and mood data, as well as the audio data. The emotion engine analyzes the audio data to identify the user's emotion. The resulting emotional state is passed along with the text input to the generative AI model.

[0901] Starting the Generative AI

[0902] The server then uses the received data to activate a generative AI model, which generates lyrics and a melody based on the specified theme, mood, and analyzed emotions. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[0903] Lyric and melody generation

[0904] The generative AI model automatically generates lyrics and melodies based on the theme, atmosphere, and emotion specified by the user. For example, if a user requests a song with a lonely autumn evening mood and the emotion engine identifies the emotion as sad, the generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will be a melancholic piano melody.

[0905] Data processing

[0906] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[0907] Sending data

[0908] The server then BASE64-encodes the generated file and sends it to the terminal as an HTTP response, which also typically includes the BASE64-encoded file data and, if necessary, metadata.

[0909] Data display and playback

[0910] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0911] User Verification

[0912] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0913] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme, mood, and emotions. This invention will further enrich users' creative activities as music lovers and help them create high-quality music even without special technical skills.

[0914] The processing flow will be explained below.

[0915] Step 1:

[0916] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[0917] Step 2:

[0918] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[0919] Step 3:

[0920] The user provides voice input to the emotion engine. For example, the user may say, "I feel a little sad today."

[0921] Step 4:

[0922] The emotion engine analyzes voice input and identifies the user's emotion. Specifically, it uses voice recognition technology to analyze the emotion "sadness."

[0923] Step 5:

[0924] The device compiles the text input data and analyzed emotion data into JSON format and sends it to the server as an HTTP request.

[0925] Step 6:

[0926] The server receives the HTTP request and parses the JSON data, extracting the theme "A song with a lonely autumn evening mood" and the emotion "Sadness."

[0927] Step 7:

[0928] The server passes the extracted data to a generative AI model, which then uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE).

[0929] Step 8:

[0930] The natural language generation model generates lyrics based on the theme "A song with a lonely autumn evening mood" and the emotion "sadness." For example, it generates lyrics like "The sky is dyed in the colors of the sunset, and a lonely wind blows..."

[0931] Step 9:

[0932] Similarly, music generation models can generate melodies that fit a given theme, mood, or emotion, such as a lonely piano melody.

[0933] Step 10:

[0934] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[0935] Step 11:

[0936] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[0937] Step 12:

[0938] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[0939] Step 13:

[0940] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[0941] Step 14:

[0942] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[0943] Through this series of steps, users can easily create and check original lyrics and melodies that match their specified theme, atmosphere, and emotions at the time.

[0944] Example 2

[0945] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0946] Conventional music generation systems often fail to respond to users' requests for specific themes and moods. Furthermore, there are no systems that can recognize users' emotions and generate music based on them, making it difficult to generate original music that perfectly matches the user's emotions. Furthermore, there is a need for a system that allows users to easily view the generated music.

[0947] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0948] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, emotion analysis means for receiving the theme and atmosphere data and analyzing the user's voice input data to identify emotions, generation means for generating lyrics and melody using a generative AI model based on the analyzed emotional state and theme and atmosphere data, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data. This allows an original piece of music to be automatically generated that reflects not only the theme and atmosphere specified by the user but also the user's emotions, and makes it easy to check the music.

[0949] The "input means" is an interface that allows the user to specify a theme or atmosphere.

[0950] The "transmission means" is a mechanism for transmitting data on the theme and atmosphere designated through the input means to the server.

[0951] "Emotion analysis means" refers to a system or algorithm for analyzing a user's voice input data to identify emotions.

[0952] "Generation means" refers to a process or mechanism for generating lyrics and melody using a generative AI model based on the analyzed emotional state and thematic or atmospheric data.

[0953] The "display means" is a mechanism for displaying or playing back the generated lyrics and melody data on a terminal.

[0954] "Generative AI models" are artificial intelligence models that automatically generate lyrics and melodies based on user input data, including natural language generation models and music generation models.

[0955] A "prompt" is input data or an instruction given to a generative AI model, and is a sentence that influences the generated results.

[0956] This invention relates to a system that not only allows a user to specify a theme and atmosphere, but also recognizes the user's emotions using emotion analysis technology, and automatically generates original lyrics and melodies based on those emotions. A specific embodiment of this system will be described.

[0957] System configuration

[0958] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion analysis engine that recognizes user emotions.

[0959] User Input

[0960] The user inputs text about a theme or mood using the provided interface on the device. For example, the user can input a theme such as "a refreshing and uplifting summer song" and express their current mood by voice using the voice input function. This input method uses hardware such as a touchscreen and microphone, and software such as a web application.

[0961] Data transmission

[0962] The device converts the input theme and mood data into JSON format. In the case of voice input, the voice data is also sent to the server for further analysis by the sentiment analysis engine. Specifically, the device creates a file called input.json and sends it to the server via an HTTP POST request.

[0963] Emotion analysis

[0964] The server analyzes the received theme and mood data, as well as the audio data. An emotion analysis engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Emotion API) analyzes the audio data to identify the user's emotion. The resulting emotional state, along with the text input, is passed to a generative AI model. This analysis is performed using a cloud-based analysis service.

[0965] Starting the Generative AI

[0966] The server then launches a generative AI model based on the received data. Specifically, a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) are used. This generation process is performed using Python scripts and the API of the generative AI model.

[0967] Specific examples

[0968] For example, if a user requests a song with a lonely autumn evening feeling and the emotion analysis engine identifies the emotion as "sadness," the generative AI model will generate the following lyrics and melody:

[0969] Example prompt sentence:

[0970] Theme: Autumn Dusk

[0971] Emotion: Sadness

[0972] Generated lyrics: "The sky is dyed in the colors of sunset, a lonely wind blows..."

[0973] Generated Melody: A melancholic piano melody

[0974] Lyric and melody generation

[0975] The generative AI model automatically generates lyrics and melodies based on the theme, mood, and emotion specified by the user. In this process, the natural language generation model generates the lyrics, and the music generation model creates the melody.

[0976] Data processing and transmission

[0977] The generated lyrics and melody are converted into the appropriate format on the server. The lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav). These files are then BASE64 encoded and sent to the terminal as an HTTP response.

[0978] Data display and playback

[0979] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[0980] User Verification

[0981] Users can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing them to easily create and check original music that matches their emotions.

[0982] This allows users to create high-quality music that matches their specified theme, atmosphere, and emotions at the time, even without detailed technical skills. The generated music also perfectly matches the user's emotions, enriching the creative activities of music lovers.

[0983] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0984] Step 1: Accepting User Input

[0985] The user uses the device to input a theme or mood as text, and also uses the voice input function to express their current mood aloud. For example, the user enters the text "A refreshing and uplifting summer song" into the device's interface and speaks "I'm feeling very excited right now" into the microphone.

[0986] Input: Theme and atmosphere text, user voice data

[0987] Output: JSON format input data, audio data file

[0988] Step 2: Sending data

[0989] The device converts the input text data into JSON format and sends it to the server along with the audio data. Specifically, it parses the text data, converts it into a JSON file, saves the audio data, and sends it to the server via an HTTP POST request.

[0990] Input: JSON format input data, audio data file

[0991] Output: HTTP POST request data sent to the server

[0992] Step 3: Receiving and analyzing themes and emotions

[0993] The server analyzes the received JSON data and voice data. In particular, it uses a sentiment analysis engine to analyze the voice data and identify the user's emotion (e.g., "excited"). In this process, the server calls the sentiment analysis API and obtains the analysis results.

[0994] Input: JSON format input data, audio data file

[0995] Output: Sentiment analysis results (text format), theme and mood data

[0996] Step 4: Launching the generative AI model

[0997] The server then launches a generative AI model based on the analysis results. Specifically, it invokes a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) to generate lyrics and a melody using the prompt sentence.

[0998] Input: Sentiment analysis results, theme and mood data

[0999] Output: Generated lyrics data, generated melody data

[1000] Step 5: Processing the generated data

[1001] The generated lyrics and melody data are converted into the appropriate format: lyrics are saved as a text file (.txt), and melodies are saved as music files (.mp3 or .wav).

[1002] Input: Generated lyrics data, generated melody data

[1003] Output: Text file and music file

[1004] Step 6: Sending data

[1005] The server encodes the generated file in BASE64 and sends it to the terminal as an HTTP response. Specifically, the server encodes the file in BASE64 format, adds it to the body of the HTTP response, and sends it to the terminal.

[1006] Input: Text file, music file

[1007] Output: BASE64 encoded data contained in the HTTP response

[1008] Step 7: Receive and decode data

[1009] The device decodes the received HTTP response and converts the BASE64-encoded data back into the original text and music files. Specifically, it decodes the BASE64 data and saves it as a file.

[1010] Input: BASE64 encoded data included in the HTTP response

[1011] Output: Original text file, music file

[1012] Step 8: View and Play Data

[1013] The device displays the lyrics as text and the melody as a playable music file, using HTML5 audio tags and text display functions.

[1014] Input: Original text file, music file

[1015] Output: Lyric text displayed, melody played

[1016] Step 9: Verify the user

[1017] The user can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing the user to check the quality of the generated music.

[1018] Input: Lyric text to be displayed, melody to be played

[1019] Output: User confirmation results (subjective evaluation)

[1020] (Application example 2)

[1021] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1022] Conventional music generation systems simply generate lyrics and melodies based on a theme or atmosphere specified by the user, and are unable to reflect the user's emotions. This makes it difficult to provide music that perfectly matches the user's current emotions. Furthermore, the music generation process is not intuitive for users, making it difficult to customize the system to meet individual needs.

[1023] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1024] In this invention, the server includes input means for the user to specify a theme and atmosphere, voice analysis means for recording and analyzing the user's voice input, transmission means for transmitting the theme and atmosphere specified through the input means and voice analysis means and analyzed emotion data to the server, generation means for receiving the theme, atmosphere, and emotion data and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data, thereby enabling the generation and playback of original lyrics and melodies that match the user's emotions.

[1025] The "input means for the user to specify the theme or atmosphere" is a device or interface function that allows the user to input the theme or atmosphere of the music piece that he or she desires in text form.

[1026] The "voice analysis means for recording and analyzing the user's voice input" is a device or function for recording the voice uttered by the user and analyzing the voice data to extract information such as emotions.

[1027] The "transmission means for transmitting data on theme, atmosphere, and emotion to the server" is a device or function for transmitting the theme and atmosphere input by the user and analyzed emotion data to the server.

[1028] A "generation means for generating lyrics and melodies using a generative AI model" is a device or function that automatically generates lyrics and melodies using an artificial intelligence model based on input data.

[1029] The "transmission means for transmitting lyric and melody data to a terminal" is a device or function for transmitting the generated lyric and melody data to a user's terminal.

[1030] The "display means for displaying or reproducing the lyric and melody data" is a device or function for visually displaying the transmitted lyric and melody data on the user terminal or reproducing it as sound.

[1031] The present invention relates to a system that automatically generates original lyrics and melodies based on a user's emotions, and displays them visually and plays them audibly on a user terminal. Specific embodiments of this system are described below.

[1032] First, the user inputs the theme and mood of the song in text using a device such as a smartphone or smart glasses, and simultaneously inputs voice to express their current emotions.

[1033] Next, the device's built-in voice analysis unit records the user's voice input and analyzes their emotions. The analysis results, along with the theme and mood data entered as text, are converted into JSON format and sent to the server.

[1034] The server generates lyrics and melodies using a generative AI model based on the received theme and mood data and the analyzed emotional data. This generation process uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE). The generated lyrics and melodies are saved in text files (.txt) and music files (.mp3 or .wav), respectively.

[1035] The server then encodes the generated lyrics and melody data into BASE64 format and sends it as an HTTP response to the user's device, where it decodes the received BASE64 data and restores it to the original text and music files.

[1036] Finally, the user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[1037] As a specific example, if a user inputs the theme "A song to cheer me up when I'm a little tired," and expresses in voice input, "I've been working all day today and I'm a little tired. But tomorrow is a day off, so I have to do my best," then based on the analysis results, a bright, uplifting song with a rhythmic melody will be generated.

[1038] An example of a prompt might be:

[1039] "A refreshing and uplifting summer song"

[1040] "I feel great today!"

[1041] This allows users to easily create and enjoy original music that perfectly matches their emotions.

[1042] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1043] Step 1:

[1044] Using a device such as a smartphone or smart glasses, the user inputs the theme and mood of the song in text, and simultaneously inputs voice to express their current emotions.

[1045] Input: Text (e.g., "I'm a little tired, so I need a song to cheer me up"), audio data (e.g., "I've been working all day today and I'm a little tired. But tomorrow is a holiday, so I have to do my best.")

[1046] Output: User-entered text and audio data of the theme and mood

[1047] Step 2:

[1048] A voice analysis means of the terminal records the user's voice input and analyzes the voice data to extract emotion data.

[1049] Input: User's voice data

[1050] Data processing: Analysis of voice data and emotion recognition (e.g., emotion classification such as happiness, sadness, fatigue, etc.)

[1051] Output: Parsed emotion data (e.g., "tired")

[1052] Step 3:

[1053] The device converts the theme and atmosphere data entered by the user and the analyzed emotional data into JSON format and sends it to the server.

[1054] Input: Theme and mood of text input, analyzed sentiment data

[1055] Data processing: Conversion to JSON format

[1056] Output: Send data in JSON format (e.g., {"theme": "Songs to cheer you up when you're tired", "emotion": "Tired"})

[1057] Step 4:

[1058] The server extracts theme, mood, and emotion data from the received JSON data, and then invokes a generative AI model to generate lyrics and a melody based on the specified theme, mood, and emotion.

[1059] Input: JSON format data (theme, mood, emotion data)

[1060] Data processing: Generate lyrics using natural language generation models (e.g., GPT-3) and melodies using music generation models (e.g., MusicVAE).

[1061] Output: Generated lyrics (text format) and melody (music file format)

[1062] Step 5:

[1063] The server encodes the generated lyrics and melody data into BASE64 format and sends it to the user's terminal as an HTTP response.

[1064] Input: Generated lyrics and melody data

[1065] Data processing: Encoding to BASE64 format and creating HTTP responses

[1066] Output: BASE64 encoded lyrics and melody data (sent as HTTP response)

[1067] Step 6:

[1068] The terminal decodes the received BASE64 encoded data and restores it to the original text and music files.

[1069] Input: BASE64 encoded lyrics and melody data

[1070] Data processing: Decoding BASE64 data

[1071] Output: Original text file (lyrics) and music file (melody)

[1072] Step 7:

[1073] The user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[1074] Input: Decoded lyrics text file and music file

[1075] Specific behavior: Display a text file and play a music file

[1076] Output: User sees lyrics visually and hears melody (e.g., lyrics displayed on device screen and melody played through speaker)

[1077] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1078] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1079] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1080] [Fourth embodiment]

[1081] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1082] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1083] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1084] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1085] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1086] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1087] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1088] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1089] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1090] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1091] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1092] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1093] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1094] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[1095] System configuration

[1096] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[1097] User Input

[1098] The user inputs the theme or mood text using the interface provided at the terminal, for example, the theme "a refreshing and uplifting summer song."

[1099] Data transmission

[1100] The device sends the entered data about the theme and mood to the server, typically using an HTTP request, with the entered text packaged in JSON format.

[1101] Receiving data and starting generation AI

[1102] The server receives and analyzes the data sent from the device. A generative AI model is then activated to generate lyrics and a melody that match the specified theme and atmosphere. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[1103] Lyric and melody generation

[1104] The generative AI model automatically generates lyrics and melodies based on user specifications. For example, it generates bright lyrics and a fun melody that matches the "refreshing and uplifting feeling of summer."

[1105] Data processing

[1106] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[1107] Sending data

[1108] The server then sends the generated file back to the terminal, usually as an HTTP response containing BASE64-encoded file data.

[1109] Data display and playback

[1110] The device decodes the received file and displays it to the user. The lyrics are displayed as text and the melody is played as a music file. The user can check and enjoy the generated song through the device interface.

[1111] Specific examples

[1112] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[1113] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[1114] The processing flow will be explained below.

[1115] Step 1:

[1116] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[1117] Step 2:

[1118] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[1119] Step 3:

[1120] The device sends the data converted to JSON format to the server using an HTTP request, using the appropriate endpoint (URL).

[1121] Step 4:

[1122] The server receives the HTTP request and parses the JSON data, extracting the theme and mood specified by the user.

[1123] Step 5:

[1124] The server then passes the received theme and atmosphere data to the generative AI model, specifically, launching a natural language generation model and a music generation model.

[1125] Step 6:

[1126] The generative AI model (natural language generation model) generates lyrics based on the theme and mood specified by the user, such as "The sky is dyed in the colors of a sunset, and a lonely wind blows."

[1127] Step 7:

[1128] Similarly, generative AI models (music generation models) can generate melodies that fit a given theme or mood, such as a lonely piano melody.

[1129] Step 8:

[1130] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[1131] Step 9:

[1132] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[1133] Step 10:

[1134] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[1135] Step 11:

[1136] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[1137] Step 12:

[1138] Through the device interface, the user can read the generated lyrics and press the play button to listen to the melody.

[1139] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme and atmosphere.

[1140] Example 1

[1141] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1142] Conventional music generation systems have had difficulty generating original lyrics and melodies that perfectly match the theme and atmosphere specified by the user. Furthermore, they lacked the means to process the generated lyrics and melodies into an appropriate format and to display and play them in a way that is easy for the user to understand. As a result, the user experience was poor and the quality of the generated music was insufficient.

[1143] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1144] In this invention, the server includes an analysis means for analyzing data on the theme and atmosphere, a generation means for generating lyrics and melody using a generative AI model, and a processing means for converting the generated lyrics and melody into an appropriate format and saving them. This makes it possible to automatically generate high-quality lyrics and melodies that perfectly match the theme and atmosphere specified by the user and provide them to the user in an appropriate format.

[1145] The "input means" is a means by which the user specifies the theme and atmosphere.

[1146] The "transmission means" is a means for transmitting data on the theme and atmosphere designated through the input means to the server.

[1147] The "generation means" is a means for receiving the theme and atmosphere data and generating lyrics and melody using a generative AI model.

[1148] "Expression means" refers to a means for transmitting the generated lyrics and melody data to a terminal.

[1149] The "display means" is a means for displaying or reproducing the lyrics and melody data.

[1150] The "processing means" is a means for converting the generated lyrics and melody into an appropriate format and saving it.

[1151] "Analysis methods" are methods for analyzing thematic and atmospheric data.

[1152] A "natural language generation model" is an algorithm or program that generates natural language sentences based on input text data.

[1153] A "music generation model" is an algorithm or program that generates music based on specified parameters and input data.

[1154] The present invention relates to a system for automatically generating high-quality original lyrics and melodies by a user specifying a theme and atmosphere. Specific embodiments of this system are described below.

[1155] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.

[1156] First, the user uses the terminal to input a specific theme or atmosphere. For example, a specific theme such as "a refreshing and uplifting summer song" can be input. An input field is displayed on the terminal, and the user inputs the theme or atmosphere into the specified field. This is called the "input means."

[1157] Next, the device sends the input data to the server. This transmission is usually done via an HTTP request. The input theme and atmosphere data is converted to JSON format and sent to the server. This is called the "transmission method."

[1158] The server receives and analyzes this transmitted data. This analysis method accurately understands the theme and mood data entered. The server then uses generative AI models to generate lyrics and a melody that match the specified theme and mood. These generative models include natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[1159] The generated lyrics and melody are then converted into the appropriate format on the server. The lyrics are saved as a text file, and the melody is saved as a music file (e.g., .mp3 or .wav). This process is called the "processing method."

[1160] The server sends the generated file to the terminal. This transmission is also usually done as an HTTP response. The transmitted data is in BASE64 encoded format.

[1161] The terminal decodes the received data and displays it to the user. The lyrics are displayed as text, and the melody is played as a music file. This process is called "display means." The user can check and enjoy the generated music through the terminal's interface.

[1162] Specific examples

[1163] For example, suppose a user wants a song with a "lonely autumn evening mood." The user inputs the theme into the device and presses the send button. The device sends the theme data to the server. The server receives the theme data and activates a generative AI model to generate lyrics and a melody. The generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will consist of a lonely piano melody. The server then converts the generated lyrics into a text file and the melody into a music file, which are then sent to the device. Finally, the user can read the generated lyrics on the device and press the play button to listen to the melody.

[1164] Prompt Sentence Examples

[1165] "Generate lyrics and a melody for a song that evokes the lonely mood of an autumn evening."

[1166] In this way, the system of the present invention allows users to easily create and enjoy high-quality original songs.

[1167] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1168] Step 1:

[1169] The user uses the terminal to input text about the theme or atmosphere. Specifically, the user enters "A refreshing and uplifting summer song" into the input field and presses the send button. The input at this point is text data about the theme or atmosphere specified by the user.

[1170] Step 2:

[1171] The device sends the text data of the theme and atmosphere entered to the server. Specifically, the device converts this data into JSON format and sends it to the server via an HTTP POST request. The input is the text data entered by the user, and the output is text data in JSON format.

[1172] Step 3:

[1173] The server receives and analyzes the data sent from the device. Specifically, the server parses the received HTTP request and extracts the "theme" field from the JSON data. The input is JSON format data, and the output is parsed text data (theme or atmosphere).

[1174] Step 4:

[1175] The server launches a generative AI model based on the analyzed theme and atmosphere. Specifically, it passes a prompt to the generative AI model, which then generates lyrics and a melody. The input is the extracted theme and atmosphere text data, and the output is the generated lyrics and melody data.

[1176] Step 5:

[1177] A generative AI model generates lyrics and a melody that fit a specified theme or atmosphere. Specifically, a natural language generation model (e.g., GPT-3) generates lyrics based on the theme, and a music generation model (e.g., MusicVAE) generates the melody. The input is a prompt to the generative AI model, and the output is text data of the lyrics and music data of the melody.

[1178] Step 6:

[1179] The server converts the generated lyrics and melody into the appropriate format and saves them. Specifically, it saves the lyrics as a text file and the melody as an MP3 or WAV file. The input is the generated lyrics and melody data, and the output is a text file and a music file.

[1180] Step 7:

[1181] The server sends the generated lyrics and melody files to the terminal. Specifically, it encodes the files in BASE64 and sends them to the terminal as an HTTP response. The input is a text file and a music file, and the output is the BASE64-encoded file data.

[1182] Step 8:

[1183] The terminal decodes the received file and displays and plays it for the user. Specifically, it decodes the lyric data into text format and displays it on the screen, and passes the melody data to the music player for playback. The input is BASE64-encoded file data, and the output is the lyrics displayed on the screen and the melody played.

[1184] (Application example 1)

[1185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1186] In today's entertainment and content distribution industries, there is a growing demand for tools that allow individual users to enjoy creative activities. However, conventional systems require specialized knowledge and an advanced system environment to automatically generate high-quality original lyrics and melodies that match a theme or atmosphere. This makes it difficult for ordinary users to easily enjoy creating music, and there is also a lack of means for sharing these creations with other users in real time. The present invention aims to solve these problems and provide a system that allows users to easily create original music and share and enjoy it.

[1187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1188] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, generation means for receiving the data on the theme and atmosphere and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, display means for displaying or playing the lyrics and melody data, and sharing means for using the generated lyrics and melody to share them with other users in real time. This enables users to easily generate high-quality original music and further share the generated music with other users in real time.

[1189] "Input means" refers to the interface that the user uses to specify the theme and atmosphere.

[1190] The "transmission means" refers to a device or system having a function of transmitting data on the theme or atmosphere designated through the input means to the server.

[1191] "Generator" means a device or system that executes a process to generate lyrics and melody using a generative AI model based on received theme and mood data.

[1192] "Display means" refers to a device or system for displaying or playing back the generated lyrics and melody data on a terminal.

[1193] "Sharing means" refers to a device or system that has the function of sharing the created lyrics and melody with other users in real time.

[1194] A "generative AI model" refers to an artificial intelligence model that generates lyrics and melodies based on a theme or atmosphere specified by the user.

[1195] "Theme" refers to a subject or concept that allows the user to specify the content and atmosphere of a piece of music.

[1196] "Atmosphere" refers to the emotional tone or mood of a song.

[1197] "Song" refers to a musical composition that combines generated lyrics and melody.

[1198] "Server" refers to a central processing unit or system for managing data transmission, reception, and generation processes.

[1199] "Terminal" refers to electronic devices such as computers, smartphones, smart glasses, and head-mounted displays that are operated by users.

[1200] The present invention relates to a system that allows a user to specify a theme or atmosphere, automatically generate high-quality original lyrics and melodies, and share them. Specific embodiments of the system are described below.

[1201] System configuration

[1202] The system consists of the following elements:

[1203] 1. Input means: Provides an interface for users to specify the theme and atmosphere. Specifically, this applies to devices such as smartphones, smart glasses, and head-mounted displays.

[1204] 2. Transmission method: The input theme and atmosphere data is sent to the server using an HTTP request, and the data is packaged in JSON format.

[1205] 3. Generation: The server generates lyrics and a melody based on the received data using a generative AI model (e.g., GPT-3 or MusicVAE). This generation process includes both natural language generation and music generation.

[1206] 4. Display: The generated lyrics and melody are displayed or played on the user's device. The lyrics are provided as text and the melody as a music file.

[1207] 5. Sharing: We provide a function to share the created music with other users in real time, so that users can instantly enjoy the music they have created with other users.

[1208] Program processing

[1209] Server side:

[1210] The server receives the HTTP request, analyzes the theme and mood data, and activates a generative AI model. This uses a natural language generation model (GPT-3) and a music generation model (MusicVAE). These models generate lyrics and a melody that match the specified theme and mood, respectively. The generated data is converted into the appropriate format (lyrics as a text file, melody as a music file), BASE64 encoded, and sent to the device.

[1211] Terminal processing:

[1212] The device decodes the data received from the server and displays or plays it according to the media format. For example, lyrics are displayed as text, and melodies are played using a music player. This allows users to check and share the generated music in real time.

[1213] Hardware and software used

[1214] Hardware:

[1215] Input devices (smartphones, smart glasses, head-mounted displays)

[1216] Server (for data processing and generation)

[1217] Output device (same as above)

[1218] software:

[1219] Natural Language Generation Model (GPT-3)

[1220] Music generation model (MusicVAE)

[1221] Communication library (requests)

[1222] Programming language (Python)

[1223] Specific examples

[1224] As a concrete example, consider a case where a user uses a smartphone to specify a theme such as "a song for having fun with friends while watching fireworks on a summer night." When the user inputs this theme and presses the send button, the device sends the input data to a server. The server receives the theme data and generates lyrics and a melody using a generative AI model. The generated lyrics and melody are then sent back to the device, where they are displayed and played in a format that the user can view. This song can also be shared with other users in real time.

[1225] Example prompt sentence:

[1226] "A song for enjoying summer nights with friends while watching fireworks"

[1227] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1228] Step 1:

[1229] The user inputs the theme and atmosphere. Specifically, the user uses a smartphone, smart glasses, or a head-mounted display to input the theme and atmosphere in text format, such as "songs for enjoying a fun summer night with friends while watching fireworks." This input generates theme and atmosphere data.

[1230] Step 2:

[1231] The device sends the input theme and atmosphere data (text information) to the server. Specifically, it uses an HTTP request to package the theme and atmosphere data in JSON format and send it to the server. The input is text data from the user, and the output is JSON format data sent to the server.

[1232] Step 3:

[1233] The server analyzes the received theme and mood data. Specifically, it analyzes the text data of the theme and mood and converts it into a format suitable for the generative AI model. This analysis process generates appropriate instructions (prompts) based on the content of the text data. The input is theme data in JSON format, and the output is the prompts to be passed to the generative AI model.

[1234] Step 4:

[1235] The server then uses the analyzed data to launch a generative AI model to generate lyrics and a melody. Specifically, it generates lyrics using a natural language generation model (GPT-3) and a music generation model (MusicVAE) to generate a melody. The input is a prompt sentence, and the output is the generated lyrics and melody data.

[1236] Step 5:

[1237] The generated lyrics and melody data are converted into the appropriate format on the server. Specifically, the lyrics are encoded as a text file (.txt) and the melody is encoded as a music file (.mp3 or .wav). The input of this step is the generated lyrics and melody data, and the output is a text file and a music file.

[1238] Step 6:

[1239] The server sends the generated lyrics and melody data to the terminal. Specifically, it encodes this data in BASE64 and sends it back to the terminal as an HTTP response. The input is a text file and a music file, and the output is sent to the terminal as BASE64-encoded data.

[1240] Step 7:

[1241] The device decodes the received files and displays them to the user. Specifically, it decodes the BASE64 encoded data back into text and music files, displays the lyrics in text format, and plays the melody using a music player. The input to this step is the BASE64 encoded data, and the output is the text displayed on the user interface and the audio played.

[1242] Step 8:

[1243] The user checks the generated song and uses the sharing function of the device to share it with other users in real time. Specifically, the user selects the sharing option and sends the generated song to the other user's device. The input of this step is the generated song data, and the output is the shared data sent to other users.

[1244] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1245] The present invention relates to a system that not only allows a user to specify a theme and atmosphere, but also uses an emotion engine to recognize the user's emotions and automatically generates original lyrics and melodies based on those emotions. Specific embodiments of this system are described below.

[1246] System configuration

[1247] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion engine that recognizes user emotions.

[1248] User Input

[1249] The user inputs the text of the theme or mood using the interface provided on the device. For example, the user can input the theme "refreshing and uplifting summer music" and then use the voice input function to express their current mood.

[1250] Data transmission

[1251] The device converts the input data about the theme and mood into JSON format. In the case of voice input, the emotion engine performs another level of analysis, so the voice data is also sent along with the input data, making it easier for the server to analyze.

[1252] Emotion analysis

[1253] The server parses the received theme and mood data, as well as the audio data. The emotion engine analyzes the audio data to identify the user's emotion. The resulting emotional state is passed along with the text input to the generative AI model.

[1254] Starting the Generative AI

[1255] The server then uses the received data to activate a generative AI model, which generates lyrics and a melody based on the specified theme, mood, and analyzed emotions. This generation process utilizes natural language generation models (e.g., GPT-3) and music generation models (e.g., MusicVAE).

[1256] Lyric and melody generation

[1257] The generative AI model automatically generates lyrics and melodies based on the theme, atmosphere, and emotion specified by the user. For example, if a user requests a song with a lonely autumn evening mood and the emotion engine identifies the emotion as sad, the generated lyrics will be something like "The sky is dyed in the colors of a sunset, a lonely wind blows..." and the melody will be a melancholic piano melody.

[1258] Data processing

[1259] The generated lyrics and melody are converted to the appropriate format on the server: lyrics are saved as a text file (.txt), and melody as a music file (.mp3 or .wav).

[1260] Sending data

[1261] The server then BASE64-encodes the generated file and sends it to the terminal as an HTTP response, which also typically includes the BASE64-encoded file data and, if necessary, metadata.

[1262] Data display and playback

[1263] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[1264] User Verification

[1265] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[1266] Through this series of steps, users can easily create and check original lyrics and melodies that fit their specified theme, mood, and emotions. This invention will further enrich users' creative activities as music lovers and help them create high-quality music even without special technical skills.

[1267] The processing flow will be explained below.

[1268] Step 1:

[1269] The user uses the terminal to input text describing a theme or mood, for example, "A song with a lonely autumn evening mood."

[1270] Step 2:

[1271] The device converts the theme and mood data you input into JSON format, which makes it easy for the server to parse.

[1272] Step 3:

[1273] The user provides voice input to the emotion engine. For example, the user may say, "I feel a little sad today."

[1274] Step 4:

[1275] The emotion engine analyzes voice input and identifies the user's emotion. Specifically, it uses voice recognition technology to analyze the emotion "sadness."

[1276] Step 5:

[1277] The device compiles the text input data and analyzed emotion data into JSON format and sends it to the server as an HTTP request.

[1278] Step 6:

[1279] The server receives the HTTP request and parses the JSON data, extracting the theme "A song with a lonely autumn evening mood" and the emotion "Sadness."

[1280] Step 7:

[1281] The server passes the extracted data to a generative AI model, which then uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE).

[1282] Step 8:

[1283] The natural language generation model generates lyrics based on the theme "A song with a lonely autumn evening mood" and the emotion "sadness." For example, it generates lyrics like "The sky is dyed in the colors of the sunset, and a lonely wind blows..."

[1284] Step 9:

[1285] Similarly, music generation models can generate melodies that fit a given theme, mood, or emotion, such as a lonely piano melody.

[1286] Step 10:

[1287] The server converts the generated lyrics and melody into the appropriate format: the lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav).

[1288] Step 11:

[1289] The server then BASE64 encodes the converted file and sends it to the terminal as an HTTP response, including metadata if necessary.

[1290] Step 12:

[1291] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files.

[1292] Step 13:

[1293] The terminal presents the user with an interface that displays the lyrics as text and provides the melody as a playable music file.

[1294] Step 14:

[1295] Through the device interface, users can read the generated lyrics and press play to listen to the melody, which will perfectly match the user's emotions.

[1296] Through this series of steps, users can easily create and check original lyrics and melodies that match their specified theme, atmosphere, and emotions at the time.

[1297] Example 2

[1298] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1299] Conventional music generation systems often fail to respond to users' requests for specific themes and moods. Furthermore, there are no systems that can recognize users' emotions and generate music based on them, making it difficult to generate original music that perfectly matches the user's emotions. Furthermore, there is a need for a system that allows users to easily view the generated music.

[1300] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1301] In this invention, the server includes input means for a user to specify a theme and atmosphere, transmission means for transmitting data on the theme and atmosphere specified through the input means to the server, emotion analysis means for receiving the theme and atmosphere data and analyzing the user's voice input data to identify emotions, generation means for generating lyrics and melody using a generative AI model based on the analyzed emotional state and theme and atmosphere data, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data. This allows an original piece of music to be automatically generated that reflects not only the theme and atmosphere specified by the user but also the user's emotions, and makes it easy to check the music.

[1302] The "input means" is an interface that allows the user to specify a theme or atmosphere.

[1303] The "transmission means" is a mechanism for transmitting data on the theme and atmosphere designated through the input means to the server.

[1304] "Emotion analysis means" refers to a system or algorithm for analyzing a user's voice input data to identify emotions.

[1305] "Generation means" refers to a process or mechanism for generating lyrics and melody using a generative AI model based on the analyzed emotional state and thematic or atmospheric data.

[1306] The "display means" is a mechanism for displaying or playing back the generated lyrics and melody data on a terminal.

[1307] "Generative AI models" are artificial intelligence models that automatically generate lyrics and melodies based on user input data, including natural language generation models and music generation models.

[1308] A "prompt" is input data or an instruction given to a generative AI model, and is a sentence that influences the generated results.

[1309] This invention relates to a system that not only allows a user to specify a theme and atmosphere, but also recognizes the user's emotions using emotion analysis technology, and automatically generates original lyrics and melodies based on those emotions. A specific embodiment of this system will be described.

[1310] System configuration

[1311] This system consists of a terminal that accepts user input, a server that processes and generates data, and a terminal that displays the generated data.It also includes an emotion analysis engine that recognizes user emotions.

[1312] User Input

[1313] The user inputs text about a theme or mood using the provided interface on the device. For example, the user can input a theme such as "a refreshing and uplifting summer song" and express their current mood by voice using the voice input function. This input method uses hardware such as a touchscreen and microphone, and software such as a web application.

[1314] Data transmission

[1315] The device converts the input theme and mood data into JSON format. In the case of voice input, the voice data is also sent to the server for further analysis by the sentiment analysis engine. Specifically, the device creates a file called input.json and sends it to the server via an HTTP POST request.

[1316] Emotion analysis

[1317] The server analyzes the received theme and mood data, as well as the audio data. An emotion analysis engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Emotion API) analyzes the audio data to identify the user's emotion. The resulting emotional state, along with the text input, is passed to a generative AI model. This analysis is performed using a cloud-based analysis service.

[1318] Starting the Generative AI

[1319] The server then launches a generative AI model based on the received data. Specifically, a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) are used. This generation process is performed using Python scripts and the API of the generative AI model.

[1320] Specific examples

[1321] For example, if a user requests a song with a lonely autumn evening feeling and the emotion analysis engine identifies the emotion as "sadness," the generative AI model will generate the following lyrics and melody:

[1322] Example prompt sentence:

[1323] Theme: Autumn Dusk

[1324] Emotion: Sadness

[1325] Generated lyrics: "The sky is dyed in the colors of sunset, a lonely wind blows..."

[1326] Generated Melody: A melancholic piano melody

[1327] Lyric and melody generation

[1328] The generative AI model automatically generates lyrics and melodies based on the theme, mood, and emotion specified by the user. In this process, the natural language generation model generates the lyrics, and the music generation model creates the melody.

[1329] Data processing and transmission

[1330] The generated lyrics and melody are converted into the appropriate format on the server. The lyrics are saved as a text file (.txt) and the melody as a music file (.mp3 or .wav). These files are then BASE64 encoded and sent to the terminal as an HTTP response.

[1331] Data display and playback

[1332] The device decodes the received HTTP response and converts the BASE64 data back into the original text and music files. The device then presents the user with an interface that displays the lyrics as text and the melody as a playable music file.

[1333] User Verification

[1334] Users can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing them to easily create and check original music that matches their emotions.

[1335] This allows users to create high-quality music that matches their specified theme, atmosphere, and emotions at the time, even without detailed technical skills. The generated music also perfectly matches the user's emotions, enriching the creative activities of music lovers.

[1336] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1337] Step 1: Accepting User Input

[1338] The user uses the device to input a theme or mood as text, and also uses the voice input function to express their current mood aloud. For example, the user enters the text "A refreshing and uplifting summer song" into the device's interface and speaks "I'm feeling very excited right now" into the microphone.

[1339] Input: Theme and atmosphere text, user voice data

[1340] Output: JSON format input data, audio data file

[1341] Step 2: Sending data

[1342] The device converts the input text data into JSON format and sends it to the server along with the audio data. Specifically, it parses the text data, converts it into a JSON file, saves the audio data, and sends it to the server via an HTTP POST request.

[1343] Input: JSON format input data, audio data file

[1344] Output: HTTP POST request data sent to the server

[1345] Step 3: Receiving and analyzing themes and emotions

[1346] The server analyzes the received JSON data and voice data. In particular, it uses a sentiment analysis engine to analyze the voice data and identify the user's emotion (e.g., "excited"). In this process, the server calls the sentiment analysis API and obtains the analysis results.

[1347] Input: JSON format input data, audio data file

[1348] Output: Sentiment analysis results (text format), theme and mood data

[1349] Step 4: Launching the generative AI model

[1350] The server then launches a generative AI model based on the analysis results. Specifically, it invokes a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE) to generate lyrics and a melody using the prompt sentence.

[1351] Input: Sentiment analysis results, theme and mood data

[1352] Output: Generated lyrics data, generated melody data

[1353] Step 5: Processing the generated data

[1354] The generated lyrics and melody data are converted into the appropriate format: lyrics are saved as a text file (.txt), and melodies are saved as music files (.mp3 or .wav).

[1355] Input: Generated lyrics data, generated melody data

[1356] Output: Text file and music file

[1357] Step 6: Sending data

[1358] The server encodes the generated file in BASE64 and sends it to the terminal as an HTTP response. Specifically, the server encodes the file in BASE64 format, adds it to the body of the HTTP response, and sends it to the terminal.

[1359] Input: Text file, music file

[1360] Output: BASE64 encoded data contained in the HTTP response

[1361] Step 7: Receive and decode data

[1362] The device decodes the received HTTP response and converts the BASE64-encoded data back into the original text and music files. Specifically, it decodes the BASE64 data and saves it as a file.

[1363] Input: BASE64 encoded data included in the HTTP response

[1364] Output: Original text file, music file

[1365] Step 8: View and Play Data

[1366] The device displays the lyrics as text and the melody as a playable music file, using HTML5 audio tags and text display functions.

[1367] Input: Original text file, music file

[1368] Output: Lyric text displayed, melody played

[1369] Step 9: Verify the user

[1370] The user can read the generated lyrics through the device interface and press the play button to listen to the melody, allowing the user to check the quality of the generated music.

[1371] Input: Lyric text to be displayed, melody to be played

[1372] Output: User confirmation results (subjective evaluation)

[1373] (Application example 2)

[1374] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1375] Conventional music generation systems simply generate lyrics and melodies based on a theme or atmosphere specified by the user, and are unable to reflect the user's emotions. This makes it difficult to provide music that perfectly matches the user's current emotions. Furthermore, the music generation process is not intuitive for users, making it difficult to customize the system to meet individual needs.

[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1377] In this invention, the server includes input means for the user to specify a theme and atmosphere, voice analysis means for recording and analyzing the user's voice input, transmission means for transmitting the theme and atmosphere specified through the input means and voice analysis means and analyzed emotion data to the server, generation means for receiving the theme, atmosphere, and emotion data and generating lyrics and melody using a generative AI model, transmission means for transmitting the generated lyrics and melody data to a terminal, and display means for displaying or playing back the lyrics and melody data, thereby enabling the generation and playback of original lyrics and melodies that match the user's emotions.

[1378] The "input means for the user to specify the theme or atmosphere" is a device or interface function that allows the user to input the theme or atmosphere of the music piece that he or she desires in text form.

[1379] The "voice analysis means for recording and analyzing the user's voice input" is a device or function for recording the voice uttered by the user and analyzing the voice data to extract information such as emotions.

[1380] The "transmission means for transmitting data on theme, atmosphere, and emotion to the server" is a device or function for transmitting the theme and atmosphere input by the user and analyzed emotion data to the server.

[1381] A "generation means for generating lyrics and melodies using a generative AI model" is a device or function that automatically generates lyrics and melodies using an artificial intelligence model based on input data.

[1382] The "transmission means for transmitting lyric and melody data to a terminal" is a device or function for transmitting the generated lyric and melody data to a user's terminal.

[1383] The "display means for displaying or reproducing the lyric and melody data" is a device or function for visually displaying the transmitted lyric and melody data on the user terminal or reproducing it as sound.

[1384] The present invention relates to a system that automatically generates original lyrics and melodies based on a user's emotions, and displays them visually and plays them audibly on a user terminal. Specific embodiments of this system are described below.

[1385] First, the user inputs the theme and mood of the song in text using a device such as a smartphone or smart glasses, and simultaneously inputs voice to express their current emotions.

[1386] Next, the device's built-in voice analysis unit records the user's voice input and analyzes their emotions. The analysis results, along with the theme and mood data entered as text, are converted into JSON format and sent to the server.

[1387] The server generates lyrics and melodies using a generative AI model based on the received theme and mood data and the analyzed emotional data. This generation process uses a natural language generation model (e.g., GPT-3) and a music generation model (e.g., MusicVAE). The generated lyrics and melodies are saved in text files (.txt) and music files (.mp3 or .wav), respectively.

[1388] The server then encodes the generated lyrics and melody data into BASE64 format and sends it as an HTTP response to the user's device, where it decodes the received BASE64 data and restores it to the original text and music files.

[1389] Finally, the user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[1390] As a specific example, if a user inputs the theme "A song to cheer me up when I'm a little tired," and expresses in voice input, "I've been working all day today and I'm a little tired. But tomorrow is a day off, so I have to do my best," then based on the analysis results, a bright, uplifting song with a rhythmic melody will be generated.

[1391] An example of a prompt might be:

[1392] "A refreshing and uplifting summer song"

[1393] "I feel great today!"

[1394] This allows users to easily create and enjoy original music that perfectly matches their emotions.

[1395] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1396] Step 1:

[1397] Using a device such as a smartphone or smart glasses, the user inputs the theme and mood of the song in text, and simultaneously inputs voice to express their current emotions.

[1398] Input: Text (e.g., "I'm a little tired, so I need a song to cheer me up"), audio data (e.g., "I've been working all day today and I'm a little tired. But tomorrow is a holiday, so I have to do my best.")

[1399] Output: User-entered text and audio data of the theme and mood

[1400] Step 2:

[1401] A voice analysis means of the terminal records the user's voice input and analyzes the voice data to extract emotion data.

[1402] Input: User's voice data

[1403] Data processing: Analysis of voice data and emotion recognition (e.g., emotion classification such as happiness, sadness, fatigue, etc.)

[1404] Output: Parsed emotion data (e.g., "tired")

[1405] Step 3:

[1406] The device converts the theme and atmosphere data entered by the user and the analyzed emotional data into JSON format and sends it to the server.

[1407] Input: Theme and mood of text input, analyzed sentiment data

[1408] Data processing: Conversion to JSON format

[1409] Output: Send data in JSON format (e.g., {"theme": "Songs to cheer you up when you're tired", "emotion": "Tired"})

[1410] Step 4:

[1411] The server extracts theme, mood, and emotion data from the received JSON data, and then invokes a generative AI model to generate lyrics and a melody based on the specified theme, mood, and emotion.

[1412] Input: JSON format data (theme, mood, emotion data)

[1413] Data processing: Generate lyrics using natural language generation models (e.g., GPT-3) and melodies using music generation models (e.g., MusicVAE).

[1414] Output: Generated lyrics (text format) and melody (music file format)

[1415] Step 5:

[1416] The server encodes the generated lyrics and melody data into BASE64 format and sends it to the user's terminal as an HTTP response.

[1417] Input: Generated lyrics and melody data

[1418] Data processing: Encoding to BASE64 format and creating HTTP responses

[1419] Output: BASE64 encoded lyrics and melody data (sent as HTTP response)

[1420] Step 6:

[1421] The terminal decodes the received BASE64 encoded data and restores it to the original text and music files.

[1422] Input: BASE64 encoded lyrics and melody data

[1423] Data processing: Decoding BASE64 data

[1424] Output: Original text file (lyrics) and music file (melody)

[1425] Step 7:

[1426] The user can visually check the generated lyrics through the display means of the terminal, and can listen to the generated melody as sound using the playback means.

[1427] Input: Decoded lyrics text file and music file

[1428] Specific behavior: Display a text file and play a music file

[1429] Output: User sees lyrics visually and hears melody (e.g., lyrics displayed on device screen and melody played through speaker)

[1430] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1431] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1432] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1433] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1434] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1435] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1436] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1437] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1438] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1439] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1440] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1441] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1442] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1443] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1444] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1445] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1446] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1447] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1448] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1449] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1450] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1451] The following is further disclosed regarding the above embodiment.

[1452] (Claim 1)

[1453] an input means for a user to specify a theme or atmosphere;

[1454] a transmission means for transmitting data on the theme or atmosphere designated through the input means to a server;

[1455] a generating means for receiving the theme and atmosphere data and generating lyrics and a melody using a generative AI model;

[1456] a transmitting means for transmitting the generated lyrics and melody data to a terminal;

[1457] a display means for displaying or reproducing the lyrics and melody data;

[1458] A system including:

[1459] (Claim 2)

[1460] 2. The system of claim 1, wherein the generation means includes a natural language generation model and a music generation model.

[1461] (Claim 3)

[1462] 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and said playback means plays back said melody as a music file.

[1463] "Example 1"

[1464] (Claim 1)

[1465] an input means for a user to specify a theme or atmosphere;

[1466] a transmission means for transmitting data on the theme or atmosphere designated through the input means to a server;

[1467] a generating means for receiving the theme and atmosphere data and generating lyrics and a melody using a generative AI model;

[1468] a transmitting means for transmitting the generated lyrics and melody data to a terminal;

[1469] display means for displaying or reproducing the lyrics and melody data;

[1470] A processing means for converting the generated lyrics and melody into an appropriate format and storing the converted lyrics and melody;

[1471] analytical means to analyze thematic and atmospheric data;

[1472] A system including:

[1473] (Claim 2)

[1474] 2. The system of claim 1, wherein the generation means includes a natural language generation model and a music generation model.

[1475] (Claim 3)

[1476] 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and said playback means plays back said melody as a music file.

[1477] "Application Example 1"

[1478] (Claim 1)

[1479] an input means for a user to specify a theme or atmosphere;

[1480] a transmission means for transmitting data on the theme or atmosphere designated through the input means to a server;

[1481] a generating means for receiving the theme and atmosphere data and generating lyrics and a melody using a generative AI model;

[1482] a transmitting means for transmitting the generated lyrics and melody data to a terminal;

[1483] display means for displaying or reproducing the lyrics and melody data;

[1484] The system includes a sharing means for using the generated lyrics and melody and sharing them with other users in real time.

[1485] (Claim 2)

[1486] 2. The system of claim 1, wherein the generation means includes a natural language generation model and a music generation model.

[1487] (Claim 3)

[1488] 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and said playback means plays back said melody as a music file.

[1489] "Example 2: Combining Emotion Engines"

[1490] (Claim 1)

[1491] an input means for a user to specify a theme or atmosphere;

[1492] a transmission means for transmitting data on the theme or atmosphere designated through the input means to a server;

[1493] emotion analysis means for receiving the theme and mood data and analyzing the user's voice input data to identify emotions;

[1494] a generating means for generating lyrics and a melody using a generative AI model based on the analyzed emotional state and the data on the theme and atmosphere;

[1495] a transmitting means for transmitting the generated lyrics and melody data to a terminal;

[1496] a display means for displaying or reproducing the lyrics and melody data;

[1497] A system including:

[1498] (Claim 2)

[1499] 2. The system of claim 1, wherein the generation means includes a natural language generation model and a music generation model.

[1500] (Claim 3)

[1501] 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and said playback means plays back said melody as a music file.

[1502] "Application example 2 when combining emotion engines"

[1503] (Claim 1)

[1504] an input means for a user to specify a theme or atmosphere;

[1505] speech analysis means for recording and analyzing a user's speech input;

[1506] a transmission means for transmitting data on the theme and atmosphere designated through the input means and the voice analysis means, as well as data on the analyzed emotions, to a server;

[1507] generating means for receiving the theme, mood, and emotion data and generating lyrics and a melody using a generative AI model;

[1508] a transmitting means for transmitting the generated lyrics and melody data to a terminal;

[1509] a display means for displaying or reproducing the lyrics and melody data;

[1510] A system including:

[1511] (Claim 2)

[1512] 2. The system of claim 1, wherein the generation means includes a natural language generation model and a music generation model.

[1513] (Claim 3)

[1514] 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and said playback means plays back said melody as a music file. [Explanation of symbols]

[1515] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. an input means for a user to specify a theme or atmosphere; a transmission means for transmitting data on the theme or atmosphere designated through the input means to a server; a generating means for receiving the theme and atmosphere data and generating lyrics and a melody using a generative AI model; a transmitting means for transmitting the generated lyrics and melody data to a terminal; a display means for displaying or reproducing the lyrics and melody data; A system including:

2. 2. The system of claim 1, wherein the generating means includes a natural language generation model and a music generation model.

3. 2. The system according to claim 1, wherein said display means displays said lyrics as a text file and includes playback means for playing said melody as a music file.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A