system

The system uses generative AI to generate and arrange songs based on user emotions and purposes, allowing for easy customization and refinement, addressing the limitations of existing systems.

JP2026063707APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing songwriting and composition systems struggle to create individual songs based on user emotions and purposes, lack customization features, and require professional knowledge for feedback and refinement.

Method used

A system that utilizes generative artificial intelligence to generate lyrics and melodies based on user input, allows for user feedback, and automatically arranges the music to match emotions and purposes.

Benefits of technology

Enables users to easily create personalized songs that align with their feelings and intended use, facilitating multiple iterations of refinement and arrangement without specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063707000001_ABST
    Figure 2026063707000001_ABST
Patent Text Reader

Abstract

This system utilizes generative artificial intelligence to efficiently generate personalized music and provides a system that flexibly modifies and arranges the music based on user feedback. [Solution] A system comprising: means for receiving emotional and purpose text data entered by a user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In existing songwriting and composition systems, it has been difficult to create individual songs based on the user's own emotions and purposes. In addition, there has been a lack of a function to customize songs according to the mood of the day and control the user's mood. Furthermore, there has been a problem that without professional knowledge, feedback and corrections for refining works cannot be effectively performed. The present invention aims to solve these problems and provide a system that efficiently generates individually specialized songs using generative artificial intelligence and flexibly modifies and arranges them based on user feedback.

Means for Solving the Problems

[0005] The present invention provides a system that includes means for receiving emotional and purpose text data entered by a user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyric and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user. This system allows users to easily create songs that match their emotions and purposes, and to revise and arrange them as many times as needed.

[0006] A "user" is an individual or group that intends to create music using this system.

[0007] "Emotions" refer to the subjective state a user is experiencing, encompassing a variety of states such as "happy," "sad," or "excited."

[0008] "Purpose" refers to the specific scenes or uses in which a user will use the music, including, for example, "a birthday celebration song" or "a cheering song."

[0009] "Text data" refers to character information entered by the user to express emotions or intentions.

[0010] "Preprocessing" refers to data processing methods used to analyze text data and extract necessary information.

[0011] "Natural language processing technology" refers to techniques for analyzing text data and extracting keywords and important phrases, and typically includes machine learning and data mining techniques.

[0012] "Generative artificial intelligence" is an artificial intelligence technology that generates creative content based on given input data.

[0013] "Lyrics" refers to the lyrical portion of a song, a combination of words designed to convey a message that aligns with the user's emotions and purpose.

[0014] "Melody" refers to the musical, melody-driven part of a song, expressing emotions and atmosphere through sound.

[0015] "Feedback" refers to the opinions and requests that users give regarding the generated lyrics and melodies.

[0016] "Revision" refers to changing the content of lyrics or melody based on user feedback.

[0017] "Arrangement" is the process of adding additional musical elements, such as instrument selection and arrangement, to the generated lyrics and melody in order to improve the overall quality of the song.

[0018] "Song data" refers to a music file that includes the completed lyrics, melody, and arrangement.

[0019] "System" refers to the entire configuration, including devices and software, that perform a series of processes, such as generating music based on the emotions and purposes input by the user, receiving feedback and making corrections, and finally providing the arranged music data. [Brief explanation of the drawing]

[0020] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. <00,00100>It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0022] First, the language used in the following description will be explained.

[0023] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0028] [First Embodiment]

[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0041] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on user input data, modifies them according to user feedback, and ultimately completes and delivers the song.

[0042] A specific embodiment of this system will be described based on the roles of the user, terminal, and server.

[0043] User roles

[0044] User

[0045] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0046] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0047] 2. Review the initial lyrics and melody provided and enter your feedback.

[0048] Example: Enter specific feedback such as, "Make this part feel more lively."

[0049] 3. Review the revised lyrics and melody again and provide further feedback if necessary.

[0050] 4. Download the completed song and use it for performances or recordings.

[0051] Terminal role

[0052] terminal

[0053] 1. Preprocessing

[0054] The system receives text data (emotion and purpose) from users and analyzes it using natural language processing technology.

[0055] This analysis extracts keywords and important phrases (e.g., "fun" = positive, "birthday celebration song" = celebration).

[0056] 2. Data transmission

[0057] The extracted keywords and phrases are sent to the server.

[0058] 3. Interface

[0059] The system displays the server's response to the user and collects and sends feedback.

[0060] 4. Data Management

[0061] It provides saving, playback, and download functions to deliver completed music data to users.

[0062] Server Role

[0063] server

[0064] 1. Idea generation

[0065] Based on the analysis data sent from the terminal, a generative AI is invoked to generate initial lyrics and melody ideas.

[0066] Example: Lyrics like "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0067] 2. Feedback Processing

[0068] We receive user feedback and use generative AI to revise the lyrics and melody again.

[0069] Example: After receiving feedback such as "Make it sound more lively," the lyrics and melody are made more vibrant.

[0070] 3. Arrangement

[0071] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[0072] 4. Data transmission

[0073] The final completed song data is sent to the device.

[0074] Specific example

[0075] For example, if a user inputs the emotion "happy" and the objective "birthday song," the system will work as follows:

[0076] 1. The terminal parses the input text and sends it to the server.

[0077] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0078] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[0079] 4. The server uses a generative AI again to revise the lyrics to a lively "Today is your special day, everyone gathers to sing a celebratory song..." and a lively melody.

[0080] 5. The server selects instruments, arranges the music, and completes the final song.

[0081] 6. The device provides the user with the completed song, which the user then uses to perform or record.

[0082] Thus, this system handles everything from generating, modifying, arranging, and delivering music tailored to the user's emotions and purpose, making it easy for users to create personalized music.

[0083] The following describes the processing flow.

[0084] Step 1:

[0085] User

[0086] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[0087] Example: Emotion = "Happiness", Purpose = "Birthday song"

[0088] Step 2:

[0089] terminal

[0090] Receive text data entered by the user.

[0091] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[0092] Example: Emotion = "Positive", Purpose = "Celebration"

[0093] Step 3:

[0094] terminal

[0095] The extracted keywords and important phrases are sent to the server.

[0096] Step 4:

[0097] server

[0098] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[0099] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0100] Step 5:

[0101] server

[0102] The generated initial lyrics and melody are sent to the device.

[0103] Step 6:

[0104] terminal

[0105] The initial lyrics and melody received from the server are presented to the user.

[0106] Step 7:

[0107] User

[0108] Review the provided lyrics and melody, and provide feedback as needed.

[0109] Example: Enter specific requests such as, "Make this part feel more lively."

[0110] Step 8:

[0111] terminal

[0112] Receive user feedback and send it to the server.

[0113] Step 9:

[0114] server

[0115] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[0116] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0117] Step 10:

[0118] server

[0119] The revised lyrics and melody are sent to the device and presented to the user again.

[0120] Step 11:

[0121] User

[0122] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[0123] Step 12:

[0124] server

[0125] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[0126] Step 13:

[0127] server

[0128] Send the completed song data to the device.

[0129] Step 14:

[0130] terminal

[0131] The completed music data is provided to the user, allowing them to download, save, play, record, and perform other operations.

[0132] This processing flow allows users to easily create specific songs tailored to their emotions and purposes, and to revise and arrange them as many times as needed based on feedback.

[0133] (Example 1)

[0134] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0135] Existing music generation systems struggle to effectively create music that aligns with the emotions and purposes input by the user. Furthermore, methods for incorporating user feedback on generated lyrics and melodies are insufficient, often resulting in final music that fails to meet user expectations. Additionally, arranging and refining music requires specialized knowledge and skills, making it difficult for users to easily generate high-quality music.

[0136] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0137] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for automatically selecting appropriate instruments and arranging the revised lyrics and melody; and means for providing the arranged song data to the user. This makes it possible to easily generate high-quality songs that match the emotions and purposes entered by the user and to revise and arrange them according to the user's feedback.

[0138] A "user" is an individual or group that wishes to use the system to generate music that matches their emotions and purpose.

[0139] "Emotions" refer to information used to express the psychological state a user is experiencing. For example, it can refer to states such as being happy, sad, or angry.

[0140] "Purpose" refers to information that clarifies the intention or use of the song. For example, it refers to a specific intended use, such as a birthday song or a wedding theme song.

[0141] "Text data" refers to string information about emotions and purposes entered by the user.

[0142] "Natural language processing technology" is a technique for analyzing text data and extracting keywords and important phrases.

[0143] A "keyword" is a particularly important word or phrase within text data. For example, "fun" or "birthday" would be considered keywords.

[0144] An "important phrase" is a part of the text data that has particular meaning in context.

[0145] "Generative artificial intelligence" refers to machine learning models or algorithms that generate new content based on input data.

[0146] "Initial lyrics and melody ideas" refers to the initial draft of lyrics and melody for a song created by a generative artificial intelligence.

[0147] "Feedback" refers to the revision requests and evaluations that users provide regarding the generated lyrics and melodies.

[0148] "Arrangement" refers to the process of selecting instruments and arranging music based on the generated lyrics and melody, thereby improving the overall quality of the song.

[0149] "Appropriate instrument selection and arrangement" refers to the process of choosing instruments that suit the atmosphere and style of the music, and determining their placement and playing methods.

[0150] "Song data" refers to the digital data of the final completed music content, including lyrics, melody, and arrangement.

[0151] A "system" is a general term for a set of devices and software that have a series of functions for generating music based on user input, making corrections and arrangements in response to user feedback, and finally providing the completed music.

[0152] Modes for carrying out the invention

[0153] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on user input data, modifies them according to user feedback, and ultimately completes and delivers the finished song. Specific embodiments of this invention will be described based on the roles of the user, terminal, and server.

[0154] User roles

[0155] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0156] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0157] 2. The user reviews the initial lyrics and melody provided and enters feedback.

[0158] Example: Enter specific feedback such as, "Make this part feel more lively."

[0159] 3. Users will review the revised lyrics and melody again and provide further feedback if necessary.

[0160] 4. Users download the completed song and use it for performances and recordings.

[0161] Terminal role

[0162] 1. The terminal receives text data (sentiment and purpose) entered by the user and analyzes it using natural language processing (NLP) techniques. This analysis uses NLP libraries such as Python's NLTK to extract important keywords and phrases.

[0163] Specific examples: "fun" = positive, "birthday celebration song" = celebratory phrases are extracted.

[0164] 2. The terminal sends the extracted keywords and phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[0165] 3. The terminal displays the response from the server to the user, collects and sends feedback. It displays the initial lyrics and melody to the user and collects new feedback from the user again.

[0166] 4. The terminal provides saving, playback, and download functions for delivering completed music data to the user. It saves the music to the file system, provides a download link, and displays an interface with playback functionality.

[0167] Server Role

[0168] 1. The server calls a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. This generation process uses APIs such as OpenAI®.

[0169] Example prompt: Generate lyrics and melody based on "happy feelings and birthday celebration songs".

[0170] 2. The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. It inputs a prompt into the generative AI and generates content based on the user's request.

[0171] 3. The server arranges the song based on the revised lyrics and melody. Specifically, it automatically selects appropriate instruments and arranges the music. Using an AI arrangement tool, it automatically selects instruments such as guitar, piano, and drums, and arranges them appropriately.

[0172] 4. The server sends the final completed music data to the terminal. It returns the completed music file to the terminal as an HTTP response.

[0173] As described above, this invention allows for the entire process from generating, modifying, arranging, and providing music tailored to the user's emotions and purpose, making it easy for users to create specialized music.

[0174] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0175] Step 1:

[0176] Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface. The entered text data includes "emotions" and "purpose."

[0177] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0178] Input: "Happy feelings", "Birthday song"

[0179] Output: Text data (emotion, purpose)

[0180] Step 2:

[0181] The terminal receives text data entered by the user and analyzes it using natural language processing (NLP) techniques. Specifically, it uses the Python NLTK library to extract keywords and important phrases from the text data.

[0182] Specific examples: "fun" → positive, "birthday song" → celebration

[0183] Input: Text data (emotion, purpose)

[0184] Output: Keywords and important phrases

[0185] Step 3:

[0186] The terminal sends extracted keywords and important phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[0187] Input: Keywords and important phrases

[0188] Output: Parsed data in JSON format (keywords, phrases)

[0189] Step 4:

[0190] The server invokes a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. The OpenAI API is used for this generation.

[0191] Example: Based on the input data, generate lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0192] Example prompt: "Happy feelings and birthday songs"

[0193] Input: Analysis data (keywords, phrases)

[0194] Output: Initial lyrics and melody

[0195] Step 5:

[0196] The terminal displays the initial lyrics and melody received from the server to the user and requests feedback. The user enters and submits their feedback.

[0197] Example: The lyrics "Today is your special day, a birthday filled with smiles..." are displayed and the melody plays. The user then provides feedback saying, "Make this part sound more lively."

[0198] Input: Early lyrics and melody

[0199] Output: User feedback

[0200] Step 6:

[0201] The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. Based on the feedback, it adjusts the prompt text and re-inputs it into the generative AI.

[0202] Specific example: Based on feedback such as "change it to sound more lively," the lyrics "Today is your special day, everyone gathers to sing a celebratory song..." are revised to a more lively melody.

[0203] Input: User feedback

[0204] Output: Revised lyrics and melody

[0205] Step 7:

[0206] The server arranges the song based on the revised lyrics and melody. Specifically, it uses an AI arrangement tool to automatically select instruments and perform the arrangement.

[0207] Specific example: Automatically select instruments such as guitar, piano, and drums, and arrange them appropriately.

[0208] Input: Modified lyrics and melody

[0209] Output: Arranged song

[0210] Step 8:

[0211] The server then sends the final completed music data to the terminal. HTTP responses are used for data transmission, returning the completed music file to the terminal.

[0212] Input: Arranged song

[0213] Output: Music data (final version)

[0214] Step 9:

[0215] The device provides saving, playback, and download functions for delivering completed music data to the user. Specifically, it saves the music to the file system, provides a download link, and displays an interface with playback capabilities.

[0216] Specific example: Click the download link to save the song and play it with your preferred media player.

[0217] Input: Music data (final version)

[0218] Output: Music files accessible to the user

[0219] (Application Example 1)

[0220] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0221] When users generate music according to their own emotions and purposes, it is difficult for users without specialized knowledge or skills to easily customize and create individually tailored music. Furthermore, providing an environment where users can stream or download the generated music in real time is not easy. To solve these problems, a consistent service is needed for generating, modifying, arranging, and delivering music tailored to the user's emotions and purposes.

[0222] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0223] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; means for providing the arranged song data to the user; and means for implementing the song generation system as a smartphone application and making the generated song available for streaming or download. This allows users to easily generate songs that match their emotions and purposes and use them in real time.

[0224] "Emotions" refer to the user's own internal feelings and moods.

[0225] "Purpose" refers to the user's intended use in a specific situation or event.

[0226] "Text data" refers to string information entered by the user.

[0227] "Natural language processing technology" refers to the technology that enables computers to understand and analyze human language.

[0228] "Keywords" refer to important words or phrases extracted from text data.

[0229] An "important phrase" refers to a meaningful sequence of words extracted from text data.

[0230] "Generative artificial intelligence" refers to artificial intelligence that has the ability to generate new information based on given data.

[0231] "Lyrics" refers to the spoken parts of a song.

[0232] "Melody" refers to a sequence of sounds in a musical piece.

[0233] An "idea" refers to an initial concept or idea.

[0234] "Feedback" refers to reactions and opinions from users.

[0235] "Modification" refers to improving initial ideas or data.

[0236] "Arrangement" refers to the selection of instruments and arrangement of music for revised lyrics and melody.

[0237] "Song data" refers to the information of a completed song.

[0238] A "smartphone application" refers to a software program that runs on a smartphone.

[0239] "Streaming distribution" refers to a method of transmitting and playing data in real time.

[0240] "Downloading" refers to saving data to a local device via a network.

[0241] One embodiment of this invention is a system that allows users to generate music according to their own emotions and purposes, and then stream or download it in real time via a smartphone application.

[0242] System Program Overview

[0243] 1. Receiving and analyzing user input

[0244] Users input their emotions and goals as text data using a smartphone application. Keywords and important phrases are extracted from the text data using natural language processing (NLP) techniques. SpaCy and NLTK are suitable NLP libraries to use.

[0245] 2. The process of creating music

[0246] The analyzed keywords and important phrases are sent to the server. The server uses generative artificial intelligence (e.g., OpenAI's GPT-4®) to generate initial lyrics and melodies based on this information. This initial generation process uses prompts such as the following:

[0247] "Please generate lyrics and melody for a fun birthday celebration song."

[0248] 3. Feedback and Corrections

[0249] The smartphone application presents the user with the initial generated lyrics and melody. The user reviews them and provides feedback. The server receives the feedback and uses generative artificial intelligence again to revise the lyrics and melody. The following prompts are used during the revision process.

[0250] "Please make this part more lively. Please revise it."

[0251] 4. Final arrangement and serving

[0252] Based on the revised lyrics and melody, the song is arranged. Automatic arrangement technology is used to select instruments and arrange the music, generating the final song data. The generated song data is provided to the user via a smartphone application. The user can stream or download this song in real time.

[0253] Implementation explanation

[0254] Hardware and software

[0255] Smartphone application: Used for user input, feedback collection, and presentation of generated music. Flutter® and React Native are suitable development frameworks.

[0256] Server: Used for data processing and execution of generative artificial intelligence. AWS® and Google® Cloud are suitable cloud platforms.

[0257] Generative artificial intelligence: Generates and modifies lyrics and melodies based on given data. OpenAI's GPT-4 is a suitable AI model to use.

[0258] Natural Language Processing (NLP) techniques: Used for analyzing user input. Suitable NLP libraries include spaCy and NLTK.

[0259] Adding specific examples

[0260] 1. Analysis of user input

[0261] The user inputs the emotion of "fun" and the objective of "birthday celebration songs." The application analyzes this input and extracts keywords and important phrases such as "fun" = positive and "birthday celebration songs" = celebration.

[0262] 2. Initial Generation and Presentation

[0263] Based on the extracted keywords, the following prompts are entered into the generative artificial intelligence.

[0264] "Please generate lyrics and melody for a fun birthday celebration song."

[0265] The initial generated lyrics and melody are presented to the user through the application.

[0266] 3. Feedback and Corrections

[0267] The user provides feedback saying, "Make this part feel more lively." The server then prompts the generative AI again with the following message.

[0268] "Please make this part more lively. Please revise it."

[0269] The revised lyrics and melody are presented to the user again.

[0270] 4. Final arrangement and serving

[0271] Based on the final feedback, the song is rearranged and the completed track is generated. The generated track can be streamed or downloaded in real time.

[0272] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0273] Step 1:

[0274] Receiving input from the user

[0275] The user uses a smartphone application to input their emotion (e.g., "happy") and purpose (e.g., "birthday celebration song") in text form. This input data is sent to the application.

[0276] Input: Text data of emotion and purpose

[0277] Output: The text data is saved in the application.

[0278] Step 2:

[0279] Analysis of text data

[0280] The terminal analyzes the received text data using natural language processing technology. Here, NLP libraries (e.g., spaCy, NLTK) are used to extract keywords and important phrases from the text data.

[0281] Input: User input text data

[0282] Data processing: Extract keywords and important phrases using an NLP library

[0283] [[ID=4|5]]Output: Extracted keywords and important phrases

[0284] Step 3: [[ID=|51]]

[0285] Sending data to the server

[0286] The terminal sends the keywords and important phrases, which are the analysis results, to the server.

[0287] Input: Extracted keywords and important phrases

[0288] Data operation: Packaging and transmission of analysis results

[0289] Output: Keywords and important phrases are sent to the server

[0290] Step 4:

[0291] Initial lyric and melody generation

[0292] The server uses a generative artificial intelligence (e.g., GPT-4 of OpenAI) to generate initial lyrics and melodies based on the received keywords and important phrases. The AI is instructed using a prompt text.

[0293] Input: Keywords and important phrases

[0294] Data processing: Input a prompt text into the generative AI model (e.g., "Please generate the lyrics and melody of a happy birthday song")

[0295] Output: Generated initial lyrics and melodies

[0296] Step 5:

[0297] Presentation of the initial generation result and receipt of feedback

[0298] The terminal presents the generated initial lyrics and melodies to the user and collects the user's feedback. The user inputs specific requests and points for improvement in text form.

[0299] Input: Generated initial lyrics and melodies

[0300] Data operation: Display to the user and receipt of feedback

[0301] Output: User feedback

[0302] Step 6:

[0303] Corrections based on feedback

[0304] Based on the feedback obtained from the user, the server uses the generative artificial intelligence again to correct the lyrics and melody. Prompts are used for the corrections.

[0305] Input: User feedback

[0306] Data processing: Input prompts into the generative AI model (e.g., "Make this part more lively. Please correct it.")

[0307] Output: Corrected lyrics and melody

[0308] Step 7:

[0309] Execution of the final arrangement

[0310] Based on the corrected lyrics and melody, the server uses automatic music arrangement technology to perform the final arrangement of the music. It automatically selects instruments and arranges the music.

[0311] Input: Corrected lyrics and melody

[0312] Data calculation: Arrangement by automatic music arrangement technology[[ID=A5]]

[0313] Output: Arranged music data

[0314] Step 8:

[0315] Provision of music data

[0316] The terminal provides the completed music data to the user. The user can stream or download this music in real time.

[0317] Input: Arranged song data

[0318] Data processing: Determining how to deliver music data (streaming or download).

[0319] Output: Music data provided to the user

[0320] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0321] This invention is a system that can generate more accurate songs by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and purpose. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[0322] The introduction of an emotion engine will enable the recognition of emotions from the user's facial expressions and voice data, allowing for the provision of more personalized and tailored music.

[0323] User roles

[0324] User

[0325] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0326] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0327] 2. You input facial expressions and voice into the system, and the emotion engine analyzes them.

[0328] 3. Review the initial lyrics and melody provided and enter your feedback.

[0329] Example: Enter specific feedback such as, "Make this part feel more lively."

[0330] 4. Review the revised lyrics and melody again and provide further feedback if necessary.

[0331] 5. Download the completed song and use it for performances or recordings.

[0332] Terminal role

[0333] terminal

[0334] 1. Preprocessing

[0335] The system receives text data, facial expressions, and voice data from the user and analyzes them using an emotion engine. This analysis allows for a more accurate recognition of the user's emotions.

[0336] This analysis extracts keywords and important phrases based on emotions and purposes.

[0337] Example: Emotion = "Positive", Purpose = "Celebration"

[0338] 2. Data transmission

[0339] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[0340] 3. Interface

[0341] The system displays the server's response to the user and collects and sends feedback.

[0342] 4. Data Management

[0343] It provides saving, playback, and download functions to deliver completed music data to users.

[0344] Server Role

[0345] server

[0346] 1. Idea generation

[0347] Based on the analysis data sent from the terminal, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[0348] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0349] 2. Feedback Processing

[0350] We receive user feedback and use generative artificial intelligence again to revise the lyrics and melody.

[0351] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0352] 3. Arrangement

[0353] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[0354] 4. Data transmission

[0355] The final completed song data is sent to the device.

[0356] Specific example

[0357] For example, if a user inputs the emotion "happy" and the objective "birthday song," and also inputs facial expressions and voice, the system will operate as follows:

[0358] 1. The device analyzes the input text, facial expressions, and voice data, and the emotion engine recognizes the emotion as "happy." Based on this, keywords and important phrases are extracted and sent to the server.

[0359] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0360] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[0361] 4. The server uses a generative AI again to revise the lyrics and melody to make them more lively.

[0362] 5. The server selects instruments, arranges the music, and completes the final song.

[0363] 6. The device provides the user with the completed music, allowing the user to download, save, play, record, and perform other operations.

[0364] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide personalized music with greater accuracy.

[0365] The following describes the processing flow.

[0366] Step 1:

[0367] User

[0368] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[0369] Example: Emotion = "Happiness", Purpose = "Birthday song"

[0370] Step 2:

[0371] User

[0372] The system receives input of facial expressions and voice. A dedicated camera and microphone are used to record facial expressions and voice tone.

[0373] Step 3:

[0374] terminal

[0375] It receives text data, facial expression data, and voice data entered by the user.

[0376] The system uses an emotion engine to analyze input facial expression and voice data to recognize the user's emotions.

[0377] Example: Analysis result = "fun"

[0378] Step 4:

[0379] terminal

[0380] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[0381] Example: Emotion = "Positive", Purpose = "Celebration"

[0382] Step 5:

[0383] terminal

[0384] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[0385] Step 6:

[0386] server

[0387] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyric and melody ideas.

[0388] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0389] Step 7:

[0390] server

[0391] The generated initial lyrics and melody are sent to the device.

[0392] Step 8:

[0393] terminal

[0394] The initial lyrics and melody received from the server are presented to the user.

[0395] Step 9:

[0396] User

[0397] Review the provided lyrics and melody, and provide feedback as needed.

[0398] Example: Enter specific requests such as, "Make this part feel more lively."

[0399] Step 10:

[0400] terminal

[0401] Receive user feedback and send it to the server.

[0402] Step 11:

[0403] server

[0404] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[0405] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0406] Step 12:

[0407] server

[0408] The revised lyrics and melody are sent to the device and presented to the user again.

[0409] Step 13:

[0410] User

[0411] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[0412] Step 14:

[0413] server

[0414] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[0415] Step 15:

[0416] server

[0417] Send the completed song data to the device.

[0418] Step 16:

[0419] terminal

[0420] The completed music data is provided to the user, allowing them to download, save, play, record, and perform other operations.

[0421] This processing flow allows users to highly customize specific songs to match their emotions and purposes through a system that combines an emotion engine and generative AI, enabling them to create more accurate and personalized songs.

[0422] (Example 2)

[0423] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0424] Conventional songwriting support systems struggled to provide sophisticated music that reflected specific emotions, even when generating songs based on the user's feelings and objectives. Furthermore, the process of revising songs while incorporating user feedback in real time was inefficient, and the automatic arrangement functions for improving song quality had limitations. As a result, there was a challenge in providing individually tailored songs.

[0425] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0426] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for using an emotion engine to analyze the text data and recognizing emotions from the user's facial expressions and voice data; means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyric and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user. This enables sophisticated song generation tailored to the user's emotions and purpose, and allows for the provision of individually tailored songs with greater accuracy.

[0427] "User" refers to a person or user who wishes to create music using this system.

[0428] "Emotions" refer to data that represents the inner feelings and intentions expressed by the user, and are primarily input through text, facial expressions, and voice data.

[0429] "Purpose" refers to data indicating the intended use or intent of the music the user wishes to create, and is entered in text format.

[0430] "Text data" refers to strings of information that users input into a system, and it includes emotions and intentions.

[0431] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate human language.

[0432] "Keywords" are important words or phrases extracted from text data and used in the creation of music.

[0433] A "key phrase" is a meaningful sequence of words extracted from text data and used in the creation of a musical piece.

[0434] An "emotion engine" refers to software or hardware that analyzes and recognizes emotions from facial expressions and voice data entered by the user.

[0435] "Generative artificial intelligence" refers to an artificial intelligence system that automatically generates content (in this case, lyrics and melody) based on given input data.

[0436] "Initial lyrics and melody" refers to the basic text and musical melody that form the basis of the song first generated by the generative artificial intelligence.

[0437] "Feedback" refers to the opinions and requests that users input into a system for corrections and improvements.

[0438] "Revised lyrics and melody" refers to the text and musical melody regenerated by the generative artificial intelligence based on user feedback.

[0439] "Arrangement" refers to the process of selecting instruments and arranging the music based on the generated lyrics and melody to complete the song.

[0440] "Song data" refers to the digital information of the final produced song, which is provided to the user.

[0441] This invention is a system that generates more accurate music by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[0442] Hardware and software configuration

[0443] User

[0444] Users interact with the system through a dedicated application or web interface. Devices such as computers, smartphones, and tablets are used.

[0445] terminal

[0446] The device will be equipped with an interface for receiving user input data (text, facial expressions, and voice).

[0447] The emotion engine used includes the Emotion API from Microsoft® Azure® Cognitive Services.

[0448] Natural language processing techniques are used to analyze the data and extract keywords and important phrases.

[0449] It will be equipped with network communication capabilities for sending analysis data to a server.

[0450] server

[0451] The server will be equipped with a function to generate lyrics and melodies using generative artificial intelligence (e.g., OpenAI's GPT-3®).

[0452] The server will be equipped with a function to receive feedback from users and then use generative artificial intelligence to revise the lyrics and melody again.

[0453] It has a function to select instruments and arrange the music based on the revised lyrics and melody, and to generate the final song data.

[0454] Program Processing Overview

[0455] 1. Input of user's emotions and purpose

[0456] Users enter their emotions and the purpose of the song in text format via a dedicated app or web interface.

[0457] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0458] 2. Input and analysis of facial expression and voice data

[0459] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[0460] The device recognizes emotions from the analysis results and extracts emotional data, such as "happy."

[0461] 3. Generation of initial lyrics and melody

[0462] The device extracts important keywords and phrases based on the extracted sentiment data and the user's input, and sends them to the server.

[0463] The server uses generative artificial intelligence to generate initial lyrics and melody.

[0464] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line.

[0465] 4. Collecting user feedback

[0466] The device displays the initial generated lyrics and melody, and the user reviews the displayed content and provides feedback.

[0467] Example: Type "Make this part sound more lively."

[0468] 5. Revise the lyrics and melody.

[0469] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody based on that feedback.

[0470] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", with a lively melody line.

[0471] 6. Arrangement and final generation of the music

[0472] Based on the revised lyrics and melody, the server selects instruments and arranges the music to improve its overall quality.

[0473] 7. Provision of the completed song

[0474] The server finally sends the completed song data to the terminal, which saves the completed song data and provides it to the user. The user can download, play, and record the song.

[0475] Example of a prompt

[0476] User: Fun

[0477] Purpose: Birthday celebration song

[0478] Feedback: "Make this part feel more lively."

[0479] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide personalized music with greater accuracy.

[0480] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0481] Step 1:

[0482] Users input emotions and the purpose of the song in text format through a dedicated application or web interface.

[0483] Input: The emotion and purpose of the song entered by the user (e.g., happy, birthday song)

[0484] Output: Input text data

[0485] Specific action: The user enters "fun" and "birthday song" into the application's input form and presses the submit button.

[0486] Step 2:

[0487] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[0488] Input: User facial expression data and voice data

[0489] Output: Analyzed emotion data (e.g., happy)

[0490] Specific operation: The user uses the camera and microphone to input a smiling expression and the voice message "Happy Birthday." The emotion engine analyzes the expression and voice and recognizes it as "happy."

[0491] Step 3:

[0492] The device extracts keywords and important phrases based on text data and analyzed sentiment data, and sends them to the server.

[0493] Input: Emotional data, objective data, user facial expression and voice analysis results

[0494] Output: Extracted keywords and important phrases (e.g., special day, birthday, smile)

[0495] Specific operation: The terminal analyzes the input "fun" and "birthday song," as well as the "fun" from the emotion engine, extracts keywords such as "special day," "birthday," and "smile," and sends them to the server.

[0496] Step 4:

[0497] The server uses generative artificial intelligence to generate initial lyrics and melodies based on keywords and important phrases.

[0498] Input: Extracted keywords and important phrases

[0499] Output: Initial lyrics and melody

[0500] Specific operation: The server inputs the received words "special day," "birthday," and "smile" as prompts into the generative artificial intelligence, and generates lyrics such as "Today is your special day, a birthday overflowing with smiles..." along with a cheerful melody.

[0501] Step 5:

[0502] The device displays the initial generated lyrics and melody and collects feedback from the user.

[0503] Input: Early lyrics and melody

[0504] Output: User feedback (e.g., "Make this part more lively")

[0505] Specific operation: The device displays the lyrics and melody "Today is your special day, a birthday filled with smiles..." to the user, and the user provides feedback such as "Make this part sound more lively."

[0506] Step 6:

[0507] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody again.

[0508] Input: User feedback, initial lyrics and melody

[0509] Output: Revised lyrics and melody

[0510] Specific operation: The device sends the feedback "Make this part sound more lively" to the server, and the server uses generative artificial intelligence to revise the lyrics to "Today is your special day, everyone gathers to sing a celebratory song..." and also changes the melody to a more lively one.

[0511] Step 7:

[0512] The server selects instruments and arranges the music based on the revised lyrics and melody, generating the final song data.

[0513] Input: Modified lyrics and melody

[0514] Output: Final music data

[0515] Specific operation: Based on the revised lyrics and melody, the server automatically selects instruments such as piano, guitar, and drums, arranges the instrumentation and creates the final song data.

[0516] Step 8:

[0517] The server ultimately sends the completed music data to the terminal, which then provides it to the user.

[0518] Input: Final song data

[0519] Output: Interface for downloading, playing, and recording music.

[0520] Specific operation: The server sends the final music data to the terminal, and the terminal displays a download link, play button, and recording function to provide the user with the music. The user can download, play, and record the music.

[0521] (Application Example 2)

[0522] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0523] Conventional music generation systems have been insufficient in generating music that aligns with users' emotions and purposes, making it difficult to provide music tailored to specific emotions or objectives. Furthermore, limited means of incorporating user feedback during the music generation process have resulted in low final music quality and user satisfaction. Additionally, the lack of emotion analysis utilizing user facial expressions and voice data has led to insufficient accuracy in emotion-based music generation.

[0524] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving emotional and purpose text data input by the user; means for analyzing the text data and image or audio data and recognizing emotions from the user's facial expressions and voice; means for analyzing using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for modifying the lyrics and melody again using the generative artificial intelligence based on the feedback; means for arranging the modified lyrics and melody into a song; and means for providing the arranged song data to the user. This enables highly accurate song generation that is tailored to the user's emotions and purpose.

[0525] "Emotional and purposeful text data" refers to text-based information entered by the user that represents the emotions and purpose of the song.

[0526] "Image or audio data" refers to input data that includes the user's facial expressions and voice.

[0527] "Means of recognizing emotions" refers to technologies that analyze image or audio data to identify the user's emotions.

[0528] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and intent.

[0529] "Keywords and important phrases" are the main words and expressions necessary for song generation, extracted from the analyzed text data.

[0530] "Generative artificial intelligence" is an artificial intelligence technology that automatically generates lyrics and melodies based on input data.

[0531] "Initial lyrics and melody ideas" refers to the initial proposals for lyrics and melodies generated by a generative artificial intelligence.

[0532] "Feedback" refers to opinions and requests for corrections and improvements provided by users.

[0533] "Revised lyrics and melody" refers to lyrics and melodies regenerated by a generative artificial intelligence system based on user feedback.

[0534] "Arranging a song" refers to the process of selecting instruments and arranging the music based on the revised lyrics and melody to complete the song.

[0535] "Arranged song data" refers to the digital data of the completed song.

[0536] This invention is a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. By incorporating an emotion engine, it can generate highly accurate and personalized songs. Specifically, it uses devices such as smartphones, smart glasses, and head-mounted displays to input the user's emotional data, and generates songs using a generative artificial intelligence system built on a server.

[0537] The server includes means for receiving emotional and target text data entered by the user, means for analyzing the user's image or voice data to recognize emotions from facial expressions and voice, and means for analyzing text data using natural language processing technology to extract keywords and important phrases. It also includes means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence, means for presenting the initial lyric and melody ideas to the user and receiving feedback, means for revising the lyrics and melody using generative artificial intelligence again based on the feedback, means for arranging the revised lyrics and melody into a song, and means for providing the arranged song data to the user.

[0538] Users input their emotions and the purpose of the song in text format using a dedicated application or web interface. They also input facial expressions and voice, which the emotion engine analyzes. For example, if a user inputs the emotion "happy" and the purpose "birthday song," and also inputs facial expressions and voice into the system, the emotion engine will recognize the emotion as "happy" and extract keywords and important phrases.

[0539] The server uses generative artificial intelligence to generate initial lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody, which are then presented to the user via the terminal. If the user provides feedback such as "Make this part more lively," the generative AI can revise the lyrics and melody again to make it more vibrant. Finally, the completed song undergoes instrument selection and arrangement, and the final song data is provided to the user.

[0540] As a concrete example, if a user logs into a virtual store, selects an emotion (e.g., "fun") and a purpose (e.g., "relaxing music"), and inputs facial expressions and voice into the app, the following prompts will be generated based on the emotion analysis results.

[0541] Example of a prompt:

[0542] Please create lyrics and a melody that express the emotion of "joy" and aims to be "relaxing music."

[0543] By sending this prompt to a generative artificial intelligence system, high-quality music tailored to the user's emotions and purpose can be automatically generated. Users can then play, purchase, and download this music within a virtual store.

[0544] As a result, it becomes possible to generate highly accurate music tailored to the user's emotions and purpose, and to provide personalized music.

[0545] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0546] Step 1:

[0547] Users input emotions and the purpose of the music in text format using a dedicated application or web interface. The entered text data of emotions and purpose is sent to the terminal and transferred to the server. The user's input (text data of emotions and purpose) serves as the starting point for analysis.

[0548] Step 2:

[0549] The device receives text data and image and audio data (facial expressions and voice) entered by the user and analyzes them using an emotion engine. Specifically, it analyzes the user's facial expressions and voice data to recognize emotions, and extracts keywords and important phrases based on this analysis. The input for this step is facial expressions and voice data, and the output is the emotion recognition result and extracted keywords.

[0550] Step 3:

[0551] The server receives the analysis results, extracted keywords, and important phrases sent from the terminal, and uses a generative artificial intelligence model to generate initial lyric and melody ideas based on them. This generation process involves creating prompt sentences using the generative AI model and generating lyrics and melodies based on them. The input is the analysis results and keywords, and the output is the initial lyric and melody drafts.

[0552] Step 4:

[0553] The server presents the user with initial lyrics and melody ideas via a terminal. The user reviews these and provides specific feedback, such as "Make this part sound more lively." In this step, the user's input is feedback, and the output is a revision instruction that reflects that feedback.

[0554] Step 5:

[0555] The server receives user feedback and uses a generative artificial intelligence model to revise the lyrics and melody again. The generative AI regenerates prompt sentences based on the feedback and produces the revised lyrics and melody. The input is the feedback, and the output is the revised lyrics and melody.

[0556] Step 6:

[0557] The server then arranges the revised lyrics and melody into a song. Specifically, it automatically selects instruments and arranges the music to improve its overall quality. The input for this step is the revised lyrics and melody, and the output is the final arranged song data.

[0558] Step 7:

[0559] The server sends the final arranged music data to the terminal and provides it to the user. The user can play, purchase, and download this music data. The input is the final arranged music data, and the output is the data provided to the user.

[0560] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0561] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0562] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0563] [Second Embodiment]

[0564] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0565] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0566] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0567] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0568] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0569] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0570] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0571] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0572] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0573] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0574] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0575] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0576] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on the user's input data, modifies them according to the user's feedback, and ultimately completes and delivers the song.

[0577] A specific embodiment of this system will be described based on the roles of the user, terminal, and server.

[0578] User roles

[0579] User

[0580] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0581] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0582] 2. Review the initial lyrics and melody provided and enter your feedback.

[0583] Example: Enter specific feedback such as, "Make this part feel more lively."

[0584] 3. Review the revised lyrics and melody again and provide further feedback if necessary.

[0585] 4. Download the completed song and use it for performances or recordings.

[0586] Terminal role

[0587] terminal

[0588] 1. Preprocessing

[0589] The system receives text data (emotion and purpose) from users and analyzes it using natural language processing technology.

[0590] This analysis extracts keywords and important phrases (e.g., "fun" = positive, "birthday celebration song" = celebration).

[0591] 2. Data transmission

[0592] The extracted keywords and phrases are sent to the server.

[0593] 3. Interface

[0594] The system displays the server's response to the user and collects and sends feedback.

[0595] 4. Data Management

[0596] It provides saving, playback, and download functions to deliver completed music data to users.

[0597] Server Role

[0598] server

[0599] 1. Idea generation

[0600] Based on the analysis data sent from the terminal, a generative AI is invoked to generate initial lyrics and melody ideas.

[0601] Example: Lyrics like "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0602] 2. Feedback Processing

[0603] We receive user feedback and use generative AI to revise the lyrics and melody again.

[0604] Example: After receiving feedback such as "Make it sound more lively," the lyrics and melody are made more vibrant.

[0605] 3. Arrangement

[0606] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[0607] 4. Data transmission

[0608] The final completed song data is sent to the device.

[0609] Specific example

[0610] For example, if a user inputs the emotion "happy" and the objective "birthday song," the system will work as follows:

[0611] 1. The terminal parses the input text and sends it to the server.

[0612] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0613] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[0614] 4. The server uses a generative AI again to revise the lyrics to a more lively "Today is your special day, everyone gathers to sing a celebratory song..." and to a more lively melody.

[0615] 5. The server selects instruments, arranges the music, and completes the final song.

[0616] 6. The device provides the user with the completed song, which the user then uses to perform or record.

[0617] Thus, this system handles everything from generating, modifying, arranging, and delivering music tailored to the user's emotions and purpose, making it easy for users to create personalized music.

[0618] The following describes the processing flow.

[0619] Step 1:

[0620] User

[0621] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[0622] Example: Emotion = "Happiness", Purpose = "Birthday song"

[0623] Step 2:

[0624] terminal

[0625] Receive text data entered by the user.

[0626] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[0627] Example: Emotion = "Positive", Purpose = "Celebration"

[0628] Step 3:

[0629] terminal

[0630] The extracted keywords and important phrases are sent to the server.

[0631] Step 4:

[0632] server

[0633] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[0634] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0635] Step 5:

[0636] server

[0637] Send the generated initial lyrics and melody to the device.

[0638] Step 6:

[0639] terminal

[0640] The initial lyrics and melody received from the server are presented to the user.

[0641] Step 7:

[0642] User

[0643] Review the provided lyrics and melody, and provide feedback as needed.

[0644] Example: Enter specific requests such as, "Make this part feel more lively."

[0645] Step 8:

[0646] terminal

[0647] Receive user feedback and send it to the server.

[0648] Step 9:

[0649] server

[0650] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[0651] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0652] Step 10:

[0653] server

[0654] The revised lyrics and melody are sent to the device and presented to the user again.

[0655] Step 11:

[0656] User

[0657] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[0658] Step 12:

[0659] server

[0660] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[0661] Step 13:

[0662] server

[0663] Send the completed song data to the device.

[0664] Step 14:

[0665] terminal

[0666] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[0667] This processing flow allows users to easily create specific songs tailored to their emotions and purposes, and to revise and arrange them as many times as needed based on feedback.

[0668] (Example 1)

[0669] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0670] Existing music generation systems struggle to effectively create music that aligns with the emotions and purposes input by the user. Furthermore, methods for incorporating user feedback on generated lyrics and melodies are insufficient, often resulting in final music that fails to meet user expectations. Additionally, arranging and refining music requires specialized knowledge and skills, making it difficult for users to easily generate high-quality music.

[0671] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0672] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for automatically selecting appropriate instruments and arranging the revised lyrics and melody; and means for providing the arranged song data to the user. This makes it possible to easily generate high-quality songs that match the emotions and purposes entered by the user and to revise and arrange them according to the user's feedback.

[0673] A "user" is an individual or group that wishes to use the system to generate music that matches their emotions and purpose.

[0674] "Emotions" refer to information used to express the psychological state a user is experiencing. For example, it can refer to states such as being happy, sad, or angry.

[0675] "Purpose" refers to information that clarifies the intention or use of the song. For example, it refers to a specific intended use, such as a birthday song or a wedding theme song.

[0676] "Text data" refers to string information about emotions and purposes entered by the user.

[0677] "Natural language processing technology" is a technique for analyzing text data and extracting keywords and important phrases.

[0678] A "keyword" is a particularly important word or phrase within text data. For example, "fun" or "birthday" would be considered keywords.

[0679] An "important phrase" is a part of the text data that has particular meaning in context.

[0680] "Generative artificial intelligence" refers to machine learning models or algorithms that generate new content based on input data.

[0681] "Initial lyrics and melody ideas" refers to the initial draft of lyrics and melody for a song created by a generative artificial intelligence.

[0682] "Feedback" refers to the revision requests and evaluations that users provide regarding the generated lyrics and melodies.

[0683] "Arrangement" refers to the process of selecting instruments and arranging music based on the generated lyrics and melody, thereby improving the overall quality of the song.

[0684] "Appropriate instrument selection and arrangement" refers to the process of choosing instruments that suit the atmosphere and style of the music, and determining their placement and playing methods.

[0685] "Song data" refers to the digital data of the final completed music content, including lyrics, melody, and arrangement.

[0686] A "system" is a general term for a set of devices and software that have a series of functions for generating music based on user input, making corrections and arrangements in response to user feedback, and finally providing the completed music.

[0687] Modes for carrying out the invention

[0688] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on user input data, modifies them according to user feedback, and ultimately completes and delivers the finished song. Specific embodiments of this invention will be described based on the roles of the user, terminal, and server.

[0689] User roles

[0690] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0691] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0692] 2. The user reviews the initial lyrics and melody provided and enters feedback.

[0693] Example: Enter specific feedback such as, "Make this part feel more lively."

[0694] 3. Users will review the revised lyrics and melody again and provide further feedback if necessary.

[0695] 4. Users download the completed song and use it for performances and recordings.

[0696] Terminal role

[0697] 1. The terminal receives text data (sentiment and purpose) entered by the user and analyzes it using natural language processing (NLP) techniques. This analysis uses NLP libraries such as Python's NLTK to extract important keywords and phrases.

[0698] Specific examples: "fun" = positive, "birthday celebration song" = celebratory phrases are extracted.

[0699] 2. The terminal sends the extracted keywords and phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[0700] 3. The terminal displays the response from the server to the user, collects and sends feedback. It displays the initial lyrics and melody to the user and collects new feedback from the user again.

[0701] 4. The terminal provides saving, playback, and download functions for delivering completed music data to the user. It saves the music to the file system, provides a download link, and displays an interface with playback functionality.

[0702] Server Role

[0703] 1. The server calls a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. OpenAI APIs and other tools are used for this generation process.

[0704] Example prompt: Generate lyrics and melody based on "happy feelings and birthday celebration songs".

[0705] 2. The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. It inputs a prompt into the generative AI and generates content based on the user's request.

[0706] 3. The server arranges the song based on the revised lyrics and melody. Specifically, it automatically selects appropriate instruments and arranges the music. Using an AI arrangement tool, it automatically selects instruments such as guitar, piano, and drums, and arranges them appropriately.

[0707] 4. The server sends the final completed music data to the terminal. It returns the completed music file to the terminal as an HTTP response.

[0708] As described above, this invention allows for the entire process from generating, modifying, arranging, and providing music tailored to the user's emotions and purpose, making it easy for users to create specialized music.

[0709] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0710] Step 1:

[0711] Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface. The entered text data includes "emotions" and "purpose."

[0712] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0713] Input: "Happy feelings", "Birthday song"

[0714] Output: Text data (emotion, purpose)

[0715] Step 2:

[0716] The terminal receives text data entered by the user and analyzes it using natural language processing (NLP) techniques. Specifically, it uses the Python NLTK library to extract keywords and important phrases from the text data.

[0717] Specific examples: "fun" → positive, "birthday song" → celebration

[0718] Input: Text data (emotion, purpose)

[0719] Output: Keywords and important phrases

[0720] Step 3:

[0721] The terminal sends extracted keywords and important phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[0722] Input: Keywords and important phrases

[0723] Output: Parsed data in JSON format (keywords, phrases)

[0724] Step 4:

[0725] The server invokes a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. The OpenAI API is used for this generation.

[0726] Example: Based on the input data, generate lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0727] Example prompt: "Happy feelings and birthday songs"

[0728] Input: Analysis data (keywords, phrases)

[0729] Output: Initial lyrics and melody

[0730] Step 5:

[0731] The terminal displays the initial lyrics and melody received from the server to the user and requests feedback. The user enters and submits their feedback.

[0732] Example: The lyrics "Today is your special day, a birthday filled with smiles..." are displayed and the melody plays. The user then provides feedback saying, "Make this part sound more lively."

[0733] Input: Early lyrics and melody

[0734] Output: User feedback

[0735] Step 6:

[0736] The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. Based on the feedback, it adjusts the prompt text and re-inputs it into the generative AI.

[0737] Specific example: Based on feedback such as "change it to sound more lively," the lyrics "Today is your special day, everyone gathers to sing a celebratory song..." are revised to a more lively melody.

[0738] Input: User feedback

[0739] Output: Revised lyrics and melody

[0740] Step 7:

[0741] The server arranges the song based on the revised lyrics and melody. Specifically, it uses an AI arrangement tool to automatically select instruments and perform the arrangement.

[0742] Specific example: Automatically select instruments such as guitar, piano, and drums, and arrange them appropriately.

[0743] Input: Modified lyrics and melody

[0744] Output: Arranged song

[0745] Step 8:

[0746] The server then sends the final completed music data to the terminal. HTTP responses are used for data transmission, returning the completed music file to the terminal.

[0747] Input: Arranged song

[0748] Output: Music data (final version)

[0749] Step 9:

[0750] The device provides saving, playback, and download functions for delivering completed music data to the user. Specifically, it saves the music to the file system, provides a download link, and displays an interface with playback capabilities.

[0751] Specific example: Click the download link to save the song and play it with your preferred media player.

[0752] Input: Music data (final version)

[0753] Output: Music files accessible to the user

[0754] (Application Example 1)

[0755] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0756] When users generate music according to their own emotions and purposes, it is difficult for users without specialized knowledge or skills to easily customize and create individually tailored music. Furthermore, providing an environment where users can stream or download the generated music in real time is not easy. To solve these problems, a consistent service is needed for generating, modifying, arranging, and delivering music tailored to the user's emotions and purposes.

[0757] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0758] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; means for providing the arranged song data to the user; and means for implementing the song generation system as a smartphone application and making the generated song available for streaming or download. This allows users to easily generate songs that match their emotions and purposes and use them in real time.

[0759] "Emotions" refer to the user's own internal feelings and moods.

[0760] "Purpose" refers to the user's intended use in a specific situation or event.

[0761] "Text data" refers to string information entered by the user.

[0762] "Natural language processing technology" refers to the technology that enables computers to understand and analyze human language.

[0763] "Keywords" refer to important words or phrases extracted from text data.

[0764] An "important phrase" refers to a meaningful sequence of words extracted from text data.

[0765] "Generative artificial intelligence" refers to artificial intelligence that has the ability to generate new information based on given data.

[0766] "Lyrics" refers to the spoken parts of a song.

[0767] "Melody" refers to a sequence of sounds in a musical piece.

[0768] An "idea" refers to an initial concept or idea.

[0769] "Feedback" refers to reactions and opinions from users.

[0770] "Modification" refers to improving initial ideas or data.

[0771] "Arrangement" refers to the selection of instruments and arrangement of music for revised lyrics and melody.

[0772] "Song data" refers to the information of a completed song.

[0773] A "smartphone application" refers to a software program that runs on a smartphone.

[0774] "Streaming distribution" refers to a method of transmitting and playing data in real time.

[0775] "Downloading" refers to saving data to a local device via a network.

[0776] One embodiment of this invention is a system that allows users to generate music according to their own emotions and purposes, and then stream or download it in real time via a smartphone application.

[0777] System Program Overview

[0778] 1. Receiving and analyzing user input

[0779] Users input their emotions and goals as text data using a smartphone application. Keywords and important phrases are extracted from the text data using natural language processing (NLP) techniques. SpaCy and NLTK are suitable NLP libraries to use.

[0780] 2. The process of creating music

[0781] The analyzed keywords and important phrases are sent to the server. The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on this information. This initial generation process uses prompts such as the following:

[0782] "Please generate lyrics and melody for a fun birthday celebration song."

[0783] 3. Feedback and Corrections

[0784] The smartphone application presents the user with the initial generated lyrics and melody. The user reviews them and provides feedback. The server receives the feedback and uses generative artificial intelligence again to revise the lyrics and melody. The following prompts are used during the revision process.

[0785] "Please make this part more lively. Please revise it."

[0786] 4. Final arrangement and serving

[0787] Based on the revised lyrics and melody, the song is arranged. Automatic arrangement technology is used to select instruments and arrange the music, generating the final song data. The generated song data is provided to the user via a smartphone application. The user can stream or download this song in real time.

[0788] Implementation explanation

[0789] Hardware and software

[0790] Smartphone application: Used for user input, feedback collection, and presentation of generated music. Flutter and React Native are suitable development frameworks.

[0791] Server: Used for data processing and the execution of generative artificial intelligence. AWS and Google Cloud are suitable cloud platforms.

[0792] Generative artificial intelligence: Generates and modifies lyrics and melodies based on given data. OpenAI's GPT-4 is a suitable AI model to use.

[0793] Natural Language Processing (NLP) techniques: Used for analyzing user input. Suitable NLP libraries include spaCy and NLTK.

[0794] Adding specific examples

[0795] 1. Analysis of user input

[0796] The user inputs the emotion of "fun" and the objective of "birthday celebration songs." The application analyzes this input and extracts keywords and important phrases such as "fun" = positive and "birthday celebration songs" = celebration.

[0797] 2. Initial Generation and Presentation

[0798] Based on the extracted keywords, the following prompts are entered into the generative artificial intelligence.

[0799] "Please generate lyrics and melody for a fun birthday celebration song."

[0800] The initial generated lyrics and melody are presented to the user through the application.

[0801] 3. Feedback and Corrections

[0802] The user provides feedback saying, "Make this part feel more lively." The server then prompts the generative AI again with the following message.

[0803] "Please make this part more lively. Please revise it."

[0804] The revised lyrics and melody are presented to the user again.

[0805] 4. Final arrangement and serving

[0806] Based on the final feedback, the song is rearranged and the completed track is generated. The generated track can be streamed or downloaded in real time.

[0807] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0808] Step 1:

[0809] Receiving input from the user

[0810] The user uses a smartphone application to input their emotions (e.g., "happy") and purpose (e.g., "birthday song") in text format. This input data is then sent to the application.

[0811] Input: Text data describing emotions and objectives

[0812] Output: Text data is saved to the application.

[0813] Step 2:

[0814] Text data analysis

[0815] The terminal analyzes the received text data using natural language processing (NLP) techniques. Here, NLP libraries (e.g., spaCy, NLTK) are used to extract keywords and important phrases from the text data.

[0816] Input: User-input text data

[0817] Data processing: Extract keywords and important phrases using an NLP library.

[0818] Output: Extracted keywords and important phrases

[0819] Step 3:

[0820] Sending data to the server

[0821] The terminal sends keywords and important phrases, which are the results of the analysis, to the server.

[0822] Input: Extracted keywords and important phrases

[0823] Data processing: Packaging and sending analysis results

[0824] Output: Keywords and important phrases are sent to the server.

[0825] Step 4:

[0826] Early lyrics and melody generation

[0827] The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on received keywords and important phrases. Instructions are given to the AI ​​using prompts.

[0828] Input: Keywords and important phrases

[0829] Data processing: Enter a prompt message into the generation AI model (e.g., "Generate lyrics and melody for a fun birthday song").

[0830] Output: Generated initial lyrics and melody

[0831] Step 5:

[0832] Presentation of initial generation results and reception of feedback

[0833] The device presents the user with the initial generated lyrics and melody and collects user feedback. The user enters specific requests and suggestions for revisions in text format.

[0834] Input: Generated initial lyrics and melody

[0835] Data processing: Display to the user and receiving feedback.

[0836] Output: User feedback

[0837] Step 6:

[0838] Corrections based on feedback

[0839] Based on the feedback received from the user, the server uses generative artificial intelligence to revise the lyrics and melody again. Prompts are used for the revision process.

[0840] Input: User feedback

[0841] Data processing: Input prompt text into the generating AI model (e.g., "Make this part sound more lively. Please revise it.")

[0842] Output: Revised lyrics and melody

[0843] Step 7:

[0844] Execution of the final arrangement

[0845] The server uses automated arrangement technology to create the final arrangement of the song based on the revised lyrics and melody. Instrument selection and arrangement are performed automatically.

[0846] Input: Modified lyrics and melody

[0847] Data processing: Arrangement using automatic arrangement technology

[0848] Output: Arranged song data

[0849] Step 8:

[0850] Music data provision

[0851] The device provides the user with the completed music data. The user can then stream or download this music in real time.

[0852] Input: Arranged song data

[0853] Data processing: Determining how to deliver music data (streaming or download).

[0854] Output: Music data provided to the user

[0855] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0856] This invention is a system that can generate more accurate songs by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and purpose. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[0857] The introduction of an emotion engine will enable the recognition of emotions from the user's facial expressions and voice data, allowing for the provision of more personalized and tailored music.

[0858] User roles

[0859] User

[0860] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[0861] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0862] 2. You input facial expressions and voice into the system, and the emotion engine analyzes them.

[0863] 3. Review the initial lyrics and melody provided and enter your feedback.

[0864] Example: Enter specific feedback such as, "Make this part feel more lively."

[0865] 4. Review the revised lyrics and melody again and provide further feedback if necessary.

[0866] 5. Download the completed song and use it for performances or recordings.

[0867] Terminal role

[0868] terminal

[0869] 1. Preprocessing

[0870] The system receives text data, facial expressions, and voice data from the user and analyzes them using an emotion engine. This analysis allows for a more accurate recognition of the user's emotions.

[0871] This analysis extracts keywords and important phrases based on emotions and purposes.

[0872] Example: Emotion = "Positive", Purpose = "Celebration"

[0873] 2. Data transmission

[0874] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[0875] 3. Interface

[0876] The system displays the server's response to the user and collects and sends feedback.

[0877] 4. Data Management

[0878] It provides saving, playback, and download functions to deliver completed music data to users.

[0879] Server Role

[0880] server

[0881] 1. Idea generation

[0882] Based on the analysis data sent from the terminal, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[0883] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0884] 2. Feedback Processing

[0885] We receive user feedback and use generative artificial intelligence again to revise the lyrics and melody.

[0886] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0887] 3. Arrangement

[0888] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[0889] 4. Data transmission

[0890] The final completed song data is sent to the device.

[0891] Specific example

[0892] For example, if a user inputs the emotion "happy" and the objective "birthday song," and also inputs facial expressions and voice, the system will operate as follows:

[0893] 1. The device analyzes the input text, facial expressions, and voice data, and the emotion engine recognizes the emotion as "happy." Based on this, keywords and important phrases are extracted and sent to the server.

[0894] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[0895] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[0896] 4. The server uses a generative AI again to revise the lyrics and melody to make them more lively.

[0897] 5. The server selects instruments, arranges the music, and completes the final song.

[0898] 6. The device provides the user with the completed music, allowing the user to download, save, play, record, and perform other operations.

[0899] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[0900] The following describes the processing flow.

[0901] Step 1:

[0902] User

[0903] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[0904] Example: Emotion = "Happiness", Purpose = "Birthday song"

[0905] Step 2:

[0906] User

[0907] The system receives input of facial expressions and voice. A dedicated camera and microphone are used to record facial expressions and voice tone.

[0908] Step 3:

[0909] terminal

[0910] It receives text data, facial expression data, and voice data entered by the user.

[0911] The system uses an emotion engine to analyze input facial expression and voice data to recognize the user's emotions.

[0912] Example: Analysis result = "fun"

[0913] Step 4:

[0914] terminal

[0915] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[0916] Example: Emotion = "Positive", Purpose = "Celebration"

[0917] Step 5:

[0918] terminal

[0919] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[0920] Step 6:

[0921] server

[0922] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyric and melody ideas.

[0923] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[0924] Step 7:

[0925] server

[0926] Send the generated initial lyrics and melody to the device.

[0927] Step 8:

[0928] terminal

[0929] The initial lyrics and melody received from the server are presented to the user.

[0930] Step 9:

[0931] User

[0932] Review the provided lyrics and melody, and provide feedback as needed.

[0933] Example: Enter specific requests such as, "Make this part feel more lively."

[0934] Step 10:

[0935] terminal

[0936] Receive user feedback and send it to the server.

[0937] Step 11:

[0938] server

[0939] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[0940] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[0941] Step 12:

[0942] server

[0943] The revised lyrics and melody are sent to the device and presented to the user again.

[0944] Step 13:

[0945] User

[0946] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[0947] Step 14:

[0948] server

[0949] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[0950] Step 15:

[0951] server

[0952] Send the completed song data to the device.

[0953] Step 16:

[0954] terminal

[0955] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[0956] This processing flow allows users to highly customize specific songs to match their emotions and purposes through a system that combines an emotion engine and generative AI, enabling them to create more accurate and personalized songs.

[0957] (Example 2)

[0958] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0959] Conventional songwriting support systems struggled to provide sophisticated music that reflected specific emotions, even when generating songs based on the user's feelings and objectives. Furthermore, the process of revising songs while incorporating user feedback in real time was inefficient, and the automatic arrangement functions for improving song quality had limitations. As a result, there was a challenge in providing individually tailored songs.

[0960] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0961] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for using an emotion engine to analyze the text data and recognizing emotions from the user's facial expressions and voice data; means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyric and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user. This enables sophisticated song generation tailored to the user's emotions and purpose, and allows for the provision of individually tailored songs with greater accuracy.

[0962] "User" refers to a person or user who wishes to create music using this system.

[0963] "Emotions" refer to data that represents the inner feelings and intentions expressed by the user, and are primarily input through text, facial expressions, and voice data.

[0964] "Purpose" refers to data indicating the intended use or intent of the music the user wishes to create, and is entered in text format.

[0965] "Text data" refers to strings of information that users input into a system, and it includes emotions and intentions.

[0966] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate human language.

[0967] "Keywords" are important words or phrases extracted from text data and used in the creation of music.

[0968] A "key phrase" is a meaningful sequence of words extracted from text data and used in the creation of a musical piece.

[0969] An "emotion engine" refers to software or hardware that analyzes and recognizes emotions from facial expressions and voice data entered by the user.

[0970] "Generative artificial intelligence" refers to an artificial intelligence system that automatically generates content (in this case, lyrics and melody) based on given input data.

[0971] "Initial lyrics and melody" refers to the basic text and musical melody that form the basis of the song first generated by the generative artificial intelligence.

[0972] "Feedback" refers to the opinions and requests that users input into a system for corrections and improvements.

[0973] "Revised lyrics and melody" refers to the text and musical melody regenerated by the generative artificial intelligence based on user feedback.

[0974] "Arrangement" refers to the process of selecting instruments and arranging the music based on the generated lyrics and melody to complete the song.

[0975] "Song data" refers to the digital information of the final produced song, which is provided to the user.

[0976] This invention is a system that generates more accurate music by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[0977] Hardware and software configuration

[0978] User

[0979] Users interact with the system through a dedicated application or web interface. Devices such as computers, smartphones, and tablets are used.

[0980] terminal

[0981] The device will be equipped with an interface for receiving user input data (text, facial expressions, and voice).

[0982] Microsoft Azure Cognitive Services' Emotion API is used as the emotion engine.

[0983] Natural language processing techniques are used to analyze the data and extract keywords and important phrases.

[0984] It will be equipped with network communication capabilities for sending analysis data to a server.

[0985] server

[0986] The server will be equipped with a function to generate lyrics and melodies using generative artificial intelligence (e.g., OpenAI's GPT-3).

[0987] The server will be equipped with a function to receive feedback from users and then use generative artificial intelligence to revise the lyrics and melody again.

[0988] It has a function to select instruments and arrange the music based on the revised lyrics and melody, and to generate the final song data.

[0989] Program Processing Overview

[0990] 1. Input of user's emotions and purpose

[0991] Users enter their emotions and the purpose of the song in text format via a dedicated app or web interface.

[0992] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[0993] 2. Input and analysis of facial expression and voice data

[0994] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[0995] The device recognizes emotions from the analysis results and extracts emotional data, such as "happy."

[0996] 3. Generation of initial lyrics and melody

[0997] The device extracts important keywords and phrases based on the extracted sentiment data and the user's input, and sends them to the server.

[0998] The server uses generative artificial intelligence to generate initial lyrics and melody.

[0999] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line.

[1000] 4. Collecting user feedback

[1001] The device displays the initial generated lyrics and melody, and the user reviews the displayed content and provides feedback.

[1002] Example: Type "Make this part sound more lively."

[1003] 5. Revise the lyrics and melody.

[1004] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody based on that feedback.

[1005] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", with a lively melody line.

[1006] 6. Arrangement and final generation of the music

[1007] Based on the revised lyrics and melody, the server selects instruments and arranges the music to improve its overall quality.

[1008] 7. Provision of the completed song

[1009] The server finally sends the completed song data to the terminal, which saves the completed song data and provides it to the user. The user can download, play, and record the song.

[1010] Example of a prompt

[1011] User: Fun

[1012] Purpose: Birthday celebration song

[1013] Feedback: "Make this part feel more lively."

[1014] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[1015] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1016] Step 1:

[1017] Users input emotions and the purpose of the song in text format through a dedicated application or web interface.

[1018] Input: The emotion and purpose of the song entered by the user (e.g., happy, birthday song)

[1019] Output: Input text data

[1020] Specific action: The user enters "fun" and "birthday song" into the application's input form and presses the submit button.

[1021] Step 2:

[1022] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[1023] Input: User facial expression data and voice data

[1024] Output: Analyzed emotion data (e.g., happy)

[1025] Specific operation: The user uses the camera and microphone to input a smiling expression and the voice message "Happy Birthday." The emotion engine analyzes the expression and voice and recognizes it as "happy."

[1026] Step 3:

[1027] The device extracts keywords and important phrases based on text data and analyzed sentiment data, and sends them to the server.

[1028] Input: Emotional data, objective data, user facial expression and voice analysis results

[1029] Output: Extracted keywords and important phrases (e.g., special day, birthday, smile)

[1030] Specific operation: The terminal analyzes the input "fun" and "birthday song," as well as the "fun" from the emotion engine, extracts keywords such as "special day," "birthday," and "smile," and sends them to the server.

[1031] Step 4:

[1032] The server uses generative artificial intelligence to generate initial lyrics and melodies based on keywords and important phrases.

[1033] Input: Extracted keywords and important phrases

[1034] Output: Initial lyrics and melody

[1035] Specific operation: The server inputs the received words "special day," "birthday," and "smile" as prompts into the generative artificial intelligence, and generates lyrics such as "Today is your special day, a birthday overflowing with smiles..." along with a cheerful melody line.

[1036] Step 5:

[1037] The device displays the initial generated lyrics and melody and collects feedback from the user.

[1038] Input: Early lyrics and melody

[1039] Output: User feedback (e.g., "Make this part more lively")

[1040] Specific operation: The device displays the lyrics and melody "Today is your special day, a birthday filled with smiles..." to the user, and the user provides feedback such as "Make this part sound more lively."

[1041] Step 6:

[1042] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody again.

[1043] Input: User feedback, initial lyrics and melody

[1044] Output: Revised lyrics and melody

[1045] Specific operation: The device sends the feedback "Make this part sound more lively" to the server, and the server uses generative artificial intelligence to revise the lyrics to "Today is your special day, everyone gathers to sing a celebratory song..." and also changes the melody to a more lively one.

[1046] Step 7:

[1047] The server selects instruments and arranges the music based on the revised lyrics and melody, generating the final song data.

[1048] Input: Modified lyrics and melody

[1049] Output: Final music data

[1050] Specific operation: Based on the revised lyrics and melody, the server automatically selects instruments such as piano, guitar, and drums, arranges the instrumentation and creates the final song data.

[1051] Step 8:

[1052] The server ultimately sends the completed music data to the terminal, which then provides it to the user.

[1053] Input: Final song data

[1054] Output: Interface for downloading, playing, and recording music.

[1055] Specific operation: The server sends the final music data to the terminal, and the terminal displays a download link, play button, and recording function to provide the user with the music. The user can download, play, and record the music.

[1056] (Application Example 2)

[1057] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1058] Conventional music generation systems have been insufficient in generating music that aligns with users' emotions and purposes, making it difficult to provide music tailored to specific emotions or objectives. Furthermore, limited means of incorporating user feedback during the music generation process have resulted in low final music quality and user satisfaction. Additionally, the lack of emotion analysis utilizing user facial expressions and voice data has led to insufficient accuracy in emotion-based music generation.

[1059] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving emotional and purpose text data input by the user; means for analyzing the text data and image or audio data and recognizing emotions from the user's facial expressions and voice; means for analyzing using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for modifying the lyrics and melody again using the generative artificial intelligence based on the feedback; means for arranging the modified lyrics and melody into a song; and means for providing the arranged song data to the user. This enables highly accurate song generation that is tailored to the user's emotions and purpose.

[1060] "Emotional and purposeful text data" refers to text-based information entered by the user that represents the emotions and purpose of the song.

[1061] "Image or audio data" refers to input data that includes the user's facial expressions and voice.

[1062] "Means of recognizing emotions" refers to technologies that analyze image or audio data to identify the user's emotions.

[1063] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and intent.

[1064] "Keywords and important phrases" are the main words and expressions necessary for song generation, extracted from the analyzed text data.

[1065] "Generative artificial intelligence" is an artificial intelligence technology that automatically generates lyrics and melodies based on input data.

[1066] "Initial lyrics and melody ideas" refers to the initial proposals for lyrics and melodies generated by a generative artificial intelligence.

[1067] "Feedback" refers to opinions and requests for corrections and improvements provided by users.

[1068] "Revised lyrics and melody" refers to lyrics and melodies regenerated by a generative artificial intelligence system based on user feedback.

[1069] "Arranging a song" refers to the process of selecting instruments and arranging the music based on the revised lyrics and melody to complete the song.

[1070] "Arranged song data" refers to the digital data of the completed song.

[1071] This invention is a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. By incorporating an emotion engine, it can generate highly accurate and personalized songs. Specifically, it uses devices such as smartphones, smart glasses, and head-mounted displays to input the user's emotional data, and generates songs using a generative artificial intelligence system built on a server.

[1072] The server includes means for receiving emotional and target text data entered by the user, means for analyzing the user's image or voice data to recognize emotions from facial expressions and voice, and means for analyzing text data using natural language processing technology to extract keywords and important phrases. It also includes means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence, means for presenting the initial lyric and melody ideas to the user and receiving feedback, means for revising the lyrics and melody using generative artificial intelligence again based on the feedback, means for arranging the revised lyrics and melody into a song, and means for providing the arranged song data to the user.

[1073] Users input their emotions and the purpose of the song in text format using a dedicated application or web interface. They also input facial expressions and voice, which the emotion engine analyzes. For example, if a user inputs the emotion "happy" and the purpose "birthday song," and also inputs facial expressions and voice into the system, the emotion engine will recognize the emotion as "happy" and extract keywords and important phrases.

[1074] The server uses generative artificial intelligence to generate initial lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody, which are then presented to the user via the terminal. If the user provides feedback such as "Make this part more lively," the generative AI can revise the lyrics and melody again to make it more vibrant. Finally, the completed song undergoes instrument selection and arrangement, and the final song data is provided to the user.

[1075] As a concrete example, if a user logs into a virtual store, selects an emotion (e.g., "fun") and a purpose (e.g., "relaxing music"), and inputs facial expressions and voice into the app, the following prompts will be generated based on the emotion analysis results.

[1076] Example of a prompt:

[1077] Please create lyrics and a melody that express the emotion of "joy" and aims to be "relaxing music."

[1078] By sending this prompt to a generative artificial intelligence system, high-quality music tailored to the user's emotions and purpose can be automatically generated. Users can then play, purchase, and download this music within a virtual store.

[1079] As a result, it becomes possible to generate highly accurate music tailored to the user's emotions and purpose, and to provide personalized music.

[1080] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1081] Step 1:

[1082] Users input emotions and the purpose of the music in text format using a dedicated application or web interface. The entered text data of emotions and purpose is sent to the terminal and transferred to the server. The user's input (text data of emotions and purpose) serves as the starting point for analysis.

[1083] Step 2:

[1084] The device receives text data and image and audio data (facial expressions and voice) entered by the user and analyzes them using an emotion engine. Specifically, it analyzes the user's facial expressions and voice data to recognize emotions, and extracts keywords and important phrases based on this analysis. The input for this step is facial expressions and voice data, and the output is the emotion recognition result and extracted keywords.

[1085] Step 3:

[1086] The server receives the analysis results, extracted keywords, and important phrases sent from the terminal, and uses a generative artificial intelligence model to generate initial lyric and melody ideas based on them. This generation process involves creating prompt sentences using the generative AI model and generating lyrics and melodies based on them. The input is the analysis results and keywords, and the output is the initial lyric and melody drafts.

[1087] Step 4:

[1088] The server presents the user with initial lyrics and melody ideas via a terminal. The user reviews these and provides specific feedback, such as "Make this part sound more lively." In this step, the user's input is feedback, and the output is a revision instruction that reflects that feedback.

[1089] Step 5:

[1090] The server receives user feedback and uses a generative artificial intelligence model to revise the lyrics and melody again. The generative AI regenerates prompt sentences based on the feedback and produces the revised lyrics and melody. The input is the feedback, and the output is the revised lyrics and melody.

[1091] Step 6:

[1092] The server then arranges the revised lyrics and melody into a song. Specifically, it automatically selects instruments and arranges the music to improve its overall quality. The input for this step is the revised lyrics and melody, and the output is the final arranged song data.

[1093] Step 7:

[1094] The server sends the final arranged music data to the terminal and provides it to the user. The user can play, purchase, and download this music data. The input is the final arranged music data, and the output is the data provided to the user.

[1095] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1096] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1097] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1098] [Third Embodiment]

[1099] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1100] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1101] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1102] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1103] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1104] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1105] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1106] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1107] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1108] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1109] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1110] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1111] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on the user's input data, modifies them according to the user's feedback, and ultimately completes and delivers the song.

[1112] A specific embodiment of this system will be described based on the roles of the user, terminal, and server.

[1113] User roles

[1114] User

[1115] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1116] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1117] 2. Review the initial lyrics and melody provided and enter your feedback.

[1118] Example: Enter specific feedback such as, "Make this part feel more lively."

[1119] 3. Review the revised lyrics and melody again and provide further feedback if necessary.

[1120] 4. Download the completed song and use it for performances or recordings.

[1121] Terminal role

[1122] terminal

[1123] 1. Preprocessing

[1124] The system receives text data (emotion and purpose) from users and analyzes it using natural language processing technology.

[1125] This analysis extracts keywords and important phrases (e.g., "fun" = positive, "birthday celebration song" = celebration).

[1126] 2. Data transmission

[1127] The extracted keywords and phrases are sent to the server.

[1128] 3. Interface

[1129] The system displays the server's response to the user and collects and sends feedback.

[1130] 4. Data Management

[1131] It provides saving, playback, and download functions to deliver completed music data to users.

[1132] Server Role

[1133] server

[1134] 1. Idea generation

[1135] Based on the analysis data sent from the terminal, a generative AI is invoked to generate initial lyrics and melody ideas.

[1136] Example: Lyrics like "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1137] 2. Feedback Processing

[1138] We receive user feedback and use generative AI to revise the lyrics and melody again.

[1139] Example: After receiving feedback such as "Make it sound more lively," the lyrics and melody are made more vibrant.

[1140] 3. Arrangement

[1141] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[1142] 4. Data transmission

[1143] The final completed song data is sent to the device.

[1144] Specific example

[1145] For example, if a user inputs the emotion "happy" and the objective "birthday song," the system will work as follows:

[1146] 1. The terminal parses the input text and sends it to the server.

[1147] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1148] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[1149] 4. The server uses a generative AI again to revise the lyrics to a more lively "Today is your special day, everyone gathers to sing a celebratory song..." and to a more lively melody.

[1150] 5. The server selects instruments, arranges the music, and completes the final song.

[1151] 6. The device provides the user with the completed song, which the user then uses to perform or record.

[1152] Thus, this system handles everything from generating, modifying, arranging, and delivering music tailored to the user's emotions and purpose, making it easy for users to create personalized music.

[1153] The following describes the processing flow.

[1154] Step 1:

[1155] User

[1156] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[1157] Example: Emotion = "Happiness", Purpose = "Birthday song"

[1158] Step 2:

[1159] terminal

[1160] Receive text data entered by the user.

[1161] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[1162] Example: Emotion = "Positive", Purpose = "Celebration"

[1163] Step 3:

[1164] terminal

[1165] The extracted keywords and important phrases are sent to the server.

[1166] Step 4:

[1167] server

[1168] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[1169] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1170] Step 5:

[1171] server

[1172] Send the generated initial lyrics and melody to the device.

[1173] Step 6:

[1174] terminal

[1175] The initial lyrics and melody received from the server are presented to the user.

[1176] Step 7:

[1177] User

[1178] Review the provided lyrics and melody, and provide feedback as needed.

[1179] Example: Enter specific requests such as, "Make this part feel more lively."

[1180] Step 8:

[1181] terminal

[1182] Receive user feedback and send it to the server.

[1183] Step 9:

[1184] server

[1185] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[1186] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[1187] Step 10:

[1188] server

[1189] The revised lyrics and melody are sent to the device and presented to the user again.

[1190] Step 11:

[1191] User

[1192] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[1193] Step 12:

[1194] server

[1195] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[1196] Step 13:

[1197] server

[1198] Send the completed song data to the device.

[1199] Step 14:

[1200] terminal

[1201] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[1202] This processing flow allows users to easily create specific songs tailored to their emotions and purposes, and to revise and arrange them as many times as needed based on feedback.

[1203] (Example 1)

[1204] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1205] Existing music generation systems struggle to effectively create music that aligns with the emotions and purposes input by the user. Furthermore, methods for incorporating user feedback on generated lyrics and melodies are insufficient, often resulting in final music that fails to meet user expectations. Additionally, arranging and refining music requires specialized knowledge and skills, making it difficult for users to easily generate high-quality music.

[1206] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1207] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for automatically selecting appropriate instruments and arranging the revised lyrics and melody; and means for providing the arranged song data to the user. This makes it possible to easily generate high-quality songs that match the emotions and purposes entered by the user and to revise and arrange them according to the user's feedback.

[1208] A "user" is an individual or group that wishes to use the system to generate music that matches their emotions and purpose.

[1209] "Emotions" refer to information used to express the psychological state a user is experiencing. For example, it can refer to states such as being happy, sad, or angry.

[1210] "Purpose" refers to information that clarifies the intention or use of the song. For example, it refers to a specific intended use, such as a birthday song or a wedding theme song.

[1211] "Text data" refers to string information about emotions and purposes entered by the user.

[1212] "Natural language processing technology" is a technique for analyzing text data and extracting keywords and important phrases.

[1213] A "keyword" is a particularly important word or phrase within text data. For example, "fun" or "birthday" would be considered keywords.

[1214] An "important phrase" is a part of the text data that has particular meaning in context.

[1215] "Generative artificial intelligence" refers to machine learning models or algorithms that generate new content based on input data.

[1216] "Initial lyrics and melody ideas" refers to the initial draft of lyrics and melody for a song created by a generative artificial intelligence.

[1217] "Feedback" refers to the revision requests and evaluations that users provide regarding the generated lyrics and melodies.

[1218] "Arrangement" refers to the process of selecting instruments and arranging music based on the generated lyrics and melody, thereby improving the overall quality of the song.

[1219] "Appropriate instrument selection and arrangement" refers to the process of choosing instruments that suit the atmosphere and style of the music, and determining their placement and playing methods.

[1220] "Song data" refers to the digital data of the final completed music content, including lyrics, melody, and arrangement.

[1221] A "system" is a general term for a set of devices and software that have a series of functions for generating music based on user input, making corrections and arrangements in response to user feedback, and finally providing the completed music.

[1222] Modes for carrying out the invention

[1223] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on user input data, modifies them according to user feedback, and ultimately completes and delivers the finished song. Specific embodiments of this invention will be described based on the roles of the user, terminal, and server.

[1224] User roles

[1225] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1226] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1227] 2. The user reviews the initial lyrics and melody provided and enters feedback.

[1228] Example: Enter specific feedback such as, "Make this part feel more lively."

[1229] 3. Users will review the revised lyrics and melody again and provide further feedback if necessary.

[1230] 4. Users download the completed song and use it for performances and recordings.

[1231] Terminal role

[1232] 1. The terminal receives text data (sentiment and purpose) entered by the user and analyzes it using natural language processing (NLP) techniques. This analysis uses NLP libraries such as Python's NLTK to extract important keywords and phrases.

[1233] Specific examples: "fun" = positive, "birthday celebration song" = celebratory phrases are extracted.

[1234] 2. The terminal sends the extracted keywords and phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[1235] 3. The terminal displays the response from the server to the user, collects and sends feedback. It displays the initial lyrics and melody to the user and collects new feedback from the user again.

[1236] 4. The terminal provides saving, playback, and download functions for delivering completed music data to the user. It saves the music to the file system, provides a download link, and displays an interface with playback functionality.

[1237] Server Role

[1238] 1. The server calls a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. OpenAI APIs and other tools are used for this generation process.

[1239] Example prompt: Generate lyrics and melody based on "happy feelings and birthday celebration songs".

[1240] 2. The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. It inputs a prompt into the generative AI and generates content based on the user's request.

[1241] 3. The server arranges the song based on the revised lyrics and melody. Specifically, it automatically selects appropriate instruments and arranges the music. Using an AI arrangement tool, it automatically selects instruments such as guitar, piano, and drums, and arranges them appropriately.

[1242] 4. The server sends the final completed music data to the terminal. It returns the completed music file to the terminal as an HTTP response.

[1243] As described above, this invention allows for the entire process from generating, modifying, arranging, and providing music tailored to the user's emotions and purpose, making it easy for users to create specialized music.

[1244] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1245] Step 1:

[1246] Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface. The entered text data includes "emotions" and "purpose."

[1247] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1248] Input: "Happy feelings", "Birthday song"

[1249] Output: Text data (emotion, purpose)

[1250] Step 2:

[1251] The terminal receives text data entered by the user and analyzes it using natural language processing (NLP) techniques. Specifically, it uses the Python NLTK library to extract keywords and important phrases from the text data.

[1252] Specific examples: "fun" → positive, "birthday song" → celebration

[1253] Input: Text data (emotion, purpose)

[1254] Output: Keywords and important phrases

[1255] Step 3:

[1256] The terminal sends extracted keywords and important phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[1257] Input: Keywords and important phrases

[1258] Output: Parsed data in JSON format (keywords, phrases)

[1259] Step 4:

[1260] The server invokes a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. The OpenAI API is used for this generation.

[1261] Example: Based on the input data, generate lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1262] Example prompt: "Happy feelings and birthday songs"

[1263] Input: Analysis data (keywords, phrases)

[1264] Output: Initial lyrics and melody

[1265] Step 5:

[1266] The terminal displays the initial lyrics and melody received from the server to the user and requests feedback. The user enters and submits their feedback.

[1267] Example: The lyrics "Today is your special day, a birthday filled with smiles..." are displayed and the melody plays. The user then provides feedback saying, "Make this part sound more lively."

[1268] Input: Early lyrics and melody

[1269] Output: User feedback

[1270] Step 6:

[1271] The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. Based on the feedback, it adjusts the prompt text and re-inputs it into the generative AI.

[1272] Specific example: Based on feedback such as "change it to sound more lively," the lyrics "Today is your special day, everyone gathers to sing a celebratory song..." are revised to a more lively melody.

[1273] Input: User feedback

[1274] Output: Revised lyrics and melody

[1275] Step 7:

[1276] The server arranges the song based on the revised lyrics and melody. Specifically, it uses an AI arrangement tool to automatically select instruments and perform the arrangement.

[1277] Specific example: Automatically select instruments such as guitar, piano, and drums, and arrange them appropriately.

[1278] Input: Modified lyrics and melody

[1279] Output: Arranged song

[1280] Step 8:

[1281] The server then sends the final completed music data to the terminal. HTTP responses are used for data transmission, returning the completed music file to the terminal.

[1282] Input: Arranged song

[1283] Output: Music data (final version)

[1284] Step 9:

[1285] The device provides saving, playback, and download functions for delivering completed music data to the user. Specifically, it saves the music to the file system, provides a download link, and displays an interface with playback capabilities.

[1286] Specific example: Click the download link to save the song and play it with your preferred media player.

[1287] Input: Music data (final version)

[1288] Output: Music files accessible to the user

[1289] (Application Example 1)

[1290] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1291] When users generate music according to their own emotions and purposes, it is difficult for users without specialized knowledge or skills to easily customize and create individually tailored music. Furthermore, providing an environment where users can stream or download the generated music in real time is not easy. To solve these problems, a consistent service is needed for generating, modifying, arranging, and delivering music tailored to the user's emotions and purposes.

[1292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1293] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; means for providing the arranged song data to the user; and means for implementing the song generation system as a smartphone application and making the generated song available for streaming or download. This allows users to easily generate songs that match their emotions and purposes and use them in real time.

[1294] "Emotions" refer to the user's own internal feelings and moods.

[1295] "Purpose" refers to the user's intended use in a specific situation or event.

[1296] "Text data" refers to string information entered by the user.

[1297] "Natural language processing technology" refers to the technology that enables computers to understand and analyze human language.

[1298] "Keywords" refer to important words or phrases extracted from text data.

[1299] An "important phrase" refers to a meaningful sequence of words extracted from text data.

[1300] "Generative artificial intelligence" refers to artificial intelligence that has the ability to generate new information based on given data.

[1301] "Lyrics" refers to the spoken parts of a song.

[1302] "Melody" refers to a sequence of sounds in a musical piece.

[1303] An "idea" refers to an initial concept or idea.

[1304] "Feedback" refers to reactions and opinions from users.

[1305] "Modification" refers to improving initial ideas or data.

[1306] "Arrangement" refers to the selection of instruments and arrangement of music for revised lyrics and melody.

[1307] "Song data" refers to the information of a completed song.

[1308] A "smartphone application" refers to a software program that runs on a smartphone.

[1309] "Streaming distribution" refers to a method of transmitting and playing data in real time.

[1310] "Downloading" refers to saving data to a local device via a network.

[1311] One embodiment of this invention is a system that allows users to generate music according to their own emotions and purposes, and then stream or download it in real time via a smartphone application.

[1312] System Program Overview

[1313] 1. Receiving and analyzing user input

[1314] Users input their emotions and goals as text data using a smartphone application. Keywords and important phrases are extracted from the text data using natural language processing (NLP) techniques. SpaCy and NLTK are suitable NLP libraries to use.

[1315] 2. The process of creating music

[1316] The analyzed keywords and important phrases are sent to the server. The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on this information. This initial generation process uses prompts such as the following:

[1317] "Please generate lyrics and melody for a fun birthday celebration song."

[1318] 3. Feedback and Corrections

[1319] The smartphone application presents the user with the initial generated lyrics and melody. The user reviews them and provides feedback. The server receives the feedback and uses generative artificial intelligence again to revise the lyrics and melody. The following prompts are used during the revision process.

[1320] "Please make this part more lively. Please revise it."

[1321] 4. Final arrangement and serving

[1322] Based on the revised lyrics and melody, the song is arranged. Automatic arrangement technology is used to select instruments and arrange the music, generating the final song data. The generated song data is provided to the user via a smartphone application. The user can stream or download this song in real time.

[1323] Implementation explanation

[1324] Hardware and software

[1325] Smartphone application: Used for user input, feedback collection, and presentation of generated music. Flutter and React Native are suitable development frameworks.

[1326] Server: Used for data processing and the execution of generative artificial intelligence. AWS and Google Cloud are suitable cloud platforms.

[1327] Generative artificial intelligence: Generates and modifies lyrics and melodies based on given data. OpenAI's GPT-4 is a suitable AI model to use.

[1328] Natural Language Processing (NLP) techniques: Used for analyzing user input. Suitable NLP libraries include spaCy and NLTK.

[1329] Adding specific examples

[1330] 1. Analysis of user input

[1331] The user inputs the emotion of "fun" and the objective of "birthday celebration songs." The application analyzes this input and extracts keywords and important phrases such as "fun" = positive and "birthday celebration songs" = celebration.

[1332] 2. Initial Generation and Presentation

[1333] Based on the extracted keywords, the following prompts are entered into the generative artificial intelligence.

[1334] "Please generate lyrics and melody for a fun birthday celebration song."

[1335] The initial generated lyrics and melody are presented to the user through the application.

[1336] 3. Feedback and Corrections

[1337] The user provides feedback saying, "Make this part feel more lively." The server then prompts the generative AI again with the following message.

[1338] "Please make this part more lively. Please revise it."

[1339] The revised lyrics and melody are presented to the user again.

[1340] 4. Final arrangement and serving

[1341] Based on the final feedback, the song is rearranged and the completed track is generated. The generated track can be streamed or downloaded in real time.

[1342] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1343] Step 1:

[1344] Receiving input from the user

[1345] The user uses a smartphone application to input their emotions (e.g., "happy") and purpose (e.g., "birthday song") in text format. This input data is then sent to the application.

[1346] Input: Text data describing emotions and objectives

[1347] Output: Text data is saved to the application.

[1348] Step 2:

[1349] Text data analysis

[1350] The terminal analyzes the received text data using natural language processing (NLP) techniques. Here, NLP libraries (e.g., spaCy, NLTK) are used to extract keywords and important phrases from the text data.

[1351] Input: User-input text data

[1352] Data processing: Extract keywords and important phrases using an NLP library.

[1353] Output: Extracted keywords and important phrases

[1354] Step 3:

[1355] Sending data to the server

[1356] The terminal sends keywords and important phrases, which are the results of the analysis, to the server.

[1357] Input: Extracted keywords and important phrases

[1358] Data processing: Packaging and sending analysis results

[1359] Output: Keywords and important phrases are sent to the server.

[1360] Step 4:

[1361] Early lyrics and melody generation

[1362] The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on received keywords and important phrases. Instructions are given to the AI ​​using prompts.

[1363] Input: Keywords and important phrases

[1364] Data processing: Enter a prompt message into the generation AI model (e.g., "Generate lyrics and melody for a fun birthday song").

[1365] Output: Generated initial lyrics and melody

[1366] Step 5:

[1367] Presentation of initial generation results and reception of feedback

[1368] The device presents the user with the initial generated lyrics and melody and collects user feedback. The user enters specific requests and suggestions for revisions in text format.

[1369] Input: Generated initial lyrics and melody

[1370] Data processing: Display to the user and receiving feedback.

[1371] Output: User feedback

[1372] Step 6:

[1373] Corrections based on feedback

[1374] Based on the feedback received from the user, the server uses generative artificial intelligence to revise the lyrics and melody again. Prompts are used for the revision process.

[1375] Input: User feedback

[1376] Data processing: Input prompt text into the generating AI model (e.g., "Make this part sound more lively. Please revise it.")

[1377] Output: Revised lyrics and melody

[1378] Step 7:

[1379] Execution of the final arrangement

[1380] The server uses automated arrangement technology to create the final arrangement of the song based on the revised lyrics and melody. Instrument selection and arrangement are performed automatically.

[1381] Input: Modified lyrics and melody

[1382] Data processing: Arrangement using automatic arrangement technology

[1383] Output: Arranged song data

[1384] Step 8:

[1385] Music data provision

[1386] The device provides the user with the completed music data. The user can then stream or download this music in real time.

[1387] Input: Arranged song data

[1388] Data processing: Determining how to deliver music data (streaming or download).

[1389] Output: Music data provided to the user

[1390] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1391] This invention is a system that can generate more accurate songs by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and purpose. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[1392] The introduction of an emotion engine will enable the recognition of emotions from the user's facial expressions and voice data, allowing for the provision of more personalized and tailored music.

[1393] User roles

[1394] User

[1395] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1396] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1397] 2. You input facial expressions and voice into the system, and the emotion engine analyzes them.

[1398] 3. Review the initial lyrics and melody provided and enter your feedback.

[1399] Example: Enter specific feedback such as, "Make this part feel more lively."

[1400] 4. Review the revised lyrics and melody again and provide further feedback if necessary.

[1401] 5. Download the completed song and use it for performances or recordings.

[1402] Terminal role

[1403] terminal

[1404] 1. Preprocessing

[1405] The system receives text data, facial expressions, and voice data from the user and analyzes them using an emotion engine. This analysis allows for a more accurate recognition of the user's emotions.

[1406] This analysis extracts keywords and important phrases based on emotions and purposes.

[1407] Example: Emotion = "Positive", Purpose = "Celebration"

[1408] 2. Data transmission

[1409] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[1410] 3. Interface

[1411] The system displays the server's response to the user and collects and sends feedback.

[1412] 4. Data Management

[1413] It provides saving, playback, and download functions to deliver completed music data to users.

[1414] Server Role

[1415] server

[1416] 1. Idea generation

[1417] Based on the analysis data sent from the terminal, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[1418] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1419] 2. Feedback Processing

[1420] We receive user feedback and use generative artificial intelligence again to revise the lyrics and melody.

[1421] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[1422] 3. Arrangement

[1423] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[1424] 4. Data transmission

[1425] The final completed song data is sent to the device.

[1426] Specific example

[1427] For example, if a user inputs the emotion "happy" and the objective "birthday song," and also inputs facial expressions and voice, the system will operate as follows:

[1428] 1. The device analyzes the input text, facial expressions, and voice data, and the emotion engine recognizes the emotion as "happy." Based on this, keywords and important phrases are extracted and sent to the server.

[1429] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1430] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[1431] 4. The server uses a generative AI again to revise the lyrics and melody to make them more lively.

[1432] 5. The server selects instruments, arranges the music, and completes the final song.

[1433] 6. The device provides the user with the completed music, allowing the user to download, save, play, record, and perform other operations.

[1434] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[1435] The following describes the processing flow.

[1436] Step 1:

[1437] User

[1438] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[1439] Example: Emotion = "Happiness", Purpose = "Birthday song"

[1440] Step 2:

[1441] User

[1442] The system receives input of facial expressions and voice. A dedicated camera and microphone are used to record facial expressions and voice tone.

[1443] Step 3:

[1444] terminal

[1445] It receives text data, facial expression data, and voice data entered by the user.

[1446] The system uses an emotion engine to analyze input facial expression and voice data to recognize the user's emotions.

[1447] Example: Analysis result = "fun"

[1448] Step 4:

[1449] terminal

[1450] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[1451] Example: Emotion = "Positive", Purpose = "Celebration"

[1452] Step 5:

[1453] terminal

[1454] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[1455] Step 6:

[1456] server

[1457] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyric and melody ideas.

[1458] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1459] Step 7:

[1460] server

[1461] Send the generated initial lyrics and melody to the device.

[1462] Step 8:

[1463] terminal

[1464] The initial lyrics and melody received from the server are presented to the user.

[1465] Step 9:

[1466] User

[1467] Review the provided lyrics and melody, and provide feedback as needed.

[1468] Example: Enter specific requests such as, "Make this part feel more lively."

[1469] Step 10:

[1470] terminal

[1471] Receive user feedback and send it to the server.

[1472] Step 11:

[1473] server

[1474] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[1475] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[1476] Step 12:

[1477] server

[1478] The revised lyrics and melody are sent to the device and presented to the user again.

[1479] Step 13:

[1480] User

[1481] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[1482] Step 14:

[1483] server

[1484] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[1485] Step 15:

[1486] server

[1487] Send the completed song data to the device.

[1488] Step 16:

[1489] terminal

[1490] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[1491] This processing flow allows users to highly customize specific songs to match their emotions and purposes through a system that combines an emotion engine and generative AI, enabling them to create more accurate and personalized songs.

[1492] (Example 2)

[1493] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1494] Conventional songwriting support systems struggled to provide sophisticated music that reflected specific emotions, even when generating songs based on the user's feelings and objectives. Furthermore, the process of revising songs while incorporating user feedback in real time was inefficient, and the automatic arrangement functions for improving song quality had limitations. As a result, there was a challenge in providing individually tailored songs.

[1495] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1496] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for using an emotion engine to analyze the text data and recognizing emotions from the user's facial expressions and voice data; means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyric and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user. This enables sophisticated song generation tailored to the user's emotions and purpose, and allows for the provision of individually tailored songs with greater accuracy.

[1497] "User" refers to a person or user who wishes to create music using this system.

[1498] "Emotions" refer to data that represents the inner feelings and intentions expressed by the user, and are primarily input through text, facial expressions, and voice data.

[1499] "Purpose" refers to data indicating the intended use or intent of the music the user wishes to create, and is entered in text format.

[1500] "Text data" refers to strings of information that users input into a system, and it includes emotions and intentions.

[1501] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate human language.

[1502] "Keywords" are important words or phrases extracted from text data and used in the creation of music.

[1503] A "key phrase" is a meaningful sequence of words extracted from text data and used in the creation of a musical piece.

[1504] An "emotion engine" refers to software or hardware that analyzes and recognizes emotions from facial expressions and voice data entered by the user.

[1505] "Generative artificial intelligence" refers to an artificial intelligence system that automatically generates content (in this case, lyrics and melody) based on given input data.

[1506] "Initial lyrics and melody" refers to the basic text and musical melody that form the basis of the song first generated by the generative artificial intelligence.

[1507] "Feedback" refers to the opinions and requests that users input into a system for corrections and improvements.

[1508] "Revised lyrics and melody" refers to the text and musical melody regenerated by the generative artificial intelligence based on user feedback.

[1509] "Arrangement" refers to the process of selecting instruments and arranging the music based on the generated lyrics and melody to complete the song.

[1510] "Song data" refers to the digital information of the final produced song, which is provided to the user.

[1511] This invention is a system that generates more accurate music by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[1512] Hardware and software configuration

[1513] User

[1514] Users interact with the system through a dedicated application or web interface. Devices such as computers, smartphones, and tablets are used.

[1515] terminal

[1516] The device will be equipped with an interface for receiving user input data (text, facial expressions, and voice).

[1517] Microsoft Azure Cognitive Services' Emotion API is used as the emotion engine.

[1518] Natural language processing techniques are used to analyze the data and extract keywords and important phrases.

[1519] It will be equipped with network communication capabilities for sending analysis data to a server.

[1520] server

[1521] The server will be equipped with a function to generate lyrics and melodies using generative artificial intelligence (e.g., OpenAI's GPT-3).

[1522] The server will be equipped with a function to receive feedback from users and then use generative artificial intelligence to revise the lyrics and melody again.

[1523] It has a function to select instruments and arrange the music based on the revised lyrics and melody, and to generate the final song data.

[1524] Program Processing Overview

[1525] 1. Input of user's emotions and purpose

[1526] Users enter their emotions and the purpose of the song in text format via a dedicated app or web interface.

[1527] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1528] 2. Input and analysis of facial expression and voice data

[1529] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[1530] The device recognizes emotions from the analysis results and extracts emotional data, such as "happy."

[1531] 3. Generation of initial lyrics and melody

[1532] The device extracts important keywords and phrases based on the extracted sentiment data and the user's input, and sends them to the server.

[1533] The server uses generative artificial intelligence to generate initial lyrics and melody.

[1534] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line.

[1535] 4. Collecting user feedback

[1536] The device displays the initial generated lyrics and melody, and the user reviews the displayed content and provides feedback.

[1537] Example: Type "Make this part sound more lively."

[1538] 5. Revise the lyrics and melody.

[1539] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody based on that feedback.

[1540] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", with a lively melody line.

[1541] 6. Arrangement and final generation of the music

[1542] Based on the revised lyrics and melody, the server selects instruments and arranges the music to improve its overall quality.

[1543] 7. Provision of the completed song

[1544] The server finally sends the completed song data to the terminal, which saves the completed song data and provides it to the user. The user can download, play, and record the song.

[1545] Example of a prompt

[1546] User: Fun

[1547] Purpose: Birthday celebration song

[1548] Feedback: "Make this part feel more lively."

[1549] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[1550] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1551] Step 1:

[1552] Users input emotions and the purpose of the song in text format through a dedicated application or web interface.

[1553] Input: The emotion and purpose of the song entered by the user (e.g., happy, birthday song)

[1554] Output: Input text data

[1555] Specific action: The user enters "fun" and "birthday song" into the application's input form and presses the submit button.

[1556] Step 2:

[1557] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[1558] Input: User facial expression data and voice data

[1559] Output: Analyzed emotion data (e.g., happy)

[1560] Specific operation: The user uses the camera and microphone to input a smiling expression and the voice message "Happy Birthday." The emotion engine analyzes the expression and voice and recognizes it as "happy."

[1561] Step 3:

[1562] The device extracts keywords and important phrases based on text data and analyzed sentiment data, and sends them to the server.

[1563] Input: Emotional data, objective data, user facial expression and voice analysis results

[1564] Output: Extracted keywords and important phrases (e.g., special day, birthday, smile)

[1565] Specific operation: The terminal analyzes the input "fun" and "birthday song," as well as the "fun" from the emotion engine, extracts keywords such as "special day," "birthday," and "smile," and sends them to the server.

[1566] Step 4:

[1567] The server uses generative artificial intelligence to generate initial lyrics and melodies based on keywords and important phrases.

[1568] Input: Extracted keywords and important phrases

[1569] Output: Initial lyrics and melody

[1570] Specific operation: The server inputs the received words "special day," "birthday," and "smile" as prompts into the generative artificial intelligence, and generates lyrics such as "Today is your special day, a birthday overflowing with smiles..." along with a cheerful melody line.

[1571] Step 5:

[1572] The device displays the initial generated lyrics and melody and collects feedback from the user.

[1573] Input: Early lyrics and melody

[1574] Output: User feedback (e.g., "Make this part more lively")

[1575] Specific operation: The device displays the lyrics and melody "Today is your special day, a birthday filled with smiles..." to the user, and the user provides feedback such as "Make this part sound more lively."

[1576] Step 6:

[1577] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody again.

[1578] Input: User feedback, initial lyrics and melody

[1579] Output: Revised lyrics and melody

[1580] Specific operation: The device sends the feedback "Make this part sound more lively" to the server, and the server uses generative artificial intelligence to revise the lyrics to "Today is your special day, everyone gathers to sing a celebratory song..." and also changes the melody to a more lively one.

[1581] Step 7:

[1582] The server selects instruments and arranges the music based on the revised lyrics and melody, generating the final song data.

[1583] Input: Modified lyrics and melody

[1584] Output: Final music data

[1585] Specific operation: Based on the revised lyrics and melody, the server automatically selects instruments such as piano, guitar, and drums, arranges the instrumentation and creates the final song data.

[1586] Step 8:

[1587] The server ultimately sends the completed music data to the terminal, which then provides it to the user.

[1588] Input: Final song data

[1589] Output: Interface for downloading, playing, and recording music.

[1590] Specific operation: The server sends the final music data to the terminal, and the terminal displays a download link, play button, and recording function to provide the user with the music. The user can download, play, and record the music.

[1591] (Application Example 2)

[1592] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1593] Conventional music generation systems have been insufficient in generating music that aligns with users' emotions and purposes, making it difficult to provide music tailored to specific emotions or objectives. Furthermore, limited means of incorporating user feedback during the music generation process have resulted in low final music quality and user satisfaction. Additionally, the lack of emotion analysis utilizing user facial expressions and voice data has led to insufficient accuracy in emotion-based music generation.

[1594] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving emotional and purpose text data input by the user; means for analyzing the text data and image or audio data and recognizing emotions from the user's facial expressions and voice; means for analyzing using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for modifying the lyrics and melody again using the generative artificial intelligence based on the feedback; means for arranging the modified lyrics and melody into a song; and means for providing the arranged song data to the user. This enables highly accurate song generation that is tailored to the user's emotions and purpose.

[1595] "Emotional and purposeful text data" refers to text-based information entered by the user that represents the emotions and purpose of the song.

[1596] "Image or audio data" refers to input data that includes the user's facial expressions and voice.

[1597] "Means of recognizing emotions" refers to technologies that analyze image or audio data to identify the user's emotions.

[1598] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and intent.

[1599] "Keywords and important phrases" are the main words and expressions necessary for song generation, extracted from the analyzed text data.

[1600] "Generative artificial intelligence" is an artificial intelligence technology that automatically generates lyrics and melodies based on input data.

[1601] "Initial lyrics and melody ideas" refers to the initial proposals for lyrics and melodies generated by a generative artificial intelligence.

[1602] "Feedback" refers to opinions and requests for corrections and improvements provided by users.

[1603] "Revised lyrics and melody" refers to lyrics and melodies regenerated by a generative artificial intelligence system based on user feedback.

[1604] "Arranging a song" refers to the process of selecting instruments and arranging the music based on the revised lyrics and melody to complete the song.

[1605] "Arranged song data" refers to the digital data of the completed song.

[1606] This invention is a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. By incorporating an emotion engine, it can generate highly accurate and personalized songs. Specifically, it uses devices such as smartphones, smart glasses, and head-mounted displays to input the user's emotional data, and generates songs using a generative artificial intelligence system built on a server.

[1607] The server includes means for receiving emotional and target text data entered by the user, means for analyzing the user's image or voice data to recognize emotions from facial expressions and voice, and means for analyzing text data using natural language processing technology to extract keywords and important phrases. It also includes means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence, means for presenting the initial lyric and melody ideas to the user and receiving feedback, means for revising the lyrics and melody using generative artificial intelligence again based on the feedback, means for arranging the revised lyrics and melody into a song, and means for providing the arranged song data to the user.

[1608] Users input their emotions and the purpose of the song in text format using a dedicated application or web interface. They also input facial expressions and voice, which the emotion engine analyzes. For example, if a user inputs the emotion "happy" and the purpose "birthday song," and also inputs facial expressions and voice into the system, the emotion engine will recognize the emotion as "happy" and extract keywords and important phrases.

[1609] The server uses generative artificial intelligence to generate initial lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody, which are then presented to the user via the terminal. If the user provides feedback such as "Make this part more lively," the generative AI can revise the lyrics and melody again to make it more vibrant. Finally, the completed song undergoes instrument selection and arrangement, and the final song data is provided to the user.

[1610] As a concrete example, if a user logs into a virtual store, selects an emotion (e.g., "fun") and a purpose (e.g., "relaxing music"), and inputs facial expressions and voice into the app, the following prompts will be generated based on the emotion analysis results.

[1611] Example of a prompt:

[1612] Please create lyrics and a melody that express the emotion of "joy" and aim to create "relaxing music."

[1613] By sending this prompt to a generative artificial intelligence system, high-quality music tailored to the user's emotions and purpose can be automatically generated. Users can then play, purchase, and download this music within a virtual store.

[1614] As a result, it becomes possible to generate highly accurate music tailored to the user's emotions and purpose, and to provide personalized music.

[1615] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1616] Step 1:

[1617] Users input emotions and the purpose of the music in text format using a dedicated application or web interface. The entered text data of emotions and purpose is sent to the terminal and transferred to the server. The user's input (text data of emotions and purpose) serves as the starting point for analysis.

[1618] Step 2:

[1619] The device receives text data and image and audio data (facial expressions and voice) entered by the user and analyzes them using an emotion engine. Specifically, it analyzes the user's facial expressions and voice data to recognize emotions, and extracts keywords and important phrases based on this analysis. The input for this step is facial expressions and voice data, and the output is the emotion recognition result and extracted keywords.

[1620] Step 3:

[1621] The server receives the analysis results, extracted keywords, and important phrases sent from the terminal, and uses a generative artificial intelligence model to generate initial lyric and melody ideas based on them. This generation process involves creating prompt sentences using the generative AI model and generating lyrics and melodies based on them. The input is the analysis results and keywords, and the output is the initial lyric and melody drafts.

[1622] Step 4:

[1623] The server presents the user with initial lyrics and melody ideas via a terminal. The user reviews these and provides specific feedback, such as "Make this part sound more lively." In this step, the user's input is feedback, and the output is a revision instruction that reflects that feedback.

[1624] Step 5:

[1625] The server receives user feedback and uses a generative artificial intelligence model to revise the lyrics and melody again. The generative AI regenerates prompt sentences based on the feedback and produces the revised lyrics and melody. The input is the feedback, and the output is the revised lyrics and melody.

[1626] Step 6:

[1627] The server then arranges the revised lyrics and melody into a song. Specifically, it automatically selects instruments and arranges the music to improve its overall quality. The input for this step is the revised lyrics and melody, and the output is the final arranged song data.

[1628] Step 7:

[1629] The server sends the final arranged music data to the terminal and provides it to the user. The user can play, purchase, and download this music data. The input is the final arranged music data, and the output is the data provided to the user.

[1630] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1631] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1632] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1633] [Fourth Embodiment]

[1634] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1635] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1636] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1637] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1638] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1639] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1640] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1641] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1642] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1643] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1644] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1645] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1646] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1647] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on the user's input data, modifies them according to the user's feedback, and ultimately completes and delivers the song.

[1648] A specific embodiment of this system will be described based on the roles of the user, terminal, and server.

[1649] User roles

[1650] User

[1651] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1652] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1653] 2. Review the initial lyrics and melody provided and enter your feedback.

[1654] Example: Enter specific feedback such as, "Make this part feel more lively."

[1655] 3. Review the revised lyrics and melody again and provide further feedback if necessary.

[1656] 4. Download the completed song and use it for performances or recordings.

[1657] Terminal role

[1658] terminal

[1659] 1. Preprocessing

[1660] The system receives text data (emotion and purpose) from users and analyzes it using natural language processing technology.

[1661] This analysis extracts keywords and important phrases (e.g., "fun" = positive, "birthday celebration song" = celebration).

[1662] 2. Data transmission

[1663] The extracted keywords and phrases are sent to the server.

[1664] 3. Interface

[1665] The system displays the server's response to the user and collects and sends feedback.

[1666] 4. Data Management

[1667] It provides saving, playback, and download functions to deliver completed music data to users.

[1668] Server Role

[1669] server

[1670] 1. Idea generation

[1671] Based on the analysis data sent from the terminal, a generative AI is invoked to generate initial lyrics and melody ideas.

[1672] Example: Lyrics like "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1673] 2. Feedback Processing

[1674] We receive user feedback and use generative AI to revise the lyrics and melody again.

[1675] Example: After receiving feedback such as "Make it sound more lively," the lyrics and melody are made more vibrant.

[1676] 3. Arrangement

[1677] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[1678] 4. Data transmission

[1679] The final completed song data is sent to the device.

[1680] Specific example

[1681] For example, if a user inputs the emotion "happy" and the objective "birthday song," the system will work as follows:

[1682] 1. The terminal parses the input text and sends it to the server.

[1683] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1684] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[1685] 4. The server uses a generative AI again to revise the lyrics to a more lively "Today is your special day, everyone gathers to sing a celebratory song..." and to a more lively melody.

[1686] 5. The server selects instruments, arranges the music, and completes the final song.

[1687] 6. The device provides the user with the completed song, which the user then uses to perform or record.

[1688] Thus, this system handles everything from generating, modifying, arranging, and delivering music tailored to the user's emotions and purpose, making it easy for users to create personalized music.

[1689] The following describes the processing flow.

[1690] Step 1:

[1691] User

[1692] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[1693] Example: Emotion = "Happiness", Purpose = "Birthday song"

[1694] Step 2:

[1695] terminal

[1696] Receive text data entered by the user.

[1697] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[1698] Example: Emotion = "Positive", Purpose = "Celebration"

[1699] Step 3:

[1700] terminal

[1701] The extracted keywords and important phrases are sent to the server.

[1702] Step 4:

[1703] server

[1704] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[1705] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1706] Step 5:

[1707] server

[1708] Send the generated initial lyrics and melody to the device.

[1709] Step 6:

[1710] terminal

[1711] The initial lyrics and melody received from the server are presented to the user.

[1712] Step 7:

[1713] User

[1714] Review the provided lyrics and melody, and provide feedback as needed.

[1715] Example: Enter specific requests such as, "Make this part feel more lively."

[1716] Step 8:

[1717] terminal

[1718] Receive user feedback and send it to the server.

[1719] Step 9:

[1720] server

[1721] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[1722] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[1723] Step 10:

[1724] server

[1725] The revised lyrics and melody are sent to the device and presented to the user again.

[1726] Step 11:

[1727] User

[1728] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[1729] Step 12:

[1730] server

[1731] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[1732] Step 13:

[1733] server

[1734] Send the completed song data to the device.

[1735] Step 14:

[1736] terminal

[1737] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[1738] This processing flow allows users to easily create specific songs tailored to their emotions and purposes, and to revise and arrange them as many times as needed based on feedback.

[1739] (Example 1)

[1740] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1741] Existing music generation systems struggle to effectively create music that aligns with the emotions and purposes input by the user. Furthermore, methods for incorporating user feedback on generated lyrics and melodies are insufficient, often resulting in final music that fails to meet user expectations. Additionally, arranging and refining music requires specialized knowledge and skills, making it difficult for users to easily generate high-quality music.

[1742] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1743] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for automatically selecting appropriate instruments and arranging the revised lyrics and melody; and means for providing the arranged song data to the user. This makes it possible to easily generate high-quality songs that match the emotions and purposes entered by the user and to revise and arrange them according to the user's feedback.

[1744] A "user" is an individual or group that wishes to use the system to generate music that matches their emotions and purpose.

[1745] "Emotions" refer to information used to express the psychological state a user is experiencing. For example, it can refer to states such as being happy, sad, or angry.

[1746] "Purpose" refers to information that clarifies the intention or use of the song. For example, it refers to a specific intended use, such as a birthday song or a wedding theme song.

[1747] "Text data" refers to string information about emotions and purposes entered by the user.

[1748] "Natural language processing technology" is a technique for analyzing text data and extracting keywords and important phrases.

[1749] A "keyword" is a particularly important word or phrase within text data. For example, "fun" or "birthday" would be considered keywords.

[1750] An "important phrase" is a part of the text data that has particular meaning in context.

[1751] "Generative artificial intelligence" refers to machine learning models or algorithms that generate new content based on input data.

[1752] "Initial lyrics and melody ideas" refers to the initial draft of lyrics and melody for a song created by a generative artificial intelligence.

[1753] "Feedback" refers to the revision requests and evaluations that users provide regarding the generated lyrics and melodies.

[1754] "Arrangement" refers to the process of selecting instruments and arranging music based on the generated lyrics and melody, thereby improving the overall quality of the song.

[1755] "Appropriate instrument selection and arrangement" refers to the process of choosing instruments that suit the atmosphere and style of the music, and determining their placement and playing methods.

[1756] "Song data" refers to the digital data of the final completed music content, including lyrics, melody, and arrangement.

[1757] A "system" is a general term for a set of devices and software that have a series of functions for generating music based on user input, making corrections and arrangements in response to user feedback, and finally providing the completed music.

[1758] Modes for carrying out the invention

[1759] This invention is a music generation system that utilizes generative AI to assist in songwriting and composition according to the user's emotions and purpose. This system has a mechanism that generates lyrics and melodies based on user input data, modifies them according to user feedback, and ultimately completes and delivers the finished song. Specific embodiments of this invention will be described based on the roles of the user, terminal, and server.

[1760] User roles

[1761] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1762] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1763] 2. The user reviews the initial lyrics and melody provided and enters feedback.

[1764] Example: Enter specific feedback such as, "Make this part feel more lively."

[1765] 3. Users will review the revised lyrics and melody again and provide further feedback if necessary.

[1766] 4. Users download the completed song and use it for performances and recordings.

[1767] Terminal role

[1768] 1. The terminal receives text data (sentiment and purpose) entered by the user and analyzes it using natural language processing (NLP) techniques. This analysis uses NLP libraries such as Python's NLTK to extract important keywords and phrases.

[1769] Specific examples: "fun" = positive, "birthday celebration song" = celebratory phrases are extracted.

[1770] 2. The terminal sends the extracted keywords and phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[1771] 3. The terminal displays the response from the server to the user, collects and sends feedback. It displays the initial lyrics and melody to the user and collects new feedback from the user again.

[1772] 4. The terminal provides saving, playback, and download functions for delivering completed music data to the user. It saves the music to the file system, provides a download link, and displays an interface with playback functionality.

[1773] Server Role

[1774] 1. The server calls a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. OpenAI APIs and other tools are used for this generation process.

[1775] Example prompt: Generate lyrics and melody based on "happy feelings and birthday celebration songs".

[1776] 2. The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. It inputs a prompt into the generative AI and generates content based on the user's request.

[1777] 3. The server arranges the song based on the revised lyrics and melody. Specifically, it automatically selects appropriate instruments and arranges the music. Using an AI arrangement tool, it automatically selects instruments such as guitar, piano, and drums, and arranges them appropriately.

[1778] 4. The server sends the final completed music data to the terminal. It returns the completed music file to the terminal as an HTTP response.

[1779] As described above, this invention allows for the entire process from generating, modifying, arranging, and providing music tailored to the user's emotions and purpose, making it easy for users to create specialized music.

[1780] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1781] Step 1:

[1782] Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface. The entered text data includes "emotions" and "purpose."

[1783] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1784] Input: "Happy feelings", "Birthday song"

[1785] Output: Text data (emotion, purpose)

[1786] Step 2:

[1787] The terminal receives text data entered by the user and analyzes it using natural language processing (NLP) techniques. Specifically, it uses the Python NLTK library to extract keywords and important phrases from the text data.

[1788] Specific examples: "fun" → positive, "birthday song" → celebration

[1789] Input: Text data (emotion, purpose)

[1790] Output: Keywords and important phrases

[1791] Step 3:

[1792] The terminal sends extracted keywords and important phrases to the server. HTTP POST requests are used for data transmission, and the analysis results are sent to the server in JSON format.

[1793] Input: Keywords and important phrases

[1794] Output: Parsed data in JSON format (keywords, phrases)

[1795] Step 4:

[1796] The server invokes a generative AI model based on the analysis data sent from the terminal to generate initial lyric and melody ideas. The OpenAI API is used for this generation.

[1797] Example: Based on the input data, generate lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1798] Example prompt: "Happy feelings and birthday songs"

[1799] Input: Analysis data (keywords, phrases)

[1800] Output: Initial lyrics and melody

[1801] Step 5:

[1802] The terminal displays the initial lyrics and melody received from the server to the user and requests feedback. The user enters and submits their feedback.

[1803] Example: The lyrics "Today is your special day, a birthday filled with smiles..." are displayed and the melody plays. The user then provides feedback saying, "Make this part sound more lively."

[1804] Input: Early lyrics and melody

[1805] Output: User feedback

[1806] Step 6:

[1807] The server receives feedback from the user and uses the generative AI model again to revise the lyrics and melody. Based on the feedback, it adjusts the prompt text and re-inputs it into the generative AI.

[1808] Specific example: Based on feedback such as "change it to sound more lively," the lyrics "Today is your special day, everyone gathers to sing a celebratory song..." are revised to a more lively melody.

[1809] Input: User feedback

[1810] Output: Revised lyrics and melody

[1811] Step 7:

[1812] The server arranges the song based on the revised lyrics and melody. Specifically, it uses an AI arrangement tool to automatically select instruments and perform the arrangement.

[1813] Specific example: Automatically select instruments such as guitar, piano, and drums, and arrange them appropriately.

[1814] Input: Modified lyrics and melody

[1815] Output: Arranged song

[1816] Step 8:

[1817] The server then sends the final completed music data to the terminal. HTTP responses are used for data transmission, returning the completed music file to the terminal.

[1818] Input: Arranged song

[1819] Output: Music data (final version)

[1820] Step 9:

[1821] The device provides saving, playback, and download functions for delivering completed music data to the user. Specifically, it saves the music to the file system, provides a download link, and displays an interface with playback capabilities.

[1822] Specific example: Click the download link to save the song and play it with your preferred media player.

[1823] Input: Music data (final version)

[1824] Output: Music files accessible to the user

[1825] (Application Example 1)

[1826] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1827] When users generate music according to their own emotions and purposes, it is difficult for users without specialized knowledge or skills to easily customize and create individually tailored music. Furthermore, providing an environment where users can stream or download the generated music in real time is not easy. To solve these problems, a consistent service is needed for generating, modifying, arranging, and delivering music tailored to the user's emotions and purposes.

[1828] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1829] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; means for providing the arranged song data to the user; and means for implementing the song generation system as a smartphone application and making the generated song available for streaming or download. This allows users to easily generate songs that match their emotions and purposes and use them in real time.

[1830] "Emotions" refer to the user's own internal feelings and moods.

[1831] "Purpose" refers to the user's intended use in a specific situation or event.

[1832] "Text data" refers to string information entered by the user.

[1833] "Natural language processing technology" refers to the technology that enables computers to understand and analyze human language.

[1834] "Keywords" refer to important words or phrases extracted from text data.

[1835] An "important phrase" refers to a meaningful sequence of words extracted from text data.

[1836] "Generative artificial intelligence" refers to artificial intelligence that has the ability to generate new information based on given data.

[1837] "Lyrics" refers to the spoken parts of a song.

[1838] "Melody" refers to a sequence of sounds in a musical piece.

[1839] An "idea" refers to an initial concept or idea.

[1840] "Feedback" refers to reactions and opinions from users.

[1841] "Modification" refers to improving initial ideas or data.

[1842] "Arrangement" refers to the selection of instruments and arrangement of music for revised lyrics and melody.

[1843] "Song data" refers to the information of a completed song.

[1844] A "smartphone application" refers to a software program that runs on a smartphone.

[1845] "Streaming distribution" refers to a method of transmitting and playing data in real time.

[1846] "Downloading" refers to saving data to a local device via a network.

[1847] One embodiment of this invention is a system that allows users to generate music according to their own emotions and purposes, and then stream or download it in real time via a smartphone application.

[1848] System Program Overview

[1849] 1. Receiving and analyzing user input

[1850] Users input their emotions and goals as text data using a smartphone application. Keywords and important phrases are extracted from the text data using natural language processing (NLP) techniques. SpaCy and NLTK are suitable NLP libraries to use.

[1851] 2. The process of creating music

[1852] The analyzed keywords and important phrases are sent to the server. The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on this information. This initial generation process uses prompts such as the following:

[1853] "Please generate lyrics and melody for a fun birthday celebration song."

[1854] 3. Feedback and Corrections

[1855] The smartphone application presents the user with the initial generated lyrics and melody. The user reviews them and provides feedback. The server receives the feedback and uses generative artificial intelligence again to revise the lyrics and melody. The following prompts are used during the revision process.

[1856] "Please make this part more lively. Please revise it."

[1857] 4. Final arrangement and serving

[1858] Based on the revised lyrics and melody, the song is arranged. Automatic arrangement technology is used to select instruments and arrange the music, generating the final song data. The generated song data is provided to the user via a smartphone application. The user can stream or download this song in real time.

[1859] Implementation explanation

[1860] Hardware and software

[1861] Smartphone application: Used for user input, feedback collection, and presentation of generated music. Flutter and React Native are suitable development frameworks.

[1862] Server: Used for data processing and the execution of generative artificial intelligence. AWS and Google Cloud are suitable cloud platforms.

[1863] Generative artificial intelligence: Generates and modifies lyrics and melodies based on given data. OpenAI's GPT-4 is a suitable AI model to use.

[1864] Natural Language Processing (NLP) techniques: Used for analyzing user input. Suitable NLP libraries include spaCy and NLTK.

[1865] Adding specific examples

[1866] 1. Analysis of user input

[1867] The user inputs the emotion of "fun" and the objective of "birthday celebration songs." The application analyzes this input and extracts keywords and important phrases such as "fun" = positive and "birthday celebration songs" = celebration.

[1868] 2. Initial Generation and Presentation

[1869] Based on the extracted keywords, the following prompts are entered into the generative artificial intelligence.

[1870] "Please generate lyrics and melody for a fun birthday celebration song."

[1871] The initial generated lyrics and melody are presented to the user through the application.

[1872] 3. Feedback and Corrections

[1873] The user provides feedback saying, "Make this part feel more lively." The server then prompts the generative AI again with the following message.

[1874] "Please make this part more lively. Please revise it."

[1875] The revised lyrics and melody are presented to the user again.

[1876] 4. Final arrangement and serving

[1877] Based on the final feedback, the song is rearranged and the completed track is generated. The generated track can be streamed or downloaded in real time.

[1878] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1879] Step 1:

[1880] Receiving input from the user

[1881] The user uses a smartphone application to input their emotions (e.g., "happy") and purpose (e.g., "birthday song") in text format. This input data is then sent to the application.

[1882] Input: Text data describing emotions and objectives

[1883] Output: Text data is saved to the application.

[1884] Step 2:

[1885] Text data analysis

[1886] The terminal analyzes the received text data using natural language processing (NLP) techniques. Here, NLP libraries (e.g., spaCy, NLTK) are used to extract keywords and important phrases from the text data.

[1887] Input: User-input text data

[1888] Data processing: Extract keywords and important phrases using an NLP library.

[1889] Output: Extracted keywords and important phrases

[1890] Step 3:

[1891] Sending data to the server

[1892] The terminal sends keywords and important phrases, which are the results of the analysis, to the server.

[1893] Input: Extracted keywords and important phrases

[1894] Data processing: Packaging and sending analysis results

[1895] Output: Keywords and important phrases are sent to the server.

[1896] Step 4:

[1897] Early lyrics and melody generation

[1898] The server uses generative artificial intelligence (e.g., OpenAI's GPT-4) to generate initial lyrics and melodies based on received keywords and important phrases. Instructions are given to the AI ​​using prompts.

[1899] Input: Keywords and important phrases

[1900] Data processing: Enter a prompt message into the generation AI model (e.g., "Generate lyrics and melody for a fun birthday song").

[1901] Output: Generated initial lyrics and melody

[1902] Step 5:

[1903] Presentation of initial generation results and reception of feedback

[1904] The device presents the user with the initial generated lyrics and melody and collects user feedback. The user enters specific requests and suggestions for revisions in text format.

[1905] Input: Generated initial lyrics and melody

[1906] Data processing: Display to the user and receiving feedback.

[1907] Output: User feedback

[1908] Step 6:

[1909] Corrections based on feedback

[1910] Based on the feedback received from the user, the server uses generative artificial intelligence to revise the lyrics and melody again. Prompts are used for the revision process.

[1911] Input: User feedback

[1912] Data processing: Input prompt text into the generating AI model (e.g., "Make this part sound more lively. Please revise it.")

[1913] Output: Revised lyrics and melody

[1914] Step 7:

[1915] Execution of the final arrangement

[1916] The server uses automated arrangement technology to create the final arrangement of the song based on the revised lyrics and melody. Instrument selection and arrangement are performed automatically.

[1917] Input: Modified lyrics and melody

[1918] Data processing: Arrangement using automatic arrangement technology

[1919] Output: Arranged song data

[1920] Step 8:

[1921] Music data provision

[1922] The device provides the user with the completed music data. The user can then stream or download this music in real time.

[1923] Input: Arranged song data

[1924] Data processing: Determining how to deliver music data (streaming or download).

[1925] Output: Music data provided to the user

[1926] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1927] This invention is a system that can generate more accurate songs by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and purpose. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[1928] The introduction of an emotion engine will enable the recognition of emotions from the user's facial expressions and voice data, allowing for the provision of more personalized and tailored music.

[1929] User roles

[1930] User

[1931] 1. Users enter their emotions and the purpose of the song in text format using a dedicated application or web interface.

[1932] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[1933] 2. You input facial expressions and voice into the system, and the emotion engine analyzes them.

[1934] 3. Review the initial lyrics and melody provided and enter your feedback.

[1935] Example: Enter specific feedback such as, "Make this part feel more lively."

[1936] 4. Review the revised lyrics and melody again and provide further feedback if necessary.

[1937] 5. Download the completed song and use it for performances or recordings.

[1938] Terminal role

[1939] terminal

[1940] 1. Preprocessing

[1941] The system receives text data, facial expressions, and voice data from the user and analyzes them using an emotion engine. This analysis allows for a more accurate recognition of the user's emotions.

[1942] This analysis extracts keywords and important phrases based on emotions and purposes.

[1943] Example: Emotion = "Positive", Purpose = "Celebration"

[1944] 2. Data transmission

[1945] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[1946] 3. Interface

[1947] The system displays the server's response to the user and collects and sends feedback.

[1948] 4. Data Management

[1949] It provides saving, playback, and download functions to deliver completed music data to users.

[1950] Server Role

[1951] server

[1952] 1. Idea generation

[1953] Based on the analysis data sent from the terminal, a generative artificial intelligence is invoked to generate initial lyrics and melody ideas.

[1954] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1955] 2. Feedback Processing

[1956] We receive user feedback and use generative artificial intelligence again to revise the lyrics and melody.

[1957] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[1958] 3. Arrangement

[1959] Based on the revised lyrics and melody, the song is arranged. Specifically, instrument selection and arrangement are done automatically to improve the overall quality of the song.

[1960] 4. Data transmission

[1961] The final completed song data is sent to the device.

[1962] Specific example

[1963] For example, if a user inputs the emotion "happy" and the objective "birthday song," and also inputs facial expressions and voice, the system will operate as follows:

[1964] 1. The device analyzes the input text, facial expressions, and voice data, and the emotion engine recognizes the emotion as "happy." Based on this, keywords and important phrases are extracted and sent to the server.

[1965] 2. The server uses a generative AI to generate the lyrics "Today is your special day, a birthday filled with smiles..." and a cheerful melody.

[1966] 3. The device displays the initial lyrics and melody to the user, and the user provides feedback such as, "Make this part sound more lively."

[1967] 4. The server uses a generative AI again to revise the lyrics and melody to make them more lively.

[1968] 5. The server selects instruments, arranges the music, and completes the final song.

[1969] 6. The device provides the user with the completed music, allowing the user to download, save, play, record, and perform other operations.

[1970] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[1971] The following describes the processing flow.

[1972] Step 1:

[1973] User

[1974] You input your emotions and the purpose of the song in text format through a dedicated application or web interface.

[1975] Example: Emotion = "Happiness", Purpose = "Birthday song"

[1976] Step 2:

[1977] User

[1978] The system receives input of facial expressions and voice. A dedicated camera and microphone are used to record facial expressions and voice tone.

[1979] Step 3:

[1980] terminal

[1981] It receives text data, facial expression data, and voice data entered by the user.

[1982] The system uses an emotion engine to analyze input facial expression and voice data to recognize the user's emotions.

[1983] Example: Analysis result = "fun"

[1984] Step 4:

[1985] terminal

[1986] Using natural language processing techniques, the system analyzes input text data and extracts keywords and important phrases based on sentiment and purpose.

[1987] Example: Emotion = "Positive", Purpose = "Celebration"

[1988] Step 5:

[1989] terminal

[1990] The extracted keywords and important phrases, along with sentiment data analyzed by the sentiment engine, are sent to the server.

[1991] Step 6:

[1992] server

[1993] Based on the received keywords and important phrases, a generative artificial intelligence is invoked to generate initial lyric and melody ideas.

[1994] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line

[1995] Step 7:

[1996] server

[1997] Send the generated initial lyrics and melody to the device.

[1998] Step 8:

[1999] terminal

[2000] The initial lyrics and melody received from the server are presented to the user.

[2001] Step 9:

[2002] User

[2003] Review the provided lyrics and melody, and provide feedback as needed.

[2004] Example: Enter specific requests such as, "Make this part feel more lively."

[2005] Step 10:

[2006] terminal

[2007] Receive user feedback and send it to the server.

[2008] Step 11:

[2009] server

[2010] Based on the previous generation results and user feedback, the generative artificial intelligence is invoked again to revise the lyrics and melody.

[2011] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", lively melody line

[2012] Step 12:

[2013] server

[2014] The revised lyrics and melody are sent to the device and presented to the user again.

[2015] Step 13:

[2016] User

[2017] Review the revised lyrics and melody provided again, and provide further feedback if necessary. If there are no particular issues, proceed to the next step.

[2018] Step 14:

[2019] server

[2020] Based on the finalized lyrics and melody, the song is arranged. Specifically, this involves selecting instruments, arranging the music, and other processes to enhance its overall quality.

[2021] Step 15:

[2022] server

[2023] Send the completed song data to the device.

[2024] Step 16:

[2025] terminal

[2026] The completed music data will be provided to the user, allowing them to download, save, play, record, and perform other operations.

[2027] This processing flow allows users to highly customize specific songs to match their emotions and purposes through a system that combines an emotion engine and generative AI, enabling them to create more accurate and personalized songs.

[2028] (Example 2)

[2029] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2030] Conventional songwriting support systems struggled to provide sophisticated music that reflected specific emotions, even when generating songs based on the user's feelings and objectives. Furthermore, the process of revising songs while incorporating user feedback in real time was inefficient, and the automatic arrangement functions for improving song quality had limitations. As a result, there was a challenge in providing individually tailored songs.

[2031] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[2032] In this invention, the server includes means for receiving emotional and purpose text data entered by the user; means for analyzing the text data using natural language processing technology and extracting keywords and important phrases; means for using an emotion engine to analyze the text data and recognizing emotions from the user's facial expressions and voice data; means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyric and melody ideas to the user and receiving feedback from the user; means for revising the lyrics and melody based on the feedback using the generative artificial intelligence again; means for arranging the revised lyrics and melody into a song; and means for providing the arranged song data to the user. This enables sophisticated song generation tailored to the user's emotions and purpose, and allows for the provision of individually tailored songs with greater accuracy.

[2033] "User" refers to a person or user who wishes to create music using this system.

[2034] "Emotions" refer to data that represents the inner feelings and intentions expressed by the user, and are primarily input through text, facial expressions, and voice data.

[2035] "Purpose" refers to data indicating the intended use or intent of the music the user wishes to create, and is entered in text format.

[2036] "Text data" refers to strings of information that users input into a system, and it includes emotions and intentions.

[2037] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate human language.

[2038] "Keywords" are important words or phrases extracted from text data and used in the creation of music.

[2039] A "key phrase" is a meaningful sequence of words extracted from text data and used in the creation of a musical piece.

[2040] An "emotion engine" refers to software or hardware that analyzes and recognizes emotions from facial expressions and voice data entered by the user.

[2041] "Generative artificial intelligence" refers to an artificial intelligence system that automatically generates content (in this case, lyrics and melody) based on given input data.

[2042] "Initial lyrics and melody" refers to the basic text and musical melody that form the basis of the song first generated by the generative artificial intelligence.

[2043] "Feedback" refers to the opinions and requests that users input into a system for corrections and improvements.

[2044] "Revised lyrics and melody" refers to the text and musical melody regenerated by the generative artificial intelligence based on user feedback.

[2045] "Arrangement" refers to the process of selecting instruments and arranging the music based on the generated lyrics and melody to complete the song.

[2046] "Song data" refers to the digital information of the final produced song, which is provided to the user.

[2047] This invention is a system that generates more accurate music by combining an emotion engine with a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. This system analyzes the user's input data with the emotion engine, modifies the generated lyrics and melody based on the user's feedback, and finally completes and delivers the song.

[2048] Hardware and software configuration

[2049] User

[2050] Users interact with the system through a dedicated application or web interface. Devices such as computers, smartphones, and tablets are used.

[2051] terminal

[2052] The device will be equipped with an interface for receiving user input data (text, facial expressions, and voice).

[2053] Microsoft Azure Cognitive Services' Emotion API is used as the emotion engine.

[2054] Natural language processing techniques are used to analyze the data and extract keywords and important phrases.

[2055] It will be equipped with network communication capabilities for sending analysis data to a server.

[2056] server

[2057] The server will be equipped with a function to generate lyrics and melodies using generative artificial intelligence (e.g., OpenAI's GPT-3).

[2058] The server will be equipped with a function to receive feedback from users and then use generative artificial intelligence to revise the lyrics and melody again.

[2059] It has a function to select instruments and arrange the music based on the revised lyrics and melody, and to generate the final song data.

[2060] Program Processing Overview

[2061] 1. Input of user's emotions and purpose

[2062] Users enter their emotions and the purpose of the song in text format via a dedicated app or web interface.

[2063] Example: Enter the emotion "happy" and the objective "birthday celebration song".

[2064] 2. Input and analysis of facial expression and voice data

[2065] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[2066] The device recognizes emotions from the analysis results and extracts emotional data, such as "happy."

[2067] 3. Generation of initial lyrics and melody

[2068] The device extracts important keywords and phrases based on the extracted sentiment data and the user's input, and sends them to the server.

[2069] The server uses generative artificial intelligence to generate initial lyrics and melody.

[2070] Example: Lyrics "Today is your special day, a birthday filled with smiles...", cheerful melody line.

[2071] 4. Collecting user feedback

[2072] The device displays the initial generated lyrics and melody, and the user reviews the displayed content and provides feedback.

[2073] Example: Type "Make this part sound more lively."

[2074] 5. Revise the lyrics and melody.

[2075] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody based on that feedback.

[2076] Example: Revised lyrics "Today is your special day, everyone gathers to sing a celebratory song...", with a lively melody line.

[2077] 6. Arrangement and final generation of the music

[2078] Based on the revised lyrics and melody, the server selects instruments and arranges the music to improve its overall quality.

[2079] 7. Provision of the completed song

[2080] The server finally sends the completed song data to the terminal, which saves the completed song data and provides it to the user. The user can download, play, and record the song.

[2081] Example of a prompt

[2082] User: Fun

[2083] Purpose: Birthday celebration song

[2084] Feedback: "Make this part feel more lively."

[2085] In this way, this system generates music tailored to the user's emotions and purpose, and by combining this with analysis of facial expressions and voice, it can provide individually tailored music with greater accuracy.

[2086] The flow of the specific processing in Example 2 will be explained using Figure 13.

[2087] Step 1:

[2088] Users input emotions and the purpose of the song in text format through a dedicated application or web interface.

[2089] Input: The emotion and purpose of the song entered by the user (e.g., happy, birthday song)

[2090] Output: Input text data

[2091] Specific action: The user enters "fun" and "birthday song" into the application's input form and presses the submit button.

[2092] Step 2:

[2093] Users input facial expressions and voices into the system, and the emotion engine analyzes this information.

[2094] Input: User facial expression data and voice data

[2095] Output: Analyzed emotion data (e.g., happy)

[2096] Specific operation: The user uses the camera and microphone to input a smiling expression and the voice message "Happy Birthday." The emotion engine analyzes the expression and voice and recognizes it as "happy."

[2097] Step 3:

[2098] The device extracts keywords and important phrases based on text data and analyzed sentiment data, and sends them to the server.

[2099] Input: Emotional data, objective data, user facial expression and voice analysis results

[2100] Output: Extracted keywords and important phrases (e.g., special day, birthday, smile)

[2101] Specific operation: The terminal analyzes the input "fun" and "birthday song," as well as the "fun" from the emotion engine, extracts keywords such as "special day," "birthday," and "smile," and sends them to the server.

[2102] Step 4:

[2103] The server uses generative artificial intelligence to generate initial lyrics and melodies based on keywords and important phrases.

[2104] Input: Extracted keywords and important phrases

[2105] Output: Initial lyrics and melody

[2106] Specific operation: The server inputs the received words "special day," "birthday," and "smile" as prompts into the generative artificial intelligence, and generates lyrics such as "Today is your special day, a birthday overflowing with smiles..." along with a cheerful melody line.

[2107] Step 5:

[2108] The device displays the initial generated lyrics and melody and collects feedback from the user.

[2109] Input: Early lyrics and melody

[2110] Output: User feedback (e.g., "Make this part more lively")

[2111] Specific operation: The device displays the lyrics and melody "Today is your special day, a birthday filled with smiles..." to the user, and the user provides feedback such as "Make this part sound more lively."

[2112] Step 6:

[2113] The device sends user feedback to the server, which then uses generative artificial intelligence to revise the lyrics and melody again.

[2114] Input: User feedback, initial lyrics and melody

[2115] Output: Revised lyrics and melody

[2116] Specific operation: The device sends the feedback "Make this part sound more lively" to the server, and the server uses generative artificial intelligence to revise the lyrics to "Today is your special day, everyone gathers to sing a celebratory song..." and also changes the melody to a more lively one.

[2117] Step 7:

[2118] The server selects instruments and arranges the music based on the revised lyrics and melody, generating the final song data.

[2119] Input: Modified lyrics and melody

[2120] Output: Final music data

[2121] Specific operation: Based on the revised lyrics and melody, the server automatically selects instruments such as piano, guitar, and drums, arranges the instrumentation and creates the final song data.

[2122] Step 8:

[2123] The server ultimately sends the completed music data to the terminal, which then provides it to the user.

[2124] Input: Final song data

[2125] Output: Interface for downloading, playing, and recording music.

[2126] Specific operation: The server sends the final music data to the terminal, and the terminal displays a download link, play button, and recording function to provide the user with the music. The user can download, play, and record the music.

[2127] (Application Example 2)

[2128] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2129] Conventional music generation systems have been insufficient in generating music that aligns with users' emotions and purposes, making it difficult to provide music tailored to specific emotions or objectives. Furthermore, limited means of incorporating user feedback during the music generation process have resulted in low final music quality and user satisfaction. Additionally, the lack of emotion analysis utilizing user facial expressions and voice data has led to insufficient accuracy in emotion-based music generation.

[2130] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving emotional and purpose text data input by the user; means for analyzing the text data and image or audio data and recognizing emotions from the user's facial expressions and voice; means for analyzing using natural language processing technology and extracting keywords and important phrases; means for generating initial lyrics and melody ideas based on the keywords and important phrases using generative artificial intelligence; means for presenting the initial lyrics and melody ideas to the user and receiving feedback from the user; means for modifying the lyrics and melody again using the generative artificial intelligence based on the feedback; means for arranging the modified lyrics and melody into a song; and means for providing the arranged song data to the user. This enables highly accurate song generation that is tailored to the user's emotions and purpose.

[2131] "Emotional and purposeful text data" refers to text-based information entered by the user that represents the emotions and purpose of the song.

[2132] "Image or audio data" refers to input data that includes the user's facial expressions and voice.

[2133] "Means of recognizing emotions" refers to technologies that analyze image or audio data to identify the user's emotions.

[2134] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and intent.

[2135] "Keywords and important phrases" are the main words and expressions necessary for song generation, extracted from the analyzed text data.

[2136] "Generative artificial intelligence" is an artificial intelligence technology that automatically generates lyrics and melodies based on input data.

[2137] "Initial lyrics and melody ideas" refers to the initial proposals for lyrics and melodies generated by a generative artificial intelligence.

[2138] "Feedback" refers to opinions and requests for corrections and improvements provided by users.

[2139] "Revised lyrics and melody" refers to lyrics and melodies regenerated by a generative artificial intelligence system based on user feedback.

[2140] "Arranging a song" refers to the process of selecting instruments and arranging the music based on the revised lyrics and melody to complete the song.

[2141] "Arranged song data" refers to the digital data of the completed song.

[2142] This invention is a generative artificial intelligence system that assists in songwriting and composition based on the user's emotions and objectives. By incorporating an emotion engine, it can generate highly accurate and personalized songs. Specifically, it uses devices such as smartphones, smart glasses, and head-mounted displays to input the user's emotional data, and generates songs using a generative artificial intelligence system built on a server.

[2143] The server includes means for receiving emotional and target text data entered by the user, means for analyzing the user's image or voice data to recognize emotions from facial expressions and voice, and means for analyzing text data using natural language processing technology to extract keywords and important phrases. It also includes means for generating initial lyric and melody ideas based on the keywords and important phrases using generative artificial intelligence, means for presenting the initial lyric and melody ideas to the user and receiving feedback, means for revising the lyrics and melody using generative artificial intelligence again based on the feedback, means for arranging the revised lyrics and melody into a song, and means for providing the arranged song data to the user.

[2144] Users input their emotions and the purpose of the song in text format using a dedicated application or web interface. They also input facial expressions and voice, which the emotion engine analyzes. For example, if a user inputs the emotion "happy" and the purpose "birthday song," and also inputs facial expressions and voice into the system, the emotion engine will recognize the emotion as "happy" and extract keywords and important phrases.

[2145] The server uses generative artificial intelligence to generate initial lyrics such as "Today is your special day, a birthday filled with smiles..." and a cheerful melody, which are then presented to the user via the terminal. If the user provides feedback such as "Make this part more lively," the generative AI can revise the lyrics and melody again to make it more vibrant. Finally, the completed song undergoes instrument selection and arrangement, and the final song data is provided to the user.

[2146] As a concrete example, if a user logs into a virtual store, selects an emotion (e.g., "fun") and a purpose (e.g., "relaxing music"), and inputs facial expressions and voice into the app, the following prompts will be generated based on the emotion analysis results.

[2147] Example of a prompt:

[2148] Please create lyrics and a melody that express the emotion of "joy" and aim to create "relaxing music."

[2149] By sending this prompt to a generative artificial intelligence system, high-quality music tailored to the user's emotions and purpose can be automatically generated. Users can then play, purchase, and download this music within a virtual store.

[2150] As a result, it becomes possible to generate highly accurate music tailored to the user's emotions and purpose, and to provide personalized music.

[2151] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2152] Step 1:

[2153] Users input emotions and the purpose of the music in text format using a dedicated application or web interface. The entered text data of emotions and purpose is sent to the terminal and transferred to the server. The user's input (text data of emotions and purpose) serves as the starting point for analysis.

[2154] Step 2:

[2155] The device receives text data and image and audio data (facial expressions and voice) entered by the user and analyzes them using an emotion engine. Specifically, it analyzes the user's facial expressions and voice data to recognize emotions, and extracts keywords and important phrases based on this analysis. The input for this step is facial expressions and voice data, and the output is the emotion recognition result and extracted keywords.

[2156] Step 3:

[2157] The server receives the analysis results, extracted keywords, and important phrases sent from the terminal, and uses a generative artificial intelligence model to generate initial lyric and melody ideas based on them. This generation process involves creating prompt sentences using the generative AI model and generating lyrics and melodies based on them. The input is the analysis results and keywords, and the output is the initial lyric and melody drafts.

[2158] Step 4:

[2159] The server presents the user with initial lyrics and melody ideas via a terminal. The user reviews these and provides specific feedback, such as "Make this part sound more lively." In this step, the user's input is feedback, and the output is a revision instruction that reflects that feedback.

[2160] Step 5:

[2161] The server receives user feedback and uses a generative artificial intelligence model to revise the lyrics and melody again. The generative AI regenerates prompt sentences based on the feedback and produces the revised lyrics and melody. The input is the feedback, and the output is the revised lyrics and melody.

[2162] Step 6:

[2163] The server then arranges the revised lyrics and melody into a song. Specifically, it automatically selects instruments and arranges the music to improve its overall quality. The input for this step is the revised lyrics and melody, and the output is the final arranged song data.

[2164] Step 7:

[2165] The server sends the final arranged music data to the terminal and provides it to the user. The user can play, purchase, and download this music data. The input is the final arranged music data, and the output is the data provided to the user.

[2166] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2167] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2168] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[2169] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2170] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[2171] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[2172] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[2173] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[2174] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[2175] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[2176] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[2177] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[2178] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[2179] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2180] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[2181] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[2182] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[2183] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[2184] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[2185] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[2186] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[2187] The following is further disclosed regarding the embodiments described above.

[2188] (Claim 1)

[2189] A means for receiving emotional and purposeful text data entered by the user,

[2190] A means for analyzing the aforementioned text data using natural language processing technology and extracting keywords and important phrases,

[2191] A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence,

[2192] A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user,

[2193] A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback,

[2194] A means of arranging the aforementioned revised lyrics and melody into a musical piece,

[2195] A means for providing the aforementioned arranged music data to the user,

[2196] A system that includes this.

[2197] (Claim 2)

[2198] The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

[2199] (Claim 3)

[2200] The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests.

[2201] "Example 1"

[2202] (Claim 1)

[2203] A means for receiving emotional and purposeful text data entered by the user,

[2204] A means for analyzing the aforementioned text data using natural language processing technology and extracting keywords and important phrases,

[2205] A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence,

[2206] A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user,

[2207] A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback,

[2208] A means of arranging the aforementioned revised lyrics and melody into a musical piece,

[2209] A means for automatically selecting appropriate instruments and arranging the aforementioned revised lyrics and melody,

[2210] A means for providing the aforementioned arranged music data to the user,

[2211] A system that includes this.

[2212] (Claim 2)

[2213] The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

[2214] (Claim 3)

[2215] The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests.

[2216] "Application Example 1"

[2217] (Claim 1)

[2218] A means for receiving emotional and purposeful text data entered by the user,

[2219] A means for analyzing the aforementioned text data using natural language processing technology and extracting keywords and important phrases,

[2220] A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence,

[2221] A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user,

[2222] A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback,

[2223] A means of arranging the aforementioned revised lyrics and melody into a musical piece,

[2224] A means for providing the aforementioned arranged music data to the user,

[2225] The aforementioned music generation system is implemented as a smartphone application, and means are provided for making the generated music available for streaming or download.

[2226] A system that includes this.

[2227] (Claim 2)

[2228] The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

[2229] (Claim 3)

[2230] The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests.

[2231] "Example 2 of combining an emotion engine"

[2232] (Claim 1)

[2233] A means for receiving emotional and purposeful text data entered by the user,

[2234] A means for analyzing the aforementioned text data using natural language processing technology and extracting keywords and important phrases,

[2235] The aforementioned text data is analyzed using an emotion engine, and a means of recognizing emotions from the user's facial expressions and voice data is provided.

[2236] A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence,

[2237] A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user,

[2238] A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback,

[2239] A means of arranging the aforementioned revised lyrics and melody into a musical piece,

[2240] A means for providing the aforementioned arranged music data to the user,

[2241] A system that includes this.

[2242] (Claim 2)

[2243] The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

[2244] (Claim 3)

[2245] The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests.

[2246] "Application example 2 when combining with an emotional engine"

[2247] (Claim 1)

[2248] A means for receiving emotional and purposeful text data entered by the user,

[2249] A means for analyzing the aforementioned text data and image or audio data, and recognizing emotions from the user's facial expressions and voice,

[2250] A means for analyzing using natural language processing technology and extracting keywords and important phrases,

[2251] A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence,

[2252] A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user,

[2253] A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback,

[2254] A means of arranging the aforementioned revised lyrics and melody into a musical piece,

[2255] A means for providing the aforementioned arranged music data to the user,

[2256] A system that includes this.

[2257] (Claim 2)

[2258] The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

[2259] (Claim 3)

[2260] The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests. [Explanation of symbols]

[2261] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving emotional and purposeful text data entered by the user, A means for analyzing the aforementioned text data using natural language processing technology and extracting keywords and important phrases, A means for generating initial lyrics and melody ideas based on the aforementioned keywords and important phrases using generative artificial intelligence, A means for presenting the aforementioned initial lyrics and melody ideas to the user and receiving feedback from the user, A means for modifying the lyrics and melody using the generative artificial intelligence again based on the aforementioned feedback, A means of arranging the aforementioned revised lyrics and melody into a musical piece, A means for providing the aforementioned arranged music data to the user, A system that includes this.

2. The system according to claim 1, wherein the lyrics and melody included in the aforementioned music data are modified after receiving user feedback multiple times.

3. The system according to claim 1, wherein the arrangement of the aforementioned song is adjusted based on the user's specific requests.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A