System

The system addresses the lack of inspiration in music production by using a database, generative AI, and emotion engine to provide personalized music theory-based suggestions, enhancing creative efficiency.

JP2026022521APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124038
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Artists and creators often face a lack of inspiration and difficulty accessing new ideas during music production, hindered by the lack of efficient systems that can provide music theory-based suggestions.

Method used

A system that collects music data from a database, analyzes it based on music theory, receives user requests, generates suggestions using a generative AI model, and displays these suggestions to users, incorporating an emotion engine to personalize the recommendations based on user emotions.

Benefits of technology

Enables artists to efficiently obtain new musical ideas and inspiration by providing personalized, emotion-based suggestions, facilitating smoother creative activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022521000001_ABST
    Figure 2026022521000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting music from a music database and analyzing the music based on music theory; means for receiving a request for music production from a user and analyzing the request; means for generating suggestions based on music theory using a generative AI model; and means for displaying the generated suggestions to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In the process of music production, artists and creators often face a lack of inspiration and difficulty accessing new ideas. These creative barriers and challenges hinder efficient music production. The purpose of this invention is to solve these problems and support artists in creating music more efficiently and creatively. [Means for solving the problem]

[0005] The present invention solves the above problems by using the following system.

[0006] The system includes a means for collecting music data from a music database and analyzing it based on music theory, a means for receiving music production requests from users and analyzing the requests, a means for generating suggestions based on music theory using a generative AI model, and a means for displaying the generated suggestions to the user. This system allows artists to gain new ideas and inspiration even when they hit a creative wall.

[0007] A "music database" is a database that stores and manages a large amount of information about songs, including elements such as melody, harmony, rhythm patterns, and lyrics.

[0008] "Music data" refers to the components of a song, including melody, harmony, rhythm patterns, lyrics, etc.

[0009] "Music theory" is an academic field that scientifically studies and explains the rules and laws regarding musical composition and progression, including the laws of melody, harmony, and rhythm.

[0010] "Means of analysis" refers to processing techniques for analyzing collected music data based on music theory and extracting useful information.

[0011] "User" refers to artists and creators who wish to create music using this music creation support system.

[0012] "Means for receiving and parsing requests" refers to the technology and processes used to receive music-making requests from users and translate them into an understandable format.

[0013] A "generative AI model" is a model trained using artificial intelligence technology, which generates suggestions by applying music theory based on large amounts of music data.

[0014] "Means for generating suggestions" refers to technology that uses a generative AI model to create specific musical suggestions in response to a user's request.

[0015] "Means for displaying suggestions to the user" refers to the interface or technology for visually and audibly displaying the generated musical suggestions to the user. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention relates to a music creation support system that solves the problems artists and creators face when creating music, such as a lack of inspiration and new ideas. The system of the present invention includes a processing means for collecting music data from a music database and analyzing it based on music theory, a processing means for receiving requests for music creation from users and analyzing the requests, a processing means for generating suggestions based on music theory using a generative AI model, and a processing means for displaying the generated suggestions to the user.

[0038] Explanation of program processing

[0039] Music data collection and analysis

[0040] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0041] Receiving a user request

[0042] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user may input a request such as "I would like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0043] Proposal generation

[0044] When the server receives a user's request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord.

[0045] View Suggestions

[0046] The terminal displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production.

[0047] Specific examples

[0048] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0049] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0050] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0051] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0052] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0053] This system allows users to efficiently obtain new musical ideas and smoothly advance their creative activities.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, and lyrics, and stores this data in a local database for analysis.

[0057] Step 2:

[0058] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the AI ​​model.

[0059] Step 3:

[0060] The terminal displays an interface for the user to input requests for music production. For example, the user may input, "I would like some new harmony suggestions."

[0061] Step 4:

[0062] The terminal receives a request from the user, analyzes the content of the request, and sends the analysis results to the server.

[0063] Step 5:

[0064] The server analyzes the request received from the device and, based on the request content, selects the appropriate process to use the generative AI model.

[0065] Step 6:

[0066] The server uses a generative AI model to generate musical suggestions based on the user's request, such as a C major seventh chord if asked for harmony suggestions.

[0067] Step 7:

[0068] The server sends the generated musical suggestions to the device, including suggestions derived from the generative AI model.

[0069] Step 8:

[0070] The terminal displays the suggestions received from the server to the user using an interface that makes it easy to provide the suggestions visually or audibly to the user.

[0071] Step 9:

[0072] Users can refer to the suggestions displayed on the device to proceed with music creation, which will help them gain new ideas and inspiration and make music creation easier.

[0073] Example 1

[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0075] In modern music creation, artists and creators often face a lack of inspiration and new ideas. Furthermore, music production requires advanced knowledge of music theory, making it difficult for users without specialized knowledge to create music efficiently. A system that solves this problem and allows users to quickly and easily obtain musical ideas is needed.

[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0077] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for displaying the generated suggestions to the user, means for the generative AI model to generate specific musical ideas based on music theory based on the input requests, and means for presenting the generated suggestions to the user's terminal in real time, thereby enabling users to efficiently obtain new musical ideas.

[0078] A "music database" is a system for managing and storing a large amount of music data in one place.

[0079] "Music data" refers to information about the elements that make up a musical work, such as melody, harmony, rhythm patterns, and lyrics.

[0080] "Music theory" is the study of the basic principles of musical structure, chords, rhythm, melody, and their combinations and progressions.

[0081] A "server" is a computer system that runs behind the scenes of the entire system, handling requests, storing data, and delivering generated offers.

[0082] "User" refers to an individual or professional who uses this system to create music.

[0083] A "request" is a user's request for a specific musical suggestion or creation from the system.

[0084] A "generative AI model" is an artificial intelligence algorithm that generates new information or suggestions based on input data.

[0085] A "specific musical idea" is a specific creative element in musical composition, such as a particular melody, harmony, rhythmic pattern, or lyrics.

[0086] A "terminal" is a device, such as a computer or smartphone, that a user uses to interact with the system.

[0087] "Real-time" refers to a state in which user input and system processing are reflected immediately without delay.

[0088] The present invention relates to a music creation support system, and specific embodiments thereof will be described below.

[0089] Music data collection and analysis

[0090] The server collects music data from a music database on the cloud. The specific software used for this is a tool that extracts data through an API. The data includes elements such as melody, harmony, rhythm patterns, and lyrics. The server then uses a Python music analysis library (e.g., music21) to analyze the collected music data based on music theory. The analysis results are stored in a database system such as MySQL or PostgreSQL. This accumulates data for training the generative AI model.

[0091] Receiving a user request

[0092] The device provides the user with an interface for inputting requests related to music production. This interface is built using a front-end framework (e.g., React, Vue.js). The user inputs requests through this interface, specifically prompt sentences such as "I would like new harmony suggestions." The input request is sent to the server in JSON format.

[0093] Proposal generation

[0094] The server analyzes the received user request and generates suggestions based on music theory via a generative AI model (e.g., GPT-3). The generative AI model has the ability to generate new musical ideas based on training data. For example, in response to a request for "new harmony suggestions," it generates specific harmonies such as a C major seventh chord. The suggestions are returned to the device in JSON format.

[0095] View Suggestions

[0096] The device displays the suggestions received from the server to the user. Specifically, it uses JavaScript to render the suggestions into an HTML view and displays it in the user interface. This allows the user to use the generated suggestions as a reference for music production.

[0097] Specific examples

[0098] For example, when a user inputs a request such as "I want new harmony suggestions" into a terminal, the following is an example of the series of processes that will be carried out.

[0099] 1. Device: The user types "I want new harmony suggestions" and presses the send button. The device sends the request in JSON format to the server.

[0100] 2. Server: The server receives the request, analyzes it, and inputs the prompt "I would like some new harmony suggestions" into the generative AI model.

[0101] 3. Server: The generative AI model generates new harmonies, such as a C major seventh chord, and returns suggestions in JSON format to the device.

[0102] 4. Terminal: The terminal receives the JSON data and displays the suggestions in the user interface.

[0103] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0104] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0106] Step 1: Collecting music data

[0107] The server collects music data from a music database. It uses an API to obtain data such as melody, harmony, rhythm patterns, and lyrics. For example, it calls the API of a cloud service and receives music data in JSON format. The input is the music data from the API, and the output is music data stored in temporary data storage on the server.

[0108] Step 2: Analyzing the music data

[0109] The server analyzes the collected music data based on music theory. Specifically, it uses a Python music analysis library (e.g., music21) to analyze harmonic structure, rhythmic patterns, etc. The analysis results are stored in a database (e.g., MySQL, PostgreSQL) as training data. The input is the music data stored in temporary data storage, and the output is the analysis results stored in the database.

[0110] Step 3: Entering User Requests

[0111] Users input their music-making requests through a terminal interface. This interface is built using a front-end framework (e.g., React, Vue.js). Specifically, they input a prompt such as "I'd like some new harmony suggestions." The input is the user's request, and the output is data sent to the server in JSON format.

[0112] Step 4: Receiving and Parsing the Request

[0113] The server receives the request sent from the device. It then analyzes the request using a web framework (e.g., Flask, Django). Based on the analysis results, it creates a prompt to input to the generative AI model. The input is the request JSON sent from the device, and the output is the prompt to input to the generative AI model.

[0114] Step 5: Proposal Generation

[0115] The server uses a generative AI model (e.g., GPT-3) to generate suggestions based on music theory. The generative AI model generates new melodies and harmonies based on the training data. For example, it generates a specific harmony such as a C major seventh chord. The input is a prompt to the generative AI model, and the output is the generated musical suggestions.

[0116] Step 6: Submit your proposal

[0117] The server sends the generated proposals to the terminal in JSON format. The input is the proposal data from the generative AI model, and the output is the JSON data sent to the terminal.

[0118] Step 7: Viewing Proposals

[0119] The terminal displays the suggestions received from the server to the user by using JavaScript to render the suggestions into an HTML view and display it in the user interface. The input is the JSON data received from the server and the output is the suggestions displayed to the user.

[0120] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0121] (Application example 1)

[0122] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0123] Conventional music production support systems have the problem of being unable to effectively solve the problem of running out of inspiration and lack of new ideas. In particular, it has been difficult to instantly provide new harmony, melody, and lyric ideas while creating music. Furthermore, because such systems are not linked to content distribution services, they have been unable to quickly provide the latest musical ideas to a wide range of users. To solve these problems, a more efficient and practical music creation support system for users is needed.

[0124] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0125] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, and means for applying the suggestions to the content distribution service, thereby enabling users to quickly and efficiently obtain new musical ideas.

[0126] A "music database" is a database that stores multiple pieces of music data and allows searching and retrieval.

[0127] "Music data" refers to the elements that make up a song, such as melody, harmony, rhythm, and lyrics.

[0128] "Music theory" is the academic knowledge of musical structure and composition, which serves as a guide for analyzing and creating musical pieces.

[0129] A "generative AI model" is a technology that uses artificial intelligence to generate new music suggestions based on user requests.

[0130] A "request" is a specific request for music production, such as a new harmony, melody, or lyrics that the user desires.

[0131] A "content distribution service" is a service that provides users with digital content such as music and videos via the Internet.

[0132] "Suggestions" are new musical ideas generated by the generative AI model to assist users in their music creation.

[0133] "Analysis" is the process of understanding a user request and analyzing its content to generate appropriate music suggestions.

[0134] This invention is a music creation support system that allows artists and creators to obtain new ideas in music production. The system includes the following means.

[0135] System Configuration

[0136] 1. A server that collects music data from a music database and analyzes it based on music theory

[0137] The server resides in the cloud and collects multiple pieces of music data (melody, harmony, rhythm, lyrics, etc.) from a music database. The collected data is analyzed based on music theory and used as training data for the generative AI model.

[0138] 2. A device that receives requests from users regarding music production and analyzes those requests.

[0139] The device provides an interface for users to input music-making requests, such as "I want new harmony suggestions." The device receives the request and sends it to the server for analysis.

[0140] 3. A server that uses generative AI models to generate suggestions based on music theory

[0141] The server analyzes the user's request and generates music theory-based suggestions based on the request using a generative AI model (e.g., GPTNeo). For example, if the user's request is for a new harmony, the server generates a specific harmony, such as a C major seventh chord.

[0142] 4. A device that displays the generated suggestions to the user.

[0143] The generated suggestions are sent from the server to the device, which then displays them to the user, who can use them as a reference to proceed with their music creation.

[0144] 5. Methods applied to content distribution services

[0145] This system is particularly linked to content distribution services, making it possible to quickly provide musical ideas to a wide range of users via the Internet, allowing users to efficiently acquire new musical ideas and support their creative activities.

[0146] Hardware and software used

[0147] Hardware

[0148] Server: Cloud server

[0149] Device: Smartphone

[0150] software

[0151] Server: Python, requests library

[0152] Generative AI model: Hugging Face transformers library, GPTNeo model

[0153] Device: Smartphone application

[0154] Specific examples

[0155] For example, if a user inputs a request such as "I want new harmony suggestions" through a smartphone application, the following processing will occur.

[0156] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0157] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0158] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0159] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0160] Prompt Sentence Examples

[0161] "Users are asking for suggestions: I'd like some new harmony suggestions."

[0162] This system allows users to efficiently eliminate inspiration drain and easily come up with new musical ideas.

[0163] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0164] Step 1:

[0165] The user uses the terminal to enter a request

[0166] A user opens a smartphone application and enters a music-making request, such as "I'd like some new harmony suggestions." The device receives this request and converts it into data to send to the server. The input is the user's request, and the output is a data packet containing that request.

[0167] Step 2:

[0168] The device sends a request to the server

[0169] The terminal sends the request received from the user to the server as an HTTP POST request. This request includes the user ID and the request content. The input is the request data from the user, and the output is the HTTP POST request sent to the server.

[0170] Step 3:

[0171] The server receives and parses the request

[0172] The server receives requests sent from the device and analyzes their contents. The analysis involves analyzing the text of the request, for example, creating a prompt appropriate for the generative AI model based on a request such as "I'd like some new harmony suggestions." The input is the HTTP POST request sent to the server, and the output is the analysis result, including the prompt text.

[0173] Step 4:

[0174] The server generates suggestions using a generative AI model

[0175] The server inputs the analysis results into a generative AI model to generate a suggestion based on music theory. For example, a generative AI model (GPTNeo) is used to generate a new suggestion for the requested harmony (e.g., a C major seventh chord). The input is a prompt based on the analysis results, and the output is a musical suggestion generated by the generative AI model.

[0176] Step 5:

[0177] The server sends the generated proposal to the device.

[0178] The server sends the generated music suggestions to the device, where the suggestions are sent as an HTTP response and formatted appropriately for the user. The input is the music suggestions generated by the AI ​​model, and the output is the HTTP response sent to the device.

[0179] Step 6:

[0180] The terminal displays the suggestions to the user.

[0181] The device receives the response from the server and displays it to the user in an appropriate format. For example, "New harmony suggestion: C major seventh chord" is displayed on the screen. The input is the HTTP response from the server, and the output is the suggestion information displayed to the user.

[0182] Step 7:

[0183] The user will use the suggestions to proceed with music production.

[0184] The user can refer to the new music suggestions displayed on the device and proceed with the composition of the music. In this process, the user can gain new inspiration and ideas. The input is the suggested information displayed on the device, and the output is the user's creative activity.

[0185] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0186] The present invention further relates to a music creation support system that combines an emotion engine that recognizes user emotions. The system includes processing means for collecting music data from a music database and analyzing it based on music theory, processing means for receiving music creation requests from users and analyzing the requests, processing means for generating suggestions based on music theory using a generative AI model, processing means for displaying the generated suggestions to the user, and an emotion engine that recognizes user emotions.

[0187] Explanation of program processing

[0188] Music data collection and analysis

[0189] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0190] Receiving a user request

[0191] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user might input, "I'd like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0192] Proposal generation

[0193] When the server receives a user request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord. It also uses an emotion engine to generate suggestions that correspond to the user's emotions.

[0194] Emotion Engine Functions

[0195] The device uses an emotion engine to recognize the user's emotions. Emotion recognition utilizes the user's facial expressions, tone of voice, and input text content. The emotion engine analyzes this information to determine the user's current emotional state.

[0196] View Suggestions

[0197] The device displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production. In particular, the emotion engine provides suggestions that correspond to the user's emotions, allowing the user to obtain more appropriate musical inspiration.

[0198] Specific examples

[0199] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0200] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server. At the same time, the emotion engine analyzes the user's facial expressions to recognize their emotions.

[0201] 2. Server: The server receives the request and analyzes it. It then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested. At the same time, emotional information from the emotion engine is taken into account.

[0202] 3. Server: Based on the information from the emotion engine, for example, if the user is relaxed, it suggests harmonies with a more relaxed atmosphere.

[0203] 4. Terminal: The terminal displays suggestions received from the server, such as "C major seventh," to the user. It also displays the analysis results of the emotion engine, making it easier for the user to understand the background of the suggestions.

[0204] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0205] The system allows users to efficiently discover new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activities.

[0206] The processing flow will be explained below.

[0207] Step 1:

[0208] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, lyrics, etc. The server then stores this data in a local database.

[0209] Step 2:

[0210] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the generative AI model.

[0211] Step 3:

[0212] The terminal provides the user with an interface for inputting requests related to music production. For example, the user may input, "I would like some new harmony suggestions."

[0213] Step 4:

[0214] The terminal receives the user's request, analyzes its contents, and sends the analysis results to the server.

[0215] Step 5:

[0216] The device uses an emotion engine to recognize the user's emotions by analyzing the user's facial expressions, tone of voice, input text, etc.

[0217] Step 6:

[0218] The server analyzes the requests received from the terminal and the emotional information provided by the emotion engine, and takes this information into account when it needs to generate musical suggestions based on the emotional state.

[0219] Step 7:

[0220] The server uses a generative AI model to generate appropriate suggestions based on music theory, for example, generating specific harmonies such as a C major seventh chord if a new harmony suggestion is requested.

[0221] Step 8:

[0222] The server generates suggestions based on the user's emotional information obtained from the emotion engine: if the user is relaxed, it will suggest harmony with a relaxing atmosphere.

[0223] Step 9:

[0224] The server sends generated musical suggestions to the device, including suggestions derived from the generative AI model and emotion-based suggestions.

[0225] Step 10:

[0226] The terminal displays the proposal received from the server to the user using an interface that presents the proposal to the user visually or audibly.

[0227] Step 11:

[0228] The user can proceed with music production by referring to the suggestions displayed on the device. By also referring to emotion-based suggestions, music production can proceed more smoothly.

[0229] Example 2

[0230] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0231] Conventional music production support systems do not take into account the user's emotional state, making it difficult to provide personalized music suggestions that reflect the user's emotions. Furthermore, their ability to automatically generate suggestions based on music theory is limited. This makes it difficult for users to gain inspiration efficiently, potentially stagnateing their creative activities.

[0232] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0233] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for using an emotion engine to recognize the user's emotions, and means for displaying the generated suggestions to the user, thereby enabling personalized music suggestions to be provided according to the user's emotional state.

[0234] A "music database" is a type of data storage that collects data related to songs, including musical elements such as melody, harmony, rhythmic patterns, and lyrics.

[0235] "Music theory" is the study of rules and principles related to musical structure, form, harmony, melody, rhythm, etc. It is the basis for analyzing music and generating suggestions.

[0236] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate suggestions for music, text, etc., particularly natural language processing and machine learning techniques.

[0237] The "emotion engine" is a system for recognizing the user's emotional state by analyzing facial expressions, tone of voice, and the content of input text.

[0238] The "server" is a central control unit that processes and stores data and provides services to client systems. In this invention, it analyzes music data and processes user requests.

[0239] "Terminal" refers to a device that allows a user to access and operate the system through an interface. This includes computers, smartphones, tablets, etc.

[0240] A "user request" is a request a user makes to the system regarding a specific piece of music production. An example would be "I'd like some new harmony suggestions."

[0241] "Suggestions" refer to musical elements and ideas that the system provides to the user using generative AI models and emotion engines, which are used as references for music production.

[0242] MODE FOR CARRYING OUT THE INVENTION

[0243] The present invention provides a system for supporting music production by recognizing the emotional state of a user. Specific embodiments for carrying out the present invention will be described below.

[0244] Music data collection and analysis

[0245] The server collects music data from a specific music database (e.g., a music library API) on the cloud. The collected data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes the acquired music data using a Python music analysis library (e.g., music21, librosa), and stores the analysis results in a database. This accumulates data for training the generative AI model.

[0246] Receiving a user request

[0247] The device provides an interface through which users can input requests for music creation via a web application (e.g., React, Vue.js). For example, a user might input, "I'd like some new harmony suggestions." The device then sends the input request to the server using WebSocket or an HTTP request.

[0248] Proposal Generation

[0249] When the server receives a user request, it first analyzes the request using an NLP library (e.g., spaCy, NLTK). It then uses a generative AI model (e.g., GPT-3) to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, it generates specific chords, such as a C major seventh chord. It also incorporates information from an emotion engine to take the user's emotional state into account.

[0250] Emotion Engine Functions

[0251] The device uses a camera module (e.g., a webcam) to capture the user's facial expressions and simultaneously accepts voice input. The device then analyzes the user's emotions using emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer). The analysis results are sent to the server in real time, and the server uses this information to determine the user's current emotional state.

[0252] View Suggestions

[0253] The server sends the generated music suggestions and the analysis results from the emotion engine to the device. The device prepares UI components (e.g., HTML, CSS, JavaScript) to display the received suggestions to the user. The user then proceeds with the music creation process based on the displayed suggestions.

[0254] Specific examples

[0255] For example, if a user inputs a request to the terminal saying, "I want new harmony suggestions," the following series of processes are carried out.

[0256] 1. Device: The user types "I want new harmony suggestions" and clicks the send button. At the same time, the camera captures facial expressions and the microphone captures audio, which are then sent to the emotion engine.

[0257] Example prompt: "I'd like some new harmony suggestions for normal times."

[0258] 2. Server: Receives the request and analyzes it using an NLP library (e.g., spaCy). Uses a generative AI model (e.g., GPT-3) to generate specific harmony suggestions, such as "C major seventh chord." Takes emotion recognition results into account to suggest a relaxing atmosphere.

[0259] Example prompt: "I'd like some new harmony suggestions for when I'm relaxing."

[0260] 3. If the emotion engine determines the user's emotion as "relaxed," the generative AI model will adjust accordingly.

[0261] 4. Terminal: Displays to the user a suggestion that combines "C major seventh" received from the server with the results of sentiment analysis.

[0262] 5. The user then proceeds with creating music that corresponds to a specific emotion based on the displayed suggestions.

[0263] The system not only helps users find musical inspiration quickly, but also streamlines their creative process by receiving personalized suggestions based on their emotional state.

[0264] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0265] Step 1: Collecting music data

[0266] The server accesses a music database on the cloud. Specifically, it retrieves music data such as melody, harmony, rhythm patterns, and lyrics through an API (e.g., a music library API). This data is stored on the server and prepared for the next analysis step.

[0267] Input: Cloud music database

[0268] Output: Music data stored on the server

[0269] ---

[0270] Step 2: Analyzing the music data

[0271] The server analyzes the stored music data using Python music analysis libraries (e.g., music21, librosa). Specifically, it extracts the melody, harmony, rhythmic patterns, and lyric characteristics of each song. The analysis results are stored in a database in JSON or CSV format and used to train the generative AI model.

[0272] Input: Music data stored on the server

[0273] Output: Analyzed music data

[0274] ---

[0275] Step 3: Receiving a user request

[0276] The device provides an interface using a web application (e.g., React, Vue.js). The user inputs a request for music creation (e.g., "I'd like some new harmony suggestions.") The device receives this request and sends it to the server via WebSocket or HTTP request.

[0277] Input: User-entered music production requests

[0278] Output: Request sent to the server

[0279] ---

[0280] Step 4: Parsing the request

[0281] When the server receives a user request, it parses it using an NLP library (e.g., spaCy, NLTK) to understand the content of the request and determine what musical suggestions would be appropriate.

[0282] Input: The user request sent to the server

[0283] Output: Parsed request content

[0284] ---

[0285] Step 5: Generate proposals

[0286] The server uses a generative AI model (e.g., GPT-3) to generate music theory-based suggestions based on the analyzed request. For example, in response to a request for "new harmony suggestions," the server generates specific chords (e.g., a C major seventh chord).

[0287] Input: Analyzed request content and music theory data

[0288] Output: Generated music suggestions

[0289] ---

[0290] Step 6: Perform emotion recognition

[0291] The device uses a camera module (e.g., webcam) to capture the user's facial expressions and a microphone to capture their voice. This information is then sent to emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer) to analyze the user's emotions. The analysis results are then sent to a server.

[0292] Input: User's facial expressions, voice data

[0293] Output: Parsed emotion data

[0294] ---

[0295] Step 7: Integrating Emotional Data

[0296] The server integrates emotional data from the emotion engine into the music suggestions generated by the generative AI model. For example, if the server determines that the user is relaxed, it generates music suggestions that correspond to that emotion (e.g., relaxing harmony).

[0297] Input: Generated music suggestions, parsed emotion data

[0298] Output: Emotion-based music suggestions

[0299] ---

[0300] Step 8: Viewing Proposals

[0301] The device displays music suggestions to the user based on the emotion received from the server. Specifically, the suggestions are visually displayed using UI components (e.g., HTML, CSS, JavaScript).

[0302] Input: Emotion-based music suggestions received from the server

[0303] Output: Music suggestions displayed to the user

[0304] ---

[0305] Step 9: Music Production Progression

[0306] The user can then proceed with the composition of the music based on the displayed musical suggestions. For example, they can create a piece of music incorporating a C major seventh chord. At the same time, they can input new requests as needed, which are reflected in the system.

[0307] Input: Music suggestions displayed to the user

[0308] Output: Finished song, next request

[0309] ---

[0310] This process flow provides personalized music suggestions based on the user's emotional state, improving the efficiency of music production.

[0311] (Application example 2)

[0312] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0313] Conventional music creation support systems generate suggestions without considering the user's emotional state, which prevents personalized suggestions for each individual user and can lead to lower user satisfaction. Furthermore, users often find it difficult to find musical inspiration that matches their emotional state, limiting their creativity. A system that can resolve these issues and provide more effective and personalized music creation support is needed.

[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0315] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, emotion analysis means for recognizing the user's emotions, and means for displaying the generated suggestions to the user. This enables personalized musical suggestions based on the user's emotional state, thereby promoting creativity.

[0316] A "music database" is a large-scale information storage device for storing, saving, and managing music data.

[0317] "Music data" refers to information related to musical components such as melody, harmony, rhythm patterns, and lyrics.

[0318] "Music theory" is a set of rules and principles for understanding and analysing musical structures and components.

[0319] A "user request" is an instruction including a request or wish given by a user to a system.

[0320] A "generative AI model" is an artificial intelligence model that generates appropriate suggestions or answers based on given data or requests.

[0321] "Emotion analysis means" refers to technology or devices for recognizing and analyzing the user's emotional state.

[0322] "Suggestions" are specific ideas and advice about music production generated by the system.

[0323] The "means for displaying to the user" is an interface or device for visually presenting the generated suggestions to the user.

[0324] The present invention is a music creation support system that is combined with emotion analysis means for recognizing the emotions of the user. To specifically implement this system, the following elements are required:

[0325] Collection and analysis of music data

[0326] The server collects music data from a music database on the cloud. This music data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes this data based on music theory and saves the analysis results. The software used for this analysis is a music theory analysis program. The analysis results are used as training data for the generative AI model.

[0327] Receiving a user request

[0328] A user inputs a request for music creation using a device (such as a smartphone or a head-mounted display). For example, the request may be specific, such as "I would like new harmony suggestions." The device receives this request and sends it to the server.

[0329] emotion recognition

[0330] The device is equipped with an emotion analysis means that recognizes the user's emotional state by analyzing the user's facial expressions, tone of voice, and the content of input text, etc. This analysis is performed using the dlib library and voice analysis software.

[0331] Proposal generation

[0332] The server combines the user request with emotional data obtained from the emotion analysis tool and generates suggestions based on music theory using a generative AI model (e.g., GPT-3). For example, if a new harmony suggestion is requested, the server generates a specific harmony suggestion and adjusts it according to the user's emotional state. The generated suggestion can be specific, such as "C major seventh chord."

[0333] View Suggestions

[0334] The generated suggestions are displayed to the user via the device, allowing the user to use the suggestions as a reference for music production. In particular, the emotion analysis means provides suggestions that match the user's emotional state, allowing the user to obtain more personalized musical inspiration.

[0335] Specific examples

[0336] For example, if a user requests "I want new harmony suggestions," the device sends this request to the server and simultaneously analyzes the user's facial expressions to determine their emotions. The server uses a generative AI model based on the request and emotional data to generate new harmony suggestions. If the user is relaxed, the server will suggest a harmony with a relaxing atmosphere.

[0337] Prompt Sentence Examples

[0338] User request analysis: I would like suggestions for new harmonies. Emotion data: {'happiness': 0.8}

[0339] The system allows users to efficiently generate new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activity.

[0340] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0341] Program processing steps

[0342] Step 1:

[0343] The server collects music data from a music database. The input data are music files from the music database, including melody, harmony, rhythm patterns, and lyrics. The server analyzes this data based on music theory. This analysis saves the music data as metadata and serves as training data for the generative AI model.

[0344] Step 2:

[0345] The user inputs a request for music production through the device. The input data is a text request entered by the user. The device sends this request to the server. The server analyzes the received request and passes the analysis results to the generative AI model.

[0346] Step 3:

[0347] The device analyzes the user's emotions. Input data includes image capture, voice input, and text input. The device uses libraries such as dlib to determine the user's emotions from facial expressions and tone of voice. This data is processed by the emotion analysis means and output as an emotional state. The output data is sent to the server as emotion data.

[0348] Step 4:

[0349] The server uses a generative AI model based on the request analysis results and emotional data to generate suggestions based on music theory. The input data are the request analysis results and emotional data. The generative AI model (e.g., GPT-3) takes these data into account to generate suggestions. The output data is a specific musical suggestion (e.g., a C major seventh chord).

[0350] Step 5:

[0351] The server sends the generated proposal to the device, which then displays it to the user. The input data is the proposal from the server, and the output data is the music proposal displayed to the user. The user can use this as a reference to proceed with their music creation.

[0352] Examples:

[0353] For example, if a user requests a new harmony suggestion and the device recognizes a relaxed facial expression, the server will use the generative AI model based on the request and emotional data to suggest a C major seventh chord. This suggestion is displayed to the user via the device, and the user can use it as a reference when creating music.

[0354] Specific operations and input / output flow

[0355] Step 1:

[0356] Input: Music files from a music database

[0357] Processing: Data analysis based on music theory

[0358] Output: Analyzed music data

[0359] How it works: The server retrieves music files from a database, analyzes the data using a music theory analysis program, and stores it as metadata.

[0360] Step 2:

[0361] Input: The request text entered by the user

[0362] Processing: Parsing the request

[0363] Output: Request analysis results

[0364] How it works: The device receives a request from the user and sends it to the server, which analyzes the request and passes the results to the generative AI model.

[0365] Step 3:

[0366] Input: User's face image, voice, text

[0367] Processing: Emotion Recognition

[0368] Output: Emotion data

[0369] How it works: The device analyzes the user's emotions using libraries such as dlib and sends the results to the server.

[0370] Step 4:

[0371] Input: Request analysis results, emotion data

[0372] Processing: Generating music suggestions with a generative AI model

[0373] Output:Music suggestions

[0374] How it works: The server uses a generative AI model to generate music suggestions based on request analysis and emotional data.

[0375] Step 5:

[0376] Input: Generated Music Suggestions

[0377] Action:View Proposal

[0378] Output: What is displayed to the user

[0379] How it works: The server sends music suggestions generated by the generative AI model to the device, which then displays them to the user.

[0380] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0381] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0382] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0383] [Second embodiment]

[0384] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0385] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0386] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0387] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0388] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0389] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0390] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0391] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0392] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0393] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0394] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0395] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0396] The present invention relates to a music creation support system that solves the problems artists and creators face when creating music, such as a lack of inspiration and new ideas. The system of the present invention includes a processing means for collecting music data from a music database and analyzing it based on music theory, a processing means for receiving requests for music creation from users and analyzing the requests, a processing means for generating suggestions based on music theory using a generative AI model, and a processing means for displaying the generated suggestions to the user.

[0397] Explanation of program processing

[0398] Music data collection and analysis

[0399] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0400] Receiving a user request

[0401] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user may input a request such as "I would like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0402] Proposal generation

[0403] When the server receives a user's request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord.

[0404] View Suggestions

[0405] The terminal displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production.

[0406] Specific examples

[0407] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0408] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0409] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0410] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0411] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0412] This system allows users to efficiently obtain new musical ideas and smoothly advance their creative activities.

[0413] The processing flow will be explained below.

[0414] Step 1:

[0415] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, and lyrics, and stores this data in a local database for analysis.

[0416] Step 2:

[0417] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the AI ​​model.

[0418] Step 3:

[0419] The terminal displays an interface for the user to input requests for music production. For example, the user may input, "I would like some new harmony suggestions."

[0420] Step 4:

[0421] The terminal receives a request from the user, analyzes the content of the request, and sends the analysis results to the server.

[0422] Step 5:

[0423] The server analyzes the request received from the device and, based on the request content, selects the appropriate process to use the generative AI model.

[0424] Step 6:

[0425] The server uses a generative AI model to generate musical suggestions based on the user's request, such as a C major seventh chord if asked for harmony suggestions.

[0426] Step 7:

[0427] The server sends the generated musical suggestions to the device, including suggestions derived from the generative AI model.

[0428] Step 8:

[0429] The terminal displays the suggestions received from the server to the user using an interface that makes it easy to provide the suggestions visually or audibly to the user.

[0430] Step 9:

[0431] Users can refer to the suggestions displayed on the device to proceed with music creation, which will help them gain new ideas and inspiration and make music creation easier.

[0432] Example 1

[0433] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0434] In modern music creation, artists and creators often face a lack of inspiration and new ideas. Furthermore, music production requires advanced knowledge of music theory, making it difficult for users without specialized knowledge to create music efficiently. A system that solves this problem and allows users to quickly and easily obtain musical ideas is needed.

[0435] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0436] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for displaying the generated suggestions to the user, means for the generative AI model to generate specific musical ideas based on music theory based on the input requests, and means for presenting the generated suggestions to the user's terminal in real time, thereby enabling users to efficiently obtain new musical ideas.

[0437] A "music database" is a system for managing and storing a large amount of music data in one place.

[0438] "Music data" refers to information about the elements that make up a musical work, such as melody, harmony, rhythm patterns, and lyrics.

[0439] "Music theory" is the study of the basic principles of musical structure, chords, rhythm, melody, and their combinations and progressions.

[0440] A "server" is a computer system that runs behind the scenes of the entire system, handling requests, storing data, and delivering generated offers.

[0441] "User" refers to an individual or professional who uses this system to create music.

[0442] A "request" is a user's request for a specific musical suggestion or creation from the system.

[0443] A "generative AI model" is an artificial intelligence algorithm that generates new information or suggestions based on input data.

[0444] A "specific musical idea" is a specific creative element in musical composition, such as a particular melody, harmony, rhythmic pattern, or lyrics.

[0445] A "terminal" is a device, such as a computer or smartphone, that a user uses to interact with the system.

[0446] "Real-time" refers to a state in which user input and system processing are reflected immediately without delay.

[0447] The present invention relates to a music creation support system, and specific embodiments thereof will be described below.

[0448] Music data collection and analysis

[0449] The server collects music data from a music database on the cloud. The specific software used for this is a tool that extracts data through an API. The data includes elements such as melody, harmony, rhythm patterns, and lyrics. The server then uses a Python music analysis library (e.g., music21) to analyze the collected music data based on music theory. The analysis results are stored in a database system such as MySQL or PostgreSQL. This accumulates data for training the generative AI model.

[0450] Receiving a user request

[0451] The device provides the user with an interface for inputting requests related to music production. This interface is built using a front-end framework (e.g., React, Vue.js). The user inputs requests through this interface, specifically prompt sentences such as "I would like new harmony suggestions." The input request is sent to the server in JSON format.

[0452] Proposal generation

[0453] The server analyzes the received user request and generates suggestions based on music theory via a generative AI model (e.g., GPT-3). The generative AI model has the ability to generate new musical ideas based on training data. For example, in response to a request for "new harmony suggestions," it generates specific harmonies such as a C major seventh chord. The suggestions are returned to the device in JSON format.

[0454] View Suggestions

[0455] The device displays the suggestions received from the server to the user. Specifically, it uses JavaScript to render the suggestions into an HTML view and displays it in the user interface. This allows the user to use the generated suggestions as a reference for music production.

[0456] Specific examples

[0457] For example, when a user inputs a request such as "I want new harmony suggestions" into a terminal, the following is an example of the series of processes that will be carried out.

[0458] 1. Device: The user types "I want new harmony suggestions" and presses the send button. The device sends the request in JSON format to the server.

[0459] 2. Server: The server receives the request, analyzes it, and inputs the prompt "I would like some new harmony suggestions" into the generative AI model.

[0460] 3. Server: The generative AI model generates new harmonies, such as a C major seventh chord, and returns suggestions in JSON format to the device.

[0461] 4. Terminal: The terminal receives the JSON data and displays the suggestions in the user interface.

[0462] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0463] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0464] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0465] Step 1: Collecting music data

[0466] The server collects music data from a music database. It uses an API to obtain data such as melody, harmony, rhythm patterns, and lyrics. For example, it calls the API of a cloud service and receives music data in JSON format. The input is the music data from the API, and the output is music data stored in temporary data storage on the server.

[0467] Step 2: Analyzing the music data

[0468] The server analyzes the collected music data based on music theory. Specifically, it uses a Python music analysis library (e.g., music21) to analyze harmonic structure, rhythmic patterns, etc. The analysis results are stored in a database (e.g., MySQL, PostgreSQL) as training data. The input is the music data stored in temporary data storage, and the output is the analysis results stored in the database.

[0469] Step 3: Entering User Requests

[0470] Users input their music-making requests through a terminal interface. This interface is built using a front-end framework (e.g., React, Vue.js). Specifically, they input a prompt such as "I'd like some new harmony suggestions." The input is the user's request, and the output is data sent to the server in JSON format.

[0471] Step 4: Receiving and Parsing the Request

[0472] The server receives the request sent from the device. It then analyzes the request using a web framework (e.g., Flask, Django). Based on the analysis results, it creates a prompt to input to the generative AI model. The input is the request JSON sent from the device, and the output is the prompt to input to the generative AI model.

[0473] Step 5: Proposal Generation

[0474] The server uses a generative AI model (e.g., GPT-3) to generate suggestions based on music theory. The generative AI model generates new melodies and harmonies based on the training data. For example, it generates a specific harmony such as a C major seventh chord. The input is a prompt to the generative AI model, and the output is the generated musical suggestions.

[0475] Step 6: Submit your proposal

[0476] The server sends the generated proposals to the terminal in JSON format. The input is the proposal data from the generative AI model, and the output is the JSON data sent to the terminal.

[0477] Step 7: Viewing Proposals

[0478] The terminal displays the suggestions received from the server to the user by using JavaScript to render the suggestions into an HTML view and display it in the user interface. The input is the JSON data received from the server and the output is the suggestions displayed to the user.

[0479] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0480] (Application example 1)

[0481] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0482] Conventional music production support systems have the problem of being unable to effectively solve the problem of running out of inspiration and lack of new ideas. In particular, it has been difficult to instantly provide new harmony, melody, and lyric ideas while creating music. Furthermore, because such systems are not linked to content distribution services, they have been unable to quickly provide the latest musical ideas to a wide range of users. To solve these problems, a more efficient and practical music creation support system for users is needed.

[0483] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0484] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, and means for applying the suggestions to the content distribution service, thereby enabling users to quickly and efficiently obtain new musical ideas.

[0485] A "music database" is a database that stores multiple pieces of music data and allows searching and retrieval.

[0486] "Music data" refers to the elements that make up a song, such as melody, harmony, rhythm, and lyrics.

[0487] "Music theory" is the academic knowledge of musical structure and composition, which serves as a guide for analyzing and creating musical pieces.

[0488] A "generative AI model" is a technology that uses artificial intelligence to generate new music suggestions based on user requests.

[0489] A "request" is a specific request for music production, such as a new harmony, melody, or lyrics that the user desires.

[0490] A "content distribution service" is a service that provides users with digital content such as music and videos via the Internet.

[0491] "Suggestions" are new musical ideas generated by the generative AI model to assist users in their music creation.

[0492] "Analysis" is the process of understanding a user request and analyzing its content to generate appropriate music suggestions.

[0493] This invention is a music creation support system that allows artists and creators to obtain new ideas in music production. The system includes the following means.

[0494] System Configuration

[0495] 1. A server that collects music data from a music database and analyzes it based on music theory

[0496] The server resides in the cloud and collects multiple pieces of music data (melody, harmony, rhythm, lyrics, etc.) from a music database. The collected data is analyzed based on music theory and used as training data for the generative AI model.

[0497] 2. A device that receives requests from users regarding music production and analyzes those requests.

[0498] The device provides an interface for users to input music-making requests, such as "I want new harmony suggestions." The device receives the request and sends it to the server for analysis.

[0499] 3. A server that uses generative AI models to generate suggestions based on music theory

[0500] The server analyzes the user's request and generates music theory-based suggestions based on the request using a generative AI model (e.g., GPTNeo). For example, if the user's request is for a new harmony, the server generates a specific harmony, such as a C major seventh chord.

[0501] 4. A device that displays the generated suggestions to the user.

[0502] The generated suggestions are sent from the server to the device, which then displays them to the user, who can use them as a reference to proceed with their music creation.

[0503] 5. Methods applied to content distribution services

[0504] This system is particularly linked to content distribution services, making it possible to quickly provide musical ideas to a wide range of users via the Internet, allowing users to efficiently acquire new musical ideas and support their creative activities.

[0505] Hardware and software used

[0506] Hardware

[0507] Server: Cloud server

[0508] Device: Smartphone

[0509] software

[0510] Server: Python, requests library

[0511] Generative AI model: Hugging Face transformers library, GPTNeo model

[0512] Device: Smartphone application

[0513] Specific examples

[0514] For example, if a user inputs a request such as "I want new harmony suggestions" through a smartphone application, the following processing will occur.

[0515] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0516] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0517] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0518] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0519] Prompt Sentence Examples

[0520] "Users are asking for suggestions: I'd like some new harmony suggestions."

[0521] This system allows users to efficiently eliminate inspiration drain and easily come up with new musical ideas.

[0522] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0523] Step 1:

[0524] The user uses the terminal to enter a request

[0525] A user opens a smartphone application and enters a music-making request, such as "I'd like some new harmony suggestions." The device receives this request and converts it into data to send to the server. The input is the user's request, and the output is a data packet containing that request.

[0526] Step 2:

[0527] The device sends a request to the server

[0528] The terminal sends the request received from the user to the server as an HTTP POST request. This request includes the user ID and the request content. The input is the request data from the user, and the output is the HTTP POST request sent to the server.

[0529] Step 3:

[0530] The server receives and parses the request

[0531] The server receives requests sent from the device and analyzes their contents. The analysis involves analyzing the text of the request, for example, creating a prompt appropriate for the generative AI model based on a request such as "I'd like some new harmony suggestions." The input is the HTTP POST request sent to the server, and the output is the analysis result, including the prompt text.

[0532] Step 4:

[0533] The server generates suggestions using a generative AI model

[0534] The server inputs the analysis results into a generative AI model to generate a suggestion based on music theory. For example, a generative AI model (GPTNeo) is used to generate a new suggestion for the requested harmony (e.g., a C major seventh chord). The input is a prompt based on the analysis results, and the output is a musical suggestion generated by the generative AI model.

[0535] Step 5:

[0536] The server sends the generated proposal to the device.

[0537] The server sends the generated music suggestions to the device, where the suggestions are sent as an HTTP response and formatted appropriately for the user. The input is the music suggestions generated by the AI ​​model, and the output is the HTTP response sent to the device.

[0538] Step 6:

[0539] The terminal displays the suggestions to the user.

[0540] The device receives the response from the server and displays it to the user in an appropriate format. For example, "New harmony suggestion: C major seventh chord" is displayed on the screen. The input is the HTTP response from the server, and the output is the suggestion information displayed to the user.

[0541] Step 7:

[0542] The user will use the suggestions to proceed with music production.

[0543] The user can refer to the new music suggestions displayed on the device and proceed with the composition of the music. In this process, the user can gain new inspiration and ideas. The input is the suggested information displayed on the device, and the output is the user's creative activity.

[0544] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0545] The present invention further relates to a music creation support system that combines an emotion engine that recognizes user emotions. The system includes processing means for collecting music data from a music database and analyzing it based on music theory, processing means for receiving music creation requests from users and analyzing the requests, processing means for generating suggestions based on music theory using a generative AI model, processing means for displaying the generated suggestions to the user, and an emotion engine that recognizes user emotions.

[0546] Explanation of program processing

[0547] Music data collection and analysis

[0548] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0549] Receiving a user request

[0550] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user might input, "I'd like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0551] Proposal generation

[0552] When the server receives a user request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord. It also uses an emotion engine to generate suggestions that correspond to the user's emotions.

[0553] Emotion Engine Functions

[0554] The device uses an emotion engine to recognize the user's emotions. Emotion recognition utilizes the user's facial expressions, tone of voice, and input text content. The emotion engine analyzes this information to determine the user's current emotional state.

[0555] View Suggestions

[0556] The device displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production. In particular, the emotion engine provides suggestions that correspond to the user's emotions, allowing the user to obtain more appropriate musical inspiration.

[0557] Specific examples

[0558] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0559] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server. At the same time, the emotion engine analyzes the user's facial expressions to recognize their emotions.

[0560] 2. Server: The server receives the request and analyzes it. It then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested. At the same time, emotional information from the emotion engine is taken into account.

[0561] 3. Server: Based on the information from the emotion engine, for example, if the user is relaxed, it suggests harmonies with a more relaxed atmosphere.

[0562] 4. Terminal: The terminal displays suggestions received from the server, such as "C major seventh," to the user. It also displays the analysis results of the emotion engine, making it easier for the user to understand the background of the suggestions.

[0563] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0564] The system allows users to efficiently discover new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activities.

[0565] The processing flow will be explained below.

[0566] Step 1:

[0567] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, lyrics, etc. The server then stores this data in a local database.

[0568] Step 2:

[0569] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the generative AI model.

[0570] Step 3:

[0571] The terminal provides the user with an interface for inputting requests related to music production. For example, the user may input, "I would like some new harmony suggestions."

[0572] Step 4:

[0573] The terminal receives the user's request, analyzes its contents, and sends the analysis results to the server.

[0574] Step 5:

[0575] The device uses an emotion engine to recognize the user's emotions by analyzing the user's facial expressions, tone of voice, input text, etc.

[0576] Step 6:

[0577] The server analyzes the requests received from the terminal and the emotional information provided by the emotion engine, and takes this information into account when it needs to generate musical suggestions based on the emotional state.

[0578] Step 7:

[0579] The server uses a generative AI model to generate appropriate suggestions based on music theory, for example, generating specific harmonies such as a C major seventh chord if a new harmony suggestion is requested.

[0580] Step 8:

[0581] The server generates suggestions based on the user's emotional information obtained from the emotion engine: if the user is relaxed, it will suggest harmony with a relaxing atmosphere.

[0582] Step 9:

[0583] The server sends generated musical suggestions to the device, including suggestions derived from the generative AI model and emotion-based suggestions.

[0584] Step 10:

[0585] The terminal displays the proposal received from the server to the user using an interface that presents the proposal to the user visually or audibly.

[0586] Step 11:

[0587] The user can proceed with music production by referring to the suggestions displayed on the device. By also referring to emotion-based suggestions, music production can proceed more smoothly.

[0588] Example 2

[0589] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0590] Conventional music production support systems do not take into account the user's emotional state, making it difficult to provide personalized music suggestions that reflect the user's emotions. Furthermore, their ability to automatically generate suggestions based on music theory is limited. This makes it difficult for users to gain inspiration efficiently, potentially stagnateing their creative activities.

[0591] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0592] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for using an emotion engine to recognize the user's emotions, and means for displaying the generated suggestions to the user, thereby enabling personalized music suggestions to be provided according to the user's emotional state.

[0593] A "music database" is a type of data storage that collects data related to songs, including musical elements such as melody, harmony, rhythmic patterns, and lyrics.

[0594] "Music theory" is the study of rules and principles related to musical structure, form, harmony, melody, rhythm, etc. It is the basis for analyzing music and generating suggestions.

[0595] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate suggestions for music, text, etc., particularly natural language processing and machine learning techniques.

[0596] The "emotion engine" is a system for recognizing the user's emotional state by analyzing facial expressions, tone of voice, and the content of input text.

[0597] The "server" is a central control unit that processes and stores data and provides services to client systems. In this invention, it analyzes music data and processes user requests.

[0598] "Terminal" refers to a device that allows a user to access and operate the system through an interface. This includes computers, smartphones, tablets, etc.

[0599] A "user request" is a request a user makes to the system regarding a specific piece of music production. An example would be "I'd like some new harmony suggestions."

[0600] "Suggestions" refer to musical elements and ideas that the system provides to the user using generative AI models and emotion engines, which are used as references for music production.

[0601] MODE FOR CARRYING OUT THE INVENTION

[0602] The present invention provides a system for supporting music production by recognizing the emotional state of a user. Specific embodiments for carrying out the present invention will be described below.

[0603] Music data collection and analysis

[0604] The server collects music data from a specific music database (e.g., a music library API) on the cloud. The collected data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes the acquired music data using a Python music analysis library (e.g., music21, librosa), and stores the analysis results in a database. This accumulates data for training the generative AI model.

[0605] Receiving a user request

[0606] The device provides an interface through which users can input requests for music creation via a web application (e.g., React, Vue.js). For example, a user might input, "I'd like some new harmony suggestions." The device then sends the input request to the server using WebSocket or an HTTP request.

[0607] Proposal Generation

[0608] When the server receives a user request, it first analyzes the request using an NLP library (e.g., spaCy, NLTK). It then uses a generative AI model (e.g., GPT-3) to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, it generates specific chords, such as a C major seventh chord. It also incorporates information from an emotion engine to take the user's emotional state into account.

[0609] Emotion Engine Functions

[0610] The device uses a camera module (e.g., a webcam) to capture the user's facial expressions and simultaneously accepts voice input. The device then analyzes the user's emotions using emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer). The analysis results are sent to the server in real time, and the server uses this information to determine the user's current emotional state.

[0611] View Suggestions

[0612] The server sends the generated music suggestions and the analysis results from the emotion engine to the device. The device prepares UI components (e.g., HTML, CSS, JavaScript) to display the received suggestions to the user. The user then proceeds with the music creation process based on the displayed suggestions.

[0613] Specific examples

[0614] For example, if a user inputs a request to the terminal saying, "I want new harmony suggestions," the following series of processes are carried out.

[0615] 1. Device: The user types "I want new harmony suggestions" and clicks the send button. At the same time, the camera captures facial expressions and the microphone captures audio, which are then sent to the emotion engine.

[0616] Example prompt: "I'd like some new harmony suggestions for normal times."

[0617] 2. Server: Receives the request and analyzes it using an NLP library (e.g., spaCy). Uses a generative AI model (e.g., GPT-3) to generate specific harmony suggestions, such as "C major seventh chord." Takes emotion recognition results into account to suggest a relaxing atmosphere.

[0618] Example prompt: "I'd like some new harmony suggestions for when I'm relaxing."

[0619] 3. If the emotion engine determines the user's emotion as "relaxed," the generative AI model will adjust accordingly.

[0620] 4. Terminal: Displays to the user a suggestion that combines "C major seventh" received from the server with the results of sentiment analysis.

[0621] 5. The user then proceeds with creating music that corresponds to a specific emotion based on the displayed suggestions.

[0622] The system not only helps users find musical inspiration quickly, but also streamlines their creative process by receiving personalized suggestions based on their emotional state.

[0623] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0624] Step 1: Collecting music data

[0625] The server accesses a music database on the cloud. Specifically, it retrieves music data such as melody, harmony, rhythm patterns, and lyrics through an API (e.g., a music library API). This data is stored on the server and prepared for the next analysis step.

[0626] Input: Cloud music database

[0627] Output: Music data stored on the server

[0628] ---

[0629] Step 2: Analyzing the music data

[0630] The server analyzes the stored music data using Python music analysis libraries (e.g., music21, librosa). Specifically, it extracts the melody, harmony, rhythmic patterns, and lyric characteristics of each song. The analysis results are stored in a database in JSON or CSV format and used to train the generative AI model.

[0631] Input: Music data stored on the server

[0632] Output: Analyzed music data

[0633] ---

[0634] Step 3: Receiving a user request

[0635] The device provides an interface using a web application (e.g., React, Vue.js). The user inputs a request for music creation (e.g., "I'd like some new harmony suggestions.") The device receives this request and sends it to the server via WebSocket or HTTP request.

[0636] Input: User-entered music production requests

[0637] Output: Request sent to the server

[0638] ---

[0639] Step 4: Parsing the request

[0640] When the server receives a user request, it parses it using an NLP library (e.g., spaCy, NLTK) to understand the content of the request and determine what musical suggestions would be appropriate.

[0641] Input: The user request sent to the server

[0642] Output: Parsed request content

[0643] ---

[0644] Step 5: Generate proposals

[0645] The server uses a generative AI model (e.g., GPT-3) to generate music theory-based suggestions based on the analyzed request. For example, in response to a request for "new harmony suggestions," the server generates specific chords (e.g., a C major seventh chord).

[0646] Input: Analyzed request content and music theory data

[0647] Output: Generated music suggestions

[0648] ---

[0649] Step 6: Perform emotion recognition

[0650] The device uses a camera module (e.g., webcam) to capture the user's facial expressions and a microphone to capture their voice. This information is then sent to emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer) to analyze the user's emotions. The analysis results are then sent to a server.

[0651] Input: User's facial expressions, voice data

[0652] Output: Parsed emotion data

[0653] ---

[0654] Step 7: Integrating Emotional Data

[0655] The server integrates emotional data from the emotion engine into the music suggestions generated by the generative AI model. For example, if the server determines that the user is relaxed, it generates music suggestions that correspond to that emotion (e.g., relaxing harmony).

[0656] Input: Generated music suggestions, parsed emotion data

[0657] Output: Emotion-based music suggestions

[0658] ---

[0659] Step 8: Viewing Proposals

[0660] The device displays music suggestions to the user based on the emotion received from the server. Specifically, the suggestions are visually displayed using UI components (e.g., HTML, CSS, JavaScript).

[0661] Input: Emotion-based music suggestions received from the server

[0662] Output: Music suggestions displayed to the user

[0663] ---

[0664] Step 9: Music Production Progression

[0665] The user can then proceed with the composition of the music based on the displayed musical suggestions. For example, they can create a piece of music incorporating a C major seventh chord. At the same time, they can input new requests as needed, which are reflected in the system.

[0666] Input: Music suggestions displayed to the user

[0667] Output: Finished song, next request

[0668] ---

[0669] This process flow provides personalized music suggestions based on the user's emotional state, improving the efficiency of music production.

[0670] (Application example 2)

[0671] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0672] Conventional music creation support systems generate suggestions without considering the user's emotional state, which prevents personalized suggestions for each individual user and can lead to lower user satisfaction. Furthermore, users often find it difficult to find musical inspiration that matches their emotional state, limiting their creativity. A system that can resolve these issues and provide more effective and personalized music creation support is needed.

[0673] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0674] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, emotion analysis means for recognizing the user's emotions, and means for displaying the generated suggestions to the user. This enables personalized musical suggestions based on the user's emotional state, thereby promoting creativity.

[0675] A "music database" is a large-scale information storage device for storing, saving, and managing music data.

[0676] "Music data" refers to information related to musical components such as melody, harmony, rhythm patterns, and lyrics.

[0677] "Music theory" is a set of rules and principles for understanding and analysing musical structures and components.

[0678] A "user request" is an instruction including a request or wish given by a user to a system.

[0679] A "generative AI model" is an artificial intelligence model that generates appropriate suggestions or answers based on given data or requests.

[0680] "Emotion analysis means" refers to technology or devices for recognizing and analyzing the user's emotional state.

[0681] "Suggestions" are specific ideas and advice about music production generated by the system.

[0682] The "means for displaying to the user" is an interface or device for visually presenting the generated suggestions to the user.

[0683] The present invention is a music creation support system that is combined with emotion analysis means for recognizing the emotions of the user. To specifically implement this system, the following elements are required:

[0684] Collection and analysis of music data

[0685] The server collects music data from a music database on the cloud. This music data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes this data based on music theory and saves the analysis results. The software used for this analysis is a music theory analysis program. The analysis results are used as training data for the generative AI model.

[0686] Receiving a user request

[0687] A user inputs a request for music creation using a device (such as a smartphone or a head-mounted display). For example, the request may be specific, such as "I would like new harmony suggestions." The device receives this request and sends it to the server.

[0688] emotion recognition

[0689] The device is equipped with an emotion analysis means that recognizes the user's emotional state by analyzing the user's facial expressions, tone of voice, and the content of input text, etc. This analysis is performed using the dlib library and voice analysis software.

[0690] Proposal generation

[0691] The server combines the user request with emotional data obtained from the emotion analysis tool and generates suggestions based on music theory using a generative AI model (e.g., GPT-3). For example, if a new harmony suggestion is requested, the server generates a specific harmony suggestion and adjusts it according to the user's emotional state. The generated suggestion can be specific, such as "C major seventh chord."

[0692] View Suggestions

[0693] The generated suggestions are displayed to the user via the device, allowing the user to use the suggestions as a reference for music production. In particular, the emotion analysis means provides suggestions that match the user's emotional state, allowing the user to obtain more personalized musical inspiration.

[0694] Specific examples

[0695] For example, if a user requests "I want new harmony suggestions," the device sends this request to the server and simultaneously analyzes the user's facial expressions to determine their emotions. The server uses a generative AI model based on the request and emotional data to generate new harmony suggestions. If the user is relaxed, the server will suggest a harmony with a relaxing atmosphere.

[0696] Prompt Sentence Examples

[0697] User request analysis: I would like suggestions for new harmonies. Emotion data: {'happiness': 0.8}

[0698] The system allows users to efficiently generate new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activity.

[0699] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0700] Program processing steps

[0701] Step 1:

[0702] The server collects music data from a music database. The input data are music files from the music database, including melody, harmony, rhythm patterns, and lyrics. The server analyzes this data based on music theory. This analysis saves the music data as metadata and serves as training data for the generative AI model.

[0703] Step 2:

[0704] The user inputs a request for music production through the device. The input data is a text request entered by the user. The device sends this request to the server. The server analyzes the received request and passes the analysis results to the generative AI model.

[0705] Step 3:

[0706] The device analyzes the user's emotions. Input data includes image capture, voice input, and text input. The device uses libraries such as dlib to determine the user's emotions from facial expressions and tone of voice. This data is processed by the emotion analysis means and output as an emotional state. The output data is sent to the server as emotion data.

[0707] Step 4:

[0708] The server uses a generative AI model based on the request analysis results and emotional data to generate suggestions based on music theory. The input data are the request analysis results and emotional data. The generative AI model (e.g., GPT-3) takes these data into account to generate suggestions. The output data is a specific musical suggestion (e.g., a C major seventh chord).

[0709] Step 5:

[0710] The server sends the generated proposal to the device, which then displays it to the user. The input data is the proposal from the server, and the output data is the music proposal displayed to the user. The user can use this as a reference to proceed with their music creation.

[0711] Examples:

[0712] For example, if a user requests a new harmony suggestion and the device recognizes a relaxed facial expression, the server will use the generative AI model based on the request and emotional data to suggest a C major seventh chord. This suggestion is displayed to the user via the device, and the user can use it as a reference when creating music.

[0713] Specific operations and input / output flow

[0714] Step 1:

[0715] Input: Music files from a music database

[0716] Processing: Data analysis based on music theory

[0717] Output: Analyzed music data

[0718] How it works: The server retrieves music files from a database, analyzes the data using a music theory analysis program, and stores it as metadata.

[0719] Step 2:

[0720] Input: The request text entered by the user

[0721] Processing: Parsing the request

[0722] Output: Request analysis results

[0723] How it works: The device receives a request from the user and sends it to the server, which analyzes the request and passes the results to the generative AI model.

[0724] Step 3:

[0725] Input: User's face image, voice, text

[0726] Processing: Emotion Recognition

[0727] Output: Emotion data

[0728] How it works: The device analyzes the user's emotions using libraries such as dlib and sends the results to the server.

[0729] Step 4:

[0730] Input: Request analysis results, emotion data

[0731] Processing: Generating music suggestions with a generative AI model

[0732] Output:Music suggestions

[0733] How it works: The server uses a generative AI model to generate music suggestions based on request analysis and emotional data.

[0734] Step 5:

[0735] Input: Generated Music Suggestions

[0736] Action:View Proposal

[0737] Output: What is displayed to the user

[0738] How it works: The server sends music suggestions generated by the generative AI model to the device, which then displays them to the user.

[0739] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0740] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0741] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0742] [Third embodiment]

[0743] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0744] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0745] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0746] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0747] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0748] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0749] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0750] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0751] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0752] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0753] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0754] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0755] The present invention relates to a music creation support system that solves the problems artists and creators face when creating music, such as a lack of inspiration and new ideas. The system of the present invention includes a processing means for collecting music data from a music database and analyzing it based on music theory, a processing means for receiving requests for music creation from users and analyzing the requests, a processing means for generating suggestions based on music theory using a generative AI model, and a processing means for displaying the generated suggestions to the user.

[0756] Explanation of program processing

[0757] Music data collection and analysis

[0758] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0759] Receiving a user request

[0760] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user may input a request such as "I would like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0761] Proposal generation

[0762] When the server receives a user's request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord.

[0763] View Suggestions

[0764] The terminal displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production.

[0765] Specific examples

[0766] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0767] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0768] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0769] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0770] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0771] This system allows users to efficiently obtain new musical ideas and smoothly advance their creative activities.

[0772] The processing flow will be explained below.

[0773] Step 1:

[0774] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, and lyrics, and stores this data in a local database for analysis.

[0775] Step 2:

[0776] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the AI ​​model.

[0777] Step 3:

[0778] The terminal displays an interface for the user to input requests for music production. For example, the user may input, "I would like some new harmony suggestions."

[0779] Step 4:

[0780] The terminal receives a request from the user, analyzes the content of the request, and sends the analysis results to the server.

[0781] Step 5:

[0782] The server analyzes the request received from the device and, based on the request content, selects the appropriate process to use the generative AI model.

[0783] Step 6:

[0784] The server uses a generative AI model to generate musical suggestions based on the user's request, such as a C major seventh chord if asked for harmony suggestions.

[0785] Step 7:

[0786] The server sends the generated musical suggestions to the device, including suggestions derived from the generative AI model.

[0787] Step 8:

[0788] The terminal displays the suggestions received from the server to the user using an interface that makes it easy to provide the suggestions visually or audibly to the user.

[0789] Step 9:

[0790] Users can refer to the suggestions displayed on the device to proceed with music creation, which will help them gain new ideas and inspiration and make music creation easier.

[0791] Example 1

[0792] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0793] In modern music creation, artists and creators often face a lack of inspiration and new ideas. Furthermore, music production requires advanced knowledge of music theory, making it difficult for users without specialized knowledge to create music efficiently. A system that solves this problem and allows users to quickly and easily obtain musical ideas is needed.

[0794] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0795] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for displaying the generated suggestions to the user, means for the generative AI model to generate specific musical ideas based on music theory based on the input requests, and means for presenting the generated suggestions to the user's terminal in real time, thereby enabling users to efficiently obtain new musical ideas.

[0796] A "music database" is a system for managing and storing a large amount of music data in one place.

[0797] "Music data" refers to information about the elements that make up a musical work, such as melody, harmony, rhythm patterns, and lyrics.

[0798] "Music theory" is the study of the basic principles of musical structure, chords, rhythm, melody, and their combinations and progressions.

[0799] A "server" is a computer system that runs behind the scenes of the entire system, handling requests, storing data, and delivering generated offers.

[0800] "User" refers to an individual or professional who uses this system to create music.

[0801] A "request" is a user's request for a specific musical suggestion or creation from the system.

[0802] A "generative AI model" is an artificial intelligence algorithm that generates new information or suggestions based on input data.

[0803] A "specific musical idea" is a specific creative element in musical composition, such as a particular melody, harmony, rhythmic pattern, or lyrics.

[0804] A "terminal" is a device, such as a computer or smartphone, that a user uses to interact with the system.

[0805] "Real-time" refers to a state in which user input and system processing are reflected immediately without delay.

[0806] The present invention relates to a music creation support system, and specific embodiments thereof will be described below.

[0807] Music data collection and analysis

[0808] The server collects music data from a music database on the cloud. The specific software used for this is a tool that extracts data through an API. The data includes elements such as melody, harmony, rhythm patterns, and lyrics. The server then uses a Python music analysis library (e.g., music21) to analyze the collected music data based on music theory. The analysis results are stored in a database system such as MySQL or PostgreSQL. This accumulates data for training the generative AI model.

[0809] Receiving a user request

[0810] The device provides the user with an interface for inputting requests related to music production. This interface is built using a front-end framework (e.g., React, Vue.js). The user inputs requests through this interface, specifically prompt sentences such as "I would like new harmony suggestions." The input request is sent to the server in JSON format.

[0811] Proposal generation

[0812] The server analyzes the received user request and generates suggestions based on music theory via a generative AI model (e.g., GPT-3). The generative AI model has the ability to generate new musical ideas based on training data. For example, in response to a request for "new harmony suggestions," it generates specific harmonies such as a C major seventh chord. The suggestions are returned to the device in JSON format.

[0813] View Suggestions

[0814] The device displays the suggestions received from the server to the user. Specifically, it uses JavaScript to render the suggestions into an HTML view and displays it in the user interface. This allows the user to use the generated suggestions as a reference for music production.

[0815] Specific examples

[0816] For example, when a user inputs a request such as "I want new harmony suggestions" into a terminal, the following is an example of the series of processes that will be carried out.

[0817] 1. Device: The user types "I want new harmony suggestions" and presses the send button. The device sends the request in JSON format to the server.

[0818] 2. Server: The server receives the request, analyzes it, and inputs the prompt "I would like some new harmony suggestions" into the generative AI model.

[0819] 3. Server: The generative AI model generates new harmonies, such as a C major seventh chord, and returns suggestions in JSON format to the device.

[0820] 4. Terminal: The terminal receives the JSON data and displays the suggestions in the user interface.

[0821] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0822] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0823] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0824] Step 1: Collecting music data

[0825] The server collects music data from a music database. It uses an API to obtain data such as melody, harmony, rhythm patterns, and lyrics. For example, it calls the API of a cloud service and receives music data in JSON format. The input is the music data from the API, and the output is music data stored in temporary data storage on the server.

[0826] Step 2: Analyzing the music data

[0827] The server analyzes the collected music data based on music theory. Specifically, it uses a Python music analysis library (e.g., music21) to analyze harmonic structure, rhythmic patterns, etc. The analysis results are stored in a database (e.g., MySQL, PostgreSQL) as training data. The input is the music data stored in temporary data storage, and the output is the analysis results stored in the database.

[0828] Step 3: Entering User Requests

[0829] Users input their music-making requests through a terminal interface. This interface is built using a front-end framework (e.g., React, Vue.js). Specifically, they input a prompt such as "I'd like some new harmony suggestions." The input is the user's request, and the output is data sent to the server in JSON format.

[0830] Step 4: Receiving and Parsing the Request

[0831] The server receives the request sent from the device. It then analyzes the request using a web framework (e.g., Flask, Django). Based on the analysis results, it creates a prompt to input to the generative AI model. The input is the request JSON sent from the device, and the output is the prompt to input to the generative AI model.

[0832] Step 5: Proposal Generation

[0833] The server uses a generative AI model (e.g., GPT-3) to generate suggestions based on music theory. The generative AI model generates new melodies and harmonies based on the training data. For example, it generates a specific harmony such as a C major seventh chord. The input is a prompt to the generative AI model, and the output is the generated musical suggestions.

[0834] Step 6: Submit your proposal

[0835] The server sends the generated proposals to the terminal in JSON format. The input is the proposal data from the generative AI model, and the output is the JSON data sent to the terminal.

[0836] Step 7: Viewing Proposals

[0837] The terminal displays the suggestions received from the server to the user by using JavaScript to render the suggestions into an HTML view and display it in the user interface. The input is the JSON data received from the server and the output is the suggestions displayed to the user.

[0838] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[0839] (Application example 1)

[0840] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0841] Conventional music production support systems have the problem of being unable to effectively solve the problem of running out of inspiration and lack of new ideas. In particular, it has been difficult to instantly provide new harmony, melody, and lyric ideas while creating music. Furthermore, because such systems are not linked to content distribution services, they have been unable to quickly provide the latest musical ideas to a wide range of users. To solve these problems, a more efficient and practical music creation support system for users is needed.

[0842] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0843] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, and means for applying the suggestions to the content distribution service, thereby enabling users to quickly and efficiently obtain new musical ideas.

[0844] A "music database" is a database that stores multiple pieces of music data and allows searching and retrieval.

[0845] "Music data" refers to the elements that make up a song, such as melody, harmony, rhythm, and lyrics.

[0846] "Music theory" is the academic knowledge of musical structure and composition, which serves as a guide for analyzing and creating musical pieces.

[0847] A "generative AI model" is a technology that uses artificial intelligence to generate new music suggestions based on user requests.

[0848] A "request" is a specific request for music production, such as a new harmony, melody, or lyrics that the user desires.

[0849] A "content distribution service" is a service that provides users with digital content such as music and videos via the Internet.

[0850] "Suggestions" are new musical ideas generated by the generative AI model to assist users in their music creation.

[0851] "Analysis" is the process of understanding a user request and analyzing its content to generate appropriate music suggestions.

[0852] This invention is a music creation support system that allows artists and creators to obtain new ideas in music production. The system includes the following means.

[0853] System Configuration

[0854] 1. A server that collects music data from a music database and analyzes it based on music theory

[0855] The server resides in the cloud and collects multiple pieces of music data (melody, harmony, rhythm, lyrics, etc.) from a music database. The collected data is analyzed based on music theory and used as training data for the generative AI model.

[0856] 2. A device that receives requests from users regarding music production and analyzes those requests.

[0857] The device provides an interface for users to input music-making requests, such as "I want new harmony suggestions." The device receives the request and sends it to the server for analysis.

[0858] 3. A server that uses generative AI models to generate suggestions based on music theory

[0859] The server analyzes the user's request and generates music theory-based suggestions based on the request using a generative AI model (e.g., GPTNeo). For example, if the user's request is for a new harmony, the server generates a specific harmony, such as a C major seventh chord.

[0860] 4. A device that displays the generated suggestions to the user.

[0861] The generated suggestions are sent from the server to the device, which then displays them to the user, who can use them as a reference to proceed with their music creation.

[0862] 5. Methods applied to content distribution services

[0863] This system is particularly linked to content distribution services, making it possible to quickly provide musical ideas to a wide range of users via the Internet, allowing users to efficiently acquire new musical ideas and support their creative activities.

[0864] Hardware and software used

[0865] Hardware

[0866] Server: Cloud server

[0867] Device: Smartphone

[0868] software

[0869] Server: Python, requests library

[0870] Generative AI model: Hugging Face transformers library, GPTNeo model

[0871] Device: Smartphone application

[0872] Specific examples

[0873] For example, if a user inputs a request such as "I want new harmony suggestions" through a smartphone application, the following processing will occur.

[0874] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[0875] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[0876] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[0877] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0878] Prompt Sentence Examples

[0879] "Users are asking for suggestions: I'd like some new harmony suggestions."

[0880] This system allows users to efficiently eliminate inspiration drain and easily come up with new musical ideas.

[0881] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0882] Step 1:

[0883] The user uses the terminal to enter a request

[0884] A user opens a smartphone application and enters a music-making request, such as "I'd like some new harmony suggestions." The device receives this request and converts it into data to send to the server. The input is the user's request, and the output is a data packet containing that request.

[0885] Step 2:

[0886] The device sends a request to the server

[0887] The terminal sends the request received from the user to the server as an HTTP POST request. This request includes the user ID and the request content. The input is the request data from the user, and the output is the HTTP POST request sent to the server.

[0888] Step 3:

[0889] The server receives and parses the request

[0890] The server receives requests sent from the device and analyzes their contents. The analysis involves analyzing the text of the request, for example, creating a prompt appropriate for the generative AI model based on a request such as "I'd like some new harmony suggestions." The input is the HTTP POST request sent to the server, and the output is the analysis result, including the prompt text.

[0891] Step 4:

[0892] The server generates suggestions using a generative AI model

[0893] The server inputs the analysis results into a generative AI model to generate a suggestion based on music theory. For example, a generative AI model (GPTNeo) is used to generate a new suggestion for the requested harmony (e.g., a C major seventh chord). The input is a prompt based on the analysis results, and the output is a musical suggestion generated by the generative AI model.

[0894] Step 5:

[0895] The server sends the generated proposal to the device.

[0896] The server sends the generated music suggestions to the device, where the suggestions are sent as an HTTP response and formatted appropriately for the user. The input is the music suggestions generated by the AI ​​model, and the output is the HTTP response sent to the device.

[0897] Step 6:

[0898] The terminal displays the suggestions to the user.

[0899] The device receives the response from the server and displays it to the user in an appropriate format. For example, "New harmony suggestion: C major seventh chord" is displayed on the screen. The input is the HTTP response from the server, and the output is the suggestion information displayed to the user.

[0900] Step 7:

[0901] The user will use the suggestions to proceed with music production.

[0902] The user can refer to the new music suggestions displayed on the device and proceed with the composition of the music. In this process, the user can gain new inspiration and ideas. The input is the suggested information displayed on the device, and the output is the user's creative activity.

[0903] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0904] The present invention further relates to a music creation support system that combines an emotion engine that recognizes user emotions. The system includes processing means for collecting music data from a music database and analyzing it based on music theory, processing means for receiving music creation requests from users and analyzing the requests, processing means for generating suggestions based on music theory using a generative AI model, processing means for displaying the generated suggestions to the user, and an emotion engine that recognizes user emotions.

[0905] Explanation of program processing

[0906] Music data collection and analysis

[0907] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[0908] Receiving a user request

[0909] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user might input, "I'd like some new harmony suggestions." The terminal receives this request and sends it to the server.

[0910] Proposal generation

[0911] When the server receives a user request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord. It also uses an emotion engine to generate suggestions that correspond to the user's emotions.

[0912] Emotion Engine Functions

[0913] The device uses an emotion engine to recognize the user's emotions. Emotion recognition utilizes the user's facial expressions, tone of voice, and input text content. The emotion engine analyzes this information to determine the user's current emotional state.

[0914] View Suggestions

[0915] The device displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production. In particular, the emotion engine provides suggestions that correspond to the user's emotions, allowing the user to obtain more appropriate musical inspiration.

[0916] Specific examples

[0917] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[0918] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server. At the same time, the emotion engine analyzes the user's facial expressions to recognize their emotions.

[0919] 2. Server: The server receives the request and analyzes it. It then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested. At the same time, emotional information from the emotion engine is taken into account.

[0920] 3. Server: Based on the information from the emotion engine, for example, if the user is relaxed, it suggests harmonies with a more relaxed atmosphere.

[0921] 4. Terminal: The terminal displays suggestions received from the server, such as "C major seventh," to the user. It also displays the analysis results of the emotion engine, making it easier for the user to understand the background of the suggestions.

[0922] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[0923] The system allows users to efficiently discover new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activities.

[0924] The processing flow will be explained below.

[0925] Step 1:

[0926] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, lyrics, etc. The server then stores this data in a local database.

[0927] Step 2:

[0928] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the generative AI model.

[0929] Step 3:

[0930] The terminal provides the user with an interface for inputting requests related to music production. For example, the user may input, "I would like some new harmony suggestions."

[0931] Step 4:

[0932] The terminal receives the user's request, analyzes its contents, and sends the analysis results to the server.

[0933] Step 5:

[0934] The device uses an emotion engine to recognize the user's emotions by analyzing the user's facial expressions, tone of voice, input text, etc.

[0935] Step 6:

[0936] The server analyzes the requests received from the terminal and the emotional information provided by the emotion engine, and takes this information into account when it needs to generate musical suggestions based on the emotional state.

[0937] Step 7:

[0938] The server uses a generative AI model to generate appropriate suggestions based on music theory, for example, generating specific harmonies such as a C major seventh chord if a new harmony suggestion is requested.

[0939] Step 8:

[0940] The server generates suggestions based on the user's emotional information obtained from the emotion engine: if the user is relaxed, it will suggest harmony with a relaxing atmosphere.

[0941] Step 9:

[0942] The server sends generated musical suggestions to the device, including suggestions derived from the generative AI model and emotion-based suggestions.

[0943] Step 10:

[0944] The terminal displays the proposal received from the server to the user using an interface that presents the proposal to the user visually or audibly.

[0945] Step 11:

[0946] The user can proceed with music production by referring to the suggestions displayed on the device. By also referring to emotion-based suggestions, music production can proceed more smoothly.

[0947] Example 2

[0948] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0949] Conventional music production support systems do not take into account the user's emotional state, making it difficult to provide personalized music suggestions that reflect the user's emotions. Furthermore, their ability to automatically generate suggestions based on music theory is limited. This makes it difficult for users to gain inspiration efficiently, potentially stagnateing their creative activities.

[0950] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0951] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for using an emotion engine to recognize the user's emotions, and means for displaying the generated suggestions to the user, thereby enabling personalized music suggestions to be provided according to the user's emotional state.

[0952] A "music database" is a type of data storage that collects data related to songs, including musical elements such as melody, harmony, rhythmic patterns, and lyrics.

[0953] "Music theory" is the study of rules and principles related to musical structure, form, harmony, melody, rhythm, etc. It is the basis for analyzing music and generating suggestions.

[0954] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate suggestions for music, text, etc., particularly natural language processing and machine learning techniques.

[0955] The "emotion engine" is a system for recognizing the user's emotional state by analyzing facial expressions, tone of voice, and the content of input text.

[0956] The "server" is a central control unit that processes and stores data and provides services to client systems. In this invention, it analyzes music data and processes user requests.

[0957] "Terminal" refers to a device that allows a user to access and operate the system through an interface. This includes computers, smartphones, tablets, etc.

[0958] A "user request" is a request a user makes to the system regarding a specific piece of music production. An example would be "I'd like some new harmony suggestions."

[0959] "Suggestions" refer to musical elements and ideas that the system provides to the user using generative AI models and emotion engines, which are used as references for music production.

[0960] MODE FOR CARRYING OUT THE INVENTION

[0961] The present invention provides a system for supporting music production by recognizing the emotional state of a user. Specific embodiments for carrying out the present invention will be described below.

[0962] Music data collection and analysis

[0963] The server collects music data from a specific music database (e.g., a music library API) on the cloud. The collected data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes the acquired music data using a Python music analysis library (e.g., music21, librosa), and stores the analysis results in a database. This accumulates data for training the generative AI model.

[0964] Receiving a user request

[0965] The device provides an interface through which users can input requests for music creation via a web application (e.g., React, Vue.js). For example, a user might input, "I'd like some new harmony suggestions." The device then sends the input request to the server using WebSocket or an HTTP request.

[0966] Proposal Generation

[0967] When the server receives a user request, it first analyzes the request using an NLP library (e.g., spaCy, NLTK). It then uses a generative AI model (e.g., GPT-3) to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, it generates specific chords, such as a C major seventh chord. It also incorporates information from an emotion engine to take the user's emotional state into account.

[0968] Emotion Engine Functions

[0969] The device uses a camera module (e.g., a webcam) to capture the user's facial expressions and simultaneously accepts voice input. The device then analyzes the user's emotions using emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer). The analysis results are sent to the server in real time, and the server uses this information to determine the user's current emotional state.

[0970] View Suggestions

[0971] The server sends the generated music suggestions and the analysis results from the emotion engine to the device. The device prepares UI components (e.g., HTML, CSS, JavaScript) to display the received suggestions to the user. The user then proceeds with the music creation process based on the displayed suggestions.

[0972] Specific examples

[0973] For example, if a user inputs a request to the terminal saying, "I want new harmony suggestions," the following series of processes are carried out.

[0974] 1. Device: The user types "I want new harmony suggestions" and clicks the send button. At the same time, the camera captures facial expressions and the microphone captures audio, which are then sent to the emotion engine.

[0975] Example prompt: "I'd like some new harmony suggestions for normal times."

[0976] 2. Server: Receives the request and analyzes it using an NLP library (e.g., spaCy). Uses a generative AI model (e.g., GPT-3) to generate specific harmony suggestions, such as "C major seventh chord." Takes emotion recognition results into account to suggest a relaxing atmosphere.

[0977] Example prompt: "I'd like some new harmony suggestions for when I'm relaxing."

[0978] 3. If the emotion engine determines the user's emotion as "relaxed," the generative AI model will adjust accordingly.

[0979] 4. Terminal: Displays to the user a suggestion that combines "C major seventh" received from the server with the results of sentiment analysis.

[0980] 5. The user then proceeds with creating music that corresponds to a specific emotion based on the displayed suggestions.

[0981] The system not only helps users find musical inspiration quickly, but also streamlines their creative process by receiving personalized suggestions based on their emotional state.

[0982] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0983] Step 1: Collecting music data

[0984] The server accesses a music database on the cloud. Specifically, it retrieves music data such as melody, harmony, rhythm patterns, and lyrics through an API (e.g., a music library API). This data is stored on the server and prepared for the next analysis step.

[0985] Input: Cloud music database

[0986] Output: Music data stored on the server

[0987] ---

[0988] Step 2: Analyzing the music data

[0989] The server analyzes the stored music data using Python music analysis libraries (e.g., music21, librosa). Specifically, it extracts the melody, harmony, rhythmic patterns, and lyric characteristics of each song. The analysis results are stored in a database in JSON or CSV format and used to train the generative AI model.

[0990] Input: Music data stored on the server

[0991] Output: Analyzed music data

[0992] ---

[0993] Step 3: Receiving a user request

[0994] The device provides an interface using a web application (e.g., React, Vue.js). The user inputs a request for music creation (e.g., "I'd like some new harmony suggestions.") The device receives this request and sends it to the server via WebSocket or HTTP request.

[0995] Input: User-entered music production requests

[0996] Output: Request sent to the server

[0997] ---

[0998] Step 4: Parsing the request

[0999] When the server receives a user request, it parses it using an NLP library (e.g., spaCy, NLTK) to understand the content of the request and determine what musical suggestions would be appropriate.

[1000] Input: The user request sent to the server

[1001] Output: Parsed request content

[1002] ---

[1003] Step 5: Generate proposals

[1004] The server uses a generative AI model (e.g., GPT-3) to generate music theory-based suggestions based on the analyzed request. For example, in response to a request for "new harmony suggestions," the server generates specific chords (e.g., a C major seventh chord).

[1005] Input: Analyzed request content and music theory data

[1006] Output: Generated music suggestions

[1007] ---

[1008] Step 6: Perform emotion recognition

[1009] The device uses a camera module (e.g., webcam) to capture the user's facial expressions and a microphone to capture their voice. This information is then sent to emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer) to analyze the user's emotions. The analysis results are then sent to a server.

[1010] Input: User's facial expressions, voice data

[1011] Output: Parsed emotion data

[1012] ---

[1013] Step 7: Integrating Emotional Data

[1014] The server integrates emotional data from the emotion engine into the music suggestions generated by the generative AI model. For example, if the server determines that the user is relaxed, it generates music suggestions that correspond to that emotion (e.g., relaxing harmony).

[1015] Input: Generated music suggestions, parsed emotion data

[1016] Output: Emotion-based music suggestions

[1017] ---

[1018] Step 8: Viewing Proposals

[1019] The device displays music suggestions to the user based on the emotion received from the server. Specifically, the suggestions are visually displayed using UI components (e.g., HTML, CSS, JavaScript).

[1020] Input: Emotion-based music suggestions received from the server

[1021] Output: Music suggestions displayed to the user

[1022] ---

[1023] Step 9: Music Production Progression

[1024] The user can then proceed with the composition of the music based on the displayed musical suggestions. For example, they can create a piece of music incorporating a C major seventh chord. At the same time, they can input new requests as needed, which are reflected in the system.

[1025] Input: Music suggestions displayed to the user

[1026] Output: Finished song, next request

[1027] ---

[1028] This process flow provides personalized music suggestions based on the user's emotional state, improving the efficiency of music production.

[1029] (Application example 2)

[1030] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1031] Conventional music creation support systems generate suggestions without considering the user's emotional state, which prevents personalized suggestions for each individual user and can lead to lower user satisfaction. Furthermore, users often find it difficult to find musical inspiration that matches their emotional state, limiting their creativity. A system that can resolve these issues and provide more effective and personalized music creation support is needed.

[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1033] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, emotion analysis means for recognizing the user's emotions, and means for displaying the generated suggestions to the user. This enables personalized musical suggestions based on the user's emotional state, thereby promoting creativity.

[1034] A "music database" is a large-scale information storage device for storing, saving, and managing music data.

[1035] "Music data" refers to information related to musical components such as melody, harmony, rhythm patterns, and lyrics.

[1036] "Music theory" is a set of rules and principles for understanding and analysing musical structures and components.

[1037] A "user request" is an instruction including a request or wish given by a user to a system.

[1038] A "generative AI model" is an artificial intelligence model that generates appropriate suggestions or answers based on given data or requests.

[1039] "Emotion analysis means" refers to technology or devices for recognizing and analyzing the user's emotional state.

[1040] "Suggestions" are specific ideas and advice about music production generated by the system.

[1041] The "means for displaying to the user" is an interface or device for visually presenting the generated suggestions to the user.

[1042] The present invention is a music creation support system that is combined with emotion analysis means for recognizing the emotions of the user. To specifically implement this system, the following elements are required:

[1043] Collection and analysis of music data

[1044] The server collects music data from a music database on the cloud. This music data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes this data based on music theory and saves the analysis results. The software used for this analysis is a music theory analysis program. The analysis results are used as training data for the generative AI model.

[1045] Receiving a user request

[1046] A user inputs a request for music creation using a device (such as a smartphone or a head-mounted display). For example, the request may be specific, such as "I would like new harmony suggestions." The device receives this request and sends it to the server.

[1047] emotion recognition

[1048] The device is equipped with an emotion analysis means that recognizes the user's emotional state by analyzing the user's facial expressions, tone of voice, and the content of input text, etc. This analysis is performed using the dlib library and voice analysis software.

[1049] Proposal generation

[1050] The server combines the user request with emotional data obtained from the emotion analysis tool and generates suggestions based on music theory using a generative AI model (e.g., GPT-3). For example, if a new harmony suggestion is requested, the server generates a specific harmony suggestion and adjusts it according to the user's emotional state. The generated suggestion can be specific, such as "C major seventh chord."

[1051] View Suggestions

[1052] The generated suggestions are displayed to the user via the device, allowing the user to use the suggestions as a reference for music production. In particular, the emotion analysis means provides suggestions that match the user's emotional state, allowing the user to obtain more personalized musical inspiration.

[1053] Specific examples

[1054] For example, if a user requests "I want new harmony suggestions," the device sends this request to the server and simultaneously analyzes the user's facial expressions to determine their emotions. The server uses a generative AI model based on the request and emotional data to generate new harmony suggestions. If the user is relaxed, the server will suggest a harmony with a relaxing atmosphere.

[1055] Prompt Sentence Examples

[1056] User request analysis: I would like suggestions for new harmonies. Emotion data: {'happiness': 0.8}

[1057] The system allows users to efficiently generate new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activity.

[1058] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1059] Program processing steps

[1060] Step 1:

[1061] The server collects music data from a music database. The input data are music files from the music database, including melody, harmony, rhythm patterns, and lyrics. The server analyzes this data based on music theory. This analysis saves the music data as metadata and serves as training data for the generative AI model.

[1062] Step 2:

[1063] The user inputs a request for music production through the device. The input data is a text request entered by the user. The device sends this request to the server. The server analyzes the received request and passes the analysis results to the generative AI model.

[1064] Step 3:

[1065] The device analyzes the user's emotions. Input data includes image capture, voice input, and text input. The device uses libraries such as dlib to determine the user's emotions from facial expressions and tone of voice. This data is processed by the emotion analysis means and output as an emotional state. The output data is sent to the server as emotion data.

[1066] Step 4:

[1067] The server uses a generative AI model based on the request analysis results and emotional data to generate suggestions based on music theory. The input data are the request analysis results and emotional data. The generative AI model (e.g., GPT-3) takes these data into account to generate suggestions. The output data is a specific musical suggestion (e.g., a C major seventh chord).

[1068] Step 5:

[1069] The server sends the generated proposal to the device, which then displays it to the user. The input data is the proposal from the server, and the output data is the music proposal displayed to the user. The user can use this as a reference to proceed with their music creation.

[1070] Examples:

[1071] For example, if a user requests a new harmony suggestion and the device recognizes a relaxed facial expression, the server will use the generative AI model based on the request and emotional data to suggest a C major seventh chord. This suggestion is displayed to the user via the device, and the user can use it as a reference when creating music.

[1072] Specific operations and input / output flow

[1073] Step 1:

[1074] Input: Music files from a music database

[1075] Processing: Data analysis based on music theory

[1076] Output: Analyzed music data

[1077] How it works: The server retrieves music files from a database, analyzes the data using a music theory analysis program, and stores it as metadata.

[1078] Step 2:

[1079] Input: The request text entered by the user

[1080] Processing: Parsing the request

[1081] Output: Request analysis results

[1082] How it works: The device receives a request from the user and sends it to the server, which analyzes the request and passes the results to the generative AI model.

[1083] Step 3:

[1084] Input: User's face image, voice, text

[1085] Processing: Emotion Recognition

[1086] Output: Emotion data

[1087] How it works: The device analyzes the user's emotions using libraries such as dlib and sends the results to the server.

[1088] Step 4:

[1089] Input: Request analysis results, emotion data

[1090] Processing: Generating music suggestions with a generative AI model

[1091] Output:Music suggestions

[1092] How it works: The server uses a generative AI model to generate music suggestions based on request analysis and emotional data.

[1093] Step 5:

[1094] Input: Generated Music Suggestions

[1095] Action:View Proposal

[1096] Output: What is displayed to the user

[1097] How it works: The server sends music suggestions generated by the generative AI model to the device, which then displays them to the user.

[1098] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1099] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1100] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1101] [Fourth embodiment]

[1102] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1103] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1104] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1105] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1106] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1107] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1108] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1109] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1110] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1111] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1112] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1113] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1114] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1115] The present invention relates to a music creation support system that solves the problems artists and creators face when creating music, such as a lack of inspiration and new ideas. The system of the present invention includes a processing means for collecting music data from a music database and analyzing it based on music theory, a processing means for receiving requests for music creation from users and analyzing the requests, a processing means for generating suggestions based on music theory using a generative AI model, and a processing means for displaying the generated suggestions to the user.

[1116] Explanation of program processing

[1117] Music data collection and analysis

[1118] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[1119] Receiving a user request

[1120] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user may input a request such as "I would like some new harmony suggestions." The terminal receives this request and sends it to the server.

[1121] Proposal generation

[1122] When the server receives a user's request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord.

[1123] View Suggestions

[1124] The terminal displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production.

[1125] Specific examples

[1126] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[1127] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[1128] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[1129] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[1130] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[1131] This system allows users to efficiently obtain new musical ideas and smoothly advance their creative activities.

[1132] The processing flow will be explained below.

[1133] Step 1:

[1134] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, and lyrics, and stores this data in a local database for analysis.

[1135] Step 2:

[1136] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the AI ​​model.

[1137] Step 3:

[1138] The terminal displays an interface for the user to input requests for music production. For example, the user may input, "I would like some new harmony suggestions."

[1139] Step 4:

[1140] The terminal receives a request from the user, analyzes the content of the request, and sends the analysis results to the server.

[1141] Step 5:

[1142] The server analyzes the request received from the device and, based on the request content, selects the appropriate process to use the generative AI model.

[1143] Step 6:

[1144] The server uses a generative AI model to generate musical suggestions based on the user's request, such as a C major seventh chord if asked for harmony suggestions.

[1145] Step 7:

[1146] The server sends the generated musical suggestions to the device, including suggestions derived from the generative AI model.

[1147] Step 8:

[1148] The terminal displays the suggestions received from the server to the user using an interface that makes it easy to provide the suggestions visually or audibly to the user.

[1149] Step 9:

[1150] Users can refer to the suggestions displayed on the device to proceed with music creation, which will help them gain new ideas and inspiration and make music creation easier.

[1151] Example 1

[1152] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1153] In modern music creation, artists and creators often face a lack of inspiration and new ideas. Furthermore, music production requires advanced knowledge of music theory, making it difficult for users without specialized knowledge to create music efficiently. A system that solves this problem and allows users to quickly and easily obtain musical ideas is needed.

[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1155] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for displaying the generated suggestions to the user, means for the generative AI model to generate specific musical ideas based on music theory based on the input requests, and means for presenting the generated suggestions to the user's terminal in real time, thereby enabling users to efficiently obtain new musical ideas.

[1156] A "music database" is a system for managing and storing a large amount of music data in one place.

[1157] "Music data" refers to information about the elements that make up a musical work, such as melody, harmony, rhythm patterns, and lyrics.

[1158] "Music theory" is the study of the basic principles of musical structure, chords, rhythm, melody, and their combinations and progressions.

[1159] A "server" is a computer system that runs behind the scenes of the entire system, handling requests, storing data, and delivering generated offers.

[1160] "User" refers to an individual or professional who uses this system to create music.

[1161] A "request" is a user's request for a specific musical suggestion or creation from the system.

[1162] A "generative AI model" is an artificial intelligence algorithm that generates new information or suggestions based on input data.

[1163] A "specific musical idea" is a specific creative element in musical composition, such as a particular melody, harmony, rhythmic pattern, or lyrics.

[1164] A "terminal" is a device, such as a computer or smartphone, that a user uses to interact with the system.

[1165] "Real-time" refers to a state in which user input and system processing are reflected immediately without delay.

[1166] The present invention relates to a music creation support system, and specific embodiments thereof will be described below.

[1167] Music data collection and analysis

[1168] The server collects music data from a music database on the cloud. The specific software used for this is a tool that extracts data through an API. The data includes elements such as melody, harmony, rhythm patterns, and lyrics. The server then uses a Python music analysis library (e.g., music21) to analyze the collected music data based on music theory. The analysis results are stored in a database system such as MySQL or PostgreSQL. This accumulates data for training the generative AI model.

[1169] Receiving a user request

[1170] The device provides the user with an interface for inputting requests related to music production. This interface is built using a front-end framework (e.g., React, Vue.js). The user inputs requests through this interface, specifically prompt sentences such as "I would like new harmony suggestions." The input request is sent to the server in JSON format.

[1171] Proposal generation

[1172] The server analyzes the received user request and generates suggestions based on music theory via a generative AI model (e.g., GPT-3). The generative AI model has the ability to generate new musical ideas based on training data. For example, in response to a request for "new harmony suggestions," it generates specific harmonies such as a C major seventh chord. The suggestions are returned to the device in JSON format.

[1173] View Suggestions

[1174] The device displays the suggestions received from the server to the user. Specifically, it uses JavaScript to render the suggestions into an HTML view and displays it in the user interface. This allows the user to use the generated suggestions as a reference for music production.

[1175] Specific examples

[1176] For example, when a user inputs a request such as "I want new harmony suggestions" into a terminal, the following is an example of the series of processes that will be carried out.

[1177] 1. Device: The user types "I want new harmony suggestions" and presses the send button. The device sends the request in JSON format to the server.

[1178] 2. Server: The server receives the request, analyzes it, and inputs the prompt "I would like some new harmony suggestions" into the generative AI model.

[1179] 3. Server: The generative AI model generates new harmonies, such as a C major seventh chord, and returns suggestions in JSON format to the device.

[1180] 4. Terminal: The terminal receives the JSON data and displays the suggestions in the user interface.

[1181] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[1182] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[1183] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1184] Step 1: Collecting music data

[1185] The server collects music data from a music database. It uses an API to obtain data such as melody, harmony, rhythm patterns, and lyrics. For example, it calls the API of a cloud service and receives music data in JSON format. The input is the music data from the API, and the output is music data stored in temporary data storage on the server.

[1186] Step 2: Analyzing the music data

[1187] The server analyzes the collected music data based on music theory. Specifically, it uses a Python music analysis library (e.g., music21) to analyze harmonic structure, rhythmic patterns, etc. The analysis results are stored in a database (e.g., MySQL, PostgreSQL) as training data. The input is the music data stored in temporary data storage, and the output is the analysis results stored in the database.

[1188] Step 3: Entering User Requests

[1189] Users input their music-making requests through a terminal interface. This interface is built using a front-end framework (e.g., React, Vue.js). Specifically, they input a prompt such as "I'd like some new harmony suggestions." The input is the user's request, and the output is data sent to the server in JSON format.

[1190] Step 4: Receiving and Parsing the Request

[1191] The server receives the request sent from the device. It then analyzes the request using a web framework (e.g., Flask, Django). Based on the analysis results, it creates a prompt to input to the generative AI model. The input is the request JSON sent from the device, and the output is the prompt to input to the generative AI model.

[1192] Step 5: Proposal Generation

[1193] The server uses a generative AI model (e.g., GPT-3) to generate suggestions based on music theory. The generative AI model generates new melodies and harmonies based on the training data. For example, it generates a specific harmony such as a C major seventh chord. The input is a prompt to the generative AI model, and the output is the generated musical suggestions.

[1194] Step 6: Submit your proposal

[1195] The server sends the generated proposals to the terminal in JSON format. The input is the proposal data from the generative AI model, and the output is the JSON data sent to the terminal.

[1196] Step 7: Viewing Proposals

[1197] The terminal displays the suggestions received from the server to the user by using JavaScript to render the suggestions into an HTML view and display it in the user interface. The input is the JSON data received from the server and the output is the suggestions displayed to the user.

[1198] In this way, the user can efficiently obtain new musical ideas and smoothly advance creative activities.

[1199] (Application example 1)

[1200] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1201] Conventional music production support systems have the problem of being unable to effectively solve the problem of running out of inspiration and lack of new ideas. In particular, it has been difficult to instantly provide new harmony, melody, and lyric ideas while creating music. Furthermore, because such systems are not linked to content distribution services, they have been unable to quickly provide the latest musical ideas to a wide range of users. To solve these problems, a more efficient and practical music creation support system for users is needed.

[1202] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1203] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, and means for applying the suggestions to the content distribution service, thereby enabling users to quickly and efficiently obtain new musical ideas.

[1204] A "music database" is a database that stores multiple pieces of music data and allows searching and retrieval.

[1205] "Music data" refers to the elements that make up a song, such as melody, harmony, rhythm, and lyrics.

[1206] "Music theory" is the academic knowledge of musical structure and composition, which serves as a guide for analyzing and creating musical pieces.

[1207] A "generative AI model" is a technology that uses artificial intelligence to generate new music suggestions based on user requests.

[1208] A "request" is a specific request for music production, such as a new harmony, melody, or lyrics that the user desires.

[1209] A "content distribution service" is a service that provides users with digital content such as music and videos via the Internet.

[1210] "Suggestions" are new musical ideas generated by the generative AI model to assist users in their music creation.

[1211] "Analysis" is the process of understanding a user request and analyzing its content to generate appropriate music suggestions.

[1212] This invention is a music creation support system that allows artists and creators to obtain new ideas in music production. The system includes the following means.

[1213] System Configuration

[1214] 1. A server that collects music data from a music database and analyzes it based on music theory

[1215] The server resides in the cloud and collects multiple pieces of music data (melody, harmony, rhythm, lyrics, etc.) from a music database. The collected data is analyzed based on music theory and used as training data for the generative AI model.

[1216] 2. A device that receives requests from users regarding music production and analyzes those requests.

[1217] The device provides an interface for users to input music-making requests, such as "I want new harmony suggestions." The device receives the request and sends it to the server for analysis.

[1218] 3. A server that uses generative AI models to generate suggestions based on music theory

[1219] The server analyzes the user's request and generates music theory-based suggestions based on the request using a generative AI model (e.g., GPTNeo). For example, if the user's request is for a new harmony, the server generates a specific harmony, such as a C major seventh chord.

[1220] 4. A device that displays the generated suggestions to the user.

[1221] The generated suggestions are sent from the server to the device, which then displays them to the user, who can use them as a reference to proceed with their music creation.

[1222] 5. Methods applied to content distribution services

[1223] This system is particularly linked to content distribution services, making it possible to quickly provide musical ideas to a wide range of users via the Internet, allowing users to efficiently acquire new musical ideas and support their creative activities.

[1224] Hardware and software used

[1225] Hardware

[1226] Server: Cloud server

[1227] Device: Smartphone

[1228] software

[1229] Server: Python, requests library

[1230] Generative AI model: Hugging Face transformers library, GPTNeo model

[1231] Device: Smartphone application

[1232] Specific examples

[1233] For example, if a user inputs a request such as "I want new harmony suggestions" through a smartphone application, the following processing will occur.

[1234] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server.

[1235] 2. Server: The server receives the request, analyzes it, and then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested.

[1236] 3. Terminal: The terminal displays the suggestion "C major seventh" received from the server to the user.

[1237] 4. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[1238] Prompt Sentence Examples

[1239] "Users are asking for suggestions: I'd like some new harmony suggestions."

[1240] This system allows users to efficiently eliminate inspiration drain and easily come up with new musical ideas.

[1241] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1242] Step 1:

[1243] The user uses the terminal to enter a request

[1244] A user opens a smartphone application and enters a music-making request, such as "I'd like some new harmony suggestions." The device receives this request and converts it into data to send to the server. The input is the user's request, and the output is a data packet containing that request.

[1245] Step 2:

[1246] The device sends a request to the server

[1247] The terminal sends the request received from the user to the server as an HTTP POST request. This request includes the user ID and the request content. The input is the request data from the user, and the output is the HTTP POST request sent to the server.

[1248] Step 3:

[1249] The server receives and parses the request

[1250] The server receives requests sent from the device and analyzes their contents. The analysis involves analyzing the text of the request, for example, creating a prompt appropriate for the generative AI model based on a request such as "I'd like some new harmony suggestions." The input is the HTTP POST request sent to the server, and the output is the analysis result, including the prompt text.

[1251] Step 4:

[1252] The server generates suggestions using a generative AI model

[1253] The server inputs the analysis results into a generative AI model to generate a suggestion based on music theory. For example, a generative AI model (GPTNeo) is used to generate a new suggestion for the requested harmony (e.g., a C major seventh chord). The input is a prompt based on the analysis results, and the output is a musical suggestion generated by the generative AI model.

[1254] Step 5:

[1255] The server sends the generated proposal to the device.

[1256] The server sends the generated music suggestions to the device, where the suggestions are sent as an HTTP response and formatted appropriately for the user. The input is the music suggestions generated by the AI ​​model, and the output is the HTTP response sent to the device.

[1257] Step 6:

[1258] The terminal displays the suggestions to the user.

[1259] The device receives the response from the server and displays it to the user in an appropriate format. For example, "New harmony suggestion: C major seventh chord" is displayed on the screen. The input is the HTTP response from the server, and the output is the suggestion information displayed to the user.

[1260] Step 7:

[1261] The user will use the suggestions to proceed with music production.

[1262] The user can refer to the new music suggestions displayed on the device and proceed with the composition of the music. In this process, the user can gain new inspiration and ideas. The input is the suggested information displayed on the device, and the output is the user's creative activity.

[1263] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1264] The present invention further relates to a music creation support system that combines an emotion engine that recognizes user emotions. The system includes processing means for collecting music data from a music database and analyzing it based on music theory, processing means for receiving music creation requests from users and analyzing the requests, processing means for generating suggestions based on music theory using a generative AI model, processing means for displaying the generated suggestions to the user, and an emotion engine that recognizes user emotions.

[1265] Explanation of program processing

[1266] Music data collection and analysis

[1267] The server collects music data from a cloud-based music database, including melody, harmony, rhythmic patterns, and lyrics. The server then analyzes this data based on music theory and saves the results. This provides training data for the generative AI model.

[1268] Receiving a user request

[1269] The terminal provides the user with an interface for inputting requests related to music creation. For example, the user might input, "I'd like some new harmony suggestions." The terminal receives this request and sends it to the server.

[1270] Proposal generation

[1271] When the server receives a user request, it analyzes the request and then uses a generative AI model to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, the server will generate a specific harmony, such as a C major seventh chord. It also uses an emotion engine to generate suggestions that correspond to the user's emotions.

[1272] Emotion Engine Functions

[1273] The device uses an emotion engine to recognize the user's emotions. Emotion recognition utilizes the user's facial expressions, tone of voice, and input text content. The emotion engine analyzes this information to determine the user's current emotional state.

[1274] View Suggestions

[1275] The device displays the suggestions received from the server to the user, allowing the user to use the generated suggestions as a reference for music production. In particular, the emotion engine provides suggestions that correspond to the user's emotions, allowing the user to obtain more appropriate musical inspiration.

[1276] Specific examples

[1277] For example, this is an example of a series of processes when a user inputs a request to the terminal saying, "I want new harmony suggestions."

[1278] 1. Device: The user types "I want new harmony suggestions." The device sends this request to the server. At the same time, the emotion engine analyzes the user's facial expressions to recognize their emotions.

[1279] 2. Server: The server receives the request and analyzes it. It then uses a generative AI model to generate new harmony suggestions based on music theory. In this case, a C major seventh chord is suggested. At the same time, emotional information from the emotion engine is taken into account.

[1280] 3. Server: Based on the information from the emotion engine, for example, if the user is relaxed, it suggests harmonies with a more relaxed atmosphere.

[1281] 4. Terminal: The terminal displays suggestions received from the server, such as "C major seventh," to the user. It also displays the analysis results of the emotion engine, making it easier for the user to understand the background of the suggestions.

[1282] 5. User: The user uses the proposed harmonies as a reference to proceed with the composition of the song.

[1283] The system allows users to efficiently discover new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activities.

[1284] The processing flow will be explained below.

[1285] Step 1:

[1286] The server collects music data from a cloud-based music database, including melody, harmony, rhythm patterns, lyrics, etc. The server then stores this data in a local database.

[1287] Step 2:

[1288] The server analyzes the collected music data based on music theory, analyzing musical components and progression patterns, and saving the data as training data for the generative AI model.

[1289] Step 3:

[1290] The terminal provides the user with an interface for inputting requests related to music production. For example, the user may input, "I would like some new harmony suggestions."

[1291] Step 4:

[1292] The terminal receives the user's request, analyzes its contents, and sends the analysis results to the server.

[1293] Step 5:

[1294] The device uses an emotion engine to recognize the user's emotions by analyzing the user's facial expressions, tone of voice, input text, etc.

[1295] Step 6:

[1296] The server analyzes the requests received from the terminal and the emotional information provided by the emotion engine, and takes this information into account when it needs to generate musical suggestions based on the emotional state.

[1297] Step 7:

[1298] The server uses a generative AI model to generate appropriate suggestions based on music theory, for example, generating specific harmonies such as a C major seventh chord if a new harmony suggestion is requested.

[1299] Step 8:

[1300] The server generates suggestions based on the user's emotional information obtained from the emotion engine: if the user is relaxed, it will suggest harmony with a relaxing atmosphere.

[1301] Step 9:

[1302] The server sends generated musical suggestions to the device, including suggestions derived from the generative AI model and emotion-based suggestions.

[1303] Step 10:

[1304] The terminal displays the proposal received from the server to the user using an interface that presents the proposal to the user visually or audibly.

[1305] Step 11:

[1306] The user can proceed with music production by referring to the suggestions displayed on the device. By also referring to emotion-based suggestions, music production can proceed more smoothly.

[1307] Example 2

[1308] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1309] Conventional music production support systems do not take into account the user's emotional state, making it difficult to provide personalized music suggestions that reflect the user's emotions. Furthermore, their ability to automatically generate suggestions based on music theory is limited. This makes it difficult for users to gain inspiration efficiently, potentially stagnateing their creative activities.

[1310] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1311] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, means for using an emotion engine to recognize the user's emotions, and means for displaying the generated suggestions to the user, thereby enabling personalized music suggestions to be provided according to the user's emotional state.

[1312] A "music database" is a type of data storage that collects data related to songs, including musical elements such as melody, harmony, rhythmic patterns, and lyrics.

[1313] "Music theory" is the study of rules and principles related to musical structure, form, harmony, melody, rhythm, etc. It is the basis for analyzing music and generating suggestions.

[1314] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate suggestions for music, text, etc., particularly natural language processing and machine learning techniques.

[1315] The "emotion engine" is a system for recognizing the user's emotional state by analyzing facial expressions, tone of voice, and the content of input text.

[1316] The "server" is a central control unit that processes and stores data and provides services to client systems. In this invention, it analyzes music data and processes user requests.

[1317] "Terminal" refers to a device that allows a user to access and operate the system through an interface. This includes computers, smartphones, tablets, etc.

[1318] A "user request" is a request a user makes to the system regarding a specific piece of music production. An example would be "I'd like some new harmony suggestions."

[1319] "Suggestions" refer to musical elements and ideas that the system provides to the user using generative AI models and emotion engines, which are used as references for music production.

[1320] MODE FOR CARRYING OUT THE INVENTION

[1321] The present invention provides a system for supporting music production by recognizing the emotional state of a user. Specific embodiments for carrying out the present invention will be described below.

[1322] Music data collection and analysis

[1323] The server collects music data from a specific music database (e.g., a music library API) on the cloud. The collected data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes the acquired music data using a Python music analysis library (e.g., music21, librosa), and stores the analysis results in a database. This accumulates data for training the generative AI model.

[1324] Receiving a user request

[1325] The device provides an interface through which users can input requests for music creation via a web application (e.g., React, Vue.js). For example, a user might input, "I'd like some new harmony suggestions." The device then sends the input request to the server using WebSocket or an HTTP request.

[1326] Proposal Generation

[1327] When the server receives a user request, it first analyzes the request using an NLP library (e.g., spaCy, NLTK). It then uses a generative AI model (e.g., GPT-3) to generate appropriate suggestions based on music theory. For example, if a new harmony suggestion is requested, it generates specific chords, such as a C major seventh chord. It also incorporates information from an emotion engine to take the user's emotional state into account.

[1328] Emotion Engine Functions

[1329] The device uses a camera module (e.g., a webcam) to capture the user's facial expressions and simultaneously accepts voice input. The device then analyzes the user's emotions using emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer). The analysis results are sent to the server in real time, and the server uses this information to determine the user's current emotional state.

[1330] View Suggestions

[1331] The server sends the generated music suggestions and the analysis results from the emotion engine to the device. The device prepares UI components (e.g., HTML, CSS, JavaScript) to display the received suggestions to the user. The user then proceeds with the music creation process based on the displayed suggestions.

[1332] Specific examples

[1333] For example, if a user inputs a request to the terminal saying, "I want new harmony suggestions," the following series of processes are carried out.

[1334] 1. Device: The user types "I want new harmony suggestions" and clicks the send button. At the same time, the camera captures facial expressions and the microphone captures audio, which are then sent to the emotion engine.

[1335] Example prompt: "I'd like some new harmony suggestions for normal times."

[1336] 2. Server: Receives the request and analyzes it using an NLP library (e.g., spaCy). Uses a generative AI model (e.g., GPT-3) to generate specific harmony suggestions, such as "C major seventh chord." Takes emotion recognition results into account to suggest a relaxing atmosphere.

[1337] Example prompt: "I'd like some new harmony suggestions for when I'm relaxing."

[1338] 3. If the emotion engine determines the user's emotion as "relaxed," the generative AI model will adjust accordingly.

[1339] 4. Terminal: Displays to the user a suggestion that combines "C major seventh" received from the server with the results of sentiment analysis.

[1340] 5. The user then proceeds with creating music that corresponds to a specific emotion based on the displayed suggestions.

[1341] The system not only helps users find musical inspiration quickly, but also streamlines their creative process by receiving personalized suggestions based on their emotional state.

[1342] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1343] Step 1: Collecting music data

[1344] The server accesses a music database on the cloud. Specifically, it retrieves music data such as melody, harmony, rhythm patterns, and lyrics through an API (e.g., a music library API). This data is stored on the server and prepared for the next analysis step.

[1345] Input: Cloud music database

[1346] Output: Music data stored on the server

[1347] ---

[1348] Step 2: Analyzing the music data

[1349] The server analyzes the stored music data using Python music analysis libraries (e.g., music21, librosa). Specifically, it extracts the melody, harmony, rhythmic patterns, and lyric characteristics of each song. The analysis results are stored in a database in JSON or CSV format and used to train the generative AI model.

[1350] Input: Music data stored on the server

[1351] Output: Analyzed music data

[1352] ---

[1353] Step 3: Receiving a user request

[1354] The device provides an interface using a web application (e.g., React, Vue.js). The user inputs a request for music creation (e.g., "I'd like some new harmony suggestions.") The device receives this request and sends it to the server via WebSocket or HTTP request.

[1355] Input: User-entered music production requests

[1356] Output: Request sent to the server

[1357] ---

[1358] Step 4: Parsing the request

[1359] When the server receives a user request, it parses it using an NLP library (e.g., spaCy, NLTK) to understand the content of the request and determine what musical suggestions would be appropriate.

[1360] Input: The user request sent to the server

[1361] Output: Parsed request content

[1362] ---

[1363] Step 5: Generate proposals

[1364] The server uses a generative AI model (e.g., GPT-3) to generate music theory-based suggestions based on the analyzed request. For example, in response to a request for "new harmony suggestions," the server generates specific chords (e.g., a C major seventh chord).

[1365] Input: Analyzed request content and music theory data

[1366] Output: Generated music suggestions

[1367] ---

[1368] Step 6: Perform emotion recognition

[1369] The device uses a camera module (e.g., webcam) to capture the user's facial expressions and a microphone to capture their voice. This information is then sent to emotion recognition software (e.g., Microsoft Azure Face API, IBM Watson Tone Analyzer) to analyze the user's emotions. The analysis results are then sent to a server.

[1370] Input: User's facial expressions, voice data

[1371] Output: Parsed emotion data

[1372] ---

[1373] Step 7: Integrating Emotional Data

[1374] The server integrates emotional data from the emotion engine into the music suggestions generated by the generative AI model. For example, if the server determines that the user is relaxed, it generates music suggestions that correspond to that emotion (e.g., relaxing harmony).

[1375] Input: Generated music suggestions, parsed emotion data

[1376] Output: Emotion-based music suggestions

[1377] ---

[1378] Step 8: Viewing Proposals

[1379] The device displays music suggestions to the user based on the emotion received from the server. Specifically, the suggestions are visually displayed using UI components (e.g., HTML, CSS, JavaScript).

[1380] Input: Emotion-based music suggestions received from the server

[1381] Output: Music suggestions displayed to the user

[1382] ---

[1383] Step 9: Music Production Progression

[1384] The user can then proceed with the composition of the music based on the displayed musical suggestions. For example, they can create a piece of music incorporating a C major seventh chord. At the same time, they can input new requests as needed, which are reflected in the system.

[1385] Input: Music suggestions displayed to the user

[1386] Output: Finished song, next request

[1387] ---

[1388] This process flow provides personalized music suggestions based on the user's emotional state, improving the efficiency of music production.

[1389] (Application example 2)

[1390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1391] Conventional music creation support systems generate suggestions without considering the user's emotional state, which prevents personalized suggestions for each individual user and can lead to lower user satisfaction. Furthermore, users often find it difficult to find musical inspiration that matches their emotional state, limiting their creativity. A system that can resolve these issues and provide more effective and personalized music creation support is needed.

[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1393] In this invention, the server includes means for collecting music data from a music database and analyzing it based on music theory, means for receiving music production requests from users and analyzing the requests, means for generating suggestions based on music theory using a generative AI model, emotion analysis means for recognizing the user's emotions, and means for displaying the generated suggestions to the user. This enables personalized musical suggestions based on the user's emotional state, thereby promoting creativity.

[1394] A "music database" is a large-scale information storage device for storing, saving, and managing music data.

[1395] "Music data" refers to information related to musical components such as melody, harmony, rhythm patterns, and lyrics.

[1396] "Music theory" is a set of rules and principles for understanding and analysing musical structures and components.

[1397] A "user request" is an instruction including a request or wish given by a user to a system.

[1398] A "generative AI model" is an artificial intelligence model that generates appropriate suggestions or answers based on given data or requests.

[1399] "Emotion analysis means" refers to technology or devices for recognizing and analyzing the user's emotional state.

[1400] "Suggestions" are specific ideas and advice about music production generated by the system.

[1401] The "means for displaying to the user" is an interface or device for visually presenting the generated suggestions to the user.

[1402] The present invention is a music creation support system that is combined with emotion analysis means for recognizing the emotions of the user. To specifically implement this system, the following elements are required:

[1403] Collection and analysis of music data

[1404] The server collects music data from a music database on the cloud. This music data includes melody, harmony, rhythm patterns, lyrics, etc. The server then analyzes this data based on music theory and saves the analysis results. The software used for this analysis is a music theory analysis program. The analysis results are used as training data for the generative AI model.

[1405] Receiving a user request

[1406] A user inputs a request for music creation using a device (such as a smartphone or a head-mounted display). For example, the request may be specific, such as "I would like new harmony suggestions." The device receives this request and sends it to the server.

[1407] emotion recognition

[1408] The device is equipped with an emotion analysis means that recognizes the user's emotional state by analyzing the user's facial expressions, tone of voice, and the content of input text, etc. This analysis is performed using the dlib library and voice analysis software.

[1409] Proposal generation

[1410] The server combines the user request with emotional data obtained from the emotion analysis tool and generates suggestions based on music theory using a generative AI model (e.g., GPT-3). For example, if a new harmony suggestion is requested, the server generates a specific harmony suggestion and adjusts it according to the user's emotional state. The generated suggestion can be specific, such as "C major seventh chord."

[1411] View Suggestions

[1412] The generated suggestions are displayed to the user via the device, allowing the user to use the suggestions as a reference for music production. In particular, the emotion analysis means provides suggestions that match the user's emotional state, allowing the user to obtain more personalized musical inspiration.

[1413] Specific examples

[1414] For example, if a user requests "I want new harmony suggestions," the device sends this request to the server and simultaneously analyzes the user's facial expressions to determine their emotions. The server uses a generative AI model based on the request and emotional data to generate new harmony suggestions. If the user is relaxed, the server will suggest a harmony with a relaxing atmosphere.

[1415] Prompt Sentence Examples

[1416] User request analysis: I would like suggestions for new harmonies. Emotion data: {'happiness': 0.8}

[1417] The system allows users to efficiently generate new musical ideas and provides more personalized, emotion-based suggestions, facilitating creative activity.

[1418] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1419] Program processing steps

[1420] Step 1:

[1421] The server collects music data from a music database. The input data are music files from the music database, including melody, harmony, rhythm patterns, and lyrics. The server analyzes this data based on music theory. This analysis saves the music data as metadata and serves as training data for the generative AI model.

[1422] Step 2:

[1423] The user inputs a request for music production through the device. The input data is a text request entered by the user. The device sends this request to the server. The server analyzes the received request and passes the analysis results to the generative AI model.

[1424] Step 3:

[1425] The device analyzes the user's emotions. Input data includes image capture, voice input, and text input. The device uses libraries such as dlib to determine the user's emotions from facial expressions and tone of voice. This data is processed by the emotion analysis means and output as an emotional state. The output data is sent to the server as emotion data.

[1426] Step 4:

[1427] The server uses a generative AI model based on the request analysis results and emotional data to generate suggestions based on music theory. The input data are the request analysis results and emotional data. The generative AI model (e.g., GPT-3) takes these data into account to generate suggestions. The output data is a specific musical suggestion (e.g., a C major seventh chord).

[1428] Step 5:

[1429] The server sends the generated proposal to the device, which then displays it to the user. The input data is the proposal from the server, and the output data is the music proposal displayed to the user. The user can use this as a reference to proceed with their music creation.

[1430] Examples:

[1431] For example, if a user requests a new harmony suggestion and the device recognizes a relaxed facial expression, the server will use the generative AI model based on the request and emotional data to suggest a C major seventh chord. This suggestion is displayed to the user via the device, and the user can use it as a reference when creating music.

[1432] Specific operations and input / output flow

[1433] Step 1:

[1434] Input: Music files from a music database

[1435] Processing: Data analysis based on music theory

[1436] Output: Analyzed music data

[1437] How it works: The server retrieves music files from a database, analyzes the data using a music theory analysis program, and stores it as metadata.

[1438] Step 2:

[1439] Input: The request text entered by the user

[1440] Processing: Parsing the request

[1441] Output: Request analysis results

[1442] How it works: The device receives a request from the user and sends it to the server, which analyzes the request and passes the results to the generative AI model.

[1443] Step 3:

[1444] Input: User's face image, voice, text

[1445] Processing: Emotion Recognition

[1446] Output: Emotion data

[1447] How it works: The device analyzes the user's emotions using libraries such as dlib and sends the results to the server.

[1448] Step 4:

[1449] Input: Request analysis results, emotion data

[1450] Processing: Generating music suggestions with a generative AI model

[1451] Output:Music suggestions

[1452] How it works: The server uses a generative AI model to generate music suggestions based on request analysis and emotional data.

[1453] Step 5:

[1454] Input: Generated Music Suggestions

[1455] Action:View Proposal

[1456] Output: What is displayed to the user

[1457] How it works: The server sends music suggestions generated by the generative AI model to the device, which then displays them to the user.

[1458] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1459] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1460] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1461] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1462] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1463] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1464] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1465] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1466] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1467] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1468] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1469] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1470] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1471] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1472] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1473] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1474] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1475] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1476] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1477] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1478] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1479] The following is further disclosed regarding the above embodiment.

[1480] (Claim 1)

[1481] A means for collecting music data from a music database and analyzing it based on music theory;

[1482] means for receiving a music production request from a user and analyzing the request;

[1483] a means for generating music theory-based suggestions using a generative AI model;

[1484] means for displaying the generated suggestions to the user;

[1485] A system including:

[1486] (Claim 2)

[1487] 10. The system of claim 1, wherein the generated suggestions are related to harmony.

[1488] (Claim 3)

[1489] 10. The system of claim 1, wherein the generated suggestions relate to song lyrics.

[1490] "Example 1"

[1491] (Claim 1)

[1492] A means for collecting music data from a music database and analyzing it based on music theory;

[1493] means for receiving a music production request from a user and analyzing the request;

[1494] a means for generating music theory-based suggestions using a generative AI model;

[1495] means for displaying the generated suggestions to the user;

[1496] a means for generating specific musical ideas based on music theory in response to the input request by the generative AI model;

[1497] means for presenting the generated proposals to a user's terminal in real time;

[1498] A system including:

[1499] (Claim 2)

[1500] 10. The system of claim 1, wherein the generated suggestions are related to harmony.

[1501] (Claim 3)

[1502] 10. The system of claim 1, wherein the generated suggestions relate to song lyrics.

[1503] "Application Example 1"

[1504] (Claim 1)

[1505] A means for collecting music data from a music database and analyzing it based on music theory;

[1506] means for receiving a music production request from a user and analyzing the request;

[1507] a means for generating music theory-based suggestions using a generative AI model;

[1508] means for displaying the generated suggestions to the user;

[1509] The means by which the proposal can be applied to content distribution services,

[1510] A system including:

[1511] (Claim 2)

[1512] 10. The system of claim 1, wherein the generated suggestions are related to harmony.

[1513] (Claim 3)

[1514] 10. The system of claim 1, wherein the generated suggestions relate to song lyrics.

[1515] "Example 2: Combining Emotion Engines"

[1516] (Claim 1)

[1517] A means for collecting music data from a music database and analyzing it based on music theory;

[1518] means for receiving a music production request from a user and analyzing the request;

[1519] a means for generating music theory-based suggestions using a generative AI model;

[1520] means for using an emotion engine to recognize the emotion of a user;

[1521] means for displaying the generated suggestions to the user;

[1522] A system including:

[1523] (Claim 2)

[1524] 10. The system of claim 1, wherein the generated suggestions are related to harmony.

[1525] (Claim 3)

[1526] 10. The system of claim 1, wherein the generated suggestions relate to song lyrics.

[1527] "Application example 2 when combining emotion engines"

[1528] (Claim 1)

[1529] A means for collecting music data from a music database and analyzing it based on music theory;

[1530] means for receiving a music production request from a user and analyzing the request;

[1531] a means for generating music theory-based suggestions using a generative AI model;

[1532] emotion analysis means for recognizing the emotion of a user;

[1533] means for displaying the generated suggestions to the user;

[1534] A system including:

[1535] (Claim 2)

[1536] 10. The system of claim 1, wherein the generated suggestions are harmonic and tailored based on the user's emotional state.

[1537] (Claim 3)

[1538] 10. The system of claim 1, wherein the generated suggestions are lyric-related and tailored based on the user's emotional state. [Explanation of symbols]

[1539] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for collecting music data from a music database and analyzing it based on music theory; means for receiving a music production request from a user and analyzing the request; a means for generating music theory-based suggestions using a generative AI model; means for displaying the generated suggestions to the user; A system including:

2. 10. The system of claim 1, wherein the generated suggestions relate to harmony.

3. The system of claim 1 , wherein the generated suggestions relate to song lyrics.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A