system

The music generation AI system addresses the inefficiencies and copyright limitations of traditional music production by enabling rapid, cost-effective, and flexible music creation.

JP2026068492APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional music production is time-consuming, costly, and requires specialized knowledge, and is limited by copyright restrictions, lacking flexibility in music use and editing.

Method used

A system utilizing music generation AI to automatically create music based on user input, allowing for rapid, cost-effective music production without copyright issues.

Benefits of technology

Enables quick and affordable music generation that is customizable and free from copyright restrictions, enhancing user flexibility and efficiency in music creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068492000001_ABST
    Figure 2026068492000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for automatically generating music based on specified music information using music generation AI, A terminal means for receiving the aforementioned music information from the user and inputting it into the music generation AI in an appropriate format, A server means for transmitting the automatically generated music to the terminal in order to provide it to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The conventional music production process required a lot of time and cost, and a large number of personnel with specialized knowledge had to be involved, so there was a problem that efficient and rapid music production could not be carried out. In addition, there was also a problem that there was a lack of flexibility in the use and editing of new music due to copyright issues.

Means for Solving the Problems

[0005] This invention provides a system that automatically generates music based on music information (e.g., title, style, keywords) input by a user, using a music generation AI. This system includes a means by which the user inputs music information from a terminal, and the server passes this information to the music generation AI, instantly generating music that meets the specified conditions and providing it to the terminal. This enables rapid and cost-effective music production and allows for the free use of music without copyright restrictions.

[0006] "Music generation AI" is a system that uses artificial intelligence technology to automatically generate music based on specified musical styles and elements.

[0007] A "musical piece" is a collection of sounds that make up music, and is a musical work consisting of melody, rhythm, chords, etc.

[0008] A "user" refers to an individual or group that operates the system and inputs music information.

[0009] A "terminal" refers to an input / output device used by users to input music information or receive generated music.

[0010] A "server" is a computer system that sends and receives data to and from the music generation AI, and manages data communication between the terminal and the music generation AI.

[0011] "Music information" refers to the information that users input to the music generation AI to generate songs, and includes, for example, the song title, style, and keywords.

[0012] "Automated generation" refers to the process of automatically generating results by machine or software without human intervention.

[0013] "Style" refers to a genre or form of music that includes musical elements such as rhythm, melody, and harmony that define the characteristics of a song.

[0014] A "keyword" refers to an important word or phrase that influences the lyrics or theme of a song. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention provides a method for users to easily generate music in a system using music generation AI. This system is implemented through a series of steps in which the user inputs music information via a terminal, sends that information to a server, and the music generation AI automatically generates a song.

[0037] First, the user launches a music generation application using their device. This device can be a computer, smartphone, or tablet. The user enters music information such as the song title, style, and lyrics keywords into the application screen. For example, if the user wants to generate a song titled "Summer Campaign," in the "Pop" style, and include keywords like "beach," "sun," and "fun" in the lyrics, they would enter this information into their device.

[0038] Next, the terminal sends the user's input information to the server. This server receives the music information, converts it into an appropriate format, and provides it to the music generation AI. The music generation AI on the server receives this information and generates a song using machine learning models and algorithms stored internally. In this process, the AI ​​automatically selects and integrates melodies, rhythms, harmonies, etc., based on the specified style and keywords.

[0039] The generated music is sent to the device via a server. Users can then view, play, and use this music in advertisements and media content on their devices. Because this process is completed quickly, it significantly reduces time and cost compared to traditional music production. Furthermore, since newly generated music by AI is not subject to existing copyright issues, it offers a high degree of freedom in its use.

[0040] This invention streamlines the music production process and allows users to quickly generate music that meets diverse needs.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user launches the music generation application using their device. The user fills in information about the song in the input form displayed on the screen. Specifically, they enter the song title, style (e.g., "Pop"), and keywords they want to include in the lyrics (e.g., "Beach," "Sun," "Fun").

[0044] Step 2:

[0045] The device verifies the music information entered by the user and converts it to the appropriate data format. After this conversion is complete, the device sends a music generation request to the server.

[0046] Step 3:

[0047] The server analyzes the music information received from the terminal and prepares it as input data for the music generation AI. The server selects a specific music generation algorithm based on the user's request and prepares to invoke the AI.

[0048] Step 4:

[0049] The music generation AI is activated and begins the music generation process based on the data received from the server. The AI ​​synthesizes a song by combining elements such as melody, rhythm, and chords to match the specified style and keywords.

[0050] Step 5:

[0051] After the music generation is complete, the server receives the music file generated by the AI ​​and verifies it for errors. It then prepares to send the verified file to the terminal.

[0052] Step 6:

[0053] The device receives generated music data sent from the server and displays it to the user in a playable format. The user can review the music and save or re-edit it as needed.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] Traditional music generation methods require specialized knowledge and a long time, making it difficult for ordinary users to easily create musical works. Furthermore, existing music is subject to copyright and other restrictions, making free use difficult. Therefore, there is a need for the development of music generation technologies that can quickly meet diverse needs.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes processing means for automatically generating musical works based on input musical information using artificial intelligence that generates music, information processing means for acquiring the musical information from the user and providing it to the artificial intelligence that generates music in an appropriate format, and data processing means for transmitting the automatically generated musical works to the information processing means in order to supply them to the user. As a result, users can generate and use original musical works in a short time without requiring specialized knowledge.

[0059] "Artificial intelligence that generates music" refers to a machine learning system equipped with algorithms and models that automatically generate musical works based on input song information.

[0060] "Song information" refers to data that users input as elements necessary for music creation, and includes title, style, lyrics keywords, and so on.

[0061] An "information processing device" refers to electronic equipment and software that converts music information received from a user into an appropriate format and provides it to artificial intelligence that generates music.

[0062] A "data processing device" is a device or system used to transmit musical works generated by artificial intelligence that performs music generation to a user.

[0063] A "musical work" refers to audio data, including melody, rhythm, and harmony, that is automatically generated by artificial intelligence used for music generation.

[0064] This invention is a system that uses artificial intelligence (generative AI model) to generate music, automatically generating musical works based on song information entered by the user. In the implementation of this system, the server, terminals, and users play important roles.

[0065] Users can launch the music generation application using devices such as computers, smartphones, and tablets. Through the application's interface, users input song information such as the song title, style, and lyric keywords. This information is the basic data necessary for generating the musical work.

[0066] The terminal has the functionality to send music information entered by the user to a server. The server receives this information and converts it into an appropriate format. The artificial intelligence that generates music, running on the server, uses the transmitted information to generate a musical work. The generation AI model utilizes neural networks and machine learning algorithms based on the input information to generate a musical work that includes melody, rhythm, and harmony. The generated musical work is then sent from the server to the terminal in the form of a digital music file.

[0067] Users can play and review the generated music on their devices. Since the generated works are copyright-free, they can be freely used in advertisements and media content.

[0068] As a concrete example of its use, if a user wants to generate a song with the theme "Summer Vacation," specifying an "Upbeat Pop" style and including keywords such as "sea," "sun," and "vacation," they would input this information as a prompt into the AI ​​model. Based on this input, the AI ​​would generate a musical work. This allows users to easily obtain original musical works without requiring specialized musical knowledge or a lengthy production process.

[0069] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0070] Step 1:

[0071] The user launches the music generation application using their device. The user accesses the application's interface and enters the song title, style, and lyrics keywords on the input screen. This information is then prepared as prompt text. The entered information becomes the basic data for generating the music.

[0072] Step 2:

[0073] The terminal processes the prompt text entered by the user and converts it into a format for transmission to the server. During this conversion process, the data format is adjusted according to a standard communication protocol. The converted data is then sent to the server.

[0074] Step 3:

[0075] The server receives data sent from the terminal. It analyzes the received data and converts it into a format that can be processed by the generative AI model. This conversion structures the input data and sets the parameters necessary for music generation. The converted data is then provided to the generative AI model.

[0076] Step 4:

[0077] The AI ​​model on the server generates musical works based on the provided data. The AI ​​model utilizes neural networks and machine learning algorithms to create and integrate melodies, rhythms, harmonies, and other elements. The generated musical works are saved as digital music files.

[0078] Step 5:

[0079] The server sends the generated music file to the terminal. The transmitted data is securely encoded to ensure data integrity during transmission.

[0080] Step 6:

[0081] The terminal receives and decodes music files sent from the server. The user can then play and verify the received music on the terminal. The generated music can be used directly according to the user's needs.

[0082] (Application Example 1)

[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] Conventional music generation systems have presented challenges for advertisers, making it difficult to quickly generate music optimized for their ads and integrate it into videos. Furthermore, there has been a lack of efficient methods for generating music themes tailored to target demographics and advertising objectives.

[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0086] In this invention, the server includes means for generating music themes tailored to advertising purposes, terminal means for receiving music information and keywords, and server means for linking the generated advertising music with a video editing function. This makes it possible for advertisers to generate music optimized for their advertisements in a short amount of time and easily conduct promotions tailored to the target audience.

[0087] "Music generation AI" refers to artificial intelligence technology that automatically generates music based on music information specified by the user.

[0088] "Terminal means" refers to devices or equipment that receive music information and advertising keywords from users and input them into a music generation AI in an appropriate format.

[0089] "Server system" refers to a computer system that has the function of sending generated music to users and further provides the generated music in conjunction with video editing functions.

[0090] "Music themes tailored to advertising objectives" refer to songs optimized for specific advertising campaigns or promotions, and these songs are generated based on the target audience and product characteristics.

[0091] The system implementing this invention includes a program for generating advertising music using a music generation AI. Users can operate the application using a device such as a smartphone or smart glasses to input music information, keywords, and target attributes necessary for an advertising campaign. The device sends this information to a server in the cloud. The server uses a music generation algorithm to input data into a music generation AI model based on specified parameters and generates music.

[0092] The hardware used includes smartphones and smart glasses for user input, while the software consists of a music generation application and a cloud server hosting the music generation AI. The server analyzes this input data and generates music using a music generation AI model (e.g., OpenAI's music-specific model). In this process, it constructs melodies, rhythms, and harmonies based on specified keywords and musical styles, enabling the rapid delivery of music suitable for the target attributes.

[0093] For example, when planning an advertising campaign for a pair of sports shoes, a user might input keywords such as "sports," "vitality," and "comfort." Based on this information, the server can generate music with an active image and provide it as music suitable for the advertising video.

[0094] Examples of prompt statements are as follows:

[0095] "Generated background music for a new sneaker ad targeting people in their 20s. Keywords: sports, vitality, comfort."

[0096] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0097] Step 1:

[0098] The device is powered on, and the user operates a music generation application to input music information and keywords necessary for the advertising campaign. This input includes music style, target audience, and advertising theme. This input information is temporarily stored on the device and prepared for transfer to the server.

[0099] Step 2:

[0100] The device sends music information entered by the user to a server in the cloud. The input data includes music style, keywords, target attributes, etc. The server receives this data and performs preprocessing such as cleanup and flagging.

[0101] Step 3:

[0102] The server inputs pre-processed data into the music generation AI model and starts the music generation process. The generation AI model constructs melodies, rhythms, and harmonies based on the given style and keywords. This process generates new music data.

[0103] Step 4:

[0104] The server prepares to send the generated music to the device. The generated music data is converted into an audio file and formatted for distribution to the user. Since this data is in a standard music format, it can be played on the device.

[0105] Step 5:

[0106] The user checks and plays the received music on their device. At this stage, they can request modifications or regeneration of the music if necessary. Based on the user's feedback, the music will be regenerated if required.

[0107] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0108] This invention is a system that combines a music generation AI and an emotion engine to automatically generate music that responds to the user's emotional state. The aim of this system is to provide more personalized music by enhancing the user experience based on emotion recognition.

[0109] First, the user launches a music generation application through their device. The device is equipped with sensors necessary for emotion recognition, such as a camera and microphone. When the user wishes to generate music, they can input their current emotions through these sensors. The emotion engine analyzes this sensor information in real time to identify the user's emotional state.

[0110] Next, the emotion engine identifies the user's emotional state and generates corresponding emotion data. For example, if the user indicates emotions such as "happy" or "relaxed," the emotion engine provides this to the music generation AI, which then selects an appropriate music style and mood. Using the emotion data received from the emotion engine, the music generation AI constructs a song with a melody, rhythm, and tempo that matches the user's emotions.

[0111] The generated music is then sent to the user's device via a server, allowing the user to immediately play and enjoy it on their device. For example, if the user is feeling "energetic," fast-paced, upbeat pop music will be generated, providing an experience that further enhances their energy.

[0112] This invention allows users to obtain personalized musical experiences tailored to their emotional state at any given time, enabling a new and unprecedented form of music consumption. By providing music that resonates with the user's emotions, this system creates even greater value in the musical experience.

[0113] The following describes the processing flow.

[0114] Step 1:

[0115] The user launches a music generation application using their device. The device prepares to collect the user's emotional data through its camera and microphone.

[0116] Step 2:

[0117] The device senses the user's facial expressions and voice in real time and sends data to the emotion engine for emotion analysis. The emotion engine analyzes the emotion data and identifies the user's emotional state.

[0118] Step 3:

[0119] The emotion engine identifies the user's emotions and generates corresponding emotion data based on the analysis results. This emotion data is then sent to the music generation AI.

[0120] Step 4:

[0121] The server receives emotional data and inputs it into the music generation AI, which then instructs the AI ​​to set a music style and mood that corresponds to those emotions.

[0122] Step 5:

[0123] The music generation AI automatically generates songs that match the specified mood and style based on emotional data. The AI ​​selects and integrates musical elements such as melody, rhythm, and tempo to complete the song.

[0124] Step 6:

[0125] The server sends the generated music to the terminal. The terminal immediately presents the received music to the user in a playable format.

[0126] Step 7:

[0127] Users can play music generated on their devices and enjoy a musical experience tailored to their emotions. They can also request the regeneration of songs as needed.

[0128] (Example 2)

[0129] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0130] In modern music generation systems, providing music optimized to a user's emotions is difficult, resulting in a challenge where users cannot obtain a musical experience that matches their mood at any given moment. With the vast amount of audio content available, there is a need for a function that can instantly generate appropriate music according to each user's emotional state.

[0131] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0132] In this invention, the server includes an emotion analysis means for analyzing the user's emotional state, a means for providing emotion data collected through an emotion recognition sensor to a generating AI model, and a means for distributing the generated music and making it playable on a terminal. This allows the user to receive music that best suits their emotions at that moment in real time.

[0133] A "music generation program" is software designed to generate music that responds to the user's emotional state.

[0134] A "sentiment analysis tool" is a system component that collects and analyzes user emotional data and selects an appropriate music style based on the analysis results.

[0135] An "emotion recognition sensor" is a device, such as a camera or microphone, used to read a user's emotions from their voice or facial expressions.

[0136] A "generative AI model" is an AI system equipped with an algorithm that automatically constructs music based on specified input data.

[0137] A "musical piece" is a sonic work that possesses a musical structure and includes elements such as melody, rhythm, and tempo.

[0138] A "server" is an internet-connected computer system that distributes generated music to user devices, making the music available to users.

[0139] "Distribution" refers to the process of sending generated music to a user's device and making it available for use.

[0140] A "musical style" is a genre or form that possesses specific musical characteristics or emotional expression.

[0141] This invention relates to a system that automatically generates music that corresponds to a user's emotions using a music generation program. This system includes emotion analysis means, a terminal equipped with an emotion recognition sensor, and a server that distributes the generated music.

[0142] The user first launches a music generation application on their device. The device uses emotion recognition sensors such as a camera and microphone to collect emotional information from the user's facial expressions and voice. Next, an emotion analysis system analyzes this data to identify the user's state. This analyzed data is then provided to the generation AI model as emotional information.

[0143] The generative AI model selects a musical style that matches the user's emotions and constructs a song. This process considers musical elements such as melody, rhythm, and tempo. The generated song is sent to the device via a server, and the user can play it in real time.

[0144] For example, if the system analyzes that the user is in a relaxed state, the generative AI model will receive a prompt to create a calm, slow-tempo song. An example of a specific prompt to the generative AI module would be, "The user is feeling relaxed, so please generate calm, moody music." In this way, users can experience music that matches their mood at that moment.

[0145] This system allows users to quickly find songs that match their emotions, enabling them to discover value in new musical experiences.

[0146] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0147] Step 1:

[0148] The user launches a music generation application on their device. The application requests access to emotion recognition sensors such as the camera and microphone. At this stage, the user's action of launching the application is considered the input. Based on this action, the device prepares the necessary sensors and gets ready to collect the user's emotion data.

[0149] Step 2:

[0150] The device collects user emotion data through emotion recognition sensors. Specifically, it analyzes the user's facial expressions with a camera and records their voice tone with a microphone. The input to this process is the user's real-time facial video and voice, and the output is initial emotion data based on this. The device then processes this data into a format that can be used in the next step.

[0151] Step 3:

[0152] The device analyzes the emotional data collected using an emotion analysis tool. The input to this analysis is the emotional data obtained in step 2, and the output is the analyzed emotional state of the user. In this step, the device uses an algorithm to identify the user's emotions and generates specific emotion labels, such as "happy" or "relaxed."

[0153] Step 4:

[0154] The device provides the analyzed emotional state to the generating AI model as a prompt. The input is the emotional label identified in step 3, and the output is a prompt to the generating AI model. Specifically, if the emotional state is identified as "relaxed," an instruction such as "Generate music with a relaxed mood" will be generated.

[0155] Step 5:

[0156] The server uses the prompt message received from the terminal to activate the AI ​​model and generate the music. The input to this process is the prompt message created in step 4, and the output is the generated music data. The AI ​​selects a musical style that matches the specified emotion and constructs the musical structure.

[0157] Step 6:

[0158] The server sends the generated music data to the terminal. At this stage, the input is the generated music data, and the output is the music data prepared for playback on the user's terminal. The server sends this data immediately so that the user can enjoy a seamless music experience.

[0159] Step 7:

[0160] Users enjoy music received from the server by playing it through their device. The input is music data sent from the server, and the output is the user's auditory experience. By enjoying music that matches their current emotional state, users can achieve a personalized musical experience.

[0161] (Application Example 2)

[0162] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0163] In modern, diverse real-world experience settings, providing musical experiences tailored to the emotional states of visitors and users is challenging. In particular, there is a need to enhance the experience by providing individually optimized music for a large number of users with diverse emotional states. This invention aims to solve this problem and improve customer satisfaction in real-world experience settings.

[0164] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0165] In this invention, the server includes means for automatically generating music based on specified music information using music generation AI, means for detecting a person's emotional state using multiple emotion analysis means and determining the music information based on the emotional state, and output means for playing music corresponding to the emotional state in a real-world setting. This makes it possible to provide an optimal music experience in real time that is tailored to the emotions of visitors and users.

[0166] "Music generation AI" is an artificial intelligence system that automatically generates music based on specified musical information.

[0167] "Emotional analysis methods" refer to techniques that analyze emotions in real time using various sensors and devices designed to detect a person's emotional state.

[0168] A "terminal device" is a device that receives music information from the user and has the functionality to input data in a format suitable for music generation AI.

[0169] "Transmission means" refers to network components for transferring automatically generated music to the user's terminal.

[0170] "Output means" refers to a device or function for playing back music generated according to an emotional state in a real-life experience setting.

[0171] This invention provides a system for improving the musical experience in real-world settings. The system includes a music generation AI, emotion analysis means, terminal means, transmission means, and output means.

[0172] The server uses a music generation AI to generate music based on specified musical information. The generation AI model understands the user's emotional state through emotion analysis and uses that data as input. Specifically, OpenCV and emotion recognition models are used to analyze emotional data obtained from the user's facial expressions and voice, and to select a musical style accordingly. This analyzed data is then used by the music generation AI to generate music that optimizes the user experience.

[0173] The terminal device is the device that performs the aforementioned emotion analysis. The user directly provides music information through this terminal, and the music generation AI generates a song based on this information. The generated song is sent from the server to the terminal by the transmission device, and can be listened to in the real-world setting.

[0174] The output device is a device placed in the real-world setting to play the generated music in real time, typically an audio system within a store. This allows for a music experience tailored to the customer's emotions; for example, if a customer is relaxed, calming music will be played.

[0175] As a concrete example, there are systems in cafes that play background music tailored to the mood of the customers. When customers are relaxed, calm acoustic music is played, and when they are looking for a lively atmosphere, upbeat jazz music is played.

[0176] Examples of prompts include, "Identify the customer's emotions from their facial expressions and generate background music to increase their desire to buy," and "Create relaxing music to play when a customer is smiling."

[0177] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0178] Step 1:

[0179] The device uses its camera and microphone to collect user facial expression and audio data. The input is the user's real-time video and audio, which is fed into an emotion analysis model. The output is the analyzed user's emotional state data.

[0180] Step 2:

[0181] The server receives and processes emotional state data sent from the terminal. Based on this data, it creates appropriate prompt statements for the generative AI model and inputs them into the music generation AI. Specifically, instructions such as "The user is currently relaxed, so please generate calming music" are created.

[0182] Step 3:

[0183] The music generation AI generates appropriate music based on prompts. The input to this process is a prompt statement and emotional state data. The output is music data that matches the user's emotions. The music generation process involves data processing of musical elements such as tempo, melody, and harmony.

[0184] Step 4:

[0185] The server sends the generated music data to the terminal. The input here is the music data, and the output is the decoded music file on the terminal.

[0186] Step 5:

[0187] The device plays the received music data through an audio system in the real-world setting. Users can experience music optimized for their environment, such as in a store or event venue. In this step, the input is music data, and the output is the actual music that is played.

[0188] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0189] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0190] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0191] [Second Embodiment]

[0192] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0193] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0194] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0195] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0196] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0198] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0199] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0200] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0201] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0202] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0203] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0204] This invention provides a method for users to easily generate music in a system using music generation AI. This system is implemented through a series of steps in which the user inputs music information via a terminal, sends that information to a server, and the music generation AI automatically generates a song.

[0205] First, the user launches a music generation application using their device. This device can be a computer, smartphone, or tablet. The user enters music information such as the song title, style, and lyrics keywords into the application screen. For example, if the user wants to generate a song titled "Summer Campaign," in the "Pop" style, and include keywords like "beach," "sun," and "fun" in the lyrics, they would enter this information into their device.

[0206] Next, the terminal sends the user's input information to the server. This server receives the music information, converts it into an appropriate format, and provides it to the music generation AI. The music generation AI on the server receives this information and generates a song using machine learning models and algorithms stored internally. In this process, the AI ​​automatically selects and integrates melodies, rhythms, harmonies, etc., based on the specified style and keywords.

[0207] The generated music is sent to the device via a server. Users can then view, play, and use this music in advertisements and media content on their devices. Because this process is completed quickly, it significantly reduces time and cost compared to traditional music production. Furthermore, since newly generated music by AI is not subject to existing copyright issues, it offers a high degree of freedom in its use.

[0208] This invention streamlines the music production process and allows users to quickly generate music that meets diverse needs.

[0209] The following describes the processing flow.

[0210] Step 1:

[0211] The user launches the music generation application using their device. The user fills in information about the song in the input form displayed on the screen. Specifically, they enter the song title, style (e.g., "Pop"), and keywords they want to include in the lyrics (e.g., "Beach," "Sun," "Fun").

[0212] Step 2:

[0213] The device verifies the music information entered by the user and converts it to the appropriate data format. After this conversion is complete, the device sends a music generation request to the server.

[0214] Step 3:

[0215] The server analyzes the music information received from the terminal and prepares it as input data for the music generation AI. The server selects a specific music generation algorithm based on the user's request and prepares to invoke the AI.

[0216] Step 4:

[0217] The music generation AI is activated and begins the music generation process based on the data received from the server. The AI ​​synthesizes a song by combining elements such as melody, rhythm, and chords to match the specified style and keywords.

[0218] Step 5:

[0219] After the music generation is complete, the server receives the music file generated by the AI ​​and verifies it for errors. It then prepares to send the verified file to the terminal.

[0220] Step 6:

[0221] The device receives generated music data sent from the server and displays it to the user in a playable format. The user can review the music and save or re-edit it as needed.

[0222] (Example 1)

[0223] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0224] Traditional music generation methods require specialized knowledge and a long time, making it difficult for ordinary users to easily create musical works. Furthermore, existing music is subject to copyright and other restrictions, making free use difficult. Therefore, there is a need for the development of music generation technologies that can quickly meet diverse needs.

[0225] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0226] In this invention, the server includes processing means for automatically generating musical works based on input musical information using artificial intelligence that generates music, information processing means for acquiring the musical information from the user and providing it to the artificial intelligence that generates music in an appropriate format, and data processing means for transmitting the automatically generated musical works to the information processing means in order to supply them to the user. As a result, users can generate and use original musical works in a short time without requiring specialized knowledge.

[0227] "Artificial intelligence that generates music" refers to a machine learning system equipped with algorithms and models that automatically generate musical works based on input song information.

[0228] "Song information" refers to data that users input as elements necessary for music creation, and includes title, style, lyrics keywords, and so on.

[0229] An "information processing device" refers to electronic equipment and software that converts music information received from a user into an appropriate format and provides it to artificial intelligence that generates music.

[0230] A "data processing device" is a device or system used to transmit musical works generated by artificial intelligence that performs music generation to a user.

[0231] A "musical work" refers to audio data, including melody, rhythm, and harmony, that is automatically generated by artificial intelligence used for music generation.

[0232] This invention is a system that uses artificial intelligence (generative AI model) to generate music, automatically generating musical works based on song information entered by the user. In the implementation of this system, the server, terminals, and users play important roles.

[0233] Users can launch the music generation application using devices such as computers, smartphones, and tablets. Through the application's interface, users input song information such as the song title, style, and lyric keywords. This information is the basic data necessary for generating the musical work.

[0234] The terminal has the functionality to send music information entered by the user to a server. The server receives this information and converts it into an appropriate format. The artificial intelligence that generates music, running on the server, uses the transmitted information to generate a musical work. The generation AI model utilizes neural networks and machine learning algorithms based on the input information to generate a musical work that includes melody, rhythm, and harmony. The generated musical work is then sent from the server to the terminal in the form of a digital music file.

[0235] Users can play and review the generated music on their devices. Since the generated works are copyright-free, they can be freely used in advertisements and media content.

[0236] As a concrete example of its use, if a user wants to generate a song with the theme "Summer Vacation," specifying an "Upbeat Pop" style and including keywords such as "sea," "sun," and "vacation," they would input this information as a prompt into the AI ​​model. Based on this input, the AI ​​would generate a musical work. This allows users to easily obtain original musical works without requiring specialized musical knowledge or a lengthy production process.

[0237] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0238] Step 1:

[0239] The user launches the music generation application using their device. The user accesses the application's interface and enters the song title, style, and lyrics keywords on the input screen. This information is then prepared as prompt text. The entered information becomes the basic data for generating the music.

[0240] Step 2:

[0241] The terminal processes the prompt text entered by the user and converts it into a format for transmission to the server. During this conversion process, the data format is adjusted according to a standard communication protocol. The converted data is then sent to the server.

[0242] Step 3:

[0243] The server receives data sent from the terminal. It analyzes the received data and converts it into a format that can be processed by the generative AI model. This conversion structures the input data and sets the parameters necessary for music generation. The converted data is then provided to the generative AI model.

[0244] Step 4:

[0245] The AI ​​model on the server generates musical works based on the provided data. The AI ​​model utilizes neural networks and machine learning algorithms to create and integrate melodies, rhythms, harmonies, and other elements. The generated musical works are saved as digital music files.

[0246] Step 5:

[0247] The server sends the generated music file to the terminal. The transmitted data is securely encoded to ensure data integrity during transmission.

[0248] Step 6:

[0249] The terminal receives and decodes music files sent from the server. The user can then play and verify the received music on the terminal. The generated music can be used directly according to the user's needs.

[0250] (Application Example 1)

[0251] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0252] Conventional music generation systems have presented challenges for advertisers, making it difficult to quickly generate music optimized for their ads and integrate it into videos. Furthermore, there has been a lack of efficient methods for generating music themes tailored to target demographics and advertising objectives.

[0253] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0254] In this invention, the server includes means for generating music themes tailored to advertising purposes, terminal means for receiving music information and keywords, and server means for linking the generated advertising music with a video editing function. This makes it possible for advertisers to generate music optimized for their advertisements in a short amount of time and easily conduct promotions tailored to the target audience.

[0255] "Music generation AI" refers to artificial intelligence technology that automatically generates music based on music information specified by the user.

[0256] "Terminal means" refers to devices or equipment that receive music information and advertising keywords from users and input them into a music generation AI in an appropriate format.

[0257] "Server system" refers to a computer system that has the function of sending generated music to users and further provides the generated music in conjunction with video editing functions.

[0258] "Music themes tailored to advertising objectives" refer to songs optimized for specific advertising campaigns or promotions, and these songs are generated based on the target audience and product characteristics.

[0259] The system implementing this invention includes a program for generating advertising music using a music generation AI. Users can operate the application using a device such as a smartphone or smart glasses to input music information, keywords, and target attributes necessary for an advertising campaign. The device sends this information to a server in the cloud. The server uses a music generation algorithm to input data into a music generation AI model based on specified parameters and generates music.

[0260] The hardware used includes smartphones and smart glasses for user input, while the software consists of a music generation application and a cloud server hosting the music generation AI. The server analyzes this input data and generates music using a music generation AI model (e.g., OpenAI's music-specific model). In this process, it constructs melodies, rhythms, and harmonies based on specified keywords and musical styles, quickly providing music suitable for the target attributes.

[0261] For example, when planning an advertising campaign for a pair of sports shoes, a user might input keywords such as "sports," "vitality," and "comfort." Based on this information, the server can generate music with an active image and provide it as music suitable for the advertising video.

[0262] Examples of prompt statements are as follows:

[0263] "Generated background music for a new sneaker ad targeting people in their 20s. Keywords: sports, vitality, comfort."

[0264] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0265] Step 1:

[0266] The device is powered on, and the user operates a music generation application to input music information and keywords necessary for the advertising campaign. This input includes music style, target audience, and advertising theme. This input information is temporarily stored on the device and prepared for transfer to the server.

[0267] Step 2:

[0268] The device sends music information entered by the user to a server in the cloud. The input data includes music style, keywords, target attributes, etc. The server receives this data and performs preprocessing such as cleanup and flagging.

[0269] Step 3:

[0270] The server inputs pre-processed data into the music generation AI model and starts the music generation process. The generation AI model constructs melodies, rhythms, and harmonies based on the given style and keywords. This process generates new music data.

[0271] Step 4:

[0272] The server prepares to send the generated music to the device. The generated music data is converted into an audio file and formatted for distribution to the user. Since this data is in a standard music format, it can be played on the device.

[0273] Step 5:

[0274] The user checks and plays the received music on their device. At this stage, they can request modifications or regeneration of the music if necessary. Based on the user's feedback, the music will be regenerated if required.

[0275] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0276] This invention is a system that combines a music generation AI and an emotion engine to automatically generate music that responds to the user's emotional state. The aim of this system is to provide more personalized music by enhancing the user experience based on emotion recognition.

[0277] First, the user launches a music generation application through the terminal. The terminal is equipped with sensors necessary for emotion recognition, such as a camera and a microphone. In a scenario where the user wishes to generate music, it is possible to input the current emotion through these sensors. The emotion engine analyzes this sensor information in real time and identifies the user's emotional state.

[0278] Next, the emotion engine identifies the user's emotional state and generates corresponding emotion data. For example, when the user shows emotions such as "happy" or "relaxed", the emotion engine provides this to the music generation AI and selects an appropriate music style and mood. Utilizing the emotion data received from the emotion engine, the music generation AI constructs a music piece with a melody, rhythm, and tempo that match the user's emotions.

[0279] After that, the generated music is transmitted to the user's terminal via the server, and the user can immediately play and enjoy it on the terminal. As an example, if the user is in an "energetic" state, fast-paced and bright pop music is generated, providing an experience that further enhances the energy.

[0280] According to the present invention, the user can obtain an individual music experience according to their emotional state at that time, enabling a new form of music consumption that has never existed before. By providing music that conforms to the emotions of the user, this system creates further value for the music experience.

[0281] The following describes the processing flow.

[0282] Step 1:

[0283] The user launches a music generation application using the terminal. The terminal prepares to collect the user's emotion data through a camera and a microphone.

[0284] Step 2:

[0285] The terminal perceives the user's expressions and voice in real time and sends data for emotion analysis to the emotion engine. The emotion engine analyzes the emotion data and identifies the user's emotional state.

[0286] Step 3:

[0287] The emotion engine identifies the user's emotion and generates corresponding emotion data based on the analysis result. The emotion data is sent to the music generation AI.

[0288] Step 4:

[0289] The server inputs the received emotion data into the music generation AI and gives an instruction to set the music style and mood corresponding to the emotion for the AI.

[0290] Step 5:

[0291] The music generation AI automatically generates a piece of music according to the specified mood and style based on the emotion data. The AI selects and integrates music elements such as melody, rhythm, and tempo to complete the music.

[0292] Step 6:

[0293] The server sends the generated music to the terminal. The terminal immediately presents the received music to the user in a playable form.

[0294] Step 7:

[0295] The user plays the music generated on the terminal and enjoys a music experience in line with their own emotions. If necessary, the user can also request the regeneration of the music.

[0296] (Example 2)

[0297] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0298] In modern music generation systems, providing music optimized to a user's emotions is difficult, resulting in a challenge where users cannot obtain a musical experience that matches their mood at any given moment. With the vast amount of audio content available, there is a need for a function that can instantly generate appropriate music according to each user's emotional state.

[0299] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0300] In this invention, the server includes an emotion analysis means for analyzing the user's emotional state, a means for providing emotion data collected through an emotion recognition sensor to a generating AI model, and a means for distributing the generated music and making it playable on a terminal. This allows the user to receive music that best suits their emotions at that moment in real time.

[0301] A "music generation program" is software designed to generate music that responds to the user's emotional state.

[0302] A "sentiment analysis tool" is a system component that collects and analyzes user emotional data and selects an appropriate music style based on the analysis results.

[0303] An "emotion recognition sensor" is a device, such as a camera or microphone, used to read a user's emotions from their voice or facial expressions.

[0304] A "generative AI model" is an AI system equipped with an algorithm that automatically constructs music based on specified input data.

[0305] A "musical piece" is a sonic work that possesses a musical structure and includes elements such as melody, rhythm, and tempo.

[0306] A "server" is a computer system connected to the Internet that distributes the generated music to user devices so that users can use the music.

[0307] "Distribution" is a process of transmitting the generated music to the user's terminal and making it available for use.

[0308] "Music style" is a genre or form with specific musical characteristics and emotional expressions.

[0309] The present invention is a system that automatically generates music corresponding to the user's emotions using a music generation program. This system includes an emotion analysis means, a terminal equipped with an emotion recognition sensor, and a server that distributes the generated music.

[0310] First, the user launches a music generation application on the terminal. The terminal uses emotion recognition sensors such as a camera and a microphone to collect the user's emotional state from the user's expression and voice. Next, the emotion analysis means analyzes this data and identifies the user's state. This analyzed data is provided as emotion information to the generation AI model.

[0311] The generation AI model selects a music style that matches the user's emotions and constructs a music piece. In this process, musical elements such as melody, rhythm, and tempo are considered. The generated music is transmitted to the terminal via the server, and the user can play it in real time.

[0312] As a specific example, when it is analyzed that the user is in a relaxed state, the generation AI model receives a prompt to create a gentle music piece with a slow tempo. An example of a specific prompt sentence to the generation AI module is "Since the user is in a relaxed mood, please generate music with a gentle mood." In this way, the user can realize a music experience that suits their emotions at that time.

[0313] This system allows users to quickly find songs that match their emotions, enabling them to discover value in new musical experiences.

[0314] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0315] Step 1:

[0316] The user launches a music generation application on their device. The application requests access to emotion recognition sensors such as the camera and microphone. At this stage, the user's action of launching the application is considered the input. Based on this action, the device prepares the necessary sensors and gets ready to collect the user's emotion data.

[0317] Step 2:

[0318] The device collects user emotion data through emotion recognition sensors. Specifically, it analyzes the user's facial expressions with a camera and records their voice tone with a microphone. The input to this process is the user's real-time facial video and voice, and the output is initial emotion data based on this. The device then processes this data into a format that can be used in the next step.

[0319] Step 3:

[0320] The device analyzes the emotional data collected using an emotion analysis tool. The input to this analysis is the emotional data obtained in step 2, and the output is the analyzed emotional state of the user. In this step, the device uses an algorithm to identify the user's emotions and generates specific emotion labels, such as "happy" or "relaxed."

[0321] Step 4:

[0322] The device provides the analyzed emotional state to the generating AI model as a prompt. The input is the emotional label identified in step 3, and the output is a prompt to the generating AI model. Specifically, if the emotional state is identified as "relaxed," an instruction such as "Generate music with a relaxed mood" will be generated.

[0323] Step 5:

[0324] The server uses the prompt message received from the terminal to activate the AI ​​model and generate the music. The input to this process is the prompt message created in step 4, and the output is the generated music data. The AI ​​selects a musical style that matches the specified emotion and constructs the musical structure.

[0325] Step 6:

[0326] The server sends the generated music data to the terminal. At this stage, the input is the generated music data, and the output is the music data prepared for playback on the user's terminal. The server sends this data immediately so that the user can enjoy a seamless music experience.

[0327] Step 7:

[0328] Users enjoy music received from the server by playing it through their device. The input is music data sent from the server, and the output is the user's auditory experience. By enjoying music that matches their current emotional state, users can achieve a personalized musical experience.

[0329] (Application Example 2)

[0330] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0331] In modern, diverse real-world experience settings, providing musical experiences tailored to the emotional states of visitors and users is challenging. In particular, there is a need to enhance the experience by providing individually optimized music for a large number of users with diverse emotional states. This invention aims to solve this problem and improve customer satisfaction in real-world experience settings.

[0332] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0333] In this invention, the server includes means for automatically generating music based on specified music information using music generation AI, means for detecting a person's emotional state using multiple emotion analysis means and determining the music information based on the emotional state, and output means for playing music corresponding to the emotional state in a real-world setting. This makes it possible to provide an optimal music experience in real time that is tailored to the emotions of visitors and users.

[0334] "Music generation AI" is an artificial intelligence system that automatically generates music based on specified musical information.

[0335] "Emotional analysis methods" refer to techniques that analyze emotions in real time using various sensors and devices designed to detect a person's emotional state.

[0336] A "terminal device" is a device that receives music information from the user and has the functionality to input data in a format suitable for music generation AI.

[0337] "Transmission means" refers to network components for transferring automatically generated music to the user's terminal.

[0338] "Output means" refers to a device or function for playing back music generated according to an emotional state in a real-life experience setting.

[0339] This invention provides a system for improving the musical experience in real-world settings. The system includes a music generation AI, emotion analysis means, terminal means, transmission means, and output means.

[0340] The server uses a music generation AI to generate music based on specified musical information. The generation AI model understands the user's emotional state through emotion analysis and uses that data as input. Specifically, OpenCV and emotion recognition models are used to analyze emotional data obtained from the user's facial expressions and voice, and to select a musical style accordingly. This analyzed data is then used by the music generation AI to generate music that optimizes the user experience.

[0341] The terminal device is the device that performs the aforementioned emotion analysis. The user directly provides music information through this terminal, and the music generation AI generates a song based on this information. The generated song is sent from the server to the terminal by the transmission device, and can be listened to in the real-world setting.

[0342] The output device is a device placed in the real-world setting to play the generated music in real time, typically an audio system within a store. This allows for a music experience tailored to the customer's emotions; for example, if a customer is relaxed, calming music will be played.

[0343] As a concrete example, there are systems in cafes that play background music tailored to the mood of the customers. When customers are relaxed, calm acoustic music is played, and when they are looking for a lively atmosphere, upbeat jazz music is played.

[0344] Examples of prompts include, "Identify the customer's emotions from their facial expressions and generate background music to increase their desire to buy," and "Create relaxing music to play when a customer is smiling."

[0345] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0346] Step 1:

[0347] The device uses its camera and microphone to collect user facial expression and audio data. The input is the user's real-time video and audio, which is fed into an emotion analysis model. The output is the analyzed user's emotional state data.

[0348] Step 2:

[0349] The server receives and processes emotional state data sent from the terminal. Based on this data, it creates appropriate prompt statements for the generative AI model and inputs them into the music generation AI. Specifically, instructions such as "The user is currently relaxed, so please generate calming music" are created.

[0350] Step 3:

[0351] The music generation AI generates appropriate music based on prompts. The input to this process is a prompt statement and emotional state data. The output is music data that matches the user's emotions. The music generation process involves data processing of musical elements such as tempo, melody, and harmony.

[0352] Step 4:

[0353] The server sends the generated music data to the terminal. The input here is the music data, and the output is the decoded music file on the terminal.

[0354] Step 5:

[0355] The device plays the received music data through an audio system in the real-world setting. Users can experience music optimized for their environment, such as in a store or event venue. In this step, the input is music data, and the output is the actual music that is played.

[0356] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0357] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0358] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0359] [Third Embodiment]

[0360] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0361] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0362] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0363] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0364] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0365] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0366] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0367] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0368] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0369] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0370] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0371] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0372] This invention provides a method for users to easily generate music in a system using music generation AI. This system is implemented through a series of steps in which the user inputs music information via a terminal, sends that information to a server, and the music generation AI automatically generates a song.

[0373] First, the user launches a music generation application using their device. This device can be a computer, smartphone, or tablet. The user enters music information such as the song title, style, and lyrics keywords into the application screen. For example, if the user wants to generate a song titled "Summer Campaign," in the "Pop" style, and include keywords like "beach," "sun," and "fun" in the lyrics, they would enter this information into their device.

[0374] Next, the terminal sends the user's input information to the server. This server receives the music information, converts it into an appropriate format, and provides it to the music generation AI. The music generation AI on the server receives this information and generates a song using machine learning models and algorithms stored internally. In this process, the AI ​​automatically selects and integrates melodies, rhythms, harmonies, etc., based on the specified style and keywords.

[0375] The generated music is sent to the device via a server. Users can then view, play, and use this music in advertisements and media content on their devices. Because this process is completed quickly, it significantly reduces time and cost compared to traditional music production. Furthermore, since newly generated music by AI is not subject to existing copyright issues, it offers a high degree of freedom in its use.

[0376] This invention streamlines the music production process and allows users to quickly generate music that meets diverse needs.

[0377] The following describes the processing flow.

[0378] Step 1:

[0379] The user launches the music generation application using their device. The user fills in information about the song in the input form displayed on the screen. Specifically, they enter the song title, style (e.g., "Pop"), and keywords they want to include in the lyrics (e.g., "Beach," "Sun," "Fun").

[0380] Step 2:

[0381] The device verifies the music information entered by the user and converts it to the appropriate data format. After this conversion is complete, the device sends a music generation request to the server.

[0382] Step 3:

[0383] The server analyzes the music information received from the terminal and prepares it as input data for the music generation AI. The server selects a specific music generation algorithm based on the user's request and prepares to invoke the AI.

[0384] Step 4:

[0385] The music generation AI is activated and begins the music generation process based on the data received from the server. The AI ​​synthesizes a song by combining elements such as melody, rhythm, and chords to match the specified style and keywords.

[0386] Step 5:

[0387] After the music generation is complete, the server receives the music file generated by the AI ​​and verifies it for errors. It then prepares to send the verified file to the terminal.

[0388] Step 6:

[0389] The device receives generated music data sent from the server and displays it to the user in a playable format. The user can review the music and save or re-edit it as needed.

[0390] (Example 1)

[0391] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0392] Traditional music generation methods require specialized knowledge and a long time, making it difficult for ordinary users to easily create musical works. Furthermore, existing music is subject to copyright and other restrictions, making free use difficult. Therefore, there is a need for the development of music generation technologies that can quickly meet diverse needs.

[0393] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0394] In this invention, the server includes processing means for automatically generating musical works based on input musical information using artificial intelligence that generates music, information processing means for acquiring the musical information from the user and providing it to the artificial intelligence that generates music in an appropriate format, and data processing means for transmitting the automatically generated musical works to the information processing means in order to supply them to the user. As a result, users can generate and use original musical works in a short time without requiring specialized knowledge.

[0395] "Artificial intelligence that generates music" refers to a machine learning system equipped with algorithms and models that automatically generate musical works based on input song information.

[0396] "Song information" refers to data that users input as elements necessary for music creation, and includes title, style, lyrics keywords, and so on.

[0397] An "information processing device" refers to electronic equipment and software that converts music information received from a user into an appropriate format and provides it to artificial intelligence that generates music.

[0398] A "data processing device" is a device or system used to transmit musical works generated by artificial intelligence that performs music generation to a user.

[0399] A "musical work" refers to audio data, including melody, rhythm, and harmony, that is automatically generated by artificial intelligence used for music generation.

[0400] This invention is a system that uses artificial intelligence (generative AI model) to generate music, automatically generating musical works based on song information entered by the user. In the implementation of this system, the server, terminals, and users play important roles.

[0401] Users can launch the music generation application using devices such as computers, smartphones, and tablets. Through the application's interface, users input song information such as the song title, style, and lyric keywords. This information is the basic data necessary for generating the musical work.

[0402] The terminal has the functionality to send music information entered by the user to a server. The server receives this information and converts it into an appropriate format. The artificial intelligence that generates music, running on the server, uses the transmitted information to generate a musical work. The generation AI model utilizes neural networks and machine learning algorithms based on the input information to generate a musical work that includes melody, rhythm, and harmony. The generated musical work is then sent from the server to the terminal in the form of a digital music file.

[0403] Users can play and review the generated music on their devices. Since the generated works are copyright-free, they can be freely used in advertisements and media content.

[0404] As a concrete example of its use, if a user wants to generate a song with the theme "Summer Vacation," specifying an "Upbeat Pop" style and including keywords such as "sea," "sun," and "vacation," they would input this information as a prompt into the AI ​​model. Based on this input, the AI ​​would generate a musical work. This allows users to easily obtain original musical works without requiring specialized musical knowledge or a lengthy production process.

[0405] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0406] Step 1:

[0407] The user launches the music generation application using their device. The user accesses the application's interface and enters the song title, style, and lyrics keywords on the input screen. This information is then prepared as prompt text. The entered information becomes the basic data for generating the music.

[0408] Step 2:

[0409] The terminal processes the prompt text entered by the user and converts it into a format for transmission to the server. During this conversion process, the data format is adjusted according to a standard communication protocol. The converted data is then sent to the server.

[0410] Step 3:

[0411] The server receives data sent from the terminal. It analyzes the received data and converts it into a format that can be processed by the generative AI model. This conversion structures the input data and sets the parameters necessary for music generation. The converted data is then provided to the generative AI model.

[0412] Step 4:

[0413] The AI ​​model on the server generates musical works based on the provided data. The AI ​​model utilizes neural networks and machine learning algorithms to create and integrate melodies, rhythms, harmonies, and other elements. The generated musical works are saved as digital music files.

[0414] Step 5:

[0415] The server sends the generated music file to the terminal. The transmitted data is securely encoded to ensure data integrity during transmission.

[0416] Step 6:

[0417] The terminal receives and decodes music files sent from the server. The user can then play and verify the received music on the terminal. The generated music can be used directly according to the user's needs.

[0418] (Application Example 1)

[0419] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0420] Conventional music generation systems have presented challenges for advertisers, making it difficult to quickly generate music optimized for their ads and integrate it into videos. Furthermore, there has been a lack of efficient methods for generating music themes tailored to target demographics and advertising objectives.

[0421] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0422] In this invention, the server includes means for generating music themes tailored to advertising purposes, terminal means for receiving music information and keywords, and server means for linking the generated advertising music with a video editing function. This makes it possible for advertisers to generate music optimized for their advertisements in a short amount of time and easily conduct promotions tailored to the target audience.

[0423] "Music generation AI" refers to artificial intelligence technology that automatically generates music based on music information specified by the user.

[0424] "Terminal means" refers to devices or equipment that receive music information and advertising keywords from users and input them into a music generation AI in an appropriate format.

[0425] "Server system" refers to a computer system that has the function of sending generated music to users and further provides the generated music in conjunction with video editing functions.

[0426] "Music themes tailored to advertising objectives" refer to songs optimized for specific advertising campaigns or promotions, and these songs are generated based on the target audience and product characteristics.

[0427] The system implementing this invention includes a program for generating advertising music using a music generation AI. Users can operate the application using a device such as a smartphone or smart glasses to input music information, keywords, and target attributes necessary for an advertising campaign. The device sends this information to a server in the cloud. The server uses a music generation algorithm to input data into a music generation AI model based on specified parameters and generates music.

[0428] The hardware used includes smartphones and smart glasses for user input, while the software consists of a music generation application and a cloud server hosting the music generation AI. The server analyzes this input data and generates music using a music generation AI model (e.g., OpenAI's music-specific model). In this process, it constructs melodies, rhythms, and harmonies based on specified keywords and musical styles, quickly providing music suitable for the target attributes.

[0429] For example, when planning an advertising campaign for a pair of sports shoes, a user might input keywords such as "sports," "vitality," and "comfort." Based on this information, the server can generate music with an active image and provide it as music suitable for the advertising video.

[0430] Examples of prompt statements are as follows:

[0431] "Generated background music for a new sneaker ad targeting people in their 20s. Keywords: sports, vitality, comfort."

[0432] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0433] Step 1:

[0434] The device is powered on, and the user operates a music generation application to input music information and keywords necessary for the advertising campaign. This input includes music style, target audience, and advertising theme. This input information is temporarily stored on the device and prepared for transfer to the server.

[0435] Step 2:

[0436] The device sends music information entered by the user to a server in the cloud. The input data includes music style, keywords, target attributes, etc. The server receives this data and performs preprocessing such as cleanup and flagging.

[0437] Step 3:

[0438] The server inputs pre-processed data into the music generation AI model and starts the music generation process. The generation AI model constructs melodies, rhythms, and harmonies based on the given style and keywords. This process generates new music data.

[0439] Step 4:

[0440] The server prepares to send the generated music to the device. The generated music data is converted into an audio file and formatted for distribution to the user. Since this data is in a standard music format, it can be played on the device.

[0441] Step 5:

[0442] The user checks and plays the received music on their device. At this stage, they can request modifications or regeneration of the music if necessary. Based on the user's feedback, the music will be regenerated if required.

[0443] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0444] This invention is a system that combines a music generation AI and an emotion engine to automatically generate music that responds to the user's emotional state. The aim of this system is to provide more personalized music by enhancing the user experience based on emotion recognition.

[0445] First, the user launches a music generation application through their device. The device is equipped with sensors necessary for emotion recognition, such as a camera and microphone. When the user wishes to generate music, they can input their current emotions through these sensors. The emotion engine analyzes this sensor information in real time to identify the user's emotional state.

[0446] Next, the emotion engine identifies the user's emotional state and generates corresponding emotion data. For example, if the user indicates emotions such as "happy" or "relaxed," the emotion engine provides this to the music generation AI, which then selects an appropriate music style and mood. Using the emotion data received from the emotion engine, the music generation AI constructs a song with a melody, rhythm, and tempo that matches the user's emotions.

[0447] The generated music is then sent to the user's device via a server, allowing the user to immediately play and enjoy it on their device. For example, if the user is feeling "energetic," fast-paced, upbeat pop music will be generated, providing an experience that further enhances their energy.

[0448] This invention allows users to obtain personalized musical experiences tailored to their emotional state at any given time, enabling a new and unprecedented form of music consumption. By providing music that resonates with the user's emotions, this system creates even greater value in the musical experience.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The user launches a music generation application using their device. The device prepares to collect the user's emotional data through its camera and microphone.

[0452] Step 2:

[0453] The device senses the user's facial expressions and voice in real time and sends data to the emotion engine for emotion analysis. The emotion engine analyzes the emotion data and identifies the user's emotional state.

[0454] Step 3:

[0455] The emotion engine identifies the user's emotions and generates corresponding emotion data based on the analysis results. This emotion data is then sent to the music generation AI.

[0456] Step 4:

[0457] The server receives emotional data and inputs it into the music generation AI, which then instructs the AI ​​to set a music style and mood that corresponds to those emotions.

[0458] Step 5:

[0459] The music generation AI automatically generates songs that match the specified mood and style based on emotional data. The AI ​​selects and integrates musical elements such as melody, rhythm, and tempo to complete the song.

[0460] Step 6:

[0461] The server sends the generated music to the terminal. The terminal immediately presents the received music to the user in a playable format.

[0462] Step 7:

[0463] Users can play music generated on their devices and enjoy a musical experience tailored to their emotions. They can also request the regeneration of songs as needed.

[0464] (Example 2)

[0465] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0466] In modern music generation systems, providing music optimized to a user's emotions is difficult, resulting in a challenge where users cannot obtain a musical experience that matches their mood at any given moment. With the vast amount of audio content available, there is a need for a function that can instantly generate appropriate music according to each user's emotional state.

[0467] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0468] In this invention, the server includes an emotion analysis means for analyzing the user's emotional state, a means for providing emotion data collected through an emotion recognition sensor to a generating AI model, and a means for distributing the generated music and making it playable on a terminal. This allows the user to receive music that best suits their emotions at that moment in real time.

[0469] A "music generation program" is software designed to generate music that responds to the user's emotional state.

[0470] A "sentiment analysis tool" is a system component that collects and analyzes user emotional data and selects an appropriate music style based on the analysis results.

[0471] An "emotion recognition sensor" is a device, such as a camera or microphone, used to read a user's emotions from their voice or facial expressions.

[0472] A "generative AI model" is an AI system equipped with an algorithm that automatically constructs music based on specified input data.

[0473] A "musical piece" is a sonic work that possesses a musical structure and includes elements such as melody, rhythm, and tempo.

[0474] A "server" is an internet-connected computer system that distributes generated music to user devices, making the music available to users.

[0475] "Distribution" refers to the process of sending generated music to a user's device and making it available for use.

[0476] A "musical style" is a genre or form that possesses specific musical characteristics or emotional expression.

[0477] This invention relates to a system that automatically generates music that corresponds to a user's emotions using a music generation program. This system includes emotion analysis means, a terminal equipped with an emotion recognition sensor, and a server that distributes the generated music.

[0478] The user first launches a music generation application on their device. The device uses emotion recognition sensors such as a camera and microphone to collect emotional information from the user's facial expressions and voice. Next, an emotion analysis system analyzes this data to identify the user's state. This analyzed data is then provided to the generation AI model as emotional information.

[0479] The generative AI model selects a musical style that matches the user's emotions and constructs a song. This process considers musical elements such as melody, rhythm, and tempo. The generated song is sent to the device via a server, and the user can play it in real time.

[0480] For example, if the system analyzes that the user is in a relaxed state, the generative AI model will receive a prompt to create a calm, slow-tempo song. An example of a specific prompt to the generative AI module would be, "The user is feeling relaxed, so please generate calm, moody music." In this way, users can experience music that matches their mood at that moment.

[0481] This system allows users to quickly find songs that match their emotions, enabling them to discover value in new musical experiences.

[0482] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0483] Step 1:

[0484] The user launches a music generation application on their device. The application requests access to emotion recognition sensors such as the camera and microphone. At this stage, the user's action of launching the application is considered the input. Based on this action, the device prepares the necessary sensors and gets ready to collect the user's emotion data.

[0485] Step 2:

[0486] The device collects user emotion data through emotion recognition sensors. Specifically, it analyzes the user's facial expressions with a camera and records their voice tone with a microphone. The input to this process is the user's real-time facial video and voice, and the output is initial emotion data based on this. The device then processes this data into a format that can be used in the next step.

[0487] Step 3:

[0488] The device analyzes the emotional data collected using an emotion analysis tool. The input to this analysis is the emotional data obtained in step 2, and the output is the analyzed emotional state of the user. In this step, the device uses an algorithm to identify the user's emotions and generates specific emotion labels, such as "happy" or "relaxed."

[0489] Step 4:

[0490] The device provides the analyzed emotional state to the generating AI model as a prompt. The input is the emotional label identified in step 3, and the output is a prompt to the generating AI model. Specifically, if the emotional state is identified as "relaxed," an instruction such as "Generate music with a relaxed mood" will be generated.

[0491] Step 5:

[0492] The server uses the prompt message received from the terminal to activate the AI ​​model and generate the music. The input to this process is the prompt message created in step 4, and the output is the generated music data. The AI ​​selects a musical style that matches the specified emotion and constructs the musical structure.

[0493] Step 6:

[0494] The server sends the generated music data to the terminal. At this stage, the input is the generated music data, and the output is the music data prepared for playback on the user's terminal. The server sends this data immediately so that the user can enjoy a seamless music experience.

[0495] Step 7:

[0496] Users enjoy music received from the server by playing it through their device. The input is music data sent from the server, and the output is the user's auditory experience. By enjoying music that matches their current emotional state, users can achieve a personalized musical experience.

[0497] (Application Example 2)

[0498] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] In modern, diverse real-world experience settings, providing musical experiences tailored to the emotional states of visitors and users is challenging. In particular, there is a need to enhance the experience by providing individually optimized music for a large number of users with diverse emotional states. This invention aims to solve this problem and improve customer satisfaction in real-world experience settings.

[0500] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0501] In this invention, the server includes means for automatically generating music based on specified music information using music generation AI, means for detecting a person's emotional state using multiple emotion analysis means and determining the music information based on the emotional state, and output means for playing music corresponding to the emotional state in a real-world setting. This makes it possible to provide an optimal music experience in real time that is tailored to the emotions of visitors and users.

[0502] "Music generation AI" is an artificial intelligence system that automatically generates music based on specified musical information.

[0503] "Emotional analysis methods" refer to techniques that analyze emotions in real time using various sensors and devices designed to detect a person's emotional state.

[0504] A "terminal device" is a device that receives music information from the user and has the functionality to input data in a format suitable for music generation AI.

[0505] "Transmission means" refers to network components for transferring automatically generated music to the user's terminal.

[0506] "Output means" refers to a device or function for playing back music generated according to an emotional state in a real-life experience setting.

[0507] This invention provides a system for improving the musical experience in real-world settings. The system includes a music generation AI, emotion analysis means, terminal means, transmission means, and output means.

[0508] The server uses a music generation AI to generate music based on specified musical information. The generation AI model understands the user's emotional state through emotion analysis and uses that data as input. Specifically, OpenCV and emotion recognition models are used to analyze emotional data obtained from the user's facial expressions and voice, and to select a musical style accordingly. This analyzed data is then used by the music generation AI to generate music that optimizes the user experience.

[0509] The terminal device is the device that performs the aforementioned emotion analysis. The user directly provides music information through this terminal, and the music generation AI generates a song based on this information. The generated song is sent from the server to the terminal by the transmission device, and can be listened to in the real-world setting.

[0510] The output device is a device placed in the real-world setting to play the generated music in real time, typically an audio system within a store. This allows for a music experience tailored to the customer's emotions; for example, if a customer is relaxed, calming music will be played.

[0511] As a concrete example, there are systems in cafes that play background music tailored to the mood of the customers. When customers are relaxed, calm acoustic music is played, and when they are looking for a lively atmosphere, upbeat jazz music is played.

[0512] Examples of prompts include, "Identify the customer's emotions from their facial expressions and generate background music to increase their desire to buy," and "Create relaxing music to play when a customer is smiling."

[0513] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0514] Step 1:

[0515] The device uses its camera and microphone to collect user facial expression and audio data. The input is the user's real-time video and audio, which is fed into an emotion analysis model. The output is the analyzed user's emotional state data.

[0516] Step 2:

[0517] The server receives and processes emotional state data sent from the terminal. Based on this data, it creates appropriate prompt statements for the generative AI model and inputs them into the music generation AI. Specifically, instructions such as "The user is currently relaxed, so please generate calming music" are created.

[0518] Step 3:

[0519] The music generation AI generates appropriate music based on prompts. The input to this process is a prompt statement and emotional state data. The output is music data that matches the user's emotions. The music generation process involves data processing of musical elements such as tempo, melody, and harmony.

[0520] Step 4:

[0521] The server sends the generated music data to the terminal. The input here is the music data, and the output is the decoded music file on the terminal.

[0522] Step 5:

[0523] The device plays the received music data through an audio system in the real-world setting. Users can experience music optimized for their environment, such as in a store or event venue. In this step, the input is music data, and the output is the actual music that is played.

[0524] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0525] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0526] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0527] [Fourth Embodiment]

[0528] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0529] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0530] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0531] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0532] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0533] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0534] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0535] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0536] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0537] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0538] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0539] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0540] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0541] This invention provides a method for users to easily generate music in a system using music generation AI. This system is implemented through a series of steps in which the user inputs music information via a terminal, sends that information to a server, and the music generation AI automatically generates a song.

[0542] First, the user launches a music generation application using their device. This device can be a computer, smartphone, or tablet. The user enters music information such as the song title, style, and lyrics keywords into the application screen. For example, if the user wants to generate a song titled "Summer Campaign," in the "Pop" style, and include keywords like "beach," "sun," and "fun" in the lyrics, they would enter this information into their device.

[0543] Next, the terminal sends the user's input information to the server. This server receives the music information, converts it into an appropriate format, and provides it to the music generation AI. The music generation AI on the server receives this information and generates a song using machine learning models and algorithms stored internally. In this process, the AI ​​automatically selects and integrates melodies, rhythms, harmonies, etc., based on the specified style and keywords.

[0544] The generated music is sent to the device via a server. Users can then view, play, and use this music in advertisements and media content on their devices. Because this process is completed quickly, it significantly reduces time and cost compared to traditional music production. Furthermore, since newly generated music by AI is not subject to existing copyright issues, it offers a high degree of freedom in its use.

[0545] This invention streamlines the music production process and allows users to quickly generate music that meets diverse needs.

[0546] The following describes the processing flow.

[0547] Step 1:

[0548] The user launches the music generation application using their device. The user fills in information about the song in the input form displayed on the screen. Specifically, they enter the song title, style (e.g., "Pop"), and keywords they want to include in the lyrics (e.g., "Beach," "Sun," "Fun").

[0549] Step 2:

[0550] The device verifies the music information entered by the user and converts it to the appropriate data format. After this conversion is complete, the device sends a music generation request to the server.

[0551] Step 3:

[0552] The server analyzes the music information received from the terminal and prepares it as input data for the music generation AI. The server selects a specific music generation algorithm based on the user's request and prepares to invoke the AI.

[0553] Step 4:

[0554] The music generation AI is activated and begins the music generation process based on the data received from the server. The AI ​​synthesizes a song by combining elements such as melody, rhythm, and chords to match the specified style and keywords.

[0555] Step 5:

[0556] After the music generation is complete, the server receives the music file generated by the AI ​​and verifies it for errors. It then prepares to send the verified file to the terminal.

[0557] Step 6:

[0558] The device receives generated music data sent from the server and displays it to the user in a playable format. The user can review the music and save or re-edit it as needed.

[0559] (Example 1)

[0560] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0561] Traditional music generation methods require specialized knowledge and a long time, making it difficult for ordinary users to easily create musical works. Furthermore, existing music is subject to copyright and other restrictions, making free use difficult. Therefore, there is a need for the development of music generation technologies that can quickly meet diverse needs.

[0562] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0563] In this invention, the server includes processing means for automatically generating musical works based on input musical information using artificial intelligence that generates music, information processing means for acquiring the musical information from the user and providing it to the artificial intelligence that generates music in an appropriate format, and data processing means for transmitting the automatically generated musical works to the information processing means in order to supply them to the user. As a result, users can generate and use original musical works in a short time without requiring specialized knowledge.

[0564] "Artificial intelligence that generates music" refers to a machine learning system equipped with algorithms and models that automatically generate musical works based on input song information.

[0565] "Song information" refers to data that users input as elements necessary for music creation, and includes title, style, lyrics keywords, and so on.

[0566] An "information processing device" refers to electronic equipment and software that converts music information received from a user into an appropriate format and provides it to artificial intelligence that generates music.

[0567] A "data processing device" is a device or system used to transmit musical works generated by artificial intelligence that performs music generation to a user.

[0568] A "musical work" refers to audio data, including melody, rhythm, and harmony, that is automatically generated by artificial intelligence used for music generation.

[0569] This invention is a system that uses artificial intelligence (generative AI model) to generate music, automatically generating musical works based on song information entered by the user. In the implementation of this system, the server, terminals, and users play important roles.

[0570] Users can launch the music generation application using devices such as computers, smartphones, and tablets. Through the application's interface, users input song information such as the song title, style, and lyric keywords. This information is the basic data necessary for generating the musical work.

[0571] The terminal has the functionality to send music information entered by the user to a server. The server receives this information and converts it into an appropriate format. The artificial intelligence that generates music, running on the server, uses the transmitted information to generate a musical work. The generation AI model utilizes neural networks and machine learning algorithms based on the input information to generate a musical work that includes melody, rhythm, and harmony. The generated musical work is then sent from the server to the terminal in the form of a digital music file.

[0572] Users can play and review the generated music on their devices. Since the generated works are copyright-free, they can be freely used in advertisements and media content.

[0573] As a concrete example of its use, if a user wants to generate a song with the theme "Summer Vacation," specifying an "Upbeat Pop" style and including keywords such as "sea," "sun," and "vacation," they would input this information as a prompt into the AI ​​model. Based on this input, the AI ​​would generate a musical work. This allows users to easily obtain original musical works without requiring specialized musical knowledge or a lengthy production process.

[0574] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0575] Step 1:

[0576] The user launches the music generation application using their device. The user accesses the application's interface and enters the song title, style, and lyrics keywords on the input screen. This information is then prepared as prompt text. The entered information becomes the basic data for generating the music.

[0577] Step 2:

[0578] The terminal processes the prompt text entered by the user and converts it into a format for transmission to the server. During this conversion process, the data format is adjusted according to a standard communication protocol. The converted data is then sent to the server.

[0579] Step 3:

[0580] The server receives data sent from the terminal. It analyzes the received data and converts it into a format that can be processed by the generative AI model. This conversion structures the input data and sets the parameters necessary for music generation. The converted data is then provided to the generative AI model.

[0581] Step 4:

[0582] The AI ​​model on the server generates musical works based on the provided data. The AI ​​model utilizes neural networks and machine learning algorithms to create and integrate melodies, rhythms, harmonies, and other elements. The generated musical works are saved as digital music files.

[0583] Step 5:

[0584] The server sends the generated music file to the terminal. The transmitted data is securely encoded to ensure data integrity during transmission.

[0585] Step 6:

[0586] The terminal receives and decodes music files sent from the server. The user can then play and verify the received music on the terminal. The generated music can be used directly according to the user's needs.

[0587] (Application Example 1)

[0588] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0589] Conventional music generation systems have presented challenges for advertisers, making it difficult to quickly generate music optimized for their ads and integrate it into videos. Furthermore, there has been a lack of efficient methods for generating music themes tailored to target demographics and advertising objectives.

[0590] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0591] In this invention, the server includes means for generating music themes tailored to advertising purposes, terminal means for receiving music information and keywords, and server means for linking the generated advertising music with a video editing function. This makes it possible for advertisers to generate music optimized for their advertisements in a short amount of time and easily conduct promotions tailored to the target audience.

[0592] "Music generation AI" refers to artificial intelligence technology that automatically generates music based on music information specified by the user.

[0593] "Terminal means" refers to devices or equipment that receive music information and advertising keywords from users and input them into a music generation AI in an appropriate format.

[0594] "Server system" refers to a computer system that has the function of sending generated music to users and further provides the generated music in conjunction with video editing functions.

[0595] "Music themes tailored to advertising objectives" refer to songs optimized for specific advertising campaigns or promotions, and these songs are generated based on the target audience and product characteristics.

[0596] The system implementing this invention includes a program for generating advertising music using a music generation AI. Users can operate the application using a device such as a smartphone or smart glasses to input music information, keywords, and target attributes necessary for an advertising campaign. The device sends this information to a server in the cloud. The server uses a music generation algorithm to input data into a music generation AI model based on specified parameters and generates music.

[0597] The hardware used includes smartphones and smart glasses for user input, while the software consists of a music generation application and a cloud server hosting the music generation AI. The server analyzes this input data and generates music using a music generation AI model (e.g., OpenAI's music-specific model). In this process, it constructs melodies, rhythms, and harmonies based on specified keywords and musical styles, quickly providing music suitable for the target attributes.

[0598] For example, when planning an advertising campaign for a pair of sports shoes, a user might input keywords such as "sports," "vitality," and "comfort." Based on this information, the server can generate music with an active image and provide it as music suitable for the advertising video.

[0599] Examples of prompt statements are as follows:

[0600] "Generated background music for a new sneaker ad targeting people in their 20s. Keywords: sports, vitality, comfort."

[0601] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0602] Step 1:

[0603] The device is powered on, and the user operates a music generation application to input music information and keywords necessary for the advertising campaign. This input includes music style, target audience, and advertising theme. This input information is temporarily stored on the device and prepared for transfer to the server.

[0604] Step 2:

[0605] The device sends music information entered by the user to a server in the cloud. The input data includes music style, keywords, target attributes, etc. The server receives this data and performs preprocessing such as cleanup and flagging.

[0606] Step 3:

[0607] The server inputs pre-processed data into the music generation AI model and starts the music generation process. The generation AI model constructs melodies, rhythms, and harmonies based on the given style and keywords. This process generates new music data.

[0608] Step 4:

[0609] The server prepares to send the generated music to the device. The generated music data is converted into an audio file and formatted for distribution to the user. Since this data is in a standard music format, it can be played on the device.

[0610] Step 5:

[0611] The user checks and plays the received music on their device. At this stage, they can request modifications or regeneration of the music if necessary. Based on the user's feedback, the music will be regenerated if required.

[0612] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0613] This invention is a system that combines a music generation AI and an emotion engine to automatically generate music that responds to the user's emotional state. The aim of this system is to provide more personalized music by enhancing the user experience based on emotion recognition.

[0614] First, the user launches a music generation application through their device. The device is equipped with sensors necessary for emotion recognition, such as a camera and microphone. When the user wishes to generate music, they can input their current emotions through these sensors. The emotion engine analyzes this sensor information in real time to identify the user's emotional state.

[0615] Next, the emotion engine identifies the user's emotional state and generates corresponding emotion data. For example, if the user indicates emotions such as "happy" or "relaxed," the emotion engine provides this to the music generation AI, which then selects an appropriate music style and mood. Using the emotion data received from the emotion engine, the music generation AI constructs a song with a melody, rhythm, and tempo that matches the user's emotions.

[0616] The generated music is then sent to the user's device via a server, allowing the user to immediately play and enjoy it on their device. For example, if the user is feeling "energetic," fast-paced, upbeat pop music will be generated, providing an experience that further enhances their energy.

[0617] This invention allows users to obtain personalized musical experiences tailored to their emotional state at any given time, enabling a new and unprecedented form of music consumption. By providing music that resonates with the user's emotions, this system creates even greater value in the musical experience.

[0618] The following describes the processing flow.

[0619] Step 1:

[0620] The user launches a music generation application using their device. The device prepares to collect the user's emotional data through its camera and microphone.

[0621] Step 2:

[0622] The device senses the user's facial expressions and voice in real time and sends data to the emotion engine for emotion analysis. The emotion engine analyzes the emotion data and identifies the user's emotional state.

[0623] Step 3:

[0624] The emotion engine identifies the user's emotions and generates corresponding emotion data based on the analysis results. This emotion data is then sent to the music generation AI.

[0625] Step 4:

[0626] The server receives emotional data and inputs it into the music generation AI, which then instructs the AI ​​to set a music style and mood that corresponds to those emotions.

[0627] Step 5:

[0628] The music generation AI automatically generates songs that match the specified mood and style based on emotional data. The AI ​​selects and integrates musical elements such as melody, rhythm, and tempo to complete the song.

[0629] Step 6:

[0630] The server sends the generated music to the terminal. The terminal immediately presents the received music to the user in a playable format.

[0631] Step 7:

[0632] Users can play music generated on their devices and enjoy a musical experience tailored to their emotions. They can also request the regeneration of songs as needed.

[0633] (Example 2)

[0634] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0635] In modern music generation systems, providing music optimized to a user's emotions is difficult, resulting in a challenge where users cannot obtain a musical experience that matches their mood at any given moment. With the vast amount of audio content available, there is a need for a function that can instantly generate appropriate music according to each user's emotional state.

[0636] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0637] In this invention, the server includes an emotion analysis means for analyzing the user's emotional state, a means for providing emotion data collected through an emotion recognition sensor to a generating AI model, and a means for distributing the generated music and making it playable on a terminal. This allows the user to receive music that best suits their emotions at that moment in real time.

[0638] A "music generation program" is software designed to generate music that responds to the user's emotional state.

[0639] A "sentiment analysis tool" is a system component that collects and analyzes user emotional data and selects an appropriate music style based on the analysis results.

[0640] An "emotion recognition sensor" is a device, such as a camera or microphone, used to read a user's emotions from their voice or facial expressions.

[0641] A "generative AI model" is an AI system equipped with an algorithm that automatically constructs music based on specified input data.

[0642] A "musical piece" is a sonic work that possesses a musical structure and includes elements such as melody, rhythm, and tempo.

[0643] A "server" is an internet-connected computer system that distributes generated music to user devices, making the music available to users.

[0644] "Distribution" refers to the process of sending generated music to a user's device and making it available for use.

[0645] A "musical style" is a genre or form that possesses specific musical characteristics or emotional expression.

[0646] This invention relates to a system that automatically generates music that corresponds to a user's emotions using a music generation program. This system includes emotion analysis means, a terminal equipped with an emotion recognition sensor, and a server that distributes the generated music.

[0647] The user first launches a music generation application on their device. The device uses emotion recognition sensors such as a camera and microphone to collect emotional information from the user's facial expressions and voice. Next, an emotion analysis system analyzes this data to identify the user's state. This analyzed data is then provided to the generation AI model as emotional information.

[0648] The generative AI model selects a musical style that matches the user's emotions and constructs a song. This process considers musical elements such as melody, rhythm, and tempo. The generated song is sent to the device via a server, and the user can play it in real time.

[0649] For example, if the system analyzes that the user is in a relaxed state, the generative AI model will receive a prompt to create a calm, slow-tempo song. An example of a specific prompt to the generative AI module would be, "The user is feeling relaxed, so please generate calm, moody music." In this way, users can experience music that matches their mood at that moment.

[0650] This system allows users to quickly find songs that match their emotions, enabling them to discover value in new musical experiences.

[0651] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0652] Step 1:

[0653] The user launches a music generation application on their device. The application requests access to emotion recognition sensors such as the camera and microphone. At this stage, the user's action of launching the application is considered the input. Based on this action, the device prepares the necessary sensors and gets ready to collect the user's emotion data.

[0654] Step 2:

[0655] The device collects user emotion data through emotion recognition sensors. Specifically, it analyzes the user's facial expressions with a camera and records their voice tone with a microphone. The input to this process is the user's real-time facial video and voice, and the output is initial emotion data based on this. The device then processes this data into a format that can be used in the next step.

[0656] Step 3:

[0657] The device analyzes the emotional data collected using an emotion analysis tool. The input to this analysis is the emotional data obtained in step 2, and the output is the analyzed emotional state of the user. In this step, the device uses an algorithm to identify the user's emotions and generates specific emotion labels, such as "happy" or "relaxed."

[0658] Step 4:

[0659] The device provides the analyzed emotional state to the generating AI model as a prompt. The input is the emotional label identified in step 3, and the output is a prompt to the generating AI model. Specifically, if the emotional state is identified as "relaxed," an instruction such as "Generate music with a relaxed mood" will be generated.

[0660] Step 5:

[0661] The server uses the prompt message received from the terminal to activate the AI ​​model and generate the music. The input to this process is the prompt message created in step 4, and the output is the generated music data. The AI ​​selects a musical style that matches the specified emotion and constructs the musical structure.

[0662] Step 6:

[0663] The server sends the generated music data to the terminal. At this stage, the input is the generated music data, and the output is the music data prepared for playback on the user's terminal. The server sends this data immediately so that the user can enjoy a seamless music experience.

[0664] Step 7:

[0665] Users enjoy music received from the server by playing it through their device. The input is music data sent from the server, and the output is the user's auditory experience. By enjoying music that matches their current emotional state, users can achieve a personalized musical experience.

[0666] (Application Example 2)

[0667] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0668] In modern, diverse real-world experience settings, providing musical experiences tailored to the emotional states of visitors and users is challenging. In particular, there is a need to enhance the experience by providing individually optimized music for a large number of users with diverse emotional states. This invention aims to solve this problem and improve customer satisfaction in real-world experience settings.

[0669] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0670] In this invention, the server includes means for automatically generating music based on specified music information using music generation AI, means for detecting a person's emotional state using multiple emotion analysis means and determining the music information based on the emotional state, and output means for playing music corresponding to the emotional state in a real-world setting. This makes it possible to provide an optimal music experience in real time that is tailored to the emotions of visitors and users.

[0671] "Music generation AI" is an artificial intelligence system that automatically generates music based on specified musical information.

[0672] "Emotional analysis methods" refer to techniques that analyze emotions in real time using various sensors and devices designed to detect a person's emotional state.

[0673] A "terminal device" is a device that receives music information from the user and has the functionality to input data in a format suitable for music generation AI.

[0674] "Transmission means" refers to network components for transferring automatically generated music to the user's terminal.

[0675] "Output means" refers to a device or function for playing back music generated according to an emotional state in a real-life experience setting.

[0676] This invention provides a system for improving the musical experience in real-world settings. The system includes a music generation AI, emotion analysis means, terminal means, transmission means, and output means.

[0677] The server uses a music generation AI to generate music based on specified musical information. The generation AI model understands the user's emotional state through emotion analysis and uses that data as input. Specifically, OpenCV and emotion recognition models are used to analyze emotional data obtained from the user's facial expressions and voice, and to select a musical style accordingly. This analyzed data is then used by the music generation AI to generate music that optimizes the user experience.

[0678] The terminal device is the device that performs the aforementioned emotion analysis. The user directly provides music information through this terminal, and the music generation AI generates a song based on this information. The generated song is sent from the server to the terminal by the transmission device, and can be listened to in the real-world setting.

[0679] The output device is a device placed in the real-world setting to play the generated music in real time, typically an audio system within a store. This allows for a music experience tailored to the customer's emotions; for example, if a customer is relaxed, calming music will be played.

[0680] As a concrete example, there are systems in cafes that play background music tailored to the mood of the customers. When customers are relaxed, calm acoustic music is played, and when they are looking for a lively atmosphere, upbeat jazz music is played.

[0681] Examples of prompts include, "Identify the customer's emotions from their facial expressions and generate background music to increase their desire to buy," and "Create relaxing music to play when a customer is smiling."

[0682] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0683] Step 1:

[0684] The device uses its camera and microphone to collect user facial expression and audio data. The input is the user's real-time video and audio, which is fed into an emotion analysis model. The output is the analyzed user's emotional state data.

[0685] Step 2:

[0686] The server receives and processes emotional state data sent from the terminal. Based on this data, it creates appropriate prompt statements for the generative AI model and inputs them into the music generation AI. Specifically, instructions such as "The user is currently relaxed, so please generate calming music" are created.

[0687] Step 3:

[0688] The music generation AI generates appropriate music based on prompts. The input to this process is a prompt statement and emotional state data. The output is music data that matches the user's emotions. The music generation process involves data processing of musical elements such as tempo, melody, and harmony.

[0689] Step 4:

[0690] The server sends the generated music data to the terminal. The input here is the music data, and the output is the decoded music file on the terminal.

[0691] Step 5:

[0692] The device plays the received music data through an audio system in the real-world setting. Users can experience music optimized for their environment, such as in a store or event venue. In this step, the input is music data, and the output is the actual music that is played.

[0693] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0694] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0695] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0696] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0697] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0698] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0699] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0700] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0701] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0702] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0703] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0704] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0705] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0706] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0707] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0708] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0709] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0710] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0711] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0712] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0713] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0714] The following is further disclosed regarding the embodiments described above.

[0715] (Claim 1)

[0716] A means for automatically generating music based on specified music information using music generation AI,

[0717] A terminal means for receiving the aforementioned music information from the user and inputting it into the music generation AI in an appropriate format,

[0718] A server means for transmitting the automatically generated music to the terminal in order to provide it to the user,

[0719] A system that includes this.

[0720] (Claim 2)

[0721] The system according to claim 1, wherein the music generation AI generates music data that reproduces a specific musical style using a predetermined algorithm.

[0722] (Claim 3)

[0723] The system according to claim 1, wherein the server means provides a plurality of music generation options based on music information received from the user, and generates a song according to the user's selection.

[0724] "Example 1"

[0725] (Claim 1)

[0726] A processing means that uses artificial intelligence to generate music and automatically generates musical works based on input music information,

[0727] Information processing device means for obtaining the aforementioned music information from a user and providing it in an appropriate format to the artificial intelligence that generates the music,

[0728] To supply the automatically generated musical works to the user, a data processing device means for transmitting data to the information processing device,

[0729] A system that includes this.

[0730] (Claim 2)

[0731] The system according to claim 1, wherein the artificial intelligence that generates music generates music information that reproduces a specific musical format using a pre-set calculation method.

[0732] (Claim 3)

[0733] The system according to claim 1, wherein the data processing means provides a plurality of music generation options based on music information received from a user, and generates a musical work according to the user's selection.

[0734] "Application Example 1"

[0735] (Claim 1)

[0736] A means that uses music generation AI to automatically generate songs based on specified music information, and also has the function of generating music themes tailored to advertising purposes,

[0737] A terminal means that receives the aforementioned music information from the user, inputs it into the music generation AI in an appropriate format, and enables the receipt of keywords based on advertising purposes,

[0738] In order to provide the automatically generated music to the user, a server means transmits it to the terminal and links the generated advertising music with a video editing function,

[0739] A system that includes this.

[0740] (Claim 2)

[0741] The system according to claim 1, wherein the music generation AI generates music data based on a specific music style and target attributes using a predetermined algorithm.

[0742] (Claim 3)

[0743] The system according to claim 1, wherein the server means provides multiple music generation options based on music information received from the user and advertising purposes, and generates music according to the user's selection and purpose.

[0744] "Example 2 of combining an emotion engine"

[0745] (Claim 1)

[0746] A music generation program includes means for analyzing the user's emotional state using emotion analysis means and generating music based on the analysis results,

[0747] The aforementioned emotion analysis means includes means for analyzing user emotion data collected through an emotion recognition sensor and providing emotion information to a generative AI model,

[0748] A means for distributing the generated music to the user via a server and for the user to play the music on their device,

[0749] A system that includes this.

[0750] (Claim 2)

[0751] The system according to claim 1, wherein the music generation program selects a pre-set music style based on the emotional information and generates a song using a generation AI model.

[0752] (Claim 3)

[0753] The system according to claim 1, wherein the server provides multiple music generation options based on user emotion data and generates music according to the user's selection.

[0754] "Application example 2 when combining with an emotional engine"

[0755] (Claim 1)

[0756] A means for automatically generating music based on specified music information using music generation AI,

[0757] A terminal means for receiving the aforementioned music information from the user and inputting it into the music generation AI in an appropriate format,

[0758] A transmission means for sending the automatically generated music to the terminal in order to provide it to the user,

[0759] A means for detecting a person's emotional state using multiple emotion analysis means and determining the music information based on the emotional state,

[0760] An output means for playing music corresponding to the aforementioned emotional state in a real-life experience setting,

[0761] A system that includes this.

[0762] (Claim 2)

[0763] The system according to claim 1, wherein the music generation AI generates music data that reproduces a specific musical style using a predetermined algorithm.

[0764] (Claim 3)

[0765] The system according to claim 1, wherein the transmission means includes means for providing a plurality of music generation options based on music information received from the user, and for generating and playing a song according to the user's selection. [Explanation of Symbols]

[0766] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for automatically generating music based on specified music information using music generation AI, A terminal means for receiving the aforementioned music information from the user and inputting it into the music generation AI in an appropriate format, A server means for transmitting the automatically generated music to the terminal in order to provide it to the user, A system that includes this.

2. The system according to claim 1, wherein the music generation AI generates music data that reproduces a specific musical style using a predetermined algorithm.

3. The system according to claim 1, wherein the server means provides a plurality of music generation options based on music information received from the user and generates a song according to the user's selection.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A