System
The system uses a generative AI model to automate the creation, recording, and sharing of music and choreography, addressing the inefficiencies of existing systems by allowing easy content creation and continuous user engagement through sequels.
Patent Information
- Application Number
- JP2024118997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing systems are time-consuming for users to create music and choreography, record and film a performance that matches the music, and share it on social media, and lack automation for generating follow-up content based on social media reactions.
A system that includes a generative AI model to automatically generate music, lyrics, and choreography, records user performances, synthesizes the recorded data with generated content, and posts it on social media, with the ability to generate sequels based on user feedback.
Enables users to easily create and share original performance content on social media, and continuously support creative activities by generating sequels based on user reactions.
Smart Images

Figure 2026017936000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The purpose of this invention is to lower the psychological barriers for young people to use short video services and make it easier for them to create and post original content, thereby creating unique entertainment and maximizing individual creativity and performance abilities. It is necessary to provide an environment where users can create freely without being bound by existing music and choreography, and to promote the creation and sharing of new content through collaboration between AI and humans. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems with a system including a means for accepting user input, a means for generating music, lyrics, and choreography using a generative model, a means for providing the generated music, lyrics, and choreography to a user terminal, a means for recording a user's performance, a means for synthesizing the recorded performance data with the generated music, and a means for providing the synthesized performance content to a user terminal. Furthermore, this system includes a means for posting the synthesized performance content to a social networking platform, and by adding a means for generating a sequel music and lyrics based on the reaction on the social networking platform, it provides an environment in which users can easily create and share original content and engage in further creative activities.
[0006] "User input" refers to the operation or instruction a user gives through an application.
[0007] A "generative model" refers to a system or algorithm that uses artificial intelligence to automatically generate new music, lyrics, and choreography.
[0008] "Music" refers to audio data that includes musical melody, rhythm, harmony, etc.
[0009] "Lyrics" refer to sentences or phrases that are sung along with a song.
[0010] "Choreography" refers to a series of actions or movements in a dance or performance.
[0011] "User terminal" refers to a device operated by a user, such as a smartphone, tablet, or computer.
[0012] "Performance Data" refers to audio and video recordings of a User singing and dancing to a song.
[0013] "Synthesis" refers to the process of combining different data to generate a single integrated piece of data or content.
[0014] "SNS Platform" refers to an online platform for social networking services, such as a website or application, through which users can post and share content.
[0015] "Sequel songs" refer to newly created musical pieces that continue the theme and style of existing songs. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention provides a system that allows users to easily create and share original music and performance content. The following describes in detail an embodiment of the system.
[0038] Music, lyrics and choreography generation
[0039] User: First launches the application and taps the "Create a new song" button, which sends a creation request from the device to the server.
[0040] Server: Upon receiving a generation request, the server requests the internally implemented generative model to generate music, lyrics, and choreography. The generative model generates an original music piece, lyrics, and choreography of approximately 30 seconds based on the input parameters.
[0041] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[0042] Server: Receives the generated data and sends it to the terminal.
[0043] Device: The received music, lyrics, and choreography data are displayed on the app screen and played.
[0044] Recording and recording user performance
[0045] User: Sing and dance freely along with the generated music.
[0046] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[0047] Combining performance and generative music
[0048] Device: Once recording is complete, performance data is uploaded to the server.
[0049] Server: Receives the performance data and combines it with the generated music data. This process is carried out using video editing software and algorithms, ultimately generating a single integrated performance content.
[0050] Server: Sends the synthesized performance content to the terminal.
[0051] Terminal: Plays and stores the received composite content so that the user can view it.
[0052] Posting to social media
[0053] User: Indicate intention to post synthesized performance content to social media.
[0054] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[0055] Sequel Music - Commercial Use
[0056] Users: Check the reaction on social media and request sequel songs and commercialization.
[0057] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[0058] Generative model: Generates new music and lyrics and sends them back to the server.
[0059] Server: Provides the generated sequel data to the device so that the user can check it.
[0060] This system allows users to easily create performance videos based on original music and share them on social media. It also encourages further creative activities based on user feedback, providing a new form of entertainment.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] The user launches the application and taps the "Create a new song" button, which sends a request from the device to the server.
[0064] Step 2:
[0065] The server receives the generation request and asks the generative model to generate the music, lyrics, and choreography.
[0066] Step 3:
[0067] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[0068] Step 4:
[0069] The server receives the generated data and transmits it to the user terminal.
[0070] Step 5:
[0071] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[0072] Step 6:
[0073] Users sing and dance along to the music, and the device records and films this performance.
[0074] Step 7:
[0075] After the recording is complete, the device uploads this performance data to the server.
[0076] Step 8:
[0077] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[0078] Step 9:
[0079] The server generates the synthesized performance content and transmits it to the terminal.
[0080] Step 10:
[0081] The device receives, plays, and stores the composite content, which the user can then view and enjoy.
[0082] Step 11:
[0083] The user selects SNS posting and wishes to post composite content.
[0084] Step 12:
[0085] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[0086] Step 13:
[0087] Users check the reaction on social media and request sequel songs and commercialization.
[0088] Step 14:
[0089] The server sends a request to generate the sequel's music and lyrics to the generative model.
[0090] Step 15:
[0091] The generative model generates new music and lyrics and sends them back to the server.
[0092] Step 16:
[0093] The server provides the generated sequel data to the user terminal, so that the user can check it.
[0094] Example 1
[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0096] Conventional systems for generating and sharing original music and performance content have had the problem that it is time-consuming for users to easily create music and choreography, record and film a performance that matches the music, and then share it on social media. Furthermore, automation of the generation of follow-up content based on reactions on social media is also insufficient. These issues need to be resolved.
[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0098] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative AI model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for allowing the user to post the combined content to a social networking service (SNS) platform, and means for generating a sequel to the music and lyrics based on the response on the SNS platform. This allows users to easily create performance videos based on original music and share them on SNS, and further enables sequels to be automatically generated based on the response.
[0099] "Means for accepting user input" refers to a mechanism for inputting information or requests from a user, thereby enabling the user to send commands to the system.
[0100] A "generative AI model" refers to an artificial intelligence system that uses machine learning algorithms to automatically generate creative content such as music, lyrics, and choreography according to specific parameters.
[0101] "Means for generating music, lyrics, and choreography" refers to a mechanism that uses a generative AI model to automatically create music, lyrics, and choreography based on specified parameters.
[0102] "Means for providing to user terminal" refers to a mechanism for transmitting the generated data or content to the terminal used by the user, thereby enabling the user to view and play the generated content.
[0103] "Means for recording and recording a user's performance" refers to a mechanism for recording a user's performance in sync with the generated music, and involves capturing video and audio using the device's camera and microphone.
[0104] "Means for synthesizing performance data with generated music" refers to a mechanism for integrating recorded or filmed performance data of users with generated music to create a single piece of content.
[0105] "Means for providing synthesized performance content to a user terminal" refers to a mechanism for transmitting synthesized performance content to a user terminal and allowing the user to view and save it.
[0106] "Means for posting to an SNS platform" refers to a mechanism for uploading and publishing content created by a user to a social networking service (SNS).
[0107] "Means for generating sequel music and lyrics based on reactions on social media platforms" refers to a mechanism for analyzing reactions and evaluations on social media and generating new music and lyrics accordingly.
[0108] This invention provides a system that allows users to easily create and share original music and performance content. This system uses a generative AI model to generate music, lyrics, and choreography based on user input, and provides them to the user's device. It can also record and film the user's performance, combine it with the generated music, and post it on a social media platform. It can also generate sequels and lyrics based on social media reactions.
[0109] The system is implemented using the following hardware and software:
[0110] 1. A means of accepting user input
[0111] A user starts the application on a device (e.g., a smartphone or tablet) and taps the "Create a new song" button. This action sends a creation request from the device to the server.
[0112] 2. A means of generating music, lyrics, and choreography using generative AI models
[0113] The server receives the generation request and asks a generative AI model (e.g., a model using a machine learning algorithm) to generate the music, lyrics, and choreography. Specific models used here include OpenAI's GPT-4 and DeepComposer.
[0114] For example, a create request might use the following prompt:
[0115] Song generation prompt: The melody line is bright, the tempo is 120 BPM, and the theme is "Summer Memories."
[0116] 3. Means for providing the generated music, lyrics, and choreography to the user terminal
[0117] The generative model automatically generates music, lyrics, and choreography and sends them back to the server. The server receives the generated data and sends it to the device. The device displays the received data on the app screen and plays it back.
[0118] 4. Means for recording and recording your performance
[0119] The user sings and dances along to the generated music, and the device records and films the performance using the built-in microphone and camera. This data is temporarily stored in the device's storage.
[0120] 5. A means of synthesizing recorded performance data with generated music
[0121] Once the device has finished recording, it uploads the performance data to the server. The server receives the performance data and combines it with the generated music data. This process is carried out using video editing software (e.g., FFmpeg or Adobe Premiere Pro) and algorithms, resulting in a single integrated performance content.
[0122] 6. Means for providing synthesized performance content to a user terminal
[0123] The server sends the synthesized performance content to the terminal, which then plays and stores the received synthesized content for the user to view.
[0124] 7. Means for enabling users to post synthetic content to social media platforms
[0125] The user indicates their intention to post the synthesized performance content to a social networking site. The device then calls a social networking site API (e.g., Twitter API or Instagram Graph API) to upload the synthesized content to the social networking site, add information such as a caption, and post it.
[0126] 8. A means of generating sequel music and lyrics based on reactions on social media platforms
[0127] Users check the reaction on social media and request a sequel song and commercialization. The server sends the sequel song and lyric generation request to the generative AI model. The generative model generates new music and lyrics and sends them back to the server. The server provides the generated sequel data to the device so that the user can view it.
[0128] This system allows users to easily create performance videos based on original music and share them on social media, and can encourage further creative activities based on user feedback, providing a new form of entertainment.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] A user launches the application and taps the "Create a new song" button. The input is the user's tap, which causes a creation request to be sent from the device to the server. Specifically, the creation request is sent as an HTTP POST request. The data included in this request is the song creation prompt and other parameters (e.g., tempo, theme).
[0132] Step 2:
[0133] The server receives the generation request and parses the request contents. The input is JSON data included in the HTTP request body. The server parses this data and converts it into an input format appropriate for the generative AI model. Specifically, it extracts the music generation prompt and parameters from the request content and converts them into a format to be passed to the generative AI model. The output is the request data sent to the generative AI model.
[0134] Step 3:
[0135] The generative AI model generates music, lyrics, and choreography based on the request data it receives. The input is the prompt and parameters sent from the server. The generative AI model applies algorithms based on these inputs to generate original music data, lyrics data, and choreography data. The output is the generated music, lyrics, and choreography data, which are sent back to the server.
[0136] Step 4:
[0137] The server receives the data returned from the generative AI model and sends it to the device. The input is the music, lyrics, and choreography data returned from the generative AI model. The server formats this data into JSON format and sends it to the device as an HTTP response. The output is the music, lyrics, and choreography data sent to the device.
[0138] Step 5:
[0139] The device displays the received data on the app screen and plays it. The input is music, lyrics, and choreography data sent from the server. The device analyzes this data, decodes and plays the music data, and displays the lyrics and choreography information on the app screen. Specifically, it plays the music using a music playback library and displays the lyrics and choreography information in a UI component. The output is audiovisual content provided to the user.
[0140] Step 6:
[0141] The user sings and dances along to the generated music. The system records and films the user's performance. The input is the user's performance, which is recorded using the device's microphone and camera. The recorded data is temporarily stored in the smartphone's storage. The output is the recorded performance data.
[0142] Step 7:
[0143] Once the device has completed recording, it uploads the performance data to the server. The input is the recorded performance data. The device sends this data to the server as an HTTP PUT request. The output is the performance data uploaded to the server.
[0144] Step 8:
[0145] The server receives the performance data and combines it with the generated music data. The input is the uploaded performance data and the generated music data. The server uses video editing software such as FFmpeg to combine these data into a single integrated performance content. Specifically, the server uses FFmpeg commands to combine the video and audio. The output is the combined performance content.
[0146] Step 9:
[0147] The server sends the synthesized performance content to the terminal. The input is the synthesized performance content. The server sends this content to the terminal as an HTTP response. The output is the synthesized performance content sent to the terminal.
[0148] Step 10:
[0149] The device plays and saves the received composite content, making it available for viewing by the user. The input is composite performance content sent from the server. The device plays this content and saves it in storage. Specifically, it plays it using a video playback library and saves it using a file system API. The output is playable composite content provided to the user.
[0150] Step 11:
[0151] The user indicates their intention to post the synthesized performance content to a social networking site. The input is the user's posting operation, and the device uses the social networking site API to upload the synthesized content to the social networking site platform. Specifically, the device sends an HTTP POST request to the social networking site platform's API to upload information such as captions along with the synthesized content. The output is the content posted to the social networking site platform.
[0152] Step 12:
[0153] Users check the reaction on social media and request a sequel song and commercialization. The input is the user's sequel request operation, which is sent to the server. The server sends a sequel song and lyric generation request to the generative AI model. The output is the request data sent to the generative AI model.
[0154] Step 13:
[0155] The generative AI model generates new music and lyrics and sends them back to the server. The input is the sequel request data, and the generative AI model generates new music and lyrics based on this. The server provides the generated data to the device so that the user can view it. Specifically, the generated data is returned to the server in JSON format, which is then sent to the device. The output is the generated sequel music and lyrics data.
[0156] (Application example 1)
[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0158] The challenge is to provide a system that allows users to easily create original music, lyrics, and choreography, synthesize audio and video recordings of performances to match them, and then efficiently post the content to a social media platform. Another challenge is to generate follow-up music and performance content based on the reaction on the social media platform, thereby continuously supporting users' creative activities.
[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0160] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for posting the generated content to a social networking service (SNS) platform, and means for checking the reaction on the SNS platform and requesting the generation of a sequel music and lyrics. This allows users to efficiently and consistently perform processes from generating original content to sharing it on SNS and further creating sequels.
[0161] definition statement
[0162] "Means for accepting user input" refers to the interface that the user operates when requesting the creation of new music, lyrics, and choreography.
[0163] "Means for generating music, lyrics, and choreography using generative models" refers to a system that uses AI models and algorithms to automatically generate original music, lyrics, and choreography based on input parameters.
[0164] "Means for providing the generated music, lyrics and choreography to the user terminal" refers to the process of transmitting the generated content from the server to the user terminal.
[0165] "Means for recording a user's performance" refers to a recording function for recording a user's performance in sync with the generated music.
[0166] "Means of combining recorded performance data with generated music" refers to the process of integrating recorded performance data with generated music data and editing it into a complete performance content.
[0167] The term "means for providing the synthesized performance content to the user terminal" refers to a process for transmitting the synthesized performance content to the user terminal and making it playable.
[0168] "Means for posting Generated Content to a Social Media Platform" means an interface for uploading Generated Performance Content to a Social Media Platform and for users to share it.
[0169] "Means for checking reactions on social media platforms and making requests to generate sequel music and lyrics" refers to a system that analyzes user reactions on social media platforms and requests the generation of sequel content based on the results.
[0170] MODE FOR CARRYING OUT THE INVENTION
[0171] This invention is a system that generates original music, lyrics, and choreography based on prompts entered by the user, integrates performance content recorded or filmed by the user with the generated content, and shares it on a social media platform.
[0172] The specific steps for implementing this system are as follows:
[0173] First, a user accesses the application using a device (smartphone or smart glasses). The user taps the "Create a new song" button to proceed to a prompt input screen. Here, the user enters a specific prompt, such as "Create an energetic pop song."
[0174] The server then receives prompts from the user, uses its internal generative AI model to automatically generate music, lyrics, and choreography based on the prompts, and sends the generated content to the device.
[0175] The device displays the generated music, lyrics, and choreography data on the application screen, allowing the user to play it. The user records and films their performance while singing and dancing to the generated music. The recorded and filmed performance data is temporarily stored on the device.
[0176] Once the user has completed recording the performance, the device uploads the performance data to the server. The server then combines the received performance data with the generated music data using video editing software. The combined performance content is then sent back to the device from the server.
[0177] The device that receives the composite content makes it playable for the user. If the user checks the generated performance content and indicates their intention to post it to a social media platform, the device calls the social media API to upload the content. The user can then add information such as a caption and finally post it to the social media platform.
[0178] The server analyzes the reactions to the content posted by users on the SNS platform, and based on the results, users can request the generation of a sequel to the song and lyrics, thereby providing continuous support for the user's creative activities.
[0179] For example, consider the case where a user enters the following prompt:
[0180] Prompt: "Generate an energetic pop song"
[0181] In this way, this invention is a system that enables users to efficiently and consistently create original content, share it on social media, and even create sequels.
[0182] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0183] Program processing flow
[0184] Step 1:
[0185] A user accesses the application using a device, taps the "Create a new song" button, and enters a prompt, which is then sent to the server.
[0186] Input: User prompt (e.g., "Generate an energetic pop song")
[0187] Output: Prompt data is sent to the server
[0188] Step 2:
[0189] The server receives the prompts and generates the music, lyrics, and choreography based on an internal generative AI model.
[0190] Input: prompt data
[0191] Data processing: Based on the input prompt, the generative AI model generates the music, lyrics, and choreography.
[0192] Output: Generated music data, lyrics data, choreography data
[0193] Step 3:
[0194] The server transmits the generated data to the terminal, which receives and displays it.
[0195] Input: Generated music data, lyrics data, choreography data
[0196] Output: Music, lyrics, and choreography displayed on the user's device
[0197] Step 4:
[0198] The user performs along with the generated music, and the device records and films this.
[0199] Input: User performance (singing, dancing)
[0200] Data processing: Activate the recording function and capture performance data
[0201] Output: Recorded and recorded performance data is temporarily saved.
[0202] Step 5:
[0203] The device uploads the recorded performance data to the server.
[0204] Input: Audio-visual performance data
[0205] Output: Performance data is sent to the server
[0206] Step 6:
[0207] The server synthesizes the received performance data with the generated music data.
[0208] Input: Performance data, generated song data
[0209] Data processing: Composition using video editing software
[0210] Output: Synthesized performance content
[0211] Step 7:
[0212] The server transmits the synthesized performance content to the terminal, which receives it and makes it playable.
[0213] Input: Synthesized performance content
[0214] Output: The composite content displayed on the device
[0215] Step 8:
[0216] The user indicates their intention to post the composite content to a social networking platform, and the device calls the social networking API to upload the content.
[0217] Input: User's posting intent, synthetic content
[0218] Data processing: Upload content using SNS API
[0219] Output: Content posted on social media platforms
[0220] Step 9:
[0221] The server analyzes the reaction on the social media platform and begins the process of generating the sequel's music and lyrics based on user requests.
[0222] Input: Reaction data from social media platforms
[0223] Data processing: A generative AI model based on feedback data generates the music and lyrics for the sequel
[0224] Output: Sequel music data, lyrics data
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] The present invention provides a system that allows users to easily create and share original music and performance content. The system also incorporates an emotion engine that can recognize a user's emotions and adjust the generated content accordingly. The following describes in detail an embodiment of the system.
[0227] Music, lyrics and choreography generation
[0228] User: First launches the application and taps the "Generate New Song" button, which sends a request from the device to the server.
[0229] Server: Upon receiving a generation request, it launches the emotion engine and collects data to recognize the user's emotion.
[0230] Emotion engine: Analyzes the user's facial expressions, tone of voice, and other emotional data from the camera and microphone to identify the emotion the user is currently feeling. The emotion engine then sends the recognized emotional information back to the server.
[0231] Server: Requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on the input parameters and emotional information.
[0232] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[0233] Server: Receives the generated data and sends it to the terminal.
[0234] Device: The received music, lyrics, and choreography data is displayed on the app screen and played, allowing the user to listen to the music, lyrics, and choreography.
[0235] Recording and recording user performance
[0236] User: Sing and dance freely along with the generated music.
[0237] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[0238] Combining performance and generative music
[0239] Device: Once recording is complete, performance data is uploaded to the server.
[0240] Server: Receives the performance data and synthesizes it with the generated music data. The synthesis process is carried out using appropriate video editing software or algorithms, and finally a single integrated performance content is generated.
[0241] Server: Sends the synthesized performance content to the terminal.
[0242] Terminal: Plays and stores the received composite content so that the user can view it.
[0243] Posting to social media
[0244] User: Indicate intention to post synthesized performance content to social media.
[0245] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[0246] Sequel Music - Commercial Use
[0247] Users: Check the reaction on social media and request sequel songs and commercialization.
[0248] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[0249] Generative model: Generates new music and lyrics and sends them back to the server.
[0250] Server: Provides the generated sequel data to the device so that the user can check it.
[0251] This system allows users to easily create performance videos based on original music and share them on social media. Furthermore, by generating content based on the user's emotional information, it is possible to provide a more personalized experience.
[0252] The processing flow will be explained below.
[0253] Step 1:
[0254] The user launches the application and taps the "Create a new song" button, which sends a song creation request from the device to the server.
[0255] Step 2:
[0256] The server receives the generation request and starts the emotion engine, which prepares to start collecting data necessary to recognize the user's emotion.
[0257] Step 3:
[0258] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and this data is sent to the emotion engine in real time.
[0259] Step 4:
[0260] The emotion engine analyzes the received data and identifies the user's emotion. For example, if the user is smiling, it recognizes "happiness," and if they are frowning, it recognizes "anxiety."
[0261] Step 5:
[0262] The emotion engine sends the identified emotion information back to the server, which collects this information and prepares it for transmission to the generative model.
[0263] Step 6:
[0264] The server requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model takes the emotional information into consideration and generates music, lyrics, and choreography that match the emotional information.
[0265] Step 7:
[0266] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[0267] Step 8:
[0268] The server receives the generated data and transmits it to the user terminal.
[0269] Step 9:
[0270] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[0271] Step 10:
[0272] The user sings and dances freely to the generated music, and the device records and films this performance.
[0273] Step 11:
[0274] Once the recording is complete, the device uploads this performance data to a server.
[0275] Step 12:
[0276] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[0277] Step 13:
[0278] The server generates the synthesized performance content and transmits it to the terminal.
[0279] Step 14:
[0280] The device plays and saves the received composite content, which the user can then view and enjoy.
[0281] Step 15:
[0282] The user selects SNS posting and wishes to post composite content.
[0283] Step 16:
[0284] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[0285] Step 17:
[0286] Users check the reaction on social media and request sequel songs and commercialization.
[0287] Step 18:
[0288] The server sends a request to generate the sequel's music and lyrics to the generative model.
[0289] Step 19:
[0290] The generative model generates new music and lyrics and sends them back to the server.
[0291] Step 20:
[0292] The server provides the generated sequel data to the user terminal, so that the user can check it.
[0293] Example 2
[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] In today's entertainment environment, there is a demand for a means for users to easily create and share original music and performance content. However, existing systems have difficulty generating content based on users' emotions, and there are insufficient methods for efficiently synthesizing generated content and sharing it on social media. Furthermore, there is a lack of easy ways to request and realize sequels or commercialization. Therefore, a system that makes it easier for users to create and share original content that responds to individual emotions is needed.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0297] In this invention, the server includes means for recognizing the user's emotional information and generating music, lyrics, and choreography, means for providing the synthesized performance content to the user terminal, and means for posting the content to an SNS platform, thereby enabling users to generate original content according to their individual emotions, efficiently synthesize the content, and share it on the SNS.
[0298] The "means for accepting user input" is an interface that allows the user to input operational instructions to the system.
[0299] "Generative model" refers to algorithms or software that use artificial intelligence to automatically generate content such as music, lyrics, and choreography.
[0300] "Means for recognizing user's emotional information" refers to an engine or software that analyzes the user's facial expressions and tone of voice via a camera or microphone to identify their current emotional state.
[0301] "Means for generating music, lyrics and choreography" refers to a device or software that generates original music, lyrics and choreography based on emotional information and input data.
[0302] "Means for providing the generated music, lyrics and choreography to the user's terminal" refers to a device or program that has the function of transferring the generated data from the server to the user's terminal and displaying and playing it.
[0303] "Means for recording a user's performance" refers to a device or function that uses a camera and microphone to record a user's performance along with a piece of music.
[0304] "Means for synthesizing recorded performance data with generated music" refers to software or a system for integrating a user's performance data with generated music and editing and generating a single composite content.
[0305] "Means for providing synthesized performance content to a user terminal" refers to a device or program for transmitting the completed synthesized content from the server to a user terminal so that it can be displayed and played.
[0306] "Means for posting to a social media platform" refers to a device or software that has the function of uploading and sharing generated content to a social media platform using an API or other means.
[0307] "Means for generating sequel music and lyrics" refers to artificial intelligence algorithms or models for generating new music and lyrics based on existing content.
[0308] The present invention relates to a system that enables users to easily create and share original music and performance content. Specific embodiments of this system will be described below.
[0309] Music, lyrics and choreography generation
[0310] First, the user launches the application and taps the "Generate a new song" button. This action sends a generation request from the device to the server. When the server receives the generation request, it launches an emotion engine and collects data to recognize the user's emotions. The emotion engine analyzes the user's facial expressions and vocal tone through the camera and microphone to identify the user's current emotion. This emotion information is sent back to the server. The server, along with the received emotion information, requests the generative model to generate music, lyrics, and choreography. The generative model generates original music, lyrics, and choreography based on the input prompt and emotion information and sends it back to the server. The server receives the generated data and sends it to the user's device. The device displays and plays this data on the app screen.
[0311] Hardware and software used
[0312] Hardware: Smartphone (camera, microphone)
[0313] Software: Emotion engine (facial expression and voice analysis software), generative AI model (music, lyrics, choreography generation)
[0314] Specific examples
[0315] For example, when a user launches the app and taps the "Generate a new song" button, the smartphone's camera and microphone are activated, and the user's facial expressions and voice are analyzed. If the user is in a happy mood, the generative model will generate an upbeat, cheerful song.
[0316] Example prompt for a generative AI model:
[0317] "The user has a smiling face and a high-pitched voice, so generate a 30-second bright and cheerful pop song."
[0318] Recording and recording user performance
[0319] The user can freely sing and dance along to the generated music. At this time, the user's device will record and record the music. The recording function will be activated and the captured data will be temporarily saved on the device.
[0320] Combining performance and generative music
[0321] Once the recording is complete, the device uploads the captured performance data to the server. The server receives the performance data and combines it with the generated music data. This combination process is carried out using appropriate video editing software or algorithms to ultimately generate a single, integrated performance content. The server then transmits the combined performance content to the device. The device then plays and stores the received content, making it available for viewing by the user.
[0322] Posting to social media
[0323] If the user indicates their intention to post the synthesized performance content to a social networking site, the device calls the social networking site API to upload the synthesized content to the social networking site platform. Through the social networking site API, the user can add information such as a caption and post the content.
[0324] Sequel Music - Commercial Use
[0325] If a user checks the reaction on social media and requests a sequel song or commercialization, the server sends a sequel song and lyric generation request to the generative model. The generative model generates new music and lyrics and sends them back to the server. The server then provides the generated sequel data to the user's device, allowing the user to check the newly generated music and lyrics.
[0326] This allows users to create original music and performance content that reflects their individual emotions and share it efficiently on social media. It also supports the continuous creation of content by responding to requests for sequels and commercialization.
[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0328] Step 1:
[0329] The user launches the application and taps the "Create New Song" button.
[0330] Specifically, the user operates the app screen on the mobile device and performs an input operation by tapping the relevant button. This operation registers the generation request on the device, and the registered information is used in the next step.
[0331] Step 2:
[0332] The device sends a request to the server.
[0333] The terminal converts the generated request into packets and sends them to the server over the Internet. The input is the user's generated request, and the output is the generated HTTP request sent to the server.
[0334] Step 3:
[0335] The server activates the emotion engine and collects data to recognize the user's emotions.
[0336] When the server receives the request, it sends a command to activate the camera and microphone to the device. At this time, the collected data is sent to the emotion engine. The input is the HTTP request data, and the output is the command to start the emotion recognition engine.
[0337] Step 4:
[0338] The emotion engine uses the device's camera and microphone to collect and analyze the user's emotional data.
[0339] Specifically, the camera captures the user's facial expressions and the microphone records their voice. These data are analyzed in real time by the emotion engine, and the user's emotional information is extracted as numerical data. The input is raw data from the camera and microphone, and the output is analyzed emotional information.
[0340] Step 5:
[0341] The server inputs emotional information into the generative model and requests it to generate music, lyrics, and choreography.
[0342] The server sends the received emotional information along with a prompt to the generative AI model. The input is the analyzed emotional information and a generation request, and the output is instruction data including the prompt to the generative model.
[0343] Step 6:
[0344] The generative model generates the music, lyrics, and choreography and sends it back to the server.
[0345] The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on prompts and emotional information. The input is the prompts and emotional data, and the output is the generated music, lyrics, and choreography data.
[0346] Step 7:
[0347] The server transmits the generated data to the terminal.
[0348] The server converts the generated data into packets and sends them to the user terminal via the Internet. The input is the generated data from the generative model, and the output is the packets sent to the user terminal.
[0349] Step 8:
[0350] The data received by the device is displayed on the app screen and played back.
[0351] The terminal analyzes the generated data and displays it as a playback component of the music and choreography. The input is the generated data from the server, and the output is the content provided visually and audibly to the user.
[0352] Step 9:
[0353] Users can freely sing and dance along to the generated music.
[0354] The user presses the play button on the device to play the music and performs along with it. This operation is input into the next recording step.
[0355] Step 10:
[0356] The device will record and temporarily store your performance.
[0357] The device's camera and microphone are again used to record the user's singing and dancing. The input is the user's performance, and the output is the temporarily stored recording.
[0358] Step 11:
[0359] The device uploads the recording data to the server.
[0360] The terminal converts the temporarily stored performance data into packets and sends them to a server via the Internet. The input is the recorded data, and the output is the packets sent to the server.
[0361] Step 12:
[0362] The server receives the performance data and synthesizes it with the generated music data.
[0363] The server launches appropriate video editing software, inputs the audio and video data and the generated music data into the program, and synthesizes them. The input is the performance data and the generated music data, and the output is the synthesized performance content.
[0364] Step 13:
[0365] The server transmits the synthesized performance content to the terminal.
[0366] The server converts the composite content back into packets and transmits them to the user terminal. The input is the composite content, and the output is the packets sent to the user terminal.
[0367] Step 14:
[0368] The terminal plays and stores the received composite content so that the user can view it.
[0369] The terminal analyzes the composite content, loads it into the playback player, and provides it to the user. The input is the composite content from the server, and the output is the played and saved content.
[0370] Step 15:
[0371] The user indicates their intention to post the synthesized performance content to a social networking site.
[0372] The user taps the "Share to SNS" button in the app to indicate their intention to post. This action becomes the input for the next step.
[0373] Step 16:
[0374] The device calls the SNS API and uploads the composite content to the SNS platform.
[0375] The device uses the SNS API to send the composite content to the SNS platform, adding necessary information such as captions and uploading it. The input is the user's intention to post and the composite content, and the output is a notification of completion of posting to the SNS platform.
[0376] (Application example 2)
[0377] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0378] Conventional music generation systems have limited the content that users can experience, making it particularly difficult to generate personalized content based on emotions. Furthermore, there is a lack of easy ways to share generated content on social media platforms. This has prevented users from fully enjoying the fun of creating unique music and performances that respond to individual emotions.
[0379] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting user input, means for analyzing the user's emotions using an emotion engine, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for synthesizing the recorded performance data with the generated music, means for providing the synthesized performance content to a user terminal, and means for generating a prompt sentence for the generative AI model based on emotion information. This makes it possible to generate personalized music and performance content based on the user's emotions and easily share the generated content on a social media platform.
[0380] "User" refers to a person who uses this system to create music or performance content.
[0381] "Means for accepting input" refers to an interface that allows a user to input a request for music composition to the system.
[0382] An "emotion engine" refers to algorithms and software that analyze emotions from a user's facial expressions and voice.
[0383] "Means for analyzing emotions" refers to functionality for identifying a user's current emotional state using an emotion engine.
[0384] "Generative model" refers to an AI model that automatically generates music, lyrics, and choreography based on specified parameters and emotional information.
[0385] "Means for generating music, lyrics, and choreography" refers to a function for generating music, lyrics, and choreography using a generative model.
[0386] "User terminal" refers to a device (smartphone, tablet, etc.) that a user uses to view the music, lyrics, and choreography they have created.
[0387] "Means for recording performance" refers to the audio and video recording functionality used to record a User's performance.
[0388] "Recorded Performance Data" means audio and video data captured of a User's performance.
[0389] "Means of synthesis" refers to the function of integrating recorded performance data with the generated music to create a single piece of content.
[0390] "Synthesized Performance Content" refers to content generated as a result of integrating recorded or filmed performance data with generated music.
[0391] "SNS Platform" refers to a social networking service used to share and post User-Generated Content.
[0392] "Means for generating prompt sentences for a generative AI model" refers to a function that creates instruction sentences to be given to a generative AI model based on emotional information.
[0393] This invention relates to a system that generates music, lyrics, and choreography based on a user's emotions, allowing the user to create and share performance content. Specific embodiments of the system are described in detail below.
[0394] Hardware and Software Used
[0395] Hardware: smartphones, tablets, servers, cameras, microphones
[0396] Software: Smartphone applications, cloud server services (e.g., Amazon Web Services, Google Cloud), emotion engines (Microsoft Azure Face API, Google Cloud Speech-to-Text API), generative AI models (OpenAI GPT, Music VAE), social media APIs (Facebook API, Twitter API)
[0397] Program processing overview
[0398] 1. Launch the application:
[0399] The user launches the smartphone app and taps the "Generate new song" button.
[0400] The app activates the camera and microphone to capture the user's facial expressions and voice for emotion analysis.
[0401] 2. Emotion analysis:
[0402] The server uses an emotion engine to analyze the user's emotions from the captured data.
[0403] The emotion engine identifies the user's emotions from facial expressions and voice and sends that information back to the server.
[0404] 3. Music, lyrics, and choreography generation:
[0405] Based on the emotional information, the server sends prompts to the generative AI model to generate music, lyrics, and choreography.
[0406] The generative AI model generates original content according to the specified content and sends it back to the server.
[0407] 4. Receiving and Viewing Generated Content:
[0408] The server transmits the generated music, lyrics, and choreography data to the user terminal.
[0409] The user device plays and displays the received data within the app.
[0410] 5. Recording and filming of performances:
[0411] The user performs along with the generated music, and the device records and films the performance.
[0412] The recorded data will be temporarily stored and uploaded to a server.
[0413] 6. Video Composition and Distribution:
[0414] The server combines the audio and video data with the generated music to generate a single performance content.
[0415] The synthesized performance content is transmitted again to the user terminal.
[0416] 7. Posting to social media:
[0417] A user selects a social media post for the composited content.
[0418] Your app uses social media APIs to post content to social media platforms.
[0419] Specific examples
[0420] For example, if a user uses this app when they are tired, the app will analyze the tired expression on the user's face and generate relaxing music and slow choreography. The user can then perform along with the music, record the performance, and post it on Instagram.
[0421] Prompt Sentence Examples
[0422] Music generation prompt:
[0423] "The user's emotion is fatigue. Please generate a 30-second relaxing piece of music. The genre should be ambient and the tempo should be slow."
[0424] Lyric generation prompt:
[0425] "The user's emotion is fatigue. Generate 30 seconds of relaxing lyrics. The theme is nature and healing."
[0426] The present invention enables personalized content generation based on a user's emotions and makes it easy to share the generated content on a social networking platform.
[0427] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0428] Step 1:
[0429] The user launches the smartphone app and taps the "Generate New Song" button, which activates the app's camera and microphone to capture the user's facial expressions and voice for emotional analysis.
[0430] Input: Tap of the "Generate New Song" button, user facial and voice data
[0431] Output: User's facial expression and voice data
[0432] Step 2:
[0433] The device sends the captured facial and voice data to the server, which then activates the emotion engine to analyze the received data.
[0434] Input: User facial and voice data
[0435] Output: A signal that triggers the emotion engine to start analysis.
[0436] Step 3:
[0437] The server uses an emotion engine to analyze the user's emotions from facial expressions and voice data. The emotion engine analyzes data such as facial expressions and voice tone and sends the results back to the server.
[0438] Input: User facial and voice data received by the emotion engine
[0439] Output: User's emotional information
[0440] Step 4:
[0441] The server generates prompts for the generative AI model based on the user's emotional information. The generated prompts include requests to generate music, lyrics, and choreography.
[0442] Input: User's emotional information
[0443] Output: Prompt sentence for the generative AI model
[0444] Step 5:
[0445] The server sends the prompt text to the generative AI model, which generates the music, lyrics, and choreography. The generative AI model generates content based on the specified prompt and sends it back to the server.
[0446] Input: Prompt for the generative AI model
[0447] Output: Generated music, lyrics, and choreography data
[0448] Step 6:
[0449] The server then sends the generated music, lyrics, and choreography data to the user's device, which then plays and displays the received data within the app.
[0450] Input: Generated music, lyrics, and choreography data
[0451] Output: Playback and display on the user's device
[0452] Step 7:
[0453] The user performs along with the generated music, and the device records and films the performance. The recorded data is temporarily stored on the device.
[0454] Input: Performance along with the music
[0455] Output: Audio and video data
[0456] Step 8:
[0457] The device uploads the recorded data to the server, which then combines the received data with the generated music.
[0458] Input: Audio and video data
[0459] Output: Synthetic content
[0460] Step 9:
[0461] The server transmits the composite content to the user terminal, which then plays and displays the received composite content.
[0462] Input: Synthetic content
[0463] Output: Playback and display on the user's device
[0464] Step 10:
[0465] The user selects a social media post for the composited content, and the app uses the social media API to post the content to the social media platform.
[0466] Input: User's social media posts selection
[0467] Output: Posting content to social media platforms
[0468] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0469] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0470] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0471] [Second embodiment]
[0472] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0473] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0474] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0475] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0476] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0477] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0478] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0479] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0480] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0481] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0482] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0483] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0484] The present invention provides a system that allows users to easily create and share original music and performance content. The following describes in detail an embodiment of the system.
[0485] Music, lyrics and choreography generation
[0486] User: First launches the application and taps the "Create a new song" button, which sends a creation request from the device to the server.
[0487] Server: Upon receiving a generation request, the server requests the internally implemented generative model to generate music, lyrics, and choreography. The generative model generates an original music piece, lyrics, and choreography of approximately 30 seconds based on the input parameters.
[0488] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[0489] Server: Receives the generated data and sends it to the terminal.
[0490] Device: The received music, lyrics, and choreography data are displayed on the app screen and played.
[0491] Recording and recording user performance
[0492] User: Sing and dance freely along with the generated music.
[0493] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[0494] Combining performance and generative music
[0495] Device: Once recording is complete, performance data is uploaded to the server.
[0496] Server: Receives the performance data and combines it with the generated music data. This process is carried out using video editing software and algorithms, ultimately generating a single integrated performance content.
[0497] Server: Sends the synthesized performance content to the terminal.
[0498] Terminal: Plays and stores the received composite content so that the user can view it.
[0499] Posting to social media
[0500] User: Indicate intention to post synthesized performance content to social media.
[0501] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[0502] Sequel Music - Commercial Use
[0503] Users: Check the reaction on social media and request sequel songs and commercialization.
[0504] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[0505] Generative model: Generates new music and lyrics and sends them back to the server.
[0506] Server: Provides the generated sequel data to the device so that the user can check it.
[0507] This system allows users to easily create performance videos based on original music and share them on social media. It also encourages further creative activities based on user feedback, providing a new form of entertainment.
[0508] The processing flow will be explained below.
[0509] Step 1:
[0510] The user launches the application and taps the "Create a new song" button, which sends a request from the device to the server.
[0511] Step 2:
[0512] The server receives the generation request and asks the generative model to generate the music, lyrics, and choreography.
[0513] Step 3:
[0514] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[0515] Step 4:
[0516] The server receives the generated data and transmits it to the user terminal.
[0517] Step 5:
[0518] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[0519] Step 6:
[0520] Users sing and dance along to the music, and the device records and films this performance.
[0521] Step 7:
[0522] After the recording is complete, the device uploads this performance data to the server.
[0523] Step 8:
[0524] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[0525] Step 9:
[0526] The server generates the synthesized performance content and transmits it to the terminal.
[0527] Step 10:
[0528] The device receives, plays, and stores the composite content, which the user can then view and enjoy.
[0529] Step 11:
[0530] The user selects SNS posting and wishes to post composite content.
[0531] Step 12:
[0532] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[0533] Step 13:
[0534] Users check the reaction on social media and request sequel songs and commercialization.
[0535] Step 14:
[0536] The server sends a request to generate the sequel's music and lyrics to the generative model.
[0537] Step 15:
[0538] The generative model generates new music and lyrics and sends them back to the server.
[0539] Step 16:
[0540] The server provides the generated sequel data to the user terminal, so that the user can check it.
[0541] Example 1
[0542] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0543] Conventional systems for generating and sharing original music and performance content have had the problem that it is time-consuming for users to easily create music and choreography, record and film a performance that matches the music, and then share it on social media. Furthermore, automation of the generation of follow-up content based on reactions on social media is also insufficient. These issues need to be resolved.
[0544] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0545] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative AI model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for allowing the user to post the combined content to a social networking service (SNS) platform, and means for generating a sequel to the music and lyrics based on the response on the SNS platform. This allows users to easily create performance videos based on original music and share them on SNS, and further enables sequels to be automatically generated based on the response.
[0546] "Means for accepting user input" refers to a mechanism for inputting information or requests from a user, thereby enabling the user to send commands to the system.
[0547] A "generative AI model" refers to an artificial intelligence system that uses machine learning algorithms to automatically generate creative content such as music, lyrics, and choreography according to specific parameters.
[0548] "Means for generating music, lyrics, and choreography" refers to a mechanism that uses a generative AI model to automatically create music, lyrics, and choreography based on specified parameters.
[0549] "Means for providing to user terminal" refers to a mechanism for transmitting the generated data or content to the terminal used by the user, thereby enabling the user to view and play the generated content.
[0550] "Means for recording and recording a user's performance" refers to a mechanism for recording a user's performance in sync with the generated music, and involves capturing video and audio using the device's camera and microphone.
[0551] "Means for synthesizing performance data with generated music" refers to a mechanism for integrating recorded or filmed performance data of users with generated music to create a single piece of content.
[0552] "Means for providing synthesized performance content to a user terminal" refers to a mechanism for transmitting synthesized performance content to a user terminal and allowing the user to view and save it.
[0553] "Means for posting to an SNS platform" refers to a mechanism for uploading and publishing content created by a user to a social networking service (SNS).
[0554] "Means for generating sequel music and lyrics based on reactions on social media platforms" refers to a mechanism for analyzing reactions and evaluations on social media and generating new music and lyrics accordingly.
[0555] This invention provides a system that allows users to easily create and share original music and performance content. This system uses a generative AI model to generate music, lyrics, and choreography based on user input, and provides them to the user's device. It can also record and film the user's performance, combine it with the generated music, and post it on a social media platform. It can also generate sequels and lyrics based on social media reactions.
[0556] The system is implemented using the following hardware and software:
[0557] 1. A means of accepting user input
[0558] A user starts the application on a device (e.g., a smartphone or tablet) and taps the "Create a new song" button. This action sends a creation request from the device to the server.
[0559] 2. A means of generating music, lyrics, and choreography using generative AI models
[0560] The server receives the generation request and asks a generative AI model (e.g., a model using a machine learning algorithm) to generate the music, lyrics, and choreography. Specific models used here include OpenAI's GPT-4 and DeepComposer.
[0561] For example, a create request might use the following prompt:
[0562] Song generation prompt: The melody line is bright, the tempo is 120 BPM, and the theme is "Summer Memories."
[0563] 3. Means for providing the generated music, lyrics, and choreography to the user terminal
[0564] The generative model automatically generates music, lyrics, and choreography and sends them back to the server. The server receives the generated data and sends it to the device. The device displays the received data on the app screen and plays it back.
[0565] 4. Means for recording and recording your performance
[0566] The user sings and dances along to the generated music, and the device records and films the performance using the built-in microphone and camera. This data is temporarily stored in the device's storage.
[0567] 5. A means of synthesizing recorded performance data with generated music
[0568] Once the device has finished recording, it uploads the performance data to the server. The server receives the performance data and combines it with the generated music data. This process is carried out using video editing software (e.g., FFmpeg or Adobe Premiere Pro) and algorithms, resulting in a single integrated performance content.
[0569] 6. Means for providing synthesized performance content to a user terminal
[0570] The server sends the synthesized performance content to the terminal, which then plays and stores the received synthesized content for the user to view.
[0571] 7. Means for enabling users to post synthetic content to social media platforms
[0572] The user indicates their intention to post the synthesized performance content to a social networking site. The device then calls a social networking site API (e.g., Twitter API or Instagram Graph API) to upload the synthesized content to the social networking site, add information such as a caption, and post it.
[0573] 8. A means of generating sequel music and lyrics based on reactions on social media platforms
[0574] Users check the reaction on social media and request a sequel song and commercialization. The server sends the sequel song and lyric generation request to the generative AI model. The generative model generates new music and lyrics and sends them back to the server. The server provides the generated sequel data to the device so that the user can view it.
[0575] This system allows users to easily create performance videos based on original music and share them on social media, and can encourage further creative activities based on user feedback, providing a new form of entertainment.
[0576] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0577] Step 1:
[0578] A user launches the application and taps the "Create a new song" button. The input is the user's tap, which causes a creation request to be sent from the device to the server. Specifically, the creation request is sent as an HTTP POST request. The data included in this request is the song creation prompt and other parameters (e.g., tempo, theme).
[0579] Step 2:
[0580] The server receives the generation request and parses the request contents. The input is JSON data included in the HTTP request body. The server parses this data and converts it into an input format appropriate for the generative AI model. Specifically, it extracts the music generation prompt and parameters from the request content and converts them into a format to be passed to the generative AI model. The output is the request data sent to the generative AI model.
[0581] Step 3:
[0582] The generative AI model generates music, lyrics, and choreography based on the request data it receives. The input is the prompt and parameters sent from the server. The generative AI model applies algorithms based on these inputs to generate original music data, lyrics data, and choreography data. The output is the generated music, lyrics, and choreography data, which are sent back to the server.
[0583] Step 4:
[0584] The server receives the data returned from the generative AI model and sends it to the device. The input is the music, lyrics, and choreography data returned from the generative AI model. The server formats this data into JSON format and sends it to the device as an HTTP response. The output is the music, lyrics, and choreography data sent to the device.
[0585] Step 5:
[0586] The device displays the received data on the app screen and plays it. The input is music, lyrics, and choreography data sent from the server. The device analyzes this data, decodes and plays the music data, and displays the lyrics and choreography information on the app screen. Specifically, it plays the music using a music playback library and displays the lyrics and choreography information in a UI component. The output is audiovisual content provided to the user.
[0587] Step 6:
[0588] The user sings and dances along to the generated music. The system records and films the user's performance. The input is the user's performance, which is recorded using the device's microphone and camera. The recorded data is temporarily stored in the smartphone's storage. The output is the recorded performance data.
[0589] Step 7:
[0590] Once the device has completed recording, it uploads the performance data to the server. The input is the recorded performance data. The device sends this data to the server as an HTTP PUT request. The output is the performance data uploaded to the server.
[0591] Step 8:
[0592] The server receives the performance data and combines it with the generated music data. The input is the uploaded performance data and the generated music data. The server uses video editing software such as FFmpeg to combine these data into a single integrated performance content. Specifically, the server uses FFmpeg commands to combine the video and audio. The output is the combined performance content.
[0593] Step 9:
[0594] The server sends the synthesized performance content to the terminal. The input is the synthesized performance content. The server sends this content to the terminal as an HTTP response. The output is the synthesized performance content sent to the terminal.
[0595] Step 10:
[0596] The device plays and saves the received composite content, making it available for viewing by the user. The input is composite performance content sent from the server. The device plays this content and saves it in storage. Specifically, it plays it using a video playback library and saves it using a file system API. The output is playable composite content provided to the user.
[0597] Step 11:
[0598] The user indicates their intention to post the synthesized performance content to a social networking site. The input is the user's posting operation, and the device uses the social networking site API to upload the synthesized content to the social networking site platform. Specifically, the device sends an HTTP POST request to the social networking site platform's API to upload information such as captions along with the synthesized content. The output is the content posted to the social networking site platform.
[0599] Step 12:
[0600] Users check the reaction on social media and request a sequel song and commercialization. The input is the user's sequel request operation, which is sent to the server. The server sends a sequel song and lyric generation request to the generative AI model. The output is the request data sent to the generative AI model.
[0601] Step 13:
[0602] The generative AI model generates new music and lyrics and sends them back to the server. The input is the sequel request data, and the generative AI model generates new music and lyrics based on this. The server provides the generated data to the device so that the user can view it. Specifically, the generated data is returned to the server in JSON format, which is then sent to the device. The output is the generated sequel music and lyrics data.
[0603] (Application example 1)
[0604] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0605] The challenge is to provide a system that allows users to easily create original music, lyrics, and choreography, synthesize audio and video recordings of performances to match them, and then efficiently post the content to a social media platform. Another challenge is to generate follow-up music and performance content based on the reaction on the social media platform, thereby continuously supporting users' creative activities.
[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0607] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for posting the generated content to a social networking service (SNS) platform, and means for checking the reaction on the SNS platform and requesting the generation of a sequel music and lyrics. This allows users to efficiently and consistently perform processes from generating original content to sharing it on SNS and further creating sequels.
[0608] definition statement
[0609] "Means for accepting user input" refers to the interface that the user operates when requesting the creation of new music, lyrics, and choreography.
[0610] "Means for generating music, lyrics, and choreography using generative models" refers to a system that uses AI models and algorithms to automatically generate original music, lyrics, and choreography based on input parameters.
[0611] "Means for providing the generated music, lyrics and choreography to the user terminal" refers to the process of transmitting the generated content from the server to the user terminal.
[0612] "Means for recording a user's performance" refers to a recording function for recording a user's performance in sync with the generated music.
[0613] "Means of combining recorded performance data with generated music" refers to the process of integrating recorded performance data with generated music data and editing it into a complete performance content.
[0614] The term "means for providing the synthesized performance content to the user terminal" refers to a process for transmitting the synthesized performance content to the user terminal and making it playable.
[0615] "Means for posting Generated Content to a Social Media Platform" means an interface for uploading Generated Performance Content to a Social Media Platform and for users to share it.
[0616] "Means for checking reactions on social media platforms and making requests to generate sequel music and lyrics" refers to a system that analyzes user reactions on social media platforms and requests the generation of sequel content based on the results.
[0617] MODE FOR CARRYING OUT THE INVENTION
[0618] This invention is a system that generates original music, lyrics, and choreography based on prompts entered by the user, integrates performance content recorded or filmed by the user with the generated content, and shares it on a social media platform.
[0619] The specific steps for implementing this system are as follows:
[0620] First, a user accesses the application using a device (smartphone or smart glasses). The user taps the "Create a new song" button to proceed to a prompt input screen. Here, the user enters a specific prompt, such as "Create an energetic pop song."
[0621] The server then receives prompts from the user, uses its internal generative AI model to automatically generate music, lyrics, and choreography based on the prompts, and sends the generated content to the device.
[0622] The device displays the generated music, lyrics, and choreography data on the application screen, allowing the user to play it. The user records and films their performance while singing and dancing to the generated music. The recorded and filmed performance data is temporarily stored on the device.
[0623] Once the user has completed recording the performance, the device uploads the performance data to the server. The server then combines the received performance data with the generated music data using video editing software. The combined performance content is then sent back to the device from the server.
[0624] The device that receives the composite content makes it playable for the user. If the user checks the generated performance content and indicates their intention to post it to a social media platform, the device calls the social media API to upload the content. The user can then add information such as a caption and finally post it to the social media platform.
[0625] The server analyzes the reactions to the content posted by users on the SNS platform, and based on the results, users can request the generation of a sequel to the song and lyrics, thereby providing continuous support for the user's creative activities.
[0626] For example, consider the case where a user enters the following prompt:
[0627] Prompt: "Generate an energetic pop song"
[0628] In this way, this invention is a system that enables users to efficiently and consistently create original content, share it on social media, and even create sequels.
[0629] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0630] Program processing flow
[0631] Step 1:
[0632] A user accesses the application using a device, taps the "Create a new song" button, and enters a prompt, which is then sent to the server.
[0633] Input: User prompt (e.g., "Generate an energetic pop song")
[0634] Output: Prompt data is sent to the server
[0635] Step 2:
[0636] The server receives the prompts and generates the music, lyrics, and choreography based on an internal generative AI model.
[0637] Input: prompt data
[0638] Data processing: Based on the input prompt, the generative AI model generates the music, lyrics, and choreography.
[0639] Output: Generated music data, lyrics data, choreography data
[0640] Step 3:
[0641] The server transmits the generated data to the terminal, which receives and displays it.
[0642] Input: Generated music data, lyrics data, choreography data
[0643] Output: Music, lyrics, and choreography displayed on the user's device
[0644] Step 4:
[0645] The user performs along with the generated music, and the device records and films this.
[0646] Input: User performance (singing, dancing)
[0647] Data processing: Activate the recording function and capture performance data
[0648] Output: Recorded and recorded performance data is temporarily saved.
[0649] Step 5:
[0650] The device uploads the recorded performance data to the server.
[0651] Input: Audio-visual performance data
[0652] Output: Performance data is sent to the server
[0653] Step 6:
[0654] The server synthesizes the received performance data with the generated music data.
[0655] Input: Performance data, generated song data
[0656] Data processing: Composition using video editing software
[0657] Output: Synthesized performance content
[0658] Step 7:
[0659] The server transmits the synthesized performance content to the terminal, which receives it and makes it playable.
[0660] Input: Synthesized performance content
[0661] Output: The composite content displayed on the device
[0662] Step 8:
[0663] The user indicates their intention to post the composite content to a social networking platform, and the device calls the social networking API to upload the content.
[0664] Input: User's posting intent, synthetic content
[0665] Data processing: Upload content using SNS API
[0666] Output: Content posted on social media platforms
[0667] Step 9:
[0668] The server analyzes the reaction on the social media platform and begins the process of generating the sequel's music and lyrics based on user requests.
[0669] Input: Reaction data from social media platforms
[0670] Data processing: A generative AI model based on feedback data generates the music and lyrics for the sequel
[0671] Output: Sequel music data, lyrics data
[0672] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0673] The present invention provides a system that allows users to easily create and share original music and performance content. The system also incorporates an emotion engine that can recognize a user's emotions and adjust the generated content accordingly. The following describes in detail an embodiment of the system.
[0674] Music, lyrics and choreography generation
[0675] User: First launches the application and taps the "Generate New Song" button, which sends a request from the device to the server.
[0676] Server: Upon receiving a generation request, it launches the emotion engine and collects data to recognize the user's emotion.
[0677] Emotion engine: Analyzes the user's facial expressions, tone of voice, and other emotional data from the camera and microphone to identify the emotion the user is currently feeling. The emotion engine then sends the recognized emotional information back to the server.
[0678] Server: Requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on the input parameters and emotional information.
[0679] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[0680] Server: Receives the generated data and sends it to the terminal.
[0681] Device: The received music, lyrics, and choreography data is displayed on the app screen and played, allowing the user to listen to the music, lyrics, and choreography.
[0682] Recording and recording user performance
[0683] User: Sing and dance freely along with the generated music.
[0684] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[0685] Combining performance and generative music
[0686] Device: Once recording is complete, performance data is uploaded to the server.
[0687] Server: Receives the performance data and synthesizes it with the generated music data. The synthesis process is carried out using appropriate video editing software or algorithms, and finally a single integrated performance content is generated.
[0688] Server: Sends the synthesized performance content to the terminal.
[0689] Terminal: Plays and stores the received composite content so that the user can view it.
[0690] Posting to social media
[0691] User: Indicate intention to post synthesized performance content to social media.
[0692] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[0693] Sequel Music - Commercial Use
[0694] Users: Check the reaction on social media and request sequel songs and commercialization.
[0695] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[0696] Generative model: Generates new music and lyrics and sends them back to the server.
[0697] Server: Provides the generated sequel data to the device so that the user can check it.
[0698] This system allows users to easily create performance videos based on original music and share them on social media. Furthermore, by generating content based on the user's emotional information, it is possible to provide a more personalized experience.
[0699] The processing flow will be explained below.
[0700] Step 1:
[0701] The user launches the application and taps the "Create a new song" button, which sends a song creation request from the device to the server.
[0702] Step 2:
[0703] The server receives the generation request and starts the emotion engine, which prepares to start collecting data necessary to recognize the user's emotion.
[0704] Step 3:
[0705] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and this data is sent to the emotion engine in real time.
[0706] Step 4:
[0707] The emotion engine analyzes the received data and identifies the user's emotion. For example, if the user is smiling, it recognizes "happiness," and if they are frowning, it recognizes "anxiety."
[0708] Step 5:
[0709] The emotion engine sends the identified emotion information back to the server, which collects this information and prepares it for transmission to the generative model.
[0710] Step 6:
[0711] The server requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model takes the emotional information into consideration and generates music, lyrics, and choreography that match the emotional information.
[0712] Step 7:
[0713] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[0714] Step 8:
[0715] The server receives the generated data and transmits it to the user terminal.
[0716] Step 9:
[0717] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[0718] Step 10:
[0719] The user sings and dances freely to the generated music, and the device records and films this performance.
[0720] Step 11:
[0721] Once the recording is complete, the device uploads this performance data to a server.
[0722] Step 12:
[0723] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[0724] Step 13:
[0725] The server generates the synthesized performance content and transmits it to the terminal.
[0726] Step 14:
[0727] The device plays and saves the received composite content, which the user can then view and enjoy.
[0728] Step 15:
[0729] The user selects SNS posting and wishes to post composite content.
[0730] Step 16:
[0731] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[0732] Step 17:
[0733] Users check the reaction on social media and request sequel songs and commercialization.
[0734] Step 18:
[0735] The server sends a request to generate the sequel's music and lyrics to the generative model.
[0736] Step 19:
[0737] The generative model generates new music and lyrics and sends them back to the server.
[0738] Step 20:
[0739] The server provides the generated sequel data to the user terminal, so that the user can check it.
[0740] Example 2
[0741] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0742] In today's entertainment environment, there is a demand for a means for users to easily create and share original music and performance content. However, existing systems have difficulty generating content based on users' emotions, and there are insufficient methods for efficiently synthesizing generated content and sharing it on social media. Furthermore, there is a lack of easy ways to request and realize sequels or commercialization. Therefore, a system that makes it easier for users to create and share original content that responds to individual emotions is needed.
[0743] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0744] In this invention, the server includes means for recognizing the user's emotional information and generating music, lyrics, and choreography, means for providing the synthesized performance content to the user terminal, and means for posting the content to an SNS platform, thereby enabling users to generate original content according to their individual emotions, efficiently synthesize the content, and share it on the SNS.
[0745] The "means for accepting user input" is an interface that allows the user to input operational instructions to the system.
[0746] "Generative model" refers to algorithms or software that use artificial intelligence to automatically generate content such as music, lyrics, and choreography.
[0747] "Means for recognizing user's emotional information" refers to an engine or software that analyzes the user's facial expressions and tone of voice via a camera or microphone to identify their current emotional state.
[0748] "Means for generating music, lyrics and choreography" refers to a device or software that generates original music, lyrics and choreography based on emotional information and input data.
[0749] "Means for providing the generated music, lyrics and choreography to the user's terminal" refers to a device or program that has the function of transferring the generated data from the server to the user's terminal and displaying and playing it.
[0750] "Means for recording a user's performance" refers to a device or function that uses a camera and microphone to record a user's performance along with a piece of music.
[0751] "Means for synthesizing recorded performance data with generated music" refers to software or a system for integrating a user's performance data with generated music and editing and generating a single composite content.
[0752] "Means for providing synthesized performance content to a user terminal" refers to a device or program for transmitting the completed synthesized content from the server to a user terminal so that it can be displayed and played.
[0753] "Means for posting to a social media platform" refers to a device or software that has the function of uploading and sharing generated content to a social media platform using an API or other means.
[0754] "Means for generating sequel music and lyrics" refers to artificial intelligence algorithms or models for generating new music and lyrics based on existing content.
[0755] The present invention relates to a system that enables users to easily create and share original music and performance content. Specific embodiments of this system will be described below.
[0756] Music, lyrics and choreography generation
[0757] First, the user launches the application and taps the "Generate a new song" button. This action sends a generation request from the device to the server. When the server receives the generation request, it launches an emotion engine and collects data to recognize the user's emotions. The emotion engine analyzes the user's facial expressions and vocal tone through the camera and microphone to identify the user's current emotion. This emotion information is sent back to the server. The server, along with the received emotion information, requests the generative model to generate music, lyrics, and choreography. The generative model generates original music, lyrics, and choreography based on the input prompt and emotion information and sends it back to the server. The server receives the generated data and sends it to the user's device. The device displays and plays this data on the app screen.
[0758] Hardware and software used
[0759] Hardware: Smartphone (camera, microphone)
[0760] Software: Emotion engine (facial expression and voice analysis software), generative AI model (music, lyrics, choreography generation)
[0761] Specific examples
[0762] For example, when a user launches the app and taps the "Generate a new song" button, the smartphone's camera and microphone are activated, and the user's facial expressions and voice are analyzed. If the user is in a happy mood, the generative model will generate an upbeat, cheerful song.
[0763] Example prompt for a generative AI model:
[0764] "The user has a smiling face and a high-pitched voice, so generate a 30-second bright and cheerful pop song."
[0765] Recording and recording user performance
[0766] The user can freely sing and dance along to the generated music. At this time, the user's device will record and record the music. The recording function will be activated and the captured data will be temporarily saved on the device.
[0767] Combining performance and generative music
[0768] Once the recording is complete, the device uploads the captured performance data to the server. The server receives the performance data and combines it with the generated music data. This combination process is carried out using appropriate video editing software or algorithms to ultimately generate a single, integrated performance content. The server then transmits the combined performance content to the device. The device then plays and stores the received content, making it available for viewing by the user.
[0769] Posting to social media
[0770] If the user indicates their intention to post the synthesized performance content to a social networking site, the device calls the social networking site API to upload the synthesized content to the social networking site platform. Through the social networking site API, the user can add information such as a caption and post the content.
[0771] Sequel Music - Commercial Use
[0772] If a user checks the reaction on social media and requests a sequel song or commercialization, the server sends a sequel song and lyric generation request to the generative model. The generative model generates new music and lyrics and sends them back to the server. The server then provides the generated sequel data to the user's device, allowing the user to check the newly generated music and lyrics.
[0773] This allows users to create original music and performance content that reflects their individual emotions and share it efficiently on social media. It also supports the continuous creation of content by responding to requests for sequels and commercialization.
[0774] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0775] Step 1:
[0776] The user launches the application and taps the "Create New Song" button.
[0777] Specifically, the user operates the app screen on the mobile device and performs an input operation by tapping the relevant button. This operation registers the generation request on the device, and the registered information is used in the next step.
[0778] Step 2:
[0779] The device sends a request to the server.
[0780] The terminal converts the generated request into packets and sends them to the server over the Internet. The input is the user's generated request, and the output is the generated HTTP request sent to the server.
[0781] Step 3:
[0782] The server activates the emotion engine and collects data to recognize the user's emotions.
[0783] When the server receives the request, it sends a command to activate the camera and microphone to the device. At this time, the collected data is sent to the emotion engine. The input is the HTTP request data, and the output is the command to start the emotion recognition engine.
[0784] Step 4:
[0785] The emotion engine uses the device's camera and microphone to collect and analyze the user's emotional data.
[0786] Specifically, the camera captures the user's facial expressions and the microphone records their voice. These data are analyzed in real time by the emotion engine, and the user's emotional information is extracted as numerical data. The input is raw data from the camera and microphone, and the output is analyzed emotional information.
[0787] Step 5:
[0788] The server inputs emotional information into the generative model and requests it to generate music, lyrics, and choreography.
[0789] The server sends the received emotional information along with a prompt to the generative AI model. The input is the analyzed emotional information and a generation request, and the output is instruction data including the prompt to the generative model.
[0790] Step 6:
[0791] The generative model generates the music, lyrics, and choreography and sends it back to the server.
[0792] The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on prompts and emotional information. The input is the prompts and emotional data, and the output is the generated music, lyrics, and choreography data.
[0793] Step 7:
[0794] The server transmits the generated data to the terminal.
[0795] The server converts the generated data into packets and sends them to the user terminal via the Internet. The input is the generated data from the generative model, and the output is the packets sent to the user terminal.
[0796] Step 8:
[0797] The data received by the device is displayed on the app screen and played back.
[0798] The terminal analyzes the generated data and displays it as a playback component of the music and choreography. The input is the generated data from the server, and the output is the content provided visually and audibly to the user.
[0799] Step 9:
[0800] Users can freely sing and dance along to the generated music.
[0801] The user presses the play button on the device to play the music and performs along with it. This operation is input into the next recording step.
[0802] Step 10:
[0803] The device will record and temporarily store your performance.
[0804] The device's camera and microphone are again used to record the user's singing and dancing. The input is the user's performance, and the output is the temporarily stored recording.
[0805] Step 11:
[0806] The device uploads the recording data to the server.
[0807] The terminal converts the temporarily stored performance data into packets and sends them to a server via the Internet. The input is the recorded data, and the output is the packets sent to the server.
[0808] Step 12:
[0809] The server receives the performance data and synthesizes it with the generated music data.
[0810] The server launches appropriate video editing software, inputs the audio and video data and the generated music data into the program, and synthesizes them. The input is the performance data and the generated music data, and the output is the synthesized performance content.
[0811] Step 13:
[0812] The server transmits the synthesized performance content to the terminal.
[0813] The server converts the composite content back into packets and transmits them to the user terminal. The input is the composite content, and the output is the packets sent to the user terminal.
[0814] Step 14:
[0815] The terminal plays and stores the received composite content so that the user can view it.
[0816] The terminal analyzes the composite content, loads it into the playback player, and provides it to the user. The input is the composite content from the server, and the output is the played and saved content.
[0817] Step 15:
[0818] The user indicates their intention to post the synthesized performance content to a social networking site.
[0819] The user taps the "Share to SNS" button in the app to indicate their intention to post. This action becomes the input for the next step.
[0820] Step 16:
[0821] The device calls the SNS API and uploads the composite content to the SNS platform.
[0822] The device uses the SNS API to send the composite content to the SNS platform, adding necessary information such as captions and uploading it. The input is the user's intention to post and the composite content, and the output is a notification of completion of posting to the SNS platform.
[0823] (Application example 2)
[0824] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0825] Conventional music generation systems have limited the content that users can experience, making it particularly difficult to generate personalized content based on emotions. Furthermore, there is a lack of easy ways to share generated content on social media platforms. This has prevented users from fully enjoying the fun of creating unique music and performances that respond to individual emotions.
[0826] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting user input, means for analyzing the user's emotions using an emotion engine, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for synthesizing the recorded performance data with the generated music, means for providing the synthesized performance content to a user terminal, and means for generating a prompt sentence for the generative AI model based on emotion information. This makes it possible to generate personalized music and performance content based on the user's emotions and easily share the generated content on a social media platform.
[0827] "User" refers to a person who uses this system to create music or performance content.
[0828] "Means for accepting input" refers to an interface that allows a user to input a request for music composition to the system.
[0829] An "emotion engine" refers to algorithms and software that analyze emotions from a user's facial expressions and voice.
[0830] "Means for analyzing emotions" refers to functionality for identifying a user's current emotional state using an emotion engine.
[0831] "Generative model" refers to an AI model that automatically generates music, lyrics, and choreography based on specified parameters and emotional information.
[0832] "Means for generating music, lyrics, and choreography" refers to a function for generating music, lyrics, and choreography using a generative model.
[0833] "User terminal" refers to a device (smartphone, tablet, etc.) that a user uses to view the music, lyrics, and choreography they have created.
[0834] "Means for recording performance" refers to the audio and video recording functionality used to record a User's performance.
[0835] "Recorded Performance Data" means audio and video data captured of a User's performance.
[0836] "Means of synthesis" refers to the function of integrating recorded performance data with the generated music to create a single piece of content.
[0837] "Synthesized Performance Content" refers to content generated as a result of integrating recorded or filmed performance data with generated music.
[0838] "SNS Platform" refers to a social networking service used to share and post User-Generated Content.
[0839] "Means for generating prompt sentences for a generative AI model" refers to a function that creates instruction sentences to be given to a generative AI model based on emotional information.
[0840] This invention relates to a system that generates music, lyrics, and choreography based on a user's emotions, allowing the user to create and share performance content. Specific embodiments of the system are described in detail below.
[0841] Hardware and Software Used
[0842] Hardware: smartphones, tablets, servers, cameras, microphones
[0843] Software: Smartphone applications, cloud server services (e.g., Amazon Web Services, Google Cloud), emotion engines (Microsoft Azure Face API, Google Cloud Speech-to-Text API), generative AI models (OpenAI GPT, Music VAE), social media APIs (Facebook API, Twitter API)
[0844] Program processing overview
[0845] 1. Launch the application:
[0846] The user launches the smartphone app and taps the "Generate new song" button.
[0847] The app activates the camera and microphone to capture the user's facial expressions and voice for emotion analysis.
[0848] 2. Emotion analysis:
[0849] The server uses an emotion engine to analyze the user's emotions from the captured data.
[0850] The emotion engine identifies the user's emotions from facial expressions and voice and sends that information back to the server.
[0851] 3. Music, lyrics, and choreography generation:
[0852] Based on the emotional information, the server sends prompts to the generative AI model to generate music, lyrics, and choreography.
[0853] The generative AI model generates original content according to the specified content and sends it back to the server.
[0854] 4. Receiving and Viewing Generated Content:
[0855] The server transmits the generated music, lyrics, and choreography data to the user terminal.
[0856] The user device plays and displays the received data within the app.
[0857] 5. Recording and filming of performances:
[0858] The user performs along with the generated music, and the device records and films the performance.
[0859] The recorded data will be temporarily stored and uploaded to a server.
[0860] 6. Video Composition and Distribution:
[0861] The server combines the audio and video data with the generated music to generate a single performance content.
[0862] The synthesized performance content is transmitted again to the user terminal.
[0863] 7. Posting to social media:
[0864] A user selects a social media post for the composited content.
[0865] Your app uses social media APIs to post content to social media platforms.
[0866] Specific examples
[0867] For example, if a user uses this app when they are tired, the app will analyze the tired expression on the user's face and generate relaxing music and slow choreography. The user can then perform along with the music, record the performance, and post it on Instagram.
[0868] Prompt Sentence Examples
[0869] Music generation prompt:
[0870] "The user's emotion is fatigue. Please generate a 30-second relaxing piece of music. The genre should be ambient and the tempo should be slow."
[0871] Lyric generation prompt:
[0872] "The user's emotion is fatigue. Generate 30 seconds of relaxing lyrics. The theme is nature and healing."
[0873] The present invention enables personalized content generation based on a user's emotions and makes it easy to share the generated content on a social networking platform.
[0874] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0875] Step 1:
[0876] The user launches the smartphone app and taps the "Generate New Song" button, which activates the app's camera and microphone to capture the user's facial expressions and voice for emotional analysis.
[0877] Input: Tap of the "Generate New Song" button, user facial and voice data
[0878] Output: User's facial expression and voice data
[0879] Step 2:
[0880] The device sends the captured facial and voice data to the server, which then activates the emotion engine to analyze the received data.
[0881] Input: User facial and voice data
[0882] Output: A signal that triggers the emotion engine to start analysis.
[0883] Step 3:
[0884] The server uses an emotion engine to analyze the user's emotions from facial expressions and voice data. The emotion engine analyzes data such as facial expressions and voice tone and sends the results back to the server.
[0885] Input: User facial and voice data received by the emotion engine
[0886] Output: User's emotional information
[0887] Step 4:
[0888] The server generates prompts for the generative AI model based on the user's emotional information. The generated prompts include requests to generate music, lyrics, and choreography.
[0889] Input: User's emotional information
[0890] Output: Prompt sentence for the generative AI model
[0891] Step 5:
[0892] The server sends the prompt text to the generative AI model, which generates the music, lyrics, and choreography. The generative AI model generates content based on the specified prompt and sends it back to the server.
[0893] Input: Prompt for the generative AI model
[0894] Output: Generated music, lyrics, and choreography data
[0895] Step 6:
[0896] The server then sends the generated music, lyrics, and choreography data to the user's device, which then plays and displays the received data within the app.
[0897] Input: Generated music, lyrics, and choreography data
[0898] Output: Playback and display on the user's device
[0899] Step 7:
[0900] The user performs along with the generated music, and the device records and films the performance. The recorded data is temporarily stored on the device.
[0901] Input: Performance along with the music
[0902] Output: Audio and video data
[0903] Step 8:
[0904] The device uploads the recorded data to the server, which then combines the received data with the generated music.
[0905] Input: Audio and video data
[0906] Output: Synthetic content
[0907] Step 9:
[0908] The server transmits the composite content to the user terminal, which then plays and displays the received composite content.
[0909] Input: Synthetic content
[0910] Output: Playback and display on the user's device
[0911] Step 10:
[0912] The user selects a social media post for the composited content, and the app uses the social media API to post the content to the social media platform.
[0913] Input: User's social media posts selection
[0914] Output: Posting content to social media platforms
[0915] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0916] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0917] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0918] [Third embodiment]
[0919] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0920] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0921] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0922] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0923] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0924] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0925] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0926] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0927] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0928] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0929] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0930] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0931] The present invention provides a system that allows users to easily create and share original music and performance content. The following describes in detail an embodiment of the system.
[0932] Music, lyrics and choreography generation
[0933] User: First launches the application and taps the "Create a new song" button, which sends a creation request from the device to the server.
[0934] Server: Upon receiving a generation request, the server requests the internally implemented generative model to generate music, lyrics, and choreography. The generative model generates an original music piece, lyrics, and choreography of approximately 30 seconds based on the input parameters.
[0935] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[0936] Server: Receives the generated data and sends it to the terminal.
[0937] Device: The received music, lyrics, and choreography data are displayed on the app screen and played.
[0938] Recording and recording user performance
[0939] User: Sing and dance freely along with the generated music.
[0940] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[0941] Combining performance and generative music
[0942] Device: Once recording is complete, performance data is uploaded to the server.
[0943] Server: Receives the performance data and combines it with the generated music data. This process is carried out using video editing software and algorithms, ultimately generating a single integrated performance content.
[0944] Server: Sends the synthesized performance content to the terminal.
[0945] Terminal: Plays and stores the received composite content so that the user can view it.
[0946] Posting to social media
[0947] User: Indicate intention to post synthesized performance content to social media.
[0948] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[0949] Sequel Music - Commercial Use
[0950] Users: Check the reaction on social media and request sequel songs and commercialization.
[0951] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[0952] Generative model: Generates new music and lyrics and sends them back to the server.
[0953] Server: Provides the generated sequel data to the device so that the user can check it.
[0954] This system allows users to easily create performance videos based on original music and share them on social media. It also encourages further creative activities based on user feedback, providing a new form of entertainment.
[0955] The processing flow will be explained below.
[0956] Step 1:
[0957] The user launches the application and taps the "Create a new song" button, which sends a request from the device to the server.
[0958] Step 2:
[0959] The server receives the generation request and asks the generative model to generate the music, lyrics, and choreography.
[0960] Step 3:
[0961] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[0962] Step 4:
[0963] The server receives the generated data and transmits it to the user terminal.
[0964] Step 5:
[0965] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[0966] Step 6:
[0967] Users sing and dance along to the music, and the device records and films this performance.
[0968] Step 7:
[0969] After the recording is complete, the device uploads this performance data to the server.
[0970] Step 8:
[0971] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[0972] Step 9:
[0973] The server generates the synthesized performance content and transmits it to the terminal.
[0974] Step 10:
[0975] The device receives, plays, and stores the composite content, which the user can then view and enjoy.
[0976] Step 11:
[0977] The user selects SNS posting and wishes to post composite content.
[0978] Step 12:
[0979] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[0980] Step 13:
[0981] Users check the reaction on social media and request sequel songs and commercialization.
[0982] Step 14:
[0983] The server sends a request to generate the sequel's music and lyrics to the generative model.
[0984] Step 15:
[0985] The generative model generates new music and lyrics and sends them back to the server.
[0986] Step 16:
[0987] The server provides the generated sequel data to the user terminal, so that the user can check it.
[0988] Example 1
[0989] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0990] Conventional systems for generating and sharing original music and performance content have had the problem that it is time-consuming for users to easily create music and choreography, record and film a performance that matches the music, and then share it on social media. Furthermore, automation of the generation of follow-up content based on reactions on social media is also insufficient. These issues need to be resolved.
[0991] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0992] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative AI model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for allowing the user to post the combined content to a social networking service (SNS) platform, and means for generating a sequel to the music and lyrics based on the response on the SNS platform. This allows users to easily create performance videos based on original music and share them on SNS, and further enables sequels to be automatically generated based on the response.
[0993] "Means for accepting user input" refers to a mechanism for inputting information or requests from a user, thereby enabling the user to send commands to the system.
[0994] A "generative AI model" refers to an artificial intelligence system that uses machine learning algorithms to automatically generate creative content such as music, lyrics, and choreography according to specific parameters.
[0995] "Means for generating music, lyrics, and choreography" refers to a mechanism that uses a generative AI model to automatically create music, lyrics, and choreography based on specified parameters.
[0996] "Means for providing to user terminal" refers to a mechanism for transmitting the generated data or content to the terminal used by the user, thereby enabling the user to view and play the generated content.
[0997] "Means for recording and recording a user's performance" refers to a mechanism for recording a user's performance in sync with the generated music, and involves capturing video and audio using the device's camera and microphone.
[0998] "Means for synthesizing performance data with generated music" refers to a mechanism for integrating recorded or filmed performance data of users with generated music to create a single piece of content.
[0999] "Means for providing synthesized performance content to a user terminal" refers to a mechanism for transmitting synthesized performance content to a user terminal and allowing the user to view and save it.
[1000] "Means for posting to an SNS platform" refers to a mechanism for uploading and publishing content created by a user to a social networking service (SNS).
[1001] "Means for generating sequel music and lyrics based on reactions on social media platforms" refers to a mechanism for analyzing reactions and evaluations on social media and generating new music and lyrics accordingly.
[1002] This invention provides a system that allows users to easily create and share original music and performance content. This system uses a generative AI model to generate music, lyrics, and choreography based on user input, and provides them to the user's device. It can also record and film the user's performance, combine it with the generated music, and post it on a social media platform. It can also generate sequels and lyrics based on social media reactions.
[1003] The system is implemented using the following hardware and software:
[1004] 1. A means of accepting user input
[1005] A user starts the application on a device (e.g., a smartphone or tablet) and taps the "Create a new song" button. This action sends a creation request from the device to the server.
[1006] 2. A means of generating music, lyrics, and choreography using generative AI models
[1007] The server receives the generation request and asks a generative AI model (e.g., a model using a machine learning algorithm) to generate the music, lyrics, and choreography. Specific models used here include OpenAI's GPT-4 and DeepComposer.
[1008] For example, a create request might use the following prompt:
[1009] Song generation prompt: The melody line is bright, the tempo is 120 BPM, and the theme is "Summer Memories."
[1010] 3. Means for providing the generated music, lyrics, and choreography to the user terminal
[1011] The generative model automatically generates music, lyrics, and choreography and sends them back to the server. The server receives the generated data and sends it to the device. The device displays the received data on the app screen and plays it back.
[1012] 4. Means for recording and recording your performance
[1013] The user sings and dances along to the generated music, and the device records and films the performance using the built-in microphone and camera. This data is temporarily stored in the device's storage.
[1014] 5. A means of synthesizing recorded performance data with generated music
[1015] Once the device has finished recording, it uploads the performance data to the server. The server receives the performance data and combines it with the generated music data. This process is carried out using video editing software (e.g., FFmpeg or Adobe Premiere Pro) and algorithms, resulting in a single integrated performance content.
[1016] 6. Means for providing synthesized performance content to a user terminal
[1017] The server sends the synthesized performance content to the terminal, which then plays and stores the received synthesized content for the user to view.
[1018] 7. Means for enabling users to post synthetic content to social media platforms
[1019] The user indicates their intention to post the synthesized performance content to a social networking site. The device then calls a social networking site API (e.g., Twitter API or Instagram Graph API) to upload the synthesized content to the social networking site, add information such as a caption, and post it.
[1020] 8. A means of generating sequel music and lyrics based on reactions on social media platforms
[1021] Users check the reaction on social media and request a sequel song and commercialization. The server sends the sequel song and lyric generation request to the generative AI model. The generative model generates new music and lyrics and sends them back to the server. The server provides the generated sequel data to the device so that the user can view it.
[1022] This system allows users to easily create performance videos based on original music and share them on social media, and can encourage further creative activities based on user feedback, providing a new form of entertainment.
[1023] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1024] Step 1:
[1025] A user launches the application and taps the "Create a new song" button. The input is the user's tap, which causes a creation request to be sent from the device to the server. Specifically, the creation request is sent as an HTTP POST request. The data included in this request is the song creation prompt and other parameters (e.g., tempo, theme).
[1026] Step 2:
[1027] The server receives the generation request and parses the request contents. The input is JSON data included in the HTTP request body. The server parses this data and converts it into an input format appropriate for the generative AI model. Specifically, it extracts the music generation prompt and parameters from the request content and converts them into a format to be passed to the generative AI model. The output is the request data sent to the generative AI model.
[1028] Step 3:
[1029] The generative AI model generates music, lyrics, and choreography based on the request data it receives. The input is the prompt and parameters sent from the server. The generative AI model applies algorithms based on these inputs to generate original music data, lyrics data, and choreography data. The output is the generated music, lyrics, and choreography data, which are sent back to the server.
[1030] Step 4:
[1031] The server receives the data returned from the generative AI model and sends it to the device. The input is the music, lyrics, and choreography data returned from the generative AI model. The server formats this data into JSON format and sends it to the device as an HTTP response. The output is the music, lyrics, and choreography data sent to the device.
[1032] Step 5:
[1033] The device displays the received data on the app screen and plays it. The input is music, lyrics, and choreography data sent from the server. The device analyzes this data, decodes and plays the music data, and displays the lyrics and choreography information on the app screen. Specifically, it plays the music using a music playback library and displays the lyrics and choreography information in a UI component. The output is audiovisual content provided to the user.
[1034] Step 6:
[1035] The user sings and dances along to the generated music. The system records and films the user's performance. The input is the user's performance, which is recorded using the device's microphone and camera. The recorded data is temporarily stored in the smartphone's storage. The output is the recorded performance data.
[1036] Step 7:
[1037] Once the device has completed recording, it uploads the performance data to the server. The input is the recorded performance data. The device sends this data to the server as an HTTP PUT request. The output is the performance data uploaded to the server.
[1038] Step 8:
[1039] The server receives the performance data and combines it with the generated music data. The input is the uploaded performance data and the generated music data. The server uses video editing software such as FFmpeg to combine these data into a single integrated performance content. Specifically, the server uses FFmpeg commands to combine the video and audio. The output is the combined performance content.
[1040] Step 9:
[1041] The server sends the synthesized performance content to the terminal. The input is the synthesized performance content. The server sends this content to the terminal as an HTTP response. The output is the synthesized performance content sent to the terminal.
[1042] Step 10:
[1043] The device plays and saves the received composite content, making it available for viewing by the user. The input is composite performance content sent from the server. The device plays this content and saves it in storage. Specifically, it plays it using a video playback library and saves it using a file system API. The output is playable composite content provided to the user.
[1044] Step 11:
[1045] The user indicates their intention to post the synthesized performance content to a social networking site. The input is the user's posting operation, and the device uses the social networking site API to upload the synthesized content to the social networking site platform. Specifically, the device sends an HTTP POST request to the social networking site platform's API to upload information such as captions along with the synthesized content. The output is the content posted to the social networking site platform.
[1046] Step 12:
[1047] Users check the reaction on social media and request a sequel song and commercialization. The input is the user's sequel request operation, which is sent to the server. The server sends a sequel song and lyric generation request to the generative AI model. The output is the request data sent to the generative AI model.
[1048] Step 13:
[1049] The generative AI model generates new music and lyrics and sends them back to the server. The input is the sequel request data, and the generative AI model generates new music and lyrics based on this. The server provides the generated data to the device so that the user can view it. Specifically, the generated data is returned to the server in JSON format, which is then sent to the device. The output is the generated sequel music and lyrics data.
[1050] (Application example 1)
[1051] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1052] The challenge is to provide a system that allows users to easily create original music, lyrics, and choreography, synthesize audio and video recordings of performances to match them, and then efficiently post the content to a social media platform. Another challenge is to generate follow-up music and performance content based on the reaction on the social media platform, thereby continuously supporting users' creative activities.
[1053] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1054] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for posting the generated content to a social networking service (SNS) platform, and means for checking the reaction on the SNS platform and requesting the generation of a sequel music and lyrics. This allows users to efficiently and consistently perform processes from generating original content to sharing it on SNS and further creating sequels.
[1055] definition statement
[1056] "Means for accepting user input" refers to the interface that the user operates when requesting the creation of new music, lyrics, and choreography.
[1057] "Means for generating music, lyrics, and choreography using generative models" refers to a system that uses AI models and algorithms to automatically generate original music, lyrics, and choreography based on input parameters.
[1058] "Means for providing the generated music, lyrics and choreography to the user terminal" refers to the process of transmitting the generated content from the server to the user terminal.
[1059] "Means for recording a user's performance" refers to a recording function for recording a user's performance in sync with the generated music.
[1060] "Means of combining recorded performance data with generated music" refers to the process of integrating recorded performance data with generated music data and editing it into a complete performance content.
[1061] The term "means for providing the synthesized performance content to the user terminal" refers to a process for transmitting the synthesized performance content to the user terminal and making it playable.
[1062] "Means for posting Generated Content to a Social Media Platform" means an interface for uploading Generated Performance Content to a Social Media Platform and for users to share it.
[1063] "Means for checking reactions on social media platforms and making requests to generate sequel music and lyrics" refers to a system that analyzes user reactions on social media platforms and requests the generation of sequel content based on the results.
[1064] MODE FOR CARRYING OUT THE INVENTION
[1065] This invention is a system that generates original music, lyrics, and choreography based on prompts entered by the user, integrates performance content recorded or filmed by the user with the generated content, and shares it on a social media platform.
[1066] The specific steps for implementing this system are as follows:
[1067] First, a user accesses the application using a device (smartphone or smart glasses). The user taps the "Create a new song" button to proceed to a prompt input screen. Here, the user enters a specific prompt, such as "Create an energetic pop song."
[1068] The server then receives prompts from the user, uses its internal generative AI model to automatically generate music, lyrics, and choreography based on the prompts, and sends the generated content to the device.
[1069] The device displays the generated music, lyrics, and choreography data on the application screen, allowing the user to play it. The user records and films their performance while singing and dancing to the generated music. The recorded and filmed performance data is temporarily stored on the device.
[1070] Once the user has completed recording the performance, the device uploads the performance data to the server. The server then combines the received performance data with the generated music data using video editing software. The combined performance content is then sent back to the device from the server.
[1071] The device that receives the composite content makes it playable for the user. If the user checks the generated performance content and indicates their intention to post it to a social media platform, the device calls the social media API to upload the content. The user can then add information such as a caption and finally post it to the social media platform.
[1072] The server analyzes the reactions to the content posted by users on the SNS platform, and based on the results, users can request the generation of a sequel to the song and lyrics, thereby providing continuous support for the user's creative activities.
[1073] For example, consider the case where a user enters the following prompt:
[1074] Prompt: "Generate an energetic pop song"
[1075] In this way, this invention is a system that enables users to efficiently and consistently create original content, share it on social media, and even create sequels.
[1076] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1077] Program processing flow
[1078] Step 1:
[1079] A user accesses the application using a device, taps the "Create a new song" button, and enters a prompt, which is then sent to the server.
[1080] Input: User prompt (e.g., "Generate an energetic pop song")
[1081] Output: Prompt data is sent to the server
[1082] Step 2:
[1083] The server receives the prompts and generates the music, lyrics, and choreography based on an internal generative AI model.
[1084] Input: prompt data
[1085] Data processing: Based on the input prompt, the generative AI model generates the music, lyrics, and choreography.
[1086] Output: Generated music data, lyrics data, choreography data
[1087] Step 3:
[1088] The server transmits the generated data to the terminal, which receives and displays it.
[1089] Input: Generated music data, lyrics data, choreography data
[1090] Output: Music, lyrics, and choreography displayed on the user's device
[1091] Step 4:
[1092] The user performs along with the generated music, and the device records and films this.
[1093] Input: User performance (singing, dancing)
[1094] Data processing: Activate the recording function and capture performance data
[1095] Output: Recorded and recorded performance data is temporarily saved.
[1096] Step 5:
[1097] The device uploads the recorded performance data to the server.
[1098] Input: Audio-visual performance data
[1099] Output: Performance data is sent to the server
[1100] Step 6:
[1101] The server synthesizes the received performance data with the generated music data.
[1102] Input: Performance data, generated song data
[1103] Data processing: Composition using video editing software
[1104] Output: Synthesized performance content
[1105] Step 7:
[1106] The server transmits the synthesized performance content to the terminal, which receives it and makes it playable.
[1107] Input: Synthesized performance content
[1108] Output: The composite content displayed on the device
[1109] Step 8:
[1110] The user indicates their intention to post the composite content to a social networking platform, and the device calls the social networking API to upload the content.
[1111] Input: User's posting intent, synthetic content
[1112] Data processing: Upload content using SNS API
[1113] Output: Content posted on social media platforms
[1114] Step 9:
[1115] The server analyzes the reaction on the social media platform and begins the process of generating the sequel's music and lyrics based on user requests.
[1116] Input: Reaction data from social media platforms
[1117] Data processing: A generative AI model based on feedback data generates the music and lyrics for the sequel
[1118] Output: Sequel music data, lyrics data
[1119] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1120] The present invention provides a system that allows users to easily create and share original music and performance content. The system also incorporates an emotion engine that can recognize a user's emotions and adjust the generated content accordingly. The following describes in detail an embodiment of the system.
[1121] Music, lyrics and choreography generation
[1122] User: First launches the application and taps the "Generate New Song" button, which sends a request from the device to the server.
[1123] Server: Upon receiving a generation request, it launches the emotion engine and collects data to recognize the user's emotion.
[1124] Emotion engine: Analyzes the user's facial expressions, tone of voice, and other emotional data from the camera and microphone to identify the emotion the user is currently feeling. The emotion engine then sends the recognized emotional information back to the server.
[1125] Server: Requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on the input parameters and emotional information.
[1126] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[1127] Server: Receives the generated data and sends it to the terminal.
[1128] Device: The received music, lyrics, and choreography data is displayed on the app screen and played, allowing the user to listen to the music, lyrics, and choreography.
[1129] Recording and recording user performance
[1130] User: Sing and dance freely along with the generated music.
[1131] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[1132] Combining performance and generative music
[1133] Device: Once recording is complete, performance data is uploaded to the server.
[1134] Server: Receives the performance data and synthesizes it with the generated music data. The synthesis process is carried out using appropriate video editing software or algorithms, and finally a single integrated performance content is generated.
[1135] Server: Sends the synthesized performance content to the terminal.
[1136] Terminal: Plays and stores the received composite content so that the user can view it.
[1137] Posting to social media
[1138] User: Indicate intention to post synthesized performance content to social media.
[1139] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[1140] Sequel Music - Commercial Use
[1141] Users: Check the reaction on social media and request sequel songs and commercialization.
[1142] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[1143] Generative model: Generates new music and lyrics and sends them back to the server.
[1144] Server: Provides the generated sequel data to the device so that the user can check it.
[1145] This system allows users to easily create performance videos based on original music and share them on social media. Furthermore, by generating content based on the user's emotional information, it is possible to provide a more personalized experience.
[1146] The processing flow will be explained below.
[1147] Step 1:
[1148] The user launches the application and taps the "Create a new song" button, which sends a song creation request from the device to the server.
[1149] Step 2:
[1150] The server receives the generation request and starts the emotion engine, which prepares to start collecting data necessary to recognize the user's emotion.
[1151] Step 3:
[1152] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and this data is sent to the emotion engine in real time.
[1153] Step 4:
[1154] The emotion engine analyzes the received data and identifies the user's emotion. For example, if the user is smiling, it recognizes "happiness," and if they are frowning, it recognizes "anxiety."
[1155] Step 5:
[1156] The emotion engine sends the identified emotion information back to the server, which collects this information and prepares it for transmission to the generative model.
[1157] Step 6:
[1158] The server requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model takes the emotional information into consideration and generates music, lyrics, and choreography that match the emotional information.
[1159] Step 7:
[1160] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[1161] Step 8:
[1162] The server receives the generated data and transmits it to the user terminal.
[1163] Step 9:
[1164] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[1165] Step 10:
[1166] The user sings and dances freely to the generated music, and the device records and films this performance.
[1167] Step 11:
[1168] Once the recording is complete, the device uploads this performance data to a server.
[1169] Step 12:
[1170] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[1171] Step 13:
[1172] The server generates the synthesized performance content and transmits it to the terminal.
[1173] Step 14:
[1174] The device plays and saves the received composite content, which the user can then view and enjoy.
[1175] Step 15:
[1176] The user selects SNS posting and wishes to post composite content.
[1177] Step 16:
[1178] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[1179] Step 17:
[1180] Users check the reaction on social media and request sequel songs and commercialization.
[1181] Step 18:
[1182] The server sends a request to generate the sequel's music and lyrics to the generative model.
[1183] Step 19:
[1184] The generative model generates new music and lyrics and sends them back to the server.
[1185] Step 20:
[1186] The server provides the generated sequel data to the user terminal, so that the user can check it.
[1187] Example 2
[1188] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1189] In today's entertainment environment, there is a demand for a means for users to easily create and share original music and performance content. However, existing systems have difficulty generating content based on users' emotions, and there are insufficient methods for efficiently synthesizing generated content and sharing it on social media. Furthermore, there is a lack of easy ways to request and realize sequels or commercialization. Therefore, a system that makes it easier for users to create and share original content that responds to individual emotions is needed.
[1190] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1191] In this invention, the server includes means for recognizing the user's emotional information and generating music, lyrics, and choreography, means for providing the synthesized performance content to the user terminal, and means for posting the content to an SNS platform, thereby enabling users to generate original content according to their individual emotions, efficiently synthesize the content, and share it on the SNS.
[1192] The "means for accepting user input" is an interface that allows the user to input operational instructions to the system.
[1193] "Generative model" refers to algorithms or software that use artificial intelligence to automatically generate content such as music, lyrics, and choreography.
[1194] "Means for recognizing user's emotional information" refers to an engine or software that analyzes the user's facial expressions and tone of voice via a camera or microphone to identify their current emotional state.
[1195] "Means for generating music, lyrics and choreography" refers to a device or software that generates original music, lyrics and choreography based on emotional information and input data.
[1196] "Means for providing the generated music, lyrics and choreography to the user's terminal" refers to a device or program that has the function of transferring the generated data from the server to the user's terminal and displaying and playing it.
[1197] "Means for recording a user's performance" refers to a device or function that uses a camera and microphone to record a user's performance along with a piece of music.
[1198] "Means for synthesizing recorded performance data with generated music" refers to software or a system for integrating a user's performance data with generated music and editing and generating a single composite content.
[1199] "Means for providing synthesized performance content to a user terminal" refers to a device or program for transmitting the completed synthesized content from the server to a user terminal so that it can be displayed and played.
[1200] "Means for posting to a social media platform" refers to a device or software that has the function of uploading and sharing generated content to a social media platform using an API or other means.
[1201] "Means for generating sequel music and lyrics" refers to artificial intelligence algorithms or models for generating new music and lyrics based on existing content.
[1202] The present invention relates to a system that enables users to easily create and share original music and performance content. Specific embodiments of this system will be described below.
[1203] Music, lyrics and choreography generation
[1204] First, the user launches the application and taps the "Generate a new song" button. This action sends a generation request from the device to the server. When the server receives the generation request, it launches an emotion engine and collects data to recognize the user's emotions. The emotion engine analyzes the user's facial expressions and vocal tone through the camera and microphone to identify the user's current emotion. This emotion information is sent back to the server. The server, along with the received emotion information, requests the generative model to generate music, lyrics, and choreography. The generative model generates original music, lyrics, and choreography based on the input prompt and emotion information and sends it back to the server. The server receives the generated data and sends it to the user's device. The device displays and plays this data on the app screen.
[1205] Hardware and software used
[1206] Hardware: Smartphone (camera, microphone)
[1207] Software: Emotion engine (facial expression and voice analysis software), generative AI model (music, lyrics, choreography generation)
[1208] Specific examples
[1209] For example, when a user launches the app and taps the "Generate a new song" button, the smartphone's camera and microphone are activated, and the user's facial expressions and voice are analyzed. If the user is in a happy mood, the generative model will generate an upbeat, cheerful song.
[1210] Example prompt for a generative AI model:
[1211] "The user has a smiling face and a high-pitched voice, so generate a 30-second bright and cheerful pop song."
[1212] Recording and recording user performance
[1213] The user can freely sing and dance along to the generated music. At this time, the user's device will record and record the music. The recording function will be activated and the captured data will be temporarily saved on the device.
[1214] Combining performance and generative music
[1215] Once the recording is complete, the device uploads the captured performance data to the server. The server receives the performance data and combines it with the generated music data. This combination process is carried out using appropriate video editing software or algorithms to ultimately generate a single, integrated performance content. The server then transmits the combined performance content to the device. The device then plays and stores the received content, making it available for viewing by the user.
[1216] Posting to social media
[1217] If the user indicates their intention to post the synthesized performance content to a social networking site, the device calls the social networking site API to upload the synthesized content to the social networking site platform. Through the social networking site API, the user can add information such as a caption and post the content.
[1218] Sequel Music - Commercial Use
[1219] If a user checks the reaction on social media and requests a sequel song or commercialization, the server sends a sequel song and lyric generation request to the generative model. The generative model generates new music and lyrics and sends them back to the server. The server then provides the generated sequel data to the user's device, allowing the user to check the newly generated music and lyrics.
[1220] This allows users to create original music and performance content that reflects their individual emotions and share it efficiently on social media. It also supports the continuous creation of content by responding to requests for sequels and commercialization.
[1221] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1222] Step 1:
[1223] The user launches the application and taps the "Create New Song" button.
[1224] Specifically, the user operates the app screen on the mobile device and performs an input operation by tapping the relevant button. This operation registers the generation request on the device, and the registered information is used in the next step.
[1225] Step 2:
[1226] The device sends a request to the server.
[1227] The terminal converts the generated request into packets and sends them to the server over the Internet. The input is the user's generated request, and the output is the generated HTTP request sent to the server.
[1228] Step 3:
[1229] The server activates the emotion engine and collects data to recognize the user's emotions.
[1230] When the server receives the request, it sends a command to activate the camera and microphone to the device. At this time, the collected data is sent to the emotion engine. The input is the HTTP request data, and the output is the command to start the emotion recognition engine.
[1231] Step 4:
[1232] The emotion engine uses the device's camera and microphone to collect and analyze the user's emotional data.
[1233] Specifically, the camera captures the user's facial expressions and the microphone records their voice. These data are analyzed in real time by the emotion engine, and the user's emotional information is extracted as numerical data. The input is raw data from the camera and microphone, and the output is analyzed emotional information.
[1234] Step 5:
[1235] The server inputs emotional information into the generative model and requests it to generate music, lyrics, and choreography.
[1236] The server sends the received emotional information along with a prompt to the generative AI model. The input is the analyzed emotional information and a generation request, and the output is instruction data including the prompt to the generative model.
[1237] Step 6:
[1238] The generative model generates the music, lyrics, and choreography and sends it back to the server.
[1239] The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on prompts and emotional information. The input is the prompts and emotional data, and the output is the generated music, lyrics, and choreography data.
[1240] Step 7:
[1241] The server transmits the generated data to the terminal.
[1242] The server converts the generated data into packets and sends them to the user terminal via the Internet. The input is the generated data from the generative model, and the output is the packets sent to the user terminal.
[1243] Step 8:
[1244] The data received by the device is displayed on the app screen and played back.
[1245] The terminal analyzes the generated data and displays it as a playback component of the music and choreography. The input is the generated data from the server, and the output is the content provided visually and audibly to the user.
[1246] Step 9:
[1247] Users can freely sing and dance along to the generated music.
[1248] The user presses the play button on the device to play the music and performs along with it. This operation is input into the next recording step.
[1249] Step 10:
[1250] The device will record and temporarily store your performance.
[1251] The device's camera and microphone are again used to record the user's singing and dancing. The input is the user's performance, and the output is the temporarily stored recording.
[1252] Step 11:
[1253] The device uploads the recording data to the server.
[1254] The terminal converts the temporarily stored performance data into packets and sends them to a server via the Internet. The input is the recorded data, and the output is the packets sent to the server.
[1255] Step 12:
[1256] The server receives the performance data and synthesizes it with the generated music data.
[1257] The server launches appropriate video editing software, inputs the audio and video data and the generated music data into the program, and synthesizes them. The input is the performance data and the generated music data, and the output is the synthesized performance content.
[1258] Step 13:
[1259] The server transmits the synthesized performance content to the terminal.
[1260] The server converts the composite content back into packets and transmits them to the user terminal. The input is the composite content, and the output is the packets sent to the user terminal.
[1261] Step 14:
[1262] The terminal plays and stores the received composite content so that the user can view it.
[1263] The terminal analyzes the composite content, loads it into the playback player, and provides it to the user. The input is the composite content from the server, and the output is the played and saved content.
[1264] Step 15:
[1265] The user indicates their intention to post the synthesized performance content to a social networking site.
[1266] The user taps the "Share to SNS" button in the app to indicate their intention to post. This action becomes the input for the next step.
[1267] Step 16:
[1268] The device calls the SNS API and uploads the composite content to the SNS platform.
[1269] The device uses the SNS API to send the composite content to the SNS platform, adding necessary information such as captions and uploading it. The input is the user's intention to post and the composite content, and the output is a notification of completion of posting to the SNS platform.
[1270] (Application example 2)
[1271] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1272] Conventional music generation systems have limited the content that users can experience, making it particularly difficult to generate personalized content based on emotions. Furthermore, there is a lack of easy ways to share generated content on social media platforms. This has prevented users from fully enjoying the fun of creating unique music and performances that respond to individual emotions.
[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting user input, means for analyzing the user's emotions using an emotion engine, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for synthesizing the recorded performance data with the generated music, means for providing the synthesized performance content to a user terminal, and means for generating a prompt sentence for the generative AI model based on emotion information. This makes it possible to generate personalized music and performance content based on the user's emotions and easily share the generated content on a social media platform.
[1274] "User" refers to a person who uses this system to create music or performance content.
[1275] "Means for accepting input" refers to an interface that allows a user to input a request for music composition to the system.
[1276] An "emotion engine" refers to algorithms and software that analyze emotions from a user's facial expressions and voice.
[1277] "Means for analyzing emotions" refers to functionality for identifying a user's current emotional state using an emotion engine.
[1278] "Generative model" refers to an AI model that automatically generates music, lyrics, and choreography based on specified parameters and emotional information.
[1279] "Means for generating music, lyrics, and choreography" refers to a function for generating music, lyrics, and choreography using a generative model.
[1280] "User terminal" refers to a device (smartphone, tablet, etc.) that a user uses to view the music, lyrics, and choreography they have created.
[1281] "Means for recording performance" refers to the audio and video recording functionality used to record a User's performance.
[1282] "Recorded Performance Data" means audio and video data captured of a User's performance.
[1283] "Means of synthesis" refers to the function of integrating recorded performance data with the generated music to create a single piece of content.
[1284] "Synthesized Performance Content" refers to content generated as a result of integrating recorded or filmed performance data with generated music.
[1285] "SNS Platform" refers to a social networking service used to share and post User-Generated Content.
[1286] "Means for generating prompt sentences for a generative AI model" refers to a function that creates instruction sentences to be given to a generative AI model based on emotional information.
[1287] This invention relates to a system that generates music, lyrics, and choreography based on a user's emotions, allowing the user to create and share performance content. Specific embodiments of the system are described in detail below.
[1288] Hardware and Software Used
[1289] Hardware: smartphones, tablets, servers, cameras, microphones
[1290] Software: Smartphone applications, cloud server services (e.g., Amazon Web Services, Google Cloud), emotion engines (Microsoft Azure Face API, Google Cloud Speech-to-Text API), generative AI models (OpenAI GPT, Music VAE), social media APIs (Facebook API, Twitter API)
[1291] Program processing overview
[1292] 1. Launch the application:
[1293] The user launches the smartphone app and taps the "Generate new song" button.
[1294] The app activates the camera and microphone to capture the user's facial expressions and voice for emotion analysis.
[1295] 2. Emotion analysis:
[1296] The server uses an emotion engine to analyze the user's emotions from the captured data.
[1297] The emotion engine identifies the user's emotions from facial expressions and voice and sends that information back to the server.
[1298] 3. Music, lyrics, and choreography generation:
[1299] Based on the emotional information, the server sends prompts to the generative AI model to generate music, lyrics, and choreography.
[1300] The generative AI model generates original content according to the specified content and sends it back to the server.
[1301] 4. Receiving and Viewing Generated Content:
[1302] The server transmits the generated music, lyrics, and choreography data to the user terminal.
[1303] The user device plays and displays the received data within the app.
[1304] 5. Recording and filming of performances:
[1305] The user performs along with the generated music, and the device records and films the performance.
[1306] The recorded data will be temporarily stored and uploaded to a server.
[1307] 6. Video Composition and Distribution:
[1308] The server combines the audio and video data with the generated music to generate a single performance content.
[1309] The synthesized performance content is transmitted again to the user terminal.
[1310] 7. Posting to social media:
[1311] A user selects a social media post for the composited content.
[1312] Your app uses social media APIs to post content to social media platforms.
[1313] Specific examples
[1314] For example, if a user uses this app when they are tired, the app will analyze the tired expression on the user's face and generate relaxing music and slow choreography. The user can then perform along with the music, record the performance, and post it on Instagram.
[1315] Prompt Sentence Examples
[1316] Music generation prompt:
[1317] "The user's emotion is fatigue. Please generate a 30-second relaxing piece of music. The genre should be ambient and the tempo should be slow."
[1318] Lyric generation prompt:
[1319] "The user's emotion is fatigue. Generate 30 seconds of relaxing lyrics. The theme is nature and healing."
[1320] The present invention enables personalized content generation based on a user's emotions and makes it easy to share the generated content on a social networking platform.
[1321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1322] Step 1:
[1323] The user launches the smartphone app and taps the "Generate New Song" button, which activates the app's camera and microphone to capture the user's facial expressions and voice for emotional analysis.
[1324] Input: Tap of the "Generate New Song" button, user facial and voice data
[1325] Output: User's facial expression and voice data
[1326] Step 2:
[1327] The device sends the captured facial and voice data to the server, which then activates the emotion engine to analyze the received data.
[1328] Input: User facial and voice data
[1329] Output: A signal that triggers the emotion engine to start analysis.
[1330] Step 3:
[1331] The server uses an emotion engine to analyze the user's emotions from facial expressions and voice data. The emotion engine analyzes data such as facial expressions and voice tone and sends the results back to the server.
[1332] Input: User facial and voice data received by the emotion engine
[1333] Output: User's emotional information
[1334] Step 4:
[1335] The server generates prompts for the generative AI model based on the user's emotional information. The generated prompts include requests to generate music, lyrics, and choreography.
[1336] Input: User's emotional information
[1337] Output: Prompt sentence for the generative AI model
[1338] Step 5:
[1339] The server sends the prompt text to the generative AI model, which generates the music, lyrics, and choreography. The generative AI model generates content based on the specified prompt and sends it back to the server.
[1340] Input: Prompt for the generative AI model
[1341] Output: Generated music, lyrics, and choreography data
[1342] Step 6:
[1343] The server then sends the generated music, lyrics, and choreography data to the user's device, which then plays and displays the received data within the app.
[1344] Input: Generated music, lyrics, and choreography data
[1345] Output: Playback and display on the user's device
[1346] Step 7:
[1347] The user performs along with the generated music, and the device records and films the performance. The recorded data is temporarily stored on the device.
[1348] Input: Performance along with the music
[1349] Output: Audio and video data
[1350] Step 8:
[1351] The device uploads the recorded data to the server, which then combines the received data with the generated music.
[1352] Input: Audio and video data
[1353] Output: Synthetic content
[1354] Step 9:
[1355] The server transmits the composite content to the user terminal, which then plays and displays the received composite content.
[1356] Input: Synthetic content
[1357] Output: Playback and display on the user's device
[1358] Step 10:
[1359] The user selects a social media post for the composited content, and the app uses the social media API to post the content to the social media platform.
[1360] Input: User's social media posts selection
[1361] Output: Posting content to social media platforms
[1362] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1363] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1364] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1365] [Fourth embodiment]
[1366] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1367] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1368] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1369] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1370] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1371] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1372] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1373] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1374] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1375] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1376] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1377] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1378] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1379] The present invention provides a system that allows users to easily create and share original music and performance content. The following describes in detail an embodiment of the system.
[1380] Music, lyrics and choreography generation
[1381] User: First launches the application and taps the "Create a new song" button, which sends a creation request from the device to the server.
[1382] Server: Upon receiving a generation request, the server requests the internally implemented generative model to generate music, lyrics, and choreography. The generative model generates an original music piece, lyrics, and choreography of approximately 30 seconds based on the input parameters.
[1383] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[1384] Server: Receives the generated data and sends it to the terminal.
[1385] Device: The received music, lyrics, and choreography data are displayed on the app screen and played.
[1386] Recording and recording user performance
[1387] User: Sing and dance freely along with the generated music.
[1388] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[1389] Combining performance and generative music
[1390] Device: Once recording is complete, performance data is uploaded to the server.
[1391] Server: Receives the performance data and combines it with the generated music data. This process is carried out using video editing software and algorithms, ultimately generating a single integrated performance content.
[1392] Server: Sends the synthesized performance content to the terminal.
[1393] Terminal: Plays and stores the received composite content so that the user can view it.
[1394] Posting to social media
[1395] User: Indicate intention to post synthesized performance content to social media.
[1396] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[1397] Sequel Music - Commercial Use
[1398] Users: Check the reaction on social media and request sequel songs and commercialization.
[1399] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[1400] Generative model: Generates new music and lyrics and sends them back to the server.
[1401] Server: Provides the generated sequel data to the device so that the user can check it.
[1402] This system allows users to easily create performance videos based on original music and share them on social media. It also encourages further creative activities based on user feedback, providing a new form of entertainment.
[1403] The processing flow will be explained below.
[1404] Step 1:
[1405] The user launches the application and taps the "Create a new song" button, which sends a request from the device to the server.
[1406] Step 2:
[1407] The server receives the generation request and asks the generative model to generate the music, lyrics, and choreography.
[1408] Step 3:
[1409] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[1410] Step 4:
[1411] The server receives the generated data and transmits it to the user terminal.
[1412] Step 5:
[1413] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[1414] Step 6:
[1415] Users sing and dance along to the music, and the device records and films this performance.
[1416] Step 7:
[1417] After the recording is complete, the device uploads this performance data to the server.
[1418] Step 8:
[1419] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[1420] Step 9:
[1421] The server generates the synthesized performance content and transmits it to the terminal.
[1422] Step 10:
[1423] The device receives, plays, and stores the composite content, which the user can then view and enjoy.
[1424] Step 11:
[1425] The user selects SNS posting and wishes to post composite content.
[1426] Step 12:
[1427] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[1428] Step 13:
[1429] Users check the reaction on social media and request sequel songs and commercialization.
[1430] Step 14:
[1431] The server sends a request to generate the sequel's music and lyrics to the generative model.
[1432] Step 15:
[1433] The generative model generates new music and lyrics and sends them back to the server.
[1434] Step 16:
[1435] The server provides the generated sequel data to the user terminal, so that the user can check it.
[1436] Example 1
[1437] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1438] Conventional systems for generating and sharing original music and performance content have had the problem that it is time-consuming for users to easily create music and choreography, record and film a performance that matches the music, and then share it on social media. Furthermore, automation of the generation of follow-up content based on reactions on social media is also insufficient. These issues need to be resolved.
[1439] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1440] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative AI model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for allowing the user to post the combined content to a social networking service (SNS) platform, and means for generating a sequel to the music and lyrics based on the response on the SNS platform. This allows users to easily create performance videos based on original music and share them on SNS, and further enables sequels to be automatically generated based on the response.
[1441] "Means for accepting user input" refers to a mechanism for inputting information or requests from a user, thereby enabling the user to send commands to the system.
[1442] A "generative AI model" refers to an artificial intelligence system that uses machine learning algorithms to automatically generate creative content such as music, lyrics, and choreography according to specific parameters.
[1443] "Means for generating music, lyrics, and choreography" refers to a mechanism that uses a generative AI model to automatically create music, lyrics, and choreography based on specified parameters.
[1444] "Means for providing to user terminal" refers to a mechanism for transmitting the generated data or content to the terminal used by the user, thereby enabling the user to view and play the generated content.
[1445] "Means for recording and recording a user's performance" refers to a mechanism for recording a user's performance in sync with the generated music, and involves capturing video and audio using the device's camera and microphone.
[1446] "Means for synthesizing performance data with generated music" refers to a mechanism for integrating recorded or filmed performance data of users with generated music to create a single piece of content.
[1447] "Means for providing synthesized performance content to a user terminal" refers to a mechanism for transmitting synthesized performance content to a user terminal and allowing the user to view and save it.
[1448] "Means for posting to an SNS platform" refers to a mechanism for uploading and publishing content created by a user to a social networking service (SNS).
[1449] "Means for generating sequel music and lyrics based on reactions on social media platforms" refers to a mechanism for analyzing reactions and evaluations on social media and generating new music and lyrics accordingly.
[1450] This invention provides a system that allows users to easily create and share original music and performance content. This system uses a generative AI model to generate music, lyrics, and choreography based on user input, and provides them to the user's device. It can also record and film the user's performance, combine it with the generated music, and post it on a social media platform. It can also generate sequels and lyrics based on social media reactions.
[1451] The system is implemented using the following hardware and software:
[1452] 1. A means of accepting user input
[1453] A user starts the application on a device (e.g., a smartphone or tablet) and taps the "Create a new song" button. This action sends a creation request from the device to the server.
[1454] 2. A means of generating music, lyrics, and choreography using generative AI models
[1455] The server receives the generation request and asks a generative AI model (e.g., a model using a machine learning algorithm) to generate the music, lyrics, and choreography. Specific models used here include OpenAI's GPT-4 and DeepComposer.
[1456] For example, a create request might use the following prompt:
[1457] Song generation prompt: The melody line is bright, the tempo is 120 BPM, and the theme is "Summer Memories."
[1458] 3. Means for providing the generated music, lyrics, and choreography to the user terminal
[1459] The generative model automatically generates music, lyrics, and choreography and sends them back to the server. The server receives the generated data and sends it to the device. The device displays the received data on the app screen and plays it back.
[1460] 4. Means for recording and recording your performance
[1461] The user sings and dances along to the generated music, and the device records and films the performance using the built-in microphone and camera. This data is temporarily stored in the device's storage.
[1462] 5. A means of synthesizing recorded performance data with generated music
[1463] Once the device has finished recording, it uploads the performance data to the server. The server receives the performance data and combines it with the generated music data. This process is carried out using video editing software (e.g., FFmpeg or Adobe Premiere Pro) and algorithms, resulting in a single integrated performance content.
[1464] 6. Means for providing synthesized performance content to a user terminal
[1465] The server sends the synthesized performance content to the terminal, which then plays and stores the received synthesized content for the user to view.
[1466] 7. Means for enabling users to post synthetic content to social media platforms
[1467] The user indicates their intention to post the synthesized performance content to a social networking site. The device then calls a social networking site API (e.g., Twitter API or Instagram Graph API) to upload the synthesized content to the social networking site, add information such as a caption, and post it.
[1468] 8. A means of generating sequel music and lyrics based on reactions on social media platforms
[1469] Users check the reaction on social media and request a sequel song and commercialization. The server sends the sequel song and lyric generation request to the generative AI model. The generative model generates new music and lyrics and sends them back to the server. The server provides the generated sequel data to the device so that the user can view it.
[1470] This system allows users to easily create performance videos based on original music and share them on social media, and can encourage further creative activities based on user feedback, providing a new form of entertainment.
[1471] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1472] Step 1:
[1473] A user launches the application and taps the "Create a new song" button. The input is the user's tap, which causes a creation request to be sent from the device to the server. Specifically, the creation request is sent as an HTTP POST request. The data included in this request is the song creation prompt and other parameters (e.g., tempo, theme).
[1474] Step 2:
[1475] The server receives the generation request and parses the request contents. The input is JSON data included in the HTTP request body. The server parses this data and converts it into an input format appropriate for the generative AI model. Specifically, it extracts the music generation prompt and parameters from the request content and converts them into a format to be passed to the generative AI model. The output is the request data sent to the generative AI model.
[1476] Step 3:
[1477] The generative AI model generates music, lyrics, and choreography based on the request data it receives. The input is the prompt and parameters sent from the server. The generative AI model applies algorithms based on these inputs to generate original music data, lyrics data, and choreography data. The output is the generated music, lyrics, and choreography data, which are sent back to the server.
[1478] Step 4:
[1479] The server receives the data returned from the generative AI model and sends it to the device. The input is the music, lyrics, and choreography data returned from the generative AI model. The server formats this data into JSON format and sends it to the device as an HTTP response. The output is the music, lyrics, and choreography data sent to the device.
[1480] Step 5:
[1481] The device displays the received data on the app screen and plays it. The input is music, lyrics, and choreography data sent from the server. The device analyzes this data, decodes and plays the music data, and displays the lyrics and choreography information on the app screen. Specifically, it plays the music using a music playback library and displays the lyrics and choreography information in a UI component. The output is audiovisual content provided to the user.
[1482] Step 6:
[1483] The user sings and dances along to the generated music. The system records and films the user's performance. The input is the user's performance, which is recorded using the device's microphone and camera. The recorded data is temporarily stored in the smartphone's storage. The output is the recorded performance data.
[1484] Step 7:
[1485] Once the device has completed recording, it uploads the performance data to the server. The input is the recorded performance data. The device sends this data to the server as an HTTP PUT request. The output is the performance data uploaded to the server.
[1486] Step 8:
[1487] The server receives the performance data and combines it with the generated music data. The input is the uploaded performance data and the generated music data. The server uses video editing software such as FFmpeg to combine these data into a single integrated performance content. Specifically, the server uses FFmpeg commands to combine the video and audio. The output is the combined performance content.
[1488] Step 9:
[1489] The server sends the synthesized performance content to the terminal. The input is the synthesized performance content. The server sends this content to the terminal as an HTTP response. The output is the synthesized performance content sent to the terminal.
[1490] Step 10:
[1491] The device plays and saves the received composite content, making it available for viewing by the user. The input is composite performance content sent from the server. The device plays this content and saves it in storage. Specifically, it plays it using a video playback library and saves it using a file system API. The output is playable composite content provided to the user.
[1492] Step 11:
[1493] The user indicates their intention to post the synthesized performance content to a social networking site. The input is the user's posting operation, and the device uses the social networking site API to upload the synthesized content to the social networking site platform. Specifically, the device sends an HTTP POST request to the social networking site platform's API to upload information such as captions along with the synthesized content. The output is the content posted to the social networking site platform.
[1494] Step 12:
[1495] Users check the reaction on social media and request a sequel song and commercialization. The input is the user's sequel request operation, which is sent to the server. The server sends a sequel song and lyric generation request to the generative AI model. The output is the request data sent to the generative AI model.
[1496] Step 13:
[1497] The generative AI model generates new music and lyrics and sends them back to the server. The input is the sequel request data, and the generative AI model generates new music and lyrics based on this. The server provides the generated data to the device so that the user can view it. Specifically, the generated data is returned to the server in JSON format, which is then sent to the device. The output is the generated sequel music and lyrics data.
[1498] (Application example 1)
[1499] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1500] The challenge is to provide a system that allows users to easily create original music, lyrics, and choreography, synthesize audio and video recordings of performances to match them, and then efficiently post the content to a social media platform. Another challenge is to generate follow-up music and performance content based on the reaction on the social media platform, thereby continuously supporting users' creative activities.
[1501] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1502] In this invention, the server includes means for accepting user input, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for combining the recorded performance data with the generated music, means for providing the combined performance content to a user terminal, means for posting the generated content to a social networking service (SNS) platform, and means for checking the reaction on the SNS platform and requesting the generation of a sequel music and lyrics. This allows users to efficiently and consistently perform processes from generating original content to sharing it on SNS and further creating sequels.
[1503] definition statement
[1504] "Means for accepting user input" refers to the interface that the user operates when requesting the creation of new music, lyrics, and choreography.
[1505] "Means for generating music, lyrics, and choreography using generative models" refers to a system that uses AI models and algorithms to automatically generate original music, lyrics, and choreography based on input parameters.
[1506] "Means for providing the generated music, lyrics and choreography to the user terminal" refers to the process of transmitting the generated content from the server to the user terminal.
[1507] "Means for recording a user's performance" refers to a recording function for recording a user's performance in sync with the generated music.
[1508] "Means of combining recorded performance data with generated music" refers to the process of integrating recorded performance data with generated music data and editing it into a complete performance content.
[1509] The term "means for providing the synthesized performance content to the user terminal" refers to a process for transmitting the synthesized performance content to the user terminal and making it playable.
[1510] "Means for posting Generated Content to a Social Media Platform" means an interface for uploading Generated Performance Content to a Social Media Platform and for users to share it.
[1511] "Means for checking reactions on social media platforms and making requests to generate sequel music and lyrics" refers to a system that analyzes user reactions on social media platforms and requests the generation of sequel content based on the results.
[1512] MODE FOR CARRYING OUT THE INVENTION
[1513] This invention is a system that generates original music, lyrics, and choreography based on prompts entered by the user, integrates performance content recorded or filmed by the user with the generated content, and shares it on a social media platform.
[1514] The specific steps for implementing this system are as follows:
[1515] First, a user accesses the application using a device (smartphone or smart glasses). The user taps the "Create a new song" button to proceed to a prompt input screen. Here, the user enters a specific prompt, such as "Create an energetic pop song."
[1516] The server then receives prompts from the user, uses its internal generative AI model to automatically generate music, lyrics, and choreography based on the prompts, and sends the generated content to the device.
[1517] The device displays the generated music, lyrics, and choreography data on the application screen, allowing the user to play it. The user records and films their performance while singing and dancing to the generated music. The recorded and filmed performance data is temporarily stored on the device.
[1518] Once the user has completed recording the performance, the device uploads the performance data to the server. The server then combines the received performance data with the generated music data using video editing software. The combined performance content is then sent back to the device from the server.
[1519] The device that receives the composite content makes it playable for the user. If the user checks the generated performance content and indicates their intention to post it to a social media platform, the device calls the social media API to upload the content. The user can then add information such as a caption and finally post it to the social media platform.
[1520] The server analyzes the reactions to the content posted by users on the SNS platform, and based on the results, users can request the generation of a sequel to the song and lyrics, thereby providing continuous support for the user's creative activities.
[1521] For example, consider the case where a user enters the following prompt:
[1522] Prompt: "Generate an energetic pop song"
[1523] In this way, this invention is a system that enables users to efficiently and consistently create original content, share it on social media, and even create sequels.
[1524] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1525] Program processing flow
[1526] Step 1:
[1527] A user accesses the application using a device, taps the "Create a new song" button, and enters a prompt, which is then sent to the server.
[1528] Input: User prompt (e.g., "Generate an energetic pop song")
[1529] Output: Prompt data is sent to the server
[1530] Step 2:
[1531] The server receives the prompts and generates the music, lyrics, and choreography based on an internal generative AI model.
[1532] Input: prompt data
[1533] Data processing: Based on the input prompt, the generative AI model generates the music, lyrics, and choreography.
[1534] Output: Generated music data, lyrics data, choreography data
[1535] Step 3:
[1536] The server transmits the generated data to the terminal, which receives and displays it.
[1537] Input: Generated music data, lyrics data, choreography data
[1538] Output: Music, lyrics, and choreography displayed on the user's device
[1539] Step 4:
[1540] The user performs along with the generated music, and the device records and films this.
[1541] Input: User performance (singing, dancing)
[1542] Data processing: Activate the recording function and capture performance data
[1543] Output: Recorded and recorded performance data is temporarily saved.
[1544] Step 5:
[1545] The device uploads the recorded performance data to the server.
[1546] Input: Audio-visual performance data
[1547] Output: Performance data is sent to the server
[1548] Step 6:
[1549] The server synthesizes the received performance data with the generated music data.
[1550] Input: Performance data, generated song data
[1551] Data processing: Composition using video editing software
[1552] Output: Synthesized performance content
[1553] Step 7:
[1554] The server transmits the synthesized performance content to the terminal, which receives it and makes it playable.
[1555] Input: Synthesized performance content
[1556] Output: The composite content displayed on the device
[1557] Step 8:
[1558] The user indicates their intention to post the composite content to a social networking platform, and the device calls the social networking API to upload the content.
[1559] Input: User's posting intent, synthetic content
[1560] Data processing: Upload content using SNS API
[1561] Output: Content posted on social media platforms
[1562] Step 9:
[1563] The server analyzes the reaction on the social media platform and begins the process of generating the sequel's music and lyrics based on user requests.
[1564] Input: Reaction data from social media platforms
[1565] Data processing: A generative AI model based on feedback data generates the music and lyrics for the sequel
[1566] Output: Sequel music data, lyrics data
[1567] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1568] The present invention provides a system that allows users to easily create and share original music and performance content. The system also incorporates an emotion engine that can recognize a user's emotions and adjust the generated content accordingly. The following describes in detail an embodiment of the system.
[1569] Music, lyrics and choreography generation
[1570] User: First launches the application and taps the "Generate New Song" button, which sends a request from the device to the server.
[1571] Server: Upon receiving a generation request, it launches the emotion engine and collects data to recognize the user's emotion.
[1572] Emotion engine: Analyzes the user's facial expressions, tone of voice, and other emotional data from the camera and microphone to identify the emotion the user is currently feeling. The emotion engine then sends the recognized emotional information back to the server.
[1573] Server: Requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on the input parameters and emotional information.
[1574] Generative model: Automatically generates music, lyrics, and choreography and sends it back to the server.
[1575] Server: Receives the generated data and sends it to the terminal.
[1576] Device: The received music, lyrics, and choreography data is displayed on the app screen and played, allowing the user to listen to the music, lyrics, and choreography.
[1577] Recording and recording user performance
[1578] User: Sing and dance freely along with the generated music.
[1579] Terminal: Records and records the user's performance. At this time, the recording function is activated and the captured data is temporarily saved.
[1580] Combining performance and generative music
[1581] Device: Once recording is complete, performance data is uploaded to the server.
[1582] Server: Receives the performance data and synthesizes it with the generated music data. The synthesis process is carried out using appropriate video editing software or algorithms, and finally a single integrated performance content is generated.
[1583] Server: Sends the synthesized performance content to the terminal.
[1584] Terminal: Plays and stores the received composite content so that the user can view it.
[1585] Posting to social media
[1586] User: Indicate intention to post synthesized performance content to social media.
[1587] Device: Calls the SNS API to upload the composite content to the SNS platform, adds information such as a caption, and posts it.
[1588] Sequel Music - Commercial Use
[1589] Users: Check the reaction on social media and request sequel songs and commercialization.
[1590] Server: Sends a request to generate the sequel's music and lyrics to the generative model.
[1591] Generative model: Generates new music and lyrics and sends them back to the server.
[1592] Server: Provides the generated sequel data to the device so that the user can check it.
[1593] This system allows users to easily create performance videos based on original music and share them on social media. Furthermore, by generating content based on the user's emotional information, it is possible to provide a more personalized experience.
[1594] The processing flow will be explained below.
[1595] Step 1:
[1596] The user launches the application and taps the "Create a new song" button, which sends a song creation request from the device to the server.
[1597] Step 2:
[1598] The server receives the generation request and starts the emotion engine, which prepares to start collecting data necessary to recognize the user's emotion.
[1599] Step 3:
[1600] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and this data is sent to the emotion engine in real time.
[1601] Step 4:
[1602] The emotion engine analyzes the received data and identifies the user's emotion. For example, if the user is smiling, it recognizes "happiness," and if they are frowning, it recognizes "anxiety."
[1603] Step 5:
[1604] The emotion engine sends the identified emotion information back to the server, which collects this information and prepares it for transmission to the generative model.
[1605] Step 6:
[1606] The server requests the generative model to generate music, lyrics, and choreography along with emotional information. The generative model takes the emotional information into consideration and generates music, lyrics, and choreography that match the emotional information.
[1607] Step 7:
[1608] The generative model generates music, lyrics, and choreography of the specified length and sends the generated data back to the server.
[1609] Step 8:
[1610] The server receives the generated data and transmits it to the user terminal.
[1611] Step 9:
[1612] The device displays the received data on the app screen and plays the music, allowing the user to listen to the music, lyrics, and choreography.
[1613] Step 10:
[1614] The user sings and dances freely to the generated music, and the device records and films this performance.
[1615] Step 11:
[1616] Once the recording is complete, the device uploads this performance data to a server.
[1617] Step 12:
[1618] The server receives the performance data and combines it with the generated music data. The combining process uses appropriate video editing software or algorithms.
[1619] Step 13:
[1620] The server generates the synthesized performance content and transmits it to the terminal.
[1621] Step 14:
[1622] The device plays and saves the received composite content, which the user can then view and enjoy.
[1623] Step 15:
[1624] The user selects SNS posting and wishes to post composite content.
[1625] Step 16:
[1626] The device calls the social media API and uploads the composite content to the social media platform, along with additional information such as captions.
[1627] Step 17:
[1628] Users check the reaction on social media and request sequel songs and commercialization.
[1629] Step 18:
[1630] The server sends a request to generate the sequel's music and lyrics to the generative model.
[1631] Step 19:
[1632] The generative model generates new music and lyrics and sends them back to the server.
[1633] Step 20:
[1634] The server provides the generated sequel data to the user terminal, so that the user can check it.
[1635] Example 2
[1636] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1637] In today's entertainment environment, there is a demand for a means for users to easily create and share original music and performance content. However, existing systems have difficulty generating content based on users' emotions, and there are insufficient methods for efficiently synthesizing generated content and sharing it on social media. Furthermore, there is a lack of easy ways to request and realize sequels or commercialization. Therefore, a system that makes it easier for users to create and share original content that responds to individual emotions is needed.
[1638] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1639] In this invention, the server includes means for recognizing the user's emotional information and generating music, lyrics, and choreography, means for providing the synthesized performance content to the user terminal, and means for posting the content to an SNS platform, thereby enabling users to generate original content according to their individual emotions, efficiently synthesize the content, and share it on the SNS.
[1640] The "means for accepting user input" is an interface that allows the user to input operational instructions to the system.
[1641] "Generative model" refers to algorithms or software that use artificial intelligence to automatically generate content such as music, lyrics, and choreography.
[1642] "Means for recognizing user's emotional information" refers to an engine or software that analyzes the user's facial expressions and tone of voice via a camera or microphone to identify their current emotional state.
[1643] "Means for generating music, lyrics and choreography" refers to a device or software that generates original music, lyrics and choreography based on emotional information and input data.
[1644] "Means for providing the generated music, lyrics and choreography to the user's terminal" refers to a device or program that has the function of transferring the generated data from the server to the user's terminal and displaying and playing it.
[1645] "Means for recording a user's performance" refers to a device or function that uses a camera and microphone to record a user's performance along with a piece of music.
[1646] "Means for synthesizing recorded performance data with generated music" refers to software or a system for integrating a user's performance data with generated music and editing and generating a single composite content.
[1647] "Means for providing synthesized performance content to a user terminal" refers to a device or program for transmitting the completed synthesized content from the server to a user terminal so that it can be displayed and played.
[1648] "Means for posting to a social media platform" refers to a device or software that has the function of uploading and sharing generated content to a social media platform using an API or other means.
[1649] "Means for generating sequel music and lyrics" refers to artificial intelligence algorithms or models for generating new music and lyrics based on existing content.
[1650] The present invention relates to a system that enables users to easily create and share original music and performance content. Specific embodiments of this system will be described below.
[1651] Music, lyrics and choreography generation
[1652] First, the user launches the application and taps the "Generate a new song" button. This action sends a generation request from the device to the server. When the server receives the generation request, it launches an emotion engine and collects data to recognize the user's emotions. The emotion engine analyzes the user's facial expressions and vocal tone through the camera and microphone to identify the user's current emotion. This emotion information is sent back to the server. The server, along with the received emotion information, requests the generative model to generate music, lyrics, and choreography. The generative model generates original music, lyrics, and choreography based on the input prompt and emotion information and sends it back to the server. The server receives the generated data and sends it to the user's device. The device displays and plays this data on the app screen.
[1653] Hardware and software used
[1654] Hardware: Smartphone (camera, microphone)
[1655] Software: Emotion engine (facial expression and voice analysis software), generative AI model (music, lyrics, choreography generation)
[1656] Specific examples
[1657] For example, when a user launches the app and taps the "Generate a new song" button, the smartphone's camera and microphone are activated, and the user's facial expressions and voice are analyzed. If the user is in a happy mood, the generative model will generate an upbeat, cheerful song.
[1658] Example prompt for a generative AI model:
[1659] "The user has a smiling face and a high-pitched voice, so generate a 30-second bright and cheerful pop song."
[1660] Recording and recording user performance
[1661] The user can freely sing and dance along to the generated music. At this time, the user's device will record and record the music. The recording function will be activated and the captured data will be temporarily saved on the device.
[1662] Combining performance and generative music
[1663] Once the recording is complete, the device uploads the captured performance data to the server. The server receives the performance data and combines it with the generated music data. This combination process is carried out using appropriate video editing software or algorithms to ultimately generate a single, integrated performance content. The server then transmits the combined performance content to the device. The device then plays and stores the received content, making it available for viewing by the user.
[1664] Posting to social media
[1665] If the user indicates their intention to post the synthesized performance content to a social networking site, the device calls the social networking site API to upload the synthesized content to the social networking site platform. Through the social networking site API, the user can add information such as a caption and post the content.
[1666] Sequel Music - Commercial Use
[1667] If a user checks the reaction on social media and requests a sequel song or commercialization, the server sends a sequel song and lyric generation request to the generative model. The generative model generates new music and lyrics and sends them back to the server. The server then provides the generated sequel data to the user's device, allowing the user to check the newly generated music and lyrics.
[1668] This allows users to create original music and performance content that reflects their individual emotions and share it efficiently on social media. It also supports the continuous creation of content by responding to requests for sequels and commercialization.
[1669] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1670] Step 1:
[1671] The user launches the application and taps the "Create New Song" button.
[1672] Specifically, the user operates the app screen on the mobile device and performs an input operation by tapping the relevant button. This operation registers the generation request on the device, and the registered information is used in the next step.
[1673] Step 2:
[1674] The device sends a request to the server.
[1675] The terminal converts the generated request into packets and sends them to the server over the Internet. The input is the user's generated request, and the output is the generated HTTP request sent to the server.
[1676] Step 3:
[1677] The server activates the emotion engine and collects data to recognize the user's emotions.
[1678] When the server receives the request, it sends a command to activate the camera and microphone to the device. At this time, the collected data is sent to the emotion engine. The input is the HTTP request data, and the output is the command to start the emotion recognition engine.
[1679] Step 4:
[1680] The emotion engine uses the device's camera and microphone to collect and analyze the user's emotional data.
[1681] Specifically, the camera captures the user's facial expressions and the microphone records their voice. These data are analyzed in real time by the emotion engine, and the user's emotional information is extracted as numerical data. The input is raw data from the camera and microphone, and the output is analyzed emotional information.
[1682] Step 5:
[1683] The server inputs emotional information into the generative model and requests it to generate music, lyrics, and choreography.
[1684] The server sends the received emotional information along with a prompt to the generative AI model. The input is the analyzed emotional information and a generation request, and the output is instruction data including the prompt to the generative model.
[1685] Step 6:
[1686] The generative model generates the music, lyrics, and choreography and sends it back to the server.
[1687] The generative model generates original music, lyrics, and choreography of approximately 30 seconds based on prompts and emotional information. The input is the prompts and emotional data, and the output is the generated music, lyrics, and choreography data.
[1688] Step 7:
[1689] The server transmits the generated data to the terminal.
[1690] The server converts the generated data into packets and sends them to the user terminal via the Internet. The input is the generated data from the generative model, and the output is the packets sent to the user terminal.
[1691] Step 8:
[1692] The data received by the device is displayed on the app screen and played back.
[1693] The terminal analyzes the generated data and displays it as a playback component of the music and choreography. The input is the generated data from the server, and the output is the content provided visually and audibly to the user.
[1694] Step 9:
[1695] Users can freely sing and dance along to the generated music.
[1696] The user presses the play button on the device to play the music and performs along with it. This operation is input into the next recording step.
[1697] Step 10:
[1698] The device will record and temporarily store your performance.
[1699] The device's camera and microphone are again used to record the user's singing and dancing. The input is the user's performance, and the output is the temporarily stored recording.
[1700] Step 11:
[1701] The device uploads the recording data to the server.
[1702] The terminal converts the temporarily stored performance data into packets and sends them to a server via the Internet. The input is the recorded data, and the output is the packets sent to the server.
[1703] Step 12:
[1704] The server receives the performance data and synthesizes it with the generated music data.
[1705] The server launches appropriate video editing software, inputs the audio and video data and the generated music data into the program, and synthesizes them. The input is the performance data and the generated music data, and the output is the synthesized performance content.
[1706] Step 13:
[1707] The server transmits the synthesized performance content to the terminal.
[1708] The server converts the composite content back into packets and transmits them to the user terminal. The input is the composite content, and the output is the packets sent to the user terminal.
[1709] Step 14:
[1710] The terminal plays and stores the received composite content so that the user can view it.
[1711] The terminal analyzes the composite content, loads it into the playback player, and provides it to the user. The input is the composite content from the server, and the output is the played and saved content.
[1712] Step 15:
[1713] The user indicates their intention to post the synthesized performance content to a social networking site.
[1714] The user taps the "Share to SNS" button in the app to indicate their intention to post. This action becomes the input for the next step.
[1715] Step 16:
[1716] The device calls the SNS API and uploads the composite content to the SNS platform.
[1717] The device uses the SNS API to send the composite content to the SNS platform, adding necessary information such as captions and uploading it. The input is the user's intention to post and the composite content, and the output is a notification of completion of posting to the SNS platform.
[1718] (Application example 2)
[1719] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1720] Conventional music generation systems have limited the content that users can experience, making it particularly difficult to generate personalized content based on emotions. Furthermore, there is a lack of easy ways to share generated content on social media platforms. This has prevented users from fully enjoying the fun of creating unique music and performances that respond to individual emotions.
[1721] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting user input, means for analyzing the user's emotions using an emotion engine, means for generating music, lyrics, and choreography using a generative model, means for providing the generated music, lyrics, and choreography to a user terminal, means for recording a user's performance, means for synthesizing the recorded performance data with the generated music, means for providing the synthesized performance content to a user terminal, and means for generating a prompt sentence for the generative AI model based on emotion information. This makes it possible to generate personalized music and performance content based on the user's emotions and easily share the generated content on a social media platform.
[1722] "User" refers to a person who uses this system to create music or performance content.
[1723] "Means for accepting input" refers to an interface that allows a user to input a request for music composition to the system.
[1724] An "emotion engine" refers to algorithms and software that analyze emotions from a user's facial expressions and voice.
[1725] "Means for analyzing emotions" refers to functionality for identifying a user's current emotional state using an emotion engine.
[1726] "Generative model" refers to an AI model that automatically generates music, lyrics, and choreography based on specified parameters and emotional information.
[1727] "Means for generating music, lyrics, and choreography" refers to a function for generating music, lyrics, and choreography using a generative model.
[1728] "User terminal" refers to a device (smartphone, tablet, etc.) that a user uses to view the music, lyrics, and choreography they have created.
[1729] "Means for recording performance" refers to the audio and video recording functionality used to record a User's performance.
[1730] "Recorded Performance Data" means audio and video data captured of a User's performance.
[1731] "Means of synthesis" refers to the function of integrating recorded performance data with the generated music to create a single piece of content.
[1732] "Synthesized Performance Content" refers to content generated as a result of integrating recorded or filmed performance data with generated music.
[1733] "SNS Platform" refers to a social networking service used to share and post User-Generated Content.
[1734] "Means for generating prompt sentences for a generative AI model" refers to a function that creates instruction sentences to be given to a generative AI model based on emotional information.
[1735] This invention relates to a system that generates music, lyrics, and choreography based on a user's emotions, allowing the user to create and share performance content. Specific embodiments of the system are described in detail below.
[1736] Hardware and Software Used
[1737] Hardware: smartphones, tablets, servers, cameras, microphones
[1738] Software: Smartphone applications, cloud server services (e.g., Amazon Web Services, Google Cloud), emotion engines (Microsoft Azure Face API, Google Cloud Speech-to-Text API), generative AI models (OpenAI GPT, Music VAE), social media APIs (Facebook API, Twitter API)
[1739] Program processing overview
[1740] 1. Launch the application:
[1741] The user launches the smartphone app and taps the "Generate new song" button.
[1742] The app activates the camera and microphone to capture the user's facial expressions and voice for emotion analysis.
[1743] 2. Emotion analysis:
[1744] The server uses an emotion engine to analyze the user's emotions from the captured data.
[1745] The emotion engine identifies the user's emotions from facial expressions and voice and sends that information back to the server.
[1746] 3. Music, lyrics, and choreography generation:
[1747] Based on the emotional information, the server sends prompts to the generative AI model to generate music, lyrics, and choreography.
[1748] The generative AI model generates original content according to the specified content and sends it back to the server.
[1749] 4. Receiving and Viewing Generated Content:
[1750] The server transmits the generated music, lyrics, and choreography data to the user terminal.
[1751] The user device plays and displays the received data within the app.
[1752] 5. Recording and filming of performances:
[1753] The user performs along with the generated music, and the device records and films the performance.
[1754] The recorded data will be temporarily stored and uploaded to a server.
[1755] 6. Video Composition and Distribution:
[1756] The server combines the audio and video data with the generated music to generate a single performance content.
[1757] The synthesized performance content is transmitted again to the user terminal.
[1758] 7. Posting to social media:
[1759] A user selects a social media post for the composited content.
[1760] Your app uses social media APIs to post content to social media platforms.
[1761] Specific examples
[1762] For example, if a user uses this app when they are tired, the app will analyze the tired expression on the user's face and generate relaxing music and slow choreography. The user can then perform along with the music, record the performance, and post it on Instagram.
[1763] Prompt Sentence Examples
[1764] Music generation prompt:
[1765] "The user's emotion is fatigue. Please generate a 30-second relaxing piece of music. The genre should be ambient and the tempo should be slow."
[1766] Lyric generation prompt:
[1767] "The user's emotion is fatigue. Generate 30 seconds of relaxing lyrics. The theme is nature and healing."
[1768] The present invention enables personalized content generation based on a user's emotions and makes it easy to share the generated content on a social networking platform.
[1769] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1770] Step 1:
[1771] The user launches the smartphone app and taps the "Generate New Song" button, which activates the app's camera and microphone to capture the user's facial expressions and voice for emotional analysis.
[1772] Input: Tap of the "Generate New Song" button, user facial and voice data
[1773] Output: User's facial expression and voice data
[1774] Step 2:
[1775] The device sends the captured facial and voice data to the server, which then activates the emotion engine to analyze the received data.
[1776] Input: User facial and voice data
[1777] Output: A signal that triggers the emotion engine to start analysis.
[1778] Step 3:
[1779] The server uses an emotion engine to analyze the user's emotions from facial expressions and voice data. The emotion engine analyzes data such as facial expressions and voice tone and sends the results back to the server.
[1780] Input: User facial and voice data received by the emotion engine
[1781] Output: User's emotional information
[1782] Step 4:
[1783] The server generates prompts for the generative AI model based on the user's emotional information. The generated prompts include requests to generate music, lyrics, and choreography.
[1784] Input: User's emotional information
[1785] Output: Prompt sentence for the generative AI model
[1786] Step 5:
[1787] The server sends the prompt text to the generative AI model, which generates the music, lyrics, and choreography. The generative AI model generates content based on the specified prompt and sends it back to the server.
[1788] Input: Prompt for the generative AI model
[1789] Output: Generated music, lyrics, and choreography data
[1790] Step 6:
[1791] The server then sends the generated music, lyrics, and choreography data to the user's device, which then plays and displays the received data within the app.
[1792] Input: Generated music, lyrics, and choreography data
[1793] Output: Playback and display on the user's device
[1794] Step 7:
[1795] The user performs along with the generated music, and the device records and films the performance. The recorded data is temporarily stored on the device.
[1796] Input: Performance along with the music
[1797] Output: Audio and video data
[1798] Step 8:
[1799] The device uploads the recorded data to the server, which then combines the received data with the generated music.
[1800] Input: Audio and video data
[1801] Output: Synthetic content
[1802] Step 9:
[1803] The server transmits the composite content to the user terminal, which then plays and displays the received composite content.
[1804] Input: Synthetic content
[1805] Output: Playback and display on the user's device
[1806] Step 10:
[1807] The user selects a social media post for the composited content, and the app uses the social media API to post the content to the social media platform.
[1808] Input: User's social media posts selection
[1809] Output: Posting content to social media platforms
[1810] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1811] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1812] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1813] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1814] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1815] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1816] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1817] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1818] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1819] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1820] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1821] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1822] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1823] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1824] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1825] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1826] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1827] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1828] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1829] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1830] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1831] The following is further disclosed regarding the above embodiment.
[1832] (Claim 1)
[1833] means for accepting user input;
[1834] A means for generating music, lyrics, and choreography using a generative model;
[1835] means for providing the generated music, lyrics and choreography to a user terminal;
[1836] a means for recording a user's performance;
[1837] A means for synthesizing the recorded performance data with the generated music;
[1838] A system including means for providing synthesized performance content to a user terminal.
[1839] (Claim 2)
[1840] 10. The system of claim 1, further comprising means for posting the synthesized performance content to a social media platform.
[1841] (Claim 3)
[1842] 10. The system of claim 1, further comprising: means for generating a sequel song and lyrics based on the response on the social media platform.
[1843] "Example 1"
[1844] (Claim 1)
[1845] means for accepting user input;
[1846] a means for generating music, lyrics, and choreography using a generative AI model;
[1847] means for providing the generated music, lyrics and choreography to a user terminal;
[1848] a means for recording a user's performance;
[1849] A means for synthesizing the recorded performance data with the generated music;
[1850] means for providing the synthesized performance content to a user terminal;
[1851] a means for enabling a user to post the composite content to a social networking platform;
[1852] The system includes a means for generating a sequel song and lyrics based on reactions on a social media platform.
[1853] (Claim 2)
[1854] 10. The system of claim 1, further comprising means for posting the synthesized performance content to a social media platform.
[1855] (Claim 3)
[1856] 10. The system of claim 1, further comprising: means for generating a sequel song and lyrics based on the response on the social media platform.
[1857] "Application Example 1"
[1858] Claims
[1859] (Claim 1)
[1860] means for accepting user input;
[1861] A means for generating music, lyrics, and choreography using a generative model;
[1862] means for providing the generated music, lyrics and choreography to a user terminal;
[1863] a means for recording a user's performance;
[1864] A means for synthesizing the recorded performance data with the generated music;
[1865] means for providing the synthesized performance content to a user terminal;
[1866] A means of posting the generated content to social media platforms;
[1867] A way to check the reaction on social media platforms and request the creation of a sequel song and lyrics,
[1868] A system including:
[1869] (Claim 2)
[1870] The system of claim 1 generates a sequel song and lyrics based on the response on a social media platform.
[1871] (Claim 3)
[1872] 10. The system of claim 1, wherein the synthesized performance content is posted to a social media platform.
[1873] "Example 2: Combining Emotion Engines"
[1874] (Claim 1)
[1875] means for accepting user input;
[1876] a means for recognizing user emotion information using a generative model and generating music, lyrics, and choreography;
[1877] means for providing the generated music, lyrics and choreography to a user terminal;
[1878] a means for recording a user's performance;
[1879] A means for synthesizing the recorded performance data with the generated music;
[1880] A system including means for providing synthesized performance content to a user terminal.
[1881] (Claim 2)
[1882] 10. The system of claim 1, further comprising: means for posting the synthesized performance content to a social media platform.
[1883] (Claim 3)
[1884] 10. The system of claim 1, further comprising: means for generating a sequel song and lyrics based on the response on the social media platform.
[1885] "Application example 2 when combining emotion engines"
[1886] (Claim 1)
[1887] means for accepting user input;
[1888] means for analyzing a user's emotions by an emotion engine;
[1889] A means for generating music, lyrics, and choreography using a generative model;
[1890] means for providing the generated music, lyrics and choreography to a user terminal;
[1891] a means for recording a user's performance;
[1892] A means for synthesizing the recorded performance data with the generated music;
[1893] means for providing the synthesized performance content to a user terminal;
[1894] A means for generating a prompt sentence for a generative AI model based on emotion information;
[1895] A system including:
[1896] (Claim 2)
[1897] 10. The system of claim 1, further comprising means for posting the synthesized performance content to a social media platform.
[1898] (Claim 3)
[1899] 10. The system of claim 1, further comprising: means for generating a sequel song and lyrics based on the response on the social media platform. [Explanation of symbols]
[1900] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for accepting user input; A means for generating music, lyrics, and choreography using a generative model; means for providing the generated music, lyrics and choreography to a user terminal; means for recording and recording a performance by a user; means for combining the recorded and filmed performance data with the generated musical composition; A system including means for providing synthesized performance content to a user terminal.
2. 10. The system of claim 1, further comprising means for posting the synthesized performance content to a social networking platform.
3. The system of claim 1 , further comprising: means for generating a sequel song and lyrics based on the response on the social media platform.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A