System

The system addresses the challenge of music composition by recording humming, generating music, and providing feedback, enabling users to create and share music without advanced skills, thus simplifying the process.

JP2026015079APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116553
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Many individuals interested in music composition lack the skills or knowledge to create music due to the technical barriers of reading music or playing instruments, and there is a limited means for receiving feedback on their compositions.

Method used

A system that records a user's humming as audio data, analyzes it to generate scales, rhythms, and melodies, converts the data into various instrument sounds, allows playback and editing, and enables sharing for feedback, with AI suggesting the next phrase or idea.

Benefits of technology

This system significantly reduces technical barriers, allowing anyone to easily compose and share music, receive feedback, and improve their compositions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015079000001_ABST
    Figure 2026015079000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for recording and storing a user's humming as audio data; means for analyzing the audio data and generating scales, rhythms, and melodies on a server; means for converting music into various instrument sounds based on the generated scales, rhythms, and melodies; means for transmitting the generated music data to a terminal to enable playback and editing; and means for uploading the music to a community and receiving feedback.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Many people who are interested in music find it difficult to compose music because they cannot read music or do not have the skills to play an instrument. As a result, many people are unable to compose music due to a lack of skill or knowledge, despite their desire to create. Furthermore, there are limited ways to receive feedback on their creations. Given this background, there is a demand for a service that removes the technical barriers to music production and allows anyone to easily enjoy composing music. [Means for solving the problem]

[0005] The present invention has a means for recording a user's humming and saving it as audio data, and a means for transmitting the recorded audio data to a server. The server has a means for analyzing the audio data and generating scales, rhythms, and melodies, and further converts this information into various instrument sounds. The generated music data is transmitted to a terminal, where it can be played and edited. The system also has a means for AI to automatically suggest the next phrase or idea for the generated music. Furthermore, users can upload their music to a community and receive feedback. This removes the technical barriers to music production, allowing anyone to easily enjoy composing music.

[0006] "Means for recording a user's humming and saving it as audio data" refers to a device or software function that allows a user to record their humming and saves the recorded audio as digital data.

[0007] "Means for transmitting the audio data to a server" refers to a communication function for transmitting recorded audio data to a remote server via a network.

[0008] "Means of analyzing audio data on a server and generating scales, rhythms, and melodies" refers to the process of extracting and generating notes, rhythms, and melody lines from humming using algorithms and AI models for analyzing audio data.

[0009] "Means for converting a musical piece into various instrument sounds based on the generated scale, rhythm, and melody" refers to a function for converting the extracted scale, rhythm, and melody information into audio data that can be played with the tones of multiple different instruments.

[0010] "Means for transmitting generated music data to a terminal and enabling playback and editing" refers to a function for transmitting music data generated by a server to a user's terminal and enabling the user to play and edit the music.

[0011] "Means for uploading music to the community and receiving feedback" refers to the ability to post music created by users to an online community and receive comments and ratings from other users.

[0012] "Means for AI to automatically suggest the next phrase or idea" refers to the process in which AI uses a music generation algorithm to automatically suggest the next development in a song or candidates for a new melody line. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention provides a system that allows a user to record a humming tune and then create, edit, and share a song based on the humming tune.

[0035] Overall system overview

[0036] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server analyzes the received audio data and generates scale, rhythm, and melody information. A song is then generated based on the sounds of the instrument selected by the user, and the generated song is sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[0037] Recording and Data Transmission

[0038] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[0039] Analysis and music generation on the server

[0040] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[0041] Next, the server converts the generated music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sound.

[0042] Sending and playing music data

[0043] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[0044] Editing features and phrase suggestions

[0045] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[0046] Community and Feedback

[0047] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[0048] Specific examples

[0049] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on the device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0050] As a result, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music.

[0051] The processing flow will be explained below.

[0052] Step 1:

[0053] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[0054] Step 2:

[0055] The device records the user's humming and saves it as audio data.

[0056] Step 3:

[0057] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[0058] Step 4:

[0059] The terminal transmits the voice data to the server via the Internet.

[0060] Step 5:

[0061] The server receives the voice data sent from the terminal and converts the voice data into a format for analysis.

[0062] Step 6:

[0063] The server uses an AI model to analyze the audio data and extract pitch, rhythm, and melody information.

[0064] Specifically, the AI ​​model analyzes audio waveforms to identify the pitch and timing of each note.

[0065] Step 7:

[0066] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[0067] For example, if the user selects piano, the note data is converted into piano sounds.

[0068] Step 8:

[0069] The server transmits the generated music data to the terminal.

[0070] Step 9:

[0071] The terminal receives the music data sent from the server and displays music playback options to the user.

[0072] Step 10:

[0073] The user presses the play button to listen to the generated music.

[0074] Step 11:

[0075] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[0076] Step 12:

[0077] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[0078] Step 13:

[0079] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[0080] Step 14:

[0081] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[0082] Step 15:

[0083] The device displays suggested phrases and adds the user's selection to the song.

[0084] Step 16:

[0085] The user presses the upload button to upload the completed song to the community.

[0086] Step 17:

[0087] The terminal transmits the completed song data to a community server, and the song is made public.

[0088] Step 18:

[0089] Other users in the community provide feedback on the published songs.

[0090] Step 19:

[0091] The server notifies the song poster of the collected feedback.

[0092] Step 20:

[0093] The user reviews the feedback and makes any necessary changes to the song.

[0094] Through the above series of steps, the present invention provides a system that allows anyone to easily enjoy composing music.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Existing music production systems require advanced musical knowledge and complex operations, making it difficult for average users to easily create music. Other issues include a lack of functionality to instantly turn a user's original melody into digital data, and a lack of systems that can easily convert it into multiple instrument sounds. Furthermore, the lack of support for users when considering the next phrase makes it difficult to smoothly compose music.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for recording a user's voice and saving it as digital data, means for transmitting the digital data to the server, means for analyzing the digital data on the server and generating a scale, rhythm, and melody, means for converting the music data into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for automatically using AI to suggest next phrase candidates for the edited music data, and means for uploading music data to a community and receiving feedback. This enables general users to easily record, generate, and edit music, convert it into multiple instrument sounds, and receive feedback from other users to smoothly compose music.

[0100] A "user" is an entity that uses the system to record humming and create, edit, and share music.

[0101] A "terminal" is a device used by a user, and has functions such as recording, playback, editing, sending, and receiving.

[0102] "Digital data" refers to a collection of bits that have been converted to store the user's humming or voice electronically.

[0103] A "server" is a computing device that receives and analyzes digital data sent from a terminal, generates music data, and provides it to the user.

[0104] A "generative AI model" is an artificial intelligence algorithm that runs on a server and extracts scales, rhythms, and melodies from audio data.

[0105] "Music data" refers to digital music data generated based on scale, rhythm, and melody information analyzed by a generative AI model.

[0106] "Instrument sound" refers to the tone or sound produced when music data is converted into the sound of a specific instrument.

[0107] "Editing" refers to the modification or addition of music data that the user performs on the generated music data, and includes changing the tempo, modifying notes, adding harmonies, and the like.

[0108] "Phrase candidates" are options for the next melody line or harmony that the generative AI model automatically suggests for the music data being edited.

[0109] "Community" is an online platform for publishing user-generated compositions and receiving feedback from other users.

[0110] "Feedback" is the opinions and ratings provided by other users of the community, and is information that helps improve and refine your songs.

[0111] The present invention is a system that allows users to record their humming and then create, edit, and share music based on that recording.

[0112] This system consists of a terminal used by the user, a server, and software for linking them.

[0113] Recording preparation and humming

[0114] The user launches the application installed on the device. The device provides the user with a record button through the interface. When the user presses the record button, the device uses the built-in microphone to record the humming sound and saves it as digital data. This digital data is temporarily stored in the device.

[0115] Sending voice data and analyzing it on the server

[0116] Once the recording is complete, the device sends the digital data to a server, for example, using the HTTPS protocol. The server then analyzes the received digital data using an AI model (generative AI model). This model extracts scale, rhythm, and melody information and structures it as music data. Specifically, the pitch and timing of each note are detected from the audio data.

[0117] Music data generation and transmission

[0118] Next, the server converts the extracted music data into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sounds. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sounds. The converted music data is then sent back to the terminal. The terminal receives this data and notifies the user.

[0119] Playing and editing songs

[0120] Users can play the generated music by pressing the play button on their device. Furthermore, they can use the editing tools to change the tempo, modify the notes, or add harmonies as needed. This editing function allows users to create more refined music.

[0121] Phrase suggestion feature

[0122] If a user is unsure of the next phrase during the editing process, they can press the suggest button. The device then sends a request to the server, which uses an AI model to generate melody line and harmony suggestions. These suggestions are then presented to the user via their device, allowing them to select their favorite phrase and continue composing.

[0123] Community Uploads and Feedback

[0124] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song, and the server collects this feedback and notifies the song poster via their device. This allows the user to improve their song based on the opinions and ratings of other users.

[0125] Specific examples

[0126] For example, consider the case where a user hums "Happy Birthday." The user records the song on their device and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the user uploads the completed song to the community and receives feedback.

[0127] Prompt Sentence Examples

[0128] An example of a prompt to be input to the generative AI model is, "Please convert the recorded humming into instrument sounds and generate the song 'Happy Birthday.' Please also accurately determine the rhythm and scale and play it as a piano." Based on this prompt, the AI ​​model analyzes the humming audio data and generates a song using the specified instrument sounds.

[0129] As described above, by using the system of the present invention, users can easily create, edit, and share music even if they do not have advanced musical knowledge.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1: Prepare to record

[0132] Input: User actions

[0133] Output: Recording ready state

[0134] Specific operation: The user launches the application on the device. The device displays the application's main screen and prepares for recording by displaying record, play, and edit buttons on the interface.

[0135] Step 2: Record your humming

[0136] Input: User's humming (audio data)

[0137] Output: Recorded digital data

[0138] Specific operation: When the user presses the record button, the device will use the built-in microphone to capture audio data in real time and save it as digital data. Once the recording is complete, the digital data will be temporarily stored on the device.

[0139] Step 3: Sending audio data

[0140] Input: Recorded digital data

[0141] Output: Notification of completion of transmission to the server

[0142] Specific operation: When the user finishes recording, the device compresses the saved audio data and sends it to the server using the HTTPS protocol. If the data is successfully sent, the server returns a successful reception response to the device.

[0143] Step 4: Analyzing the audio data

[0144] Input: Transmitted digital data

[0145] Output: Scale, rhythm, melody information

[0146] How it works: The server passes the received audio data to a generative AI model, which then analyzes the data for pitch and timing of each note, extracting information about the scale, rhythm, and melody.

[0147] Step 5: Generate music data

[0148] Input: scale, rhythm, melody information

[0149] Output: Music data converted into specific instrument sounds

[0150] Specific operation: The server converts the analyzed music data into the instrument sound selected by the user. For example, if the user selects piano, the AI ​​model converts the note data into piano sound. The converted music data is saved in a file format.

[0151] Step 6: Send your music

[0152] Input: Generated music data

[0153] Output: Notification of completion of transmission to the terminal

[0154] Specific operation: The server sends the converted music data back to the device. The device receives the data and displays "Music created" in the notification bar. The music data is also saved in the application.

[0155] Step 7: Play and edit your song

[0156] Input: Generated music data

[0157] Output: Played songs and edited song data

[0158] Specific operation: The user presses the play button to play a song. The device plays the specified song data, and when the user switches to edit mode, they can change the tempo, modify notes, add harmonies, etc. The edited song data is updated in real time.

[0159] Step 8: Use the Phrase Suggestion Feature

[0160] Input: Request to server and current music data

[0161] Output: Suggested phrase candidates

[0162] Specific operation: When the user presses the suggest button, the device sends the current song data to the server and requests suggestions for the next phrase. The server-side AI model generates multiple melody and harmony suggestions and sends them to the device. The user then selects from the suggested phrase suggestions through the device interface.

[0163] Step 9: Upload to the Community

[0164] Input: Completed song data

[0165] Output: Upload completion notification to the community

[0166] Specific operation: When a user presses the button to upload a song to the community, the device sends the song data to the server and makes it available within the community. The server stores the song and makes it accessible to other users. When the upload is successful, a notification is displayed on the device.

[0167] Step 10: Get feedback

[0168] Input: Feedback from other users

[0169] Output: Feedback notification and feedback content

[0170] What it does: Other users in the community add comments and ratings to uploaded songs. The server collects this feedback and sends it to the uploader's device as a notification. By tapping the notification, the user can view the feedback details.

[0171] As described above, by using this system, users can easily create, edit, and share music based on their humming.

[0172] (Application example 1)

[0173] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0174] In conventional music production systems, it was difficult for users without musical knowledge or skills to create and distribute music and share it with other users. They also lacked the functionality to receive feedback on the music they created and incorporate improvements. Furthermore, there was no system that provided users with the next phrase or idea during the music generation process based on humming, which increased the time and effort required for music production.

[0175] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0176] In this invention, the server includes means for recording a user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for uploading the generated music data to a content distribution service and sharing it with other users, means for receiving feedback from other users, and means for AI to automatically suggest the next phrase or idea. This allows even users with no musical knowledge or skills to easily create music from their humming, share it with other users, and receive feedback. Furthermore, AI's suggestions of the next phrase or idea significantly reduce the effort and time required for music production.

[0177] "User" refers to an individual or organization that uses the system to record humming, create music, and share it.

[0178] "Humming" refers to an informal singing voice with a melody or rhythm that is hummed by a user.

[0179] "Audio data" refers to data that is a digital recording of the user's humming.

[0180] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[0181] A "scale" is an arrangement of pitches that indicate the pitch of a sound, and is an element that makes up the melody of a piece of music.

[0182] "Rhythm" refers to the pattern of the placement and spacing of musical notes in time, which forms the tempo and beat of a piece of music.

[0183] A "melody" is the main theme of a piece of music, which is formed by combining scales and rhythms.

[0184] "Instrument sounds" are sounds produced by a particular instrument and are used to reproduce the scale, rhythm, and melody generated from humming.

[0185] A "terminal" refers to an electronic device used by a user, such as a computer, smartphone, or tablet.

[0186] A "content distribution service" is a platform for sharing and distributing music and other media content between users over the Internet.

[0187] "Feedback" means comments and ratings provided by other users within the community.

[0188] The present invention is a system that allows users to record their humming and then create, edit, and share music based on the audio data. The system mainly includes means for data communication between a terminal and a server and for data analysis. Specific embodiments of the present invention are described below.

[0189] Recording and Data Transmission

[0190] The user records their humming using an application installed on the device. At this time, the audio data is captured through the device's microphone and saved as digital data. After recording is complete, the audio data is sent from the device to the server. The software used includes a requests library for making HTTP requests.

[0191] Analysis and music generation on the server

[0192] The server receives the audio data sent from the device and analyzes the scale, rhythm, and melody information using an AI model. This analysis includes detecting the pitch and timing of musical notes from the audio data and structuring it as music data. Next, based on the analyzed music data, it converts it into the sound of an instrument selected by the user (e.g., piano). This conversion process involves analyzing the audio using an AI model and generating acoustic data. Audio analysis on the server is performed using an audio processing library (e.g., librosa).

[0193] Sending and playing music data

[0194] The generated music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. An audio data manipulation library (e.g., pydub) is used for playback. While listening to the music, the user can edit it by changing the tempo, modifying notes, adding harmonies, and so on.

[0195] Editing features and phrase suggestions

[0196] When editing a song on a device, if a user is unsure of the next phrase, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and presents them to the user via the device. This significantly reduces the effort required for music production.

[0197] Community and Feedback

[0198] The completed song is uploaded by the user to a community on the content distribution service. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. The user can then receive the feedback and make corrections to the song.

[0199] Examples of specific examples and prompts

[0200] For example, if a user hums "Happy Birthday," the device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0201] An example of a prompt for a generative AI model is:

[0202] "User sang 'Happy Birthday' with a humming voice. Convert this recording into a piano music file."

[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0204] Step 1:

[0205] The user uses the terminal to record a hum.

[0206] Input: User's humming

[0207] Data processing: Capture and store audio data digitally through the device's microphone

[0208] Output: Digital audio data

[0209] Step 2:

[0210] The terminal transmits the recorded voice data to the server.

[0211] Input: Digital audio data

[0212] Data processing: Upload audio data to the server using an HTTP request

[0213] Output: Audio data stored on the server

[0214] Step 3:

[0215] The server receives the audio data and begins analyzing it.

[0216] Input: Audio data stored on the server

[0217] Data processing: Using audio analysis algorithms to analyze and extract pitch, rhythm, and melodic information

[0218] Output: Scale information, rhythm information, melody information

[0219] Step 4:

[0220] The server converts the analyzed music data into the sound of an instrument selected by the user.

[0221] Input: Scale information, rhythm information, melody information, and the type of instrument sound selected by the user

[0222] Data processing: Using a generative AI model to convert note data into selected instrument sounds

[0223] Output: Music data converted into instrument sounds

[0224] Step 5:

[0225] The server transmits the generated music data to the terminal.

[0226] Input: Music data converted into instrument sounds

[0227] Data processing: Send music data to the device using HTTP responses

[0228] Output: Song data stored on the device

[0229] Step 6:

[0230] The terminal allows the user to play the music data.

[0231] Input: Song data stored on the device

[0232] Data processing: Playing audio data using the pydub library

[0233] Output: The song played by the user

[0234] Step 7:

[0235] The user edits the music as needed.

[0236] Input: Song data to edit

[0237] Data processing: Editing operations such as changing the tempo, correcting notes, and adding harmonies can be performed through the application.

[0238] Output: Edited song data

[0239] Step 8:

[0240] The user uploads the completed song to a content distribution service.

[0241] Input: Completed song data

[0242] Data processing: Upload music data using HTTP requests and make it available to the community

[0243] Output: Songs published on content distribution services

[0244] Step 9:

[0245] Other users in the community provide feedback on the published songs.

[0246] Input: Published songs

[0247] Data Processing: Post a comment or rating

[0248] Output: Feedback provided

[0249] Step 10:

[0250] The server collects feedback from other users and notifies the song poster via the terminal.

[0251] Input: Feedback from other users

[0252] Data processing: Collect feedback data and notify song submitters

[0253] Output: Feedback notification received by song submitter

[0254] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0255] The following describes in detail an embodiment of the present invention: The present invention is a system that allows a user to record a humming tune and then create, edit, and share music based on the humming tune. It also incorporates an emotion engine that recognizes the user's emotional state and applies it to the creation and editing of music.

[0256] Overall system overview

[0257] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server then analyzes the received audio data and generates scale, rhythm, and melody information. The server then uses an emotion engine to analyze the user's emotions and automatically adjusts the mood and tempo of the music based on the emotion recognition results. A song is then generated based on the sounds of the instruments selected by the user and sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[0258] Recording and Data Transmission

[0259] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[0260] Analysis and music generation on the server

[0261] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[0262] The server then uses an emotion engine to analyze the user's emotional state. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and other biometric signals. Based on the emotion recognition results, the server automatically adjusts the mood and tempo of the music.

[0263] Furthermore, the extracted music data is converted into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the scale information is converted into note data corresponding to the respective instrument sound.

[0264] Sending and playing music data

[0265] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[0266] Editing features and phrase suggestions

[0267] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[0268] Community and Feedback

[0269] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[0270] Specific examples

[0271] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion, and if it recognizes it as "joy," for example, it speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0272] In this way, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music by generating music that takes the user's emotions into consideration.

[0273] The processing flow will be explained below.

[0274] Step 1:

[0275] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[0276] Step 2:

[0277] The device records the user's humming and saves it as audio data.

[0278] Step 3:

[0279] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[0280] Step 4:

[0281] The terminal transmits the voice data to the server via the Internet.

[0282] Step 5:

[0283] The server receives the voice data sent from the device and converts it into a format for analysis, which makes it easier to analyze.

[0284] Step 6:

[0285] The server uses AI models to analyze the audio data and extract pitch, rhythm, and melody information.

[0286] Example: Detecting pitch and beat from a recorded humming and converting it into musical note data.

[0287] Step 7:

[0288] Users provide emotional data, such as facial expressions and tone of voice, through their device's camera and microphone.

[0289] The device acquires the emotion data and transmits it to the server.

[0290] Step 8:

[0291] The server uses an emotion engine to analyze the received emotion data, thereby recognizing the user's emotional state (e.g., joy, sadness).

[0292] Example: AI analyzes facial expressions and recognizes that if the user is smiling, it is expressing "joy."

[0293] Step 9:

[0294] The server automatically adjusts the mood and tempo of the music based on the emotion recognition results. For example, if it recognizes "joy," it will speed up the tempo or add a more upbeat arrangement.

[0295] Step 10:

[0296] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[0297] For example, if the user selects piano, the note data is converted into piano sounds.

[0298] Step 11:

[0299] The server transmits the generated music data to the terminal.

[0300] Step 12:

[0301] The terminal receives the music data sent from the server and displays music playback options to the user.

[0302] Step 13:

[0303] The user presses the play button to listen to the generated music.

[0304] Step 14:

[0305] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[0306] Step 15:

[0307] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[0308] Step 16:

[0309] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[0310] Step 17:

[0311] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[0312] Step 18:

[0313] The device displays suggested phrases and adds the user's selection to the song.

[0314] Step 19:

[0315] The user presses the upload button to upload the completed song to the community.

[0316] Step 20:

[0317] The terminal transmits the completed song data to a community server, and the song is made public.

[0318] Step 21:

[0319] Other users in the community provide feedback on the published songs.

[0320] Step 22:

[0321] The server notifies the song poster of the collected feedback.

[0322] Step 23:

[0323] The user reviews the feedback and makes any necessary changes to the song.

[0324] Through the above series of steps, the present invention allows anyone to easily enjoy composing music, and by taking the user's emotions into consideration when generating music, it provides a more personal music-making experience.

[0325] Example 2

[0326] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0327] Conventional music generation systems require users to manually edit music, often requiring technical knowledge. Furthermore, they are unable to generate music based on the user's emotional state, making it difficult to generate optimal music that matches individual emotions. Furthermore, there is no function to automatically suggest the next phrase or idea for a generated piece of music. There is a need for a system that can resolve these issues and allow users to generate music more easily and based on their emotions.

[0328] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data using a generative AI model to generate a scale, rhythm, and melody, means for analyzing the user's emotional state and automatically adjusting the mood and tempo of the music, and means for automatically suggesting the user's next phrase or idea using the generative AI model via a suggestion button. This allows users to easily generate music without technical knowledge, and the music is optimized according to the user's emotional state, enabling more satisfying music production. Furthermore, the automatic suggestion of the next phrase or idea can support the user's creativity.

[0329] "User" means an individual who uses the System to record humming and create, edit, and share music.

[0330] "Terminal" refers to a communication device used by a user to record humming, including a smartphone, tablet, etc.

[0331] A "server" is a computer system that receives audio data sent by a user, analyzes it, and creates music.

[0332] "Audio data" refers to information stored in digital form of a user's humming.

[0333] A "generative AI model" is an artificial intelligence algorithm that analyzes audio data to generate scales, rhythms, and melodies.

[0334] A "scale" refers to the arrangement of pitches that make up the melody of a piece of music.

[0335] "Rhythm" refers to the pattern of timing and duration of notes in a piece of music.

[0336] "Melody" refers to the main theme of a piece of music that is formed by combining scales and rhythms.

[0337] The "emotion engine" is an algorithm that analyzes the user's emotional state and adjusts the mood and tempo of the music based on the results.

[0338] "Instrumental sounds" refers to sounds produced by different musical instruments such as piano, guitar, violin, drums, etc.

[0339] "Music Data" means music information in digital form that has been generated through an analysis and conversion process.

[0340] "Suggestion button" refers to an interface element that a user uses to request an automatic suggestion of the next phrase or idea.

[0341] "Community" means the online platform where users can upload their created Music and receive feedback from other users.

[0342] "Feedback" means ratings and opinions provided by other users within the Community regarding the Generated Song.

[0343] The system of the present invention allows users to record their humming and then use it to create, edit, and share music, and is characterized by using an emotion engine to apply the user's emotional state to the music creation. Below, we will explain in detail how this system is implemented.

[0344] Recording and Data Transmission

[0345] Users record their humming using an application installed on their smartphone, tablet, or other device. When they tap the record button, the device captures the audio data through the built-in microphone and saves it as digital data in PCM or AAC format. Once recording is complete, the device sends the audio data to a server via the Internet.

[0346] Hardware used: Smartphone, tablet

[0347] Software used: Recording application

[0348] Analysis and music generation on the server

[0349] The server receives the audio data sent from the device and uses a generative AI model to analyze the audio data and extract musical scale, rhythm, and melody information. This analysis involves using a speech recognition algorithm to detect the pitch and timing of musical notes.

[0350] The server then analyzes the user's emotional state using an emotion engine, which analyzes the tone of the recorded voice and facial expression images provided by the user, and automatically adjusts the mood and tempo of the music based on the emotion recognition results.

[0351] Finally, the server converts the extracted music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If another instrument (e.g., guitar or violin) is selected, the information is converted into note data corresponding to the instrument sound.

[0352] Hardware used: Server

[0353] Software used: Generative AI model, emotion engine

[0354] Sending and playing music data

[0355] The server then sends the generated music data back to the terminal, which receives it and displays it on its user interface. The user can listen to the music by pressing the play button.

[0356] Hardware used: Server, terminal

[0357] Software used: Music playback application

[0358] Editing features and phrase suggestions

[0359] Users can edit songs through the application, for example, by changing the tempo, modifying specific notes, or adding harmonies. If users are unsure of the next phrase during editing, they can press the suggestion button. This causes the server to use a generative AI model to generate several melody line and harmony candidates and suggest them to the user via their device.

[0360] Hardware used: Server, terminal

[0361] Software used: Music editing application, generative AI model

[0362] Community and Feedback

[0363] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device, allowing the user to receive opinions and ratings from other users and make corrections to the song.

[0364] Hardware used: Server, terminal

[0365] Software used: Community Platform

[0366] Specific examples

[0367] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server uses a generative AI model to analyze the scale and extract the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion and, if it recognizes it as "joy," for example, speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0368] Example prompt sentence:

[0369] "I'm humming Happy Birthday. Identify the user's emotion as joy, speed up the tempo, and emphasize the bright tone. Finally, convert it into a piano sound."

[0370] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0371] Step 1:

[0372] The user launches the device's recording application and taps the record button. The device uses the built-in microphone to capture the user's humming and saves it as digital audio data in, for example, PCM or AAC format. The input is the user's humming, and the output is the captured audio data. This recording is saved as a temporary file in the device's storage.

[0373] Specific behavior:

[0374] User taps the record button

[0375] The device captures audio through the microphone

[0376] The device generates and stores digital audio data

[0377] Step 2:

[0378] After recording is complete, the device sends the audio data to the server via the Internet. The input is the stored audio data, and the output is the audio data sent to the server. The data is transferred securely using the HTTP protocol or WebSocket.

[0379] Specific behavior:

[0380] The device is waiting for audio data

[0381] The device sends the voice data to the server

[0382] Step 3:

[0383] The server analyzes the received audio data. First, it uses a generative AI model to extract scale, rhythm, and melody from the audio data. The input is audio data, and the output is structured musical data. It uses a speech recognition algorithm (e.g., a deep learning model) to detect the pitch and timing of each note and stores this in a database format.

[0384] Specific behavior:

[0385] The server receives the audio data

[0386] The server applies generative AI models to extract scale, rhythm, and melody

[0387] The server structures and stores music data

[0388] Step 4:

[0389] The server then analyzes the user's emotional state using an emotion engine. The emotion engine recognizes emotions by analyzing voice tone and provided facial expression images. The input is voice data and facial expression data, and the output is the recognized emotional state. The emotion engine uses machine learning algorithms to determine emotions and incorporates this information into the music data.

[0390] Specific behavior:

[0391] The server applies the emotion engine

[0392] The server analyzes voice tone and facial expression data

[0393] The server recognizes and records the user's emotional state.

[0394] Step 5:

[0395] The server converts the generated music data into the instrument sounds selected by the user. If the user selects piano, it is converted into piano sounds. The input is structured music data and instrument selection information, and the output is music data converted into specific instrument sounds. The instrument sounds are synthesized using sound fonts and audio sampling.

[0396] Specific behavior:

[0397] User selects instrument

[0398] The server converts the music data into instrument sounds

[0399] Step 6:

[0400] The server sends the converted music data to the terminal. The input is the music data, and the output is the music data sent to the terminal. The data is transferred via a secure communication channel.

[0401] Specific behavior:

[0402] The server waits for music data

[0403] The server sends the music data to the device

[0404] Step 7:

[0405] The terminal displays the received music data on the user interface, and the user can listen to the music by pressing the play button. The input is the received music data, and the output is the played music.

[0406] Specific behavior:

[0407] The device receives the music data

[0408] User taps the play button

[0409] The device plays the music

[0410] Step 8:

[0411] The user edits the music through the application, changing the tempo, modifying specific notes, adding harmonies, etc. The input is the music data and editing operations, and the output is the edited music data.

[0412] Specific behavior:

[0413] User uses editing functions

[0414] The device reflects the edited content in the song data

[0415] Step 9:

[0416] When a user is unsure of the next phrase, they press the suggest button. The server uses a generative AI model to generate several melody line and harmony candidates and suggests them to the user via the device. The input is a suggestion request, and the output is the suggested melody or harmony.

[0417] Specific behavior:

[0418] User taps the suggest button

[0419] Server generates candidates

[0420] Your device will display suggestions

[0421] Step 10:

[0422] A completed song is uploaded to the community by the user. Other users in the community can provide feedback on this song. The input is the completed song data, and the output is the feedback from the user. The server collects the feedback and notifies the song uploader via their device.

[0423] Specific behavior:

[0424] Users upload songs to the community

[0425] Other users provide feedback

[0426] Server collects and notifies feedback

[0427] (Application example 2)

[0428] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0429] Previous technologies have not provided sufficient concrete methods for improving work efficiency and worker morale in factories. Furthermore, systems that automatically generate music suited to the work environment have not been able to combine it with emotion recognition technology. Therefore, there is a need for technology that automatically generates music suited to the work environment while taking into account the user's emotions.

[0430] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for converting music into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music suitable for the work environment, means for playing the background music on robots in the factory, and means for uploading music to the community and receiving feedback. This makes it possible to generate appropriate background music based on the user's humming and emotions, thereby improving work efficiency and morale in the factory.

[0431] definition statement

[0432] The "means for recording a user's humming and saving it as audio data" is a device or program that captures the audio of a user humming as digital data and saves it for later processing.

[0433] The "means for transmitting the audio data to the server" is a system that transfers the recorded audio data to a remote server via the Internet or a local network.

[0434] The "means for analyzing audio data on a server and generating scales, rhythms, and melodies" refers to a server system that performs processing to extract highly accurate musical elements based on audio data.

[0435] The "means for converting a piece of music into various instrument sounds based on the generated scale, rhythm, and melody" is a system that uses the analyzed musical elements to convert note data into different instrument sounds selected by the user.

[0436] "Means for transmitting generated music data to a terminal and enabling playback and editing" refers to a system for transferring music data generated on a server to a user's operating terminal, allowing the data to be played back and further edited.

[0437] "Means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music appropriate for the work environment" refers to a system that analyzes the user's emotional state and automatically adjusts the tempo and tone of the music to an appropriate level based on the results.

[0438] The "means for playing the background music by a robot in a factory" refers to a robot system used to play the generated background music in a physical space.

[0439] "Means for uploading music to a community and receiving feedback" is a system that allows users to share music they have created with an online community and collect ratings and comments from other users.

[0440] MODE FOR CARRYING OUT THE INVENTION

[0441] The embodiment of the present invention is a music generation system aimed at improving work efficiency and worker morale in a factory. This system allows a user to record a humming tune and then generates, edits, and plays appropriate background music based on that humming. The system also recognizes the user's emotions and reflects them in the music it generates, providing music that is optimal for the work environment.

[0442] The server processes and calculates data using the following hardware and software: The "librosa" library is used for analyzing audio data, the "EmotionEngine" is used for emotion recognition, and the "MusicGenerator" is used for music generation.

[0443] System operation explanation

[0444] Humming recording and data transmission

[0445] A user records their humming using a terminal equipped with a recording function. The recorded audio data is saved as digital data by the terminal. This digital data is then transmitted to a server via a network.

[0446] Analysis and music generation on the server

[0447] The server receives the transmitted audio data and uses the librosa library to analyze the pitch and timing of each note, extracting scale, rhythm, and melody information. It then uses the Emotion Engine to analyze the user's emotions. Based on this emotional analysis, the tempo and timbre of the music are automatically adjusted.

[0448] Based on the analyzed musical scale information, the "Music Generator" is used to convert it into an appropriate instrument sound. At this stage, the note data is converted based on the instrument sound selected by the user (e.g. piano, guitar, drums, etc.).

[0449] Sending and playing music data

[0450] The generated music data is then sent back to the terminal. The terminal displays the received music data, and the user can listen to the music by pressing the play button. The user can also edit the music by changing the tempo or correcting the notes. In particular, it is possible to use robots in factories to play background music.

[0451] Community Features and Feedback

[0452] Users can upload their created compositions to the community and receive feedback from other users, allowing for further improvements and suggestions for new ideas.

[0453] Examples and prompts

[0454] As a specific example, let us consider a case where a user working in a factory hums, saying, "I want to concentrate, so I want some calming music." The tempo and tone of the music generated from this humming are adjusted appropriately based on the tone of the user's voice and emotional analysis.

[0455] Prompt Sentence Examples

[0456] Prompt: "I need some calming music to help me concentrate."

[0457] In this way, the present invention provides a music generation system that aims to improve work efficiency and morale in factories. Furthermore, by generating music that takes emotions into consideration, it is possible to provide appropriate music that matches the user's psychological state.

[0458] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0459] System program processing flow

[0460] Step 1:

[0461] The user uses the device to record their humming. When they press the record button, the audio is captured through the device's microphone and saved as digital audio data.

[0462] Input: User's humming (audio)

[0463] Output: Digital audio data (.wav, .mp3, etc.)

[0464] Step 2:

[0465] The stored digital audio data is sent from the device to the server, where it is transferred securely and quickly using a network protocol for transmission (e.g., HTTP, FTP).

[0466] Input: Digital audio data

[0467] Output: Digital audio data received by the server

[0468] Step 3:

[0469] The server analyzes the received audio data and generates scale, rhythm, and melody information using the audio data analysis library "librosa," which detects the pitch and timing of each note.

[0470] Input: Digital audio data received by the server

[0471] Output: Scale, rhythm, melody information (structured data)

[0472] Step 4:

[0473] Based on the analyzed scale, rhythm, and melody information, the note data is converted into musical instrument sounds. To correspond to the instrument sounds selected by the user (e.g., piano, guitar), the note data is converted using "MusicGenerator."

[0474] Input: Scale, rhythm, melody information, and selected instrument information

[0475] Output: Musical note data converted into instrument sounds (music data)

[0476] Step 5:

[0477] The server analyzes the user's emotions and reflects them in the generated music data. Using the emotion recognition engine "EmotionEngine," it analyzes the user's emotional state from the audio data. Based on that emotional state, the tempo and tone of the music are automatically adjusted.

[0478] Input: Digital voice data and analyzed emotional information

[0479] Output: Music data with tempo and tone adjusted to match the emotion

[0480] Step 6:

[0481] The generated and adjusted music data is sent back to the device. The server sends the data to the device via a transfer protocol.

[0482] Input: Music data generated and adjusted on the server

[0483] Output: Music data received on the device

[0484] Step 7:

[0485] The device plays and edits the received music data. The user can listen to the created music by pressing the play button. The edit button can also be used to change the tempo or modify the notes.

[0486] Input: Music data received on the device

[0487] Output: Played and edited song

[0488] Step 8:

[0489] Users can then play the final edited song as background music on a robot in the factory, which has a built-in speaker and plays the music based on the song data.

[0490] Input: Song data edited on the device

[0491] Output: Background music played in your work environment

[0492] Step 9:

[0493] The completed songs are uploaded to the community by the users, and other users provide feedback on the uploaded songs, which is collected by the server and sent to the device.

[0494] Input: Completed song data

[0495] Output: Songs shared with the community and given feedback

[0496] Specific prompt examples:

[0497] Prompt: "I need some calming music to help me concentrate."

[0498] In this way, a system is provided that generates music that reflects the user's humming or emotional state and plays it as appropriate background music in the factory.

[0499] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0500] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0501] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0502] [Second embodiment]

[0503] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0504] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0505] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0506] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0507] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0509] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0510] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0511] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0512] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0513] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0514] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0515] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention provides a system that allows a user to record a humming tune and then create, edit, and share a song based on the humming tune.

[0516] Overall system overview

[0517] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server analyzes the received audio data and generates scale, rhythm, and melody information. A song is then generated based on the sounds of the instrument selected by the user, and the generated song is sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[0518] Recording and Data Transmission

[0519] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[0520] Analysis and music generation on the server

[0521] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[0522] Next, the server converts the generated music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sound.

[0523] Sending and playing music data

[0524] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[0525] Editing features and phrase suggestions

[0526] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[0527] Community and Feedback

[0528] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[0529] Specific examples

[0530] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on the device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0531] As a result, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music.

[0532] The processing flow will be explained below.

[0533] Step 1:

[0534] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[0535] Step 2:

[0536] The device records the user's humming and saves it as audio data.

[0537] Step 3:

[0538] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[0539] Step 4:

[0540] The terminal transmits the voice data to the server via the Internet.

[0541] Step 5:

[0542] The server receives the voice data sent from the terminal and converts the voice data into a format for analysis.

[0543] Step 6:

[0544] The server uses an AI model to analyze the audio data and extract pitch, rhythm, and melody information.

[0545] Specifically, the AI ​​model analyzes audio waveforms to identify the pitch and timing of each note.

[0546] Step 7:

[0547] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[0548] For example, if the user selects piano, the note data is converted into piano sounds.

[0549] Step 8:

[0550] The server transmits the generated music data to the terminal.

[0551] Step 9:

[0552] The terminal receives the music data sent from the server and displays music playback options to the user.

[0553] Step 10:

[0554] The user presses the play button to listen to the generated music.

[0555] Step 11:

[0556] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[0557] Step 12:

[0558] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[0559] Step 13:

[0560] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[0561] Step 14:

[0562] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[0563] Step 15:

[0564] The device displays suggested phrases and adds the user's selection to the song.

[0565] Step 16:

[0566] The user presses the upload button to upload the completed song to the community.

[0567] Step 17:

[0568] The terminal transmits the completed song data to a community server, and the song is made public.

[0569] Step 18:

[0570] Other users in the community provide feedback on the published songs.

[0571] Step 19:

[0572] The server notifies the song poster of the collected feedback.

[0573] Step 20:

[0574] The user reviews the feedback and makes any necessary changes to the song.

[0575] Through the above series of steps, the present invention provides a system that allows anyone to easily enjoy composing music.

[0576] Example 1

[0577] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0578] Existing music production systems require advanced musical knowledge and complex operations, making it difficult for average users to easily create music. Other issues include a lack of functionality to instantly turn a user's original melody into digital data, and a lack of systems that can easily convert it into multiple instrument sounds. Furthermore, the lack of support for users when considering the next phrase makes it difficult to smoothly compose music.

[0579] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0580] In this invention, the server includes means for recording a user's voice and saving it as digital data, means for transmitting the digital data to the server, means for analyzing the digital data on the server and generating a scale, rhythm, and melody, means for converting the music data into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for automatically using AI to suggest next phrase candidates for the edited music data, and means for uploading music data to a community and receiving feedback. This enables general users to easily record, generate, and edit music, convert it into multiple instrument sounds, and receive feedback from other users to smoothly compose music.

[0581] A "user" is an entity that uses the system to record humming and create, edit, and share music.

[0582] A "terminal" is a device used by a user, and has functions such as recording, playback, editing, sending, and receiving.

[0583] "Digital data" refers to a collection of bits that have been converted to store the user's humming or voice electronically.

[0584] A "server" is a computing device that receives and analyzes digital data sent from a terminal, generates music data, and provides it to the user.

[0585] A "generative AI model" is an artificial intelligence algorithm that runs on a server and extracts scales, rhythms, and melodies from audio data.

[0586] "Music data" refers to digital music data generated based on scale, rhythm, and melody information analyzed by a generative AI model.

[0587] "Instrument sound" refers to the tone or sound produced when music data is converted into the sound of a specific instrument.

[0588] "Editing" refers to the modification or addition of music data that the user performs on the generated music data, and includes changing the tempo, modifying notes, adding harmonies, and the like.

[0589] "Phrase candidates" are options for the next melody line or harmony that the generative AI model automatically suggests for the music data being edited.

[0590] "Community" is an online platform for publishing user-generated compositions and receiving feedback from other users.

[0591] "Feedback" is the opinions and ratings provided by other users of the community, and is information that helps improve and refine your songs.

[0592] The present invention is a system that allows users to record their humming and then create, edit, and share music based on that recording.

[0593] This system consists of a terminal used by the user, a server, and software for linking them.

[0594] Recording preparation and humming

[0595] The user launches the application installed on the device. The device provides the user with a record button through the interface. When the user presses the record button, the device uses the built-in microphone to record the humming sound and saves it as digital data. This digital data is temporarily stored in the device.

[0596] Sending voice data and analyzing it on the server

[0597] Once the recording is complete, the device sends the digital data to a server, for example, using the HTTPS protocol. The server then analyzes the received digital data using an AI model (generative AI model). This model extracts scale, rhythm, and melody information and structures it as music data. Specifically, the pitch and timing of each note are detected from the audio data.

[0598] Music data generation and transmission

[0599] Next, the server converts the extracted music data into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sounds. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sounds. The converted music data is then sent back to the terminal. The terminal receives this data and notifies the user.

[0600] Playing and editing songs

[0601] Users can play the generated music by pressing the play button on their device. Furthermore, they can use the editing tools to change the tempo, modify the notes, or add harmonies as needed. This editing function allows users to create more refined music.

[0602] Phrase suggestion feature

[0603] If a user is unsure of the next phrase during the editing process, they can press the suggest button. The device then sends a request to the server, which uses an AI model to generate melody line and harmony suggestions. These suggestions are then presented to the user via their device, allowing them to select their favorite phrase and continue composing.

[0604] Community Uploads and Feedback

[0605] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song, and the server collects this feedback and notifies the song poster via their device. This allows the user to improve their song based on the opinions and ratings of other users.

[0606] Specific examples

[0607] For example, consider the case where a user hums "Happy Birthday." The user records the song on their device and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the user uploads the completed song to the community and receives feedback.

[0608] Prompt Sentence Examples

[0609] An example of a prompt to be input to the generative AI model is, "Please convert the recorded humming into instrument sounds and generate the song 'Happy Birthday.' Please also accurately determine the rhythm and scale and play it as a piano." Based on this prompt, the AI ​​model analyzes the humming audio data and generates a song using the specified instrument sounds.

[0610] As described above, by using the system of the present invention, users can easily create, edit, and share music even if they do not have advanced musical knowledge.

[0611] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0612] Step 1: Prepare to record

[0613] Input: User actions

[0614] Output: Recording ready state

[0615] Specific operation: The user launches the application on the device. The device displays the application's main screen and prepares for recording by displaying record, play, and edit buttons on the interface.

[0616] Step 2: Record your humming

[0617] Input: User's humming (audio data)

[0618] Output: Recorded digital data

[0619] Specific operation: When the user presses the record button, the device will use the built-in microphone to capture audio data in real time and save it as digital data. Once the recording is complete, the digital data will be temporarily stored on the device.

[0620] Step 3: Sending audio data

[0621] Input: Recorded digital data

[0622] Output: Notification of completion of transmission to the server

[0623] Specific operation: When the user finishes recording, the device compresses the saved audio data and sends it to the server using the HTTPS protocol. If the data is successfully sent, the server returns a successful reception response to the device.

[0624] Step 4: Analyzing the audio data

[0625] Input: Transmitted digital data

[0626] Output: Scale, rhythm, melody information

[0627] How it works: The server passes the received audio data to a generative AI model, which then analyzes the data for pitch and timing of each note, extracting information about the scale, rhythm, and melody.

[0628] Step 5: Generate music data

[0629] Input: scale, rhythm, melody information

[0630] Output: Music data converted into specific instrument sounds

[0631] Specific operation: The server converts the analyzed music data into the instrument sound selected by the user. For example, if the user selects piano, the AI ​​model converts the note data into piano sound. The converted music data is saved in a file format.

[0632] Step 6: Send your music

[0633] Input: Generated music data

[0634] Output: Notification of completion of transmission to the terminal

[0635] Specific operation: The server sends the converted music data back to the device. The device receives the data and displays "Music created" in the notification bar. The music data is also saved in the application.

[0636] Step 7: Play and edit your song

[0637] Input: Generated music data

[0638] Output: Played songs and edited song data

[0639] Specific operation: The user presses the play button to play a song. The device plays the specified song data, and when the user switches to edit mode, they can change the tempo, modify notes, add harmonies, etc. The edited song data is updated in real time.

[0640] Step 8: Use the Phrase Suggestion Feature

[0641] Input: Request to server and current music data

[0642] Output: Suggested phrase candidates

[0643] Specific operation: When the user presses the suggest button, the device sends the current song data to the server and requests suggestions for the next phrase. The server-side AI model generates multiple melody and harmony suggestions and sends them to the device. The user then selects from the suggested phrase suggestions through the device interface.

[0644] Step 9: Upload to the Community

[0645] Input: Completed song data

[0646] Output: Upload completion notification to the community

[0647] Specific operation: When a user presses the button to upload a song to the community, the device sends the song data to the server and makes it available within the community. The server stores the song and makes it accessible to other users. When the upload is successful, a notification is displayed on the device.

[0648] Step 10: Get feedback

[0649] Input: Feedback from other users

[0650] Output: Feedback notification and feedback content

[0651] What it does: Other users in the community add comments and ratings to uploaded songs. The server collects this feedback and sends it to the uploader's device as a notification. By tapping the notification, the user can view the feedback details.

[0652] As described above, by using this system, users can easily create, edit, and share music based on their humming.

[0653] (Application example 1)

[0654] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0655] In conventional music production systems, it was difficult for users without musical knowledge or skills to create and distribute music and share it with other users. They also lacked the functionality to receive feedback on the music they created and incorporate improvements. Furthermore, there was no system that provided users with the next phrase or idea during the music generation process based on humming, which increased the time and effort required for music production.

[0656] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0657] In this invention, the server includes means for recording a user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for uploading the generated music data to a content distribution service and sharing it with other users, means for receiving feedback from other users, and means for AI to automatically suggest the next phrase or idea. This allows even users with no musical knowledge or skills to easily create music from their humming, share it with other users, and receive feedback. Furthermore, AI's suggestions of the next phrase or idea significantly reduce the effort and time required for music production.

[0658] "User" refers to an individual or organization that uses the system to record humming, create music, and share it.

[0659] "Humming" refers to an informal singing voice with a melody or rhythm that is hummed by a user.

[0660] "Audio data" refers to data that is a digital recording of the user's humming.

[0661] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[0662] A "scale" is an arrangement of pitches that indicate the pitch of a sound, and is an element that makes up the melody of a piece of music.

[0663] "Rhythm" refers to the pattern of the placement and spacing of musical notes in time, which forms the tempo and beat of a piece of music.

[0664] A "melody" is the main theme of a piece of music, which is formed by combining scales and rhythms.

[0665] "Instrument sounds" are sounds produced by a particular instrument and are used to reproduce the scale, rhythm, and melody generated from humming.

[0666] A "terminal" refers to an electronic device used by a user, such as a computer, smartphone, or tablet.

[0667] A "content distribution service" is a platform for sharing and distributing music and other media content between users over the Internet.

[0668] "Feedback" means comments and ratings provided by other users within the community.

[0669] The present invention is a system that allows users to record their humming and then create, edit, and share music based on the audio data. The system mainly includes means for data communication between a terminal and a server and for data analysis. Specific embodiments of the present invention are described below.

[0670] Recording and Data Transmission

[0671] The user records their humming using an application installed on the device. At this time, the audio data is captured through the device's microphone and saved as digital data. After recording is complete, the audio data is sent from the device to the server. The software used includes a requests library for making HTTP requests.

[0672] Analysis and music generation on the server

[0673] The server receives the audio data sent from the device and analyzes the scale, rhythm, and melody information using an AI model. This analysis includes detecting the pitch and timing of musical notes from the audio data and structuring it as music data. Next, based on the analyzed music data, it converts it into the sound of an instrument selected by the user (e.g., piano). This conversion process involves analyzing the audio using an AI model and generating acoustic data. Audio analysis on the server is performed using an audio processing library (e.g., librosa).

[0674] Sending and playing music data

[0675] The generated music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. An audio data manipulation library (e.g., pydub) is used for playback. While listening to the music, the user can edit it by changing the tempo, modifying notes, adding harmonies, and so on.

[0676] Editing features and phrase suggestions

[0677] When editing a song on a device, if a user is unsure of the next phrase, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and presents them to the user via the device. This significantly reduces the effort required for music production.

[0678] Community and Feedback

[0679] The completed song is uploaded by the user to a community on the content distribution service. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. The user can then receive the feedback and make corrections to the song.

[0680] Examples of specific examples and prompts

[0681] For example, if a user hums "Happy Birthday," the device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0682] An example of a prompt for a generative AI model is:

[0683] "User sang 'Happy Birthday' with a humming voice. Convert this recording into a piano music file."

[0684] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0685] Step 1:

[0686] The user uses the terminal to record a hum.

[0687] Input: User's humming

[0688] Data processing: Capture and store audio data digitally through the device's microphone

[0689] Output: Digital audio data

[0690] Step 2:

[0691] The terminal transmits the recorded voice data to the server.

[0692] Input: Digital audio data

[0693] Data processing: Upload audio data to the server using an HTTP request

[0694] Output: Audio data stored on the server

[0695] Step 3:

[0696] The server receives the audio data and begins analyzing it.

[0697] Input: Audio data stored on the server

[0698] Data processing: Using audio analysis algorithms to analyze and extract pitch, rhythm, and melodic information

[0699] Output: Scale information, rhythm information, melody information

[0700] Step 4:

[0701] The server converts the analyzed music data into the sound of an instrument selected by the user.

[0702] Input: Scale information, rhythm information, melody information, and the type of instrument sound selected by the user

[0703] Data processing: Using a generative AI model to convert note data into selected instrument sounds

[0704] Output: Music data converted into instrument sounds

[0705] Step 5:

[0706] The server transmits the generated music data to the terminal.

[0707] Input: Music data converted into instrument sounds

[0708] Data processing: Send music data to the device using HTTP responses

[0709] Output: Song data stored on the device

[0710] Step 6:

[0711] The terminal allows the user to play the music data.

[0712] Input: Song data stored on the device

[0713] Data processing: Playing audio data using the pydub library

[0714] Output: The song played by the user

[0715] Step 7:

[0716] The user edits the music as needed.

[0717] Input: Song data to edit

[0718] Data processing: Editing operations such as changing the tempo, correcting notes, and adding harmonies can be performed through the application.

[0719] Output: Edited song data

[0720] Step 8:

[0721] The user uploads the completed song to a content distribution service.

[0722] Input: Completed song data

[0723] Data processing: Upload music data using HTTP requests and make it available to the community

[0724] Output: Songs published on content distribution services

[0725] Step 9:

[0726] Other users in the community provide feedback on the published songs.

[0727] Input: Published songs

[0728] Data Processing: Post a comment or rating

[0729] Output: Feedback provided

[0730] Step 10:

[0731] The server collects feedback from other users and notifies the song poster via the terminal.

[0732] Input: Feedback from other users

[0733] Data processing: Collect feedback data and notify song submitters

[0734] Output: Feedback notification received by song submitter

[0735] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0736] The following describes in detail an embodiment of the present invention: The present invention is a system that allows a user to record a humming tune and then create, edit, and share music based on the humming tune. It also incorporates an emotion engine that recognizes the user's emotional state and applies it to the creation and editing of music.

[0737] Overall system overview

[0738] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server then analyzes the received audio data and generates scale, rhythm, and melody information. The server then uses an emotion engine to analyze the user's emotions and automatically adjusts the mood and tempo of the music based on the emotion recognition results. A song is then generated based on the sounds of the instruments selected by the user and sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[0739] Recording and Data Transmission

[0740] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[0741] Analysis and music generation on the server

[0742] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[0743] The server then uses an emotion engine to analyze the user's emotional state. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and other biometric signals. Based on the emotion recognition results, the server automatically adjusts the mood and tempo of the music.

[0744] Furthermore, the extracted music data is converted into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the scale information is converted into note data corresponding to the respective instrument sound.

[0745] Sending and playing music data

[0746] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[0747] Editing features and phrase suggestions

[0748] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[0749] Community and Feedback

[0750] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[0751] Specific examples

[0752] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion, and if it recognizes it as "joy," for example, it speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0753] In this way, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music by generating music that takes the user's emotions into consideration.

[0754] The processing flow will be explained below.

[0755] Step 1:

[0756] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[0757] Step 2:

[0758] The device records the user's humming and saves it as audio data.

[0759] Step 3:

[0760] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[0761] Step 4:

[0762] The terminal transmits the voice data to the server via the Internet.

[0763] Step 5:

[0764] The server receives the voice data sent from the device and converts it into a format for analysis, which makes it easier to analyze.

[0765] Step 6:

[0766] The server uses AI models to analyze the audio data and extract pitch, rhythm, and melody information.

[0767] Example: Detecting pitch and beat from a recorded humming and converting it into musical note data.

[0768] Step 7:

[0769] Users provide emotional data, such as facial expressions and tone of voice, through their device's camera and microphone.

[0770] The device acquires the emotion data and transmits it to the server.

[0771] Step 8:

[0772] The server uses an emotion engine to analyze the received emotion data, thereby recognizing the user's emotional state (e.g., joy, sadness).

[0773] Example: AI analyzes facial expressions and recognizes that if the user is smiling, it is expressing "joy."

[0774] Step 9:

[0775] The server automatically adjusts the mood and tempo of the music based on the emotion recognition results. For example, if it recognizes "joy," it will speed up the tempo or add a more upbeat arrangement.

[0776] Step 10:

[0777] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[0778] For example, if the user selects piano, the note data is converted into piano sounds.

[0779] Step 11:

[0780] The server transmits the generated music data to the terminal.

[0781] Step 12:

[0782] The terminal receives the music data sent from the server and displays music playback options to the user.

[0783] Step 13:

[0784] The user presses the play button to listen to the generated music.

[0785] Step 14:

[0786] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[0787] Step 15:

[0788] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[0789] Step 16:

[0790] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[0791] Step 17:

[0792] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[0793] Step 18:

[0794] The device displays suggested phrases and adds the user's selection to the song.

[0795] Step 19:

[0796] The user presses the upload button to upload the completed song to the community.

[0797] Step 20:

[0798] The terminal transmits the completed song data to a community server, and the song is made public.

[0799] Step 21:

[0800] Other users in the community provide feedback on the published songs.

[0801] Step 22:

[0802] The server notifies the song poster of the collected feedback.

[0803] Step 23:

[0804] The user reviews the feedback and makes any necessary changes to the song.

[0805] Through the above series of steps, the present invention allows anyone to easily enjoy composing music, and by taking the user's emotions into consideration when generating music, it provides a more personal music-making experience.

[0806] Example 2

[0807] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0808] Conventional music generation systems require users to manually edit music, often requiring technical knowledge. Furthermore, they are unable to generate music based on the user's emotional state, making it difficult to generate optimal music that matches individual emotions. Furthermore, there is no function to automatically suggest the next phrase or idea for a generated piece of music. There is a need for a system that can resolve these issues and allow users to generate music more easily and based on their emotions.

[0809] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data using a generative AI model to generate a scale, rhythm, and melody, means for analyzing the user's emotional state and automatically adjusting the mood and tempo of the music, and means for automatically suggesting the user's next phrase or idea using the generative AI model via a suggestion button. This allows users to easily generate music without technical knowledge, and the music is optimized according to the user's emotional state, enabling more satisfying music production. Furthermore, the automatic suggestion of the next phrase or idea can support the user's creativity.

[0810] "User" means an individual who uses the System to record humming and create, edit, and share music.

[0811] "Terminal" refers to a communication device used by a user to record humming, including a smartphone, tablet, etc.

[0812] A "server" is a computer system that receives audio data sent by a user, analyzes it, and creates music.

[0813] "Audio data" refers to information stored in digital form of a user's humming.

[0814] A "generative AI model" is an artificial intelligence algorithm that analyzes audio data to generate scales, rhythms, and melodies.

[0815] A "scale" refers to the arrangement of pitches that make up the melody of a piece of music.

[0816] "Rhythm" refers to the pattern of timing and duration of notes in a piece of music.

[0817] "Melody" refers to the main theme of a piece of music that is formed by combining scales and rhythms.

[0818] The "emotion engine" is an algorithm that analyzes the user's emotional state and adjusts the mood and tempo of the music based on the results.

[0819] "Instrumental sounds" refers to sounds produced by different musical instruments such as piano, guitar, violin, drums, etc.

[0820] "Music Data" means music information in digital form that has been generated through an analysis and conversion process.

[0821] "Suggestion button" refers to an interface element that a user uses to request an automatic suggestion of the next phrase or idea.

[0822] "Community" means the online platform where users can upload their created Music and receive feedback from other users.

[0823] "Feedback" means ratings and opinions provided by other users within the Community regarding the Generated Song.

[0824] The system of the present invention allows users to record their humming and then use it to create, edit, and share music, and is characterized by using an emotion engine to apply the user's emotional state to the music creation. Below, we will explain in detail how this system is implemented.

[0825] Recording and Data Transmission

[0826] Users record their humming using an application installed on their smartphone, tablet, or other device. When they tap the record button, the device captures the audio data through the built-in microphone and saves it as digital data in PCM or AAC format. Once recording is complete, the device sends the audio data to a server via the Internet.

[0827] Hardware used: Smartphone, tablet

[0828] Software used: Recording application

[0829] Analysis and music generation on the server

[0830] The server receives the audio data sent from the device and uses a generative AI model to analyze the audio data and extract musical scale, rhythm, and melody information. This analysis involves using a speech recognition algorithm to detect the pitch and timing of musical notes.

[0831] The server then analyzes the user's emotional state using an emotion engine, which analyzes the tone of the recorded voice and facial expression images provided by the user, and automatically adjusts the mood and tempo of the music based on the emotion recognition results.

[0832] Finally, the server converts the extracted music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If another instrument (e.g., guitar or violin) is selected, the information is converted into note data corresponding to the instrument sound.

[0833] Hardware used: Server

[0834] Software used: Generative AI model, emotion engine

[0835] Sending and playing music data

[0836] The server then sends the generated music data back to the terminal, which receives it and displays it on its user interface. The user can listen to the music by pressing the play button.

[0837] Hardware used: Server, terminal

[0838] Software used: Music playback application

[0839] Editing features and phrase suggestions

[0840] Users can edit songs through the application, for example, by changing the tempo, modifying specific notes, or adding harmonies. If users are unsure of the next phrase during editing, they can press the suggestion button. This causes the server to use a generative AI model to generate several melody line and harmony candidates and suggest them to the user via their device.

[0841] Hardware used: Server, terminal

[0842] Software used: Music editing application, generative AI model

[0843] Community and Feedback

[0844] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device, allowing the user to receive opinions and ratings from other users and make corrections to the song.

[0845] Hardware used: Server, terminal

[0846] Software used: Community Platform

[0847] Specific examples

[0848] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server uses a generative AI model to analyze the scale and extract the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion and, if it recognizes it as "joy," for example, speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[0849] Example prompt sentence:

[0850] "I'm humming Happy Birthday. Identify the user's emotion as joy, speed up the tempo, and emphasize the bright tone. Finally, convert it into a piano sound."

[0851] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0852] Step 1:

[0853] The user launches the device's recording application and taps the record button. The device uses the built-in microphone to capture the user's humming and saves it as digital audio data in, for example, PCM or AAC format. The input is the user's humming, and the output is the captured audio data. This recording is saved as a temporary file in the device's storage.

[0854] Specific behavior:

[0855] User taps the record button

[0856] The device captures audio through the microphone

[0857] The device generates and stores digital audio data

[0858] Step 2:

[0859] After recording is complete, the device sends the audio data to the server via the Internet. The input is the stored audio data, and the output is the audio data sent to the server. The data is transferred securely using the HTTP protocol or WebSocket.

[0860] Specific behavior:

[0861] The device is waiting for audio data

[0862] The device sends the voice data to the server

[0863] Step 3:

[0864] The server analyzes the received audio data. First, it uses a generative AI model to extract scale, rhythm, and melody from the audio data. The input is audio data, and the output is structured musical data. It uses a speech recognition algorithm (e.g., a deep learning model) to detect the pitch and timing of each note and stores this in a database format.

[0865] Specific behavior:

[0866] The server receives the audio data

[0867] The server applies generative AI models to extract scale, rhythm, and melody

[0868] The server structures and stores music data

[0869] Step 4:

[0870] The server then analyzes the user's emotional state using an emotion engine. The emotion engine recognizes emotions by analyzing voice tone and provided facial expression images. The input is voice data and facial expression data, and the output is the recognized emotional state. The emotion engine uses machine learning algorithms to determine emotions and incorporates this information into the music data.

[0871] Specific behavior:

[0872] The server applies the emotion engine

[0873] The server analyzes voice tone and facial expression data

[0874] The server recognizes and records the user's emotional state.

[0875] Step 5:

[0876] The server converts the generated music data into the instrument sounds selected by the user. If the user selects piano, it is converted into piano sounds. The input is structured music data and instrument selection information, and the output is music data converted into specific instrument sounds. The instrument sounds are synthesized using sound fonts and audio sampling.

[0877] Specific behavior:

[0878] User selects instrument

[0879] The server converts the music data into instrument sounds

[0880] Step 6:

[0881] The server sends the converted music data to the terminal. The input is the music data, and the output is the music data sent to the terminal. The data is transferred via a secure communication channel.

[0882] Specific behavior:

[0883] The server waits for music data

[0884] The server sends the music data to the device

[0885] Step 7:

[0886] The terminal displays the received music data on the user interface, and the user can listen to the music by pressing the play button. The input is the received music data, and the output is the played music.

[0887] Specific behavior:

[0888] The device receives the music data

[0889] User taps the play button

[0890] The device plays the music

[0891] Step 8:

[0892] The user edits the music through the application, changing the tempo, modifying specific notes, adding harmonies, etc. The input is the music data and editing operations, and the output is the edited music data.

[0893] Specific behavior:

[0894] User uses editing functions

[0895] The device reflects the edited content in the song data

[0896] Step 9:

[0897] When a user is unsure of the next phrase, they press the suggest button. The server uses a generative AI model to generate several melody line and harmony candidates and suggests them to the user via the device. The input is a suggestion request, and the output is the suggested melody or harmony.

[0898] Specific behavior:

[0899] User taps the suggest button

[0900] Server generates candidates

[0901] Your device will display suggestions

[0902] Step 10:

[0903] A completed song is uploaded to the community by the user. Other users in the community can provide feedback on this song. The input is the completed song data, and the output is the feedback from the user. The server collects the feedback and notifies the song uploader via their device.

[0904] Specific behavior:

[0905] Users upload songs to the community

[0906] Other users provide feedback

[0907] Server collects and notifies feedback

[0908] (Application example 2)

[0909] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0910] Previous technologies have not provided sufficient concrete methods for improving work efficiency and worker morale in factories. Furthermore, systems that automatically generate music suited to the work environment have not been able to combine it with emotion recognition technology. Therefore, there is a need for technology that automatically generates music suited to the work environment while taking into account the user's emotions.

[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for converting music into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music suitable for the work environment, means for playing the background music on robots in the factory, and means for uploading music to the community and receiving feedback. This makes it possible to generate appropriate background music based on the user's humming and emotions, thereby improving work efficiency and morale in the factory.

[0912] definition statement

[0913] The "means for recording a user's humming and saving it as audio data" is a device or program that captures the audio of a user humming as digital data and saves it for later processing.

[0914] The "means for transmitting the audio data to the server" is a system that transfers the recorded audio data to a remote server via the Internet or a local network.

[0915] The "means for analyzing audio data on a server and generating scales, rhythms, and melodies" refers to a server system that performs processing to extract highly accurate musical elements based on audio data.

[0916] The "means for converting a piece of music into various instrument sounds based on the generated scale, rhythm, and melody" is a system that uses the analyzed musical elements to convert note data into different instrument sounds selected by the user.

[0917] "Means for transmitting generated music data to a terminal and enabling playback and editing" refers to a system for transferring music data generated on a server to a user's operating terminal, allowing the data to be played back and further edited.

[0918] "Means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music appropriate for the work environment" refers to a system that analyzes the user's emotional state and automatically adjusts the tempo and tone of the music to an appropriate level based on the results.

[0919] The "means for playing the background music by a robot in a factory" refers to a robot system used to play the generated background music in a physical space.

[0920] "Means for uploading music to a community and receiving feedback" is a system that allows users to share music they have created with an online community and collect ratings and comments from other users.

[0921] MODE FOR CARRYING OUT THE INVENTION

[0922] The embodiment of the present invention is a music generation system aimed at improving work efficiency and worker morale in a factory. This system allows a user to record a humming tune and then generates, edits, and plays appropriate background music based on that humming. The system also recognizes the user's emotions and reflects them in the music it generates, providing music that is optimal for the work environment.

[0923] The server processes and calculates data using the following hardware and software: The "librosa" library is used for analyzing audio data, the "EmotionEngine" is used for emotion recognition, and the "MusicGenerator" is used for music generation.

[0924] System operation explanation

[0925] Humming recording and data transmission

[0926] A user records their humming using a terminal equipped with a recording function. The recorded audio data is saved as digital data by the terminal. This digital data is then transmitted to a server via a network.

[0927] Analysis and music generation on the server

[0928] The server receives the transmitted audio data and uses the librosa library to analyze the pitch and timing of each note, extracting scale, rhythm, and melody information. It then uses the Emotion Engine to analyze the user's emotions. Based on this emotional analysis, the tempo and timbre of the music are automatically adjusted.

[0929] Based on the analyzed musical scale information, the "Music Generator" is used to convert it into an appropriate instrument sound. At this stage, the note data is converted based on the instrument sound selected by the user (e.g. piano, guitar, drums, etc.).

[0930] Sending and playing music data

[0931] The generated music data is then sent back to the terminal. The terminal displays the received music data, and the user can listen to the music by pressing the play button. The user can also edit the music by changing the tempo or correcting the notes. In particular, it is possible to use robots in factories to play background music.

[0932] Community Features and Feedback

[0933] Users can upload their created compositions to the community and receive feedback from other users, allowing for further improvements and suggestions for new ideas.

[0934] Examples and prompts

[0935] As a specific example, let us consider a case where a user working in a factory hums, saying, "I want to concentrate, so I want some calming music." The tempo and tone of the music generated from this humming are adjusted appropriately based on the tone of the user's voice and emotional analysis.

[0936] Prompt Sentence Examples

[0937] Prompt: "I need some calming music to help me concentrate."

[0938] In this way, the present invention provides a music generation system that aims to improve work efficiency and morale in factories. Furthermore, by generating music that takes emotions into consideration, it is possible to provide appropriate music that matches the user's psychological state.

[0939] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0940] System program processing flow

[0941] Step 1:

[0942] The user uses the device to record their humming. When they press the record button, the audio is captured through the device's microphone and saved as digital audio data.

[0943] Input: User's humming (audio)

[0944] Output: Digital audio data (.wav, .mp3, etc.)

[0945] Step 2:

[0946] The stored digital audio data is sent from the device to the server, where it is transferred securely and quickly using a network protocol for transmission (e.g., HTTP, FTP).

[0947] Input: Digital audio data

[0948] Output: Digital audio data received by the server

[0949] Step 3:

[0950] The server analyzes the received audio data and generates scale, rhythm, and melody information using the audio data analysis library "librosa," which detects the pitch and timing of each note.

[0951] Input: Digital audio data received by the server

[0952] Output: Scale, rhythm, melody information (structured data)

[0953] Step 4:

[0954] Based on the analyzed scale, rhythm, and melody information, the note data is converted into musical instrument sounds. To correspond to the instrument sounds selected by the user (e.g., piano, guitar), the note data is converted using "MusicGenerator."

[0955] Input: Scale, rhythm, melody information, and selected instrument information

[0956] Output: Musical note data converted into instrument sounds (music data)

[0957] Step 5:

[0958] The server analyzes the user's emotions and reflects them in the generated music data. Using the emotion recognition engine "EmotionEngine," it analyzes the user's emotional state from the audio data. Based on that emotional state, the tempo and tone of the music are automatically adjusted.

[0959] Input: Digital voice data and analyzed emotional information

[0960] Output: Music data with tempo and tone adjusted to match the emotion

[0961] Step 6:

[0962] The generated and adjusted music data is sent back to the device. The server sends the data to the device via a transfer protocol.

[0963] Input: Music data generated and adjusted on the server

[0964] Output: Music data received on the device

[0965] Step 7:

[0966] The device plays and edits the received music data. The user can listen to the created music by pressing the play button. The edit button can also be used to change the tempo or modify the notes.

[0967] Input: Music data received on the device

[0968] Output: Played and edited song

[0969] Step 8:

[0970] Users can then play the final edited song as background music on a robot in the factory, which has a built-in speaker and plays the music based on the song data.

[0971] Input: Song data edited on the device

[0972] Output: Background music played in your work environment

[0973] Step 9:

[0974] The completed songs are uploaded to the community by the users, and other users provide feedback on the uploaded songs, which is collected by the server and sent to the device.

[0975] Input: Completed song data

[0976] Output: Songs shared with the community and given feedback

[0977] Specific prompt examples:

[0978] Prompt: "I need some calming music to help me concentrate."

[0979] In this way, a system is provided that generates music that reflects the user's humming or emotional state and plays it as appropriate background music in the factory.

[0980] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0981] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0982] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0983] [Third embodiment]

[0984] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0985] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0986] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0987] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0988] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0989] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0990] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0991] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0992] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0993] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0994] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0995] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0996] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention provides a system that allows a user to record a humming tune and then create, edit, and share a song based on the humming tune.

[0997] Overall system overview

[0998] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server analyzes the received audio data and generates scale, rhythm, and melody information. A song is then generated based on the sounds of the instrument selected by the user, and the generated song is sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[0999] Recording and Data Transmission

[1000] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[1001] Analysis and music generation on the server

[1002] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[1003] Next, the server converts the generated music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sound.

[1004] Sending and playing music data

[1005] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[1006] Editing features and phrase suggestions

[1007] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[1008] Community and Feedback

[1009] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[1010] Specific examples

[1011] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on the device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1012] As a result, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music.

[1013] The processing flow will be explained below.

[1014] Step 1:

[1015] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[1016] Step 2:

[1017] The device records the user's humming and saves it as audio data.

[1018] Step 3:

[1019] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[1020] Step 4:

[1021] The terminal transmits the voice data to the server via the Internet.

[1022] Step 5:

[1023] The server receives the voice data sent from the terminal and converts the voice data into a format for analysis.

[1024] Step 6:

[1025] The server uses an AI model to analyze the audio data and extract pitch, rhythm, and melody information.

[1026] Specifically, the AI ​​model analyzes audio waveforms to identify the pitch and timing of each note.

[1027] Step 7:

[1028] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[1029] For example, if the user selects piano, the note data is converted into piano sounds.

[1030] Step 8:

[1031] The server transmits the generated music data to the terminal.

[1032] Step 9:

[1033] The terminal receives the music data sent from the server and displays music playback options to the user.

[1034] Step 10:

[1035] The user presses the play button to listen to the generated music.

[1036] Step 11:

[1037] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[1038] Step 12:

[1039] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[1040] Step 13:

[1041] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[1042] Step 14:

[1043] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[1044] Step 15:

[1045] The device displays suggested phrases and adds the user's selection to the song.

[1046] Step 16:

[1047] The user presses the upload button to upload the completed song to the community.

[1048] Step 17:

[1049] The terminal transmits the completed song data to a community server, and the song is made public.

[1050] Step 18:

[1051] Other users in the community provide feedback on the published songs.

[1052] Step 19:

[1053] The server notifies the song poster of the collected feedback.

[1054] Step 20:

[1055] The user reviews the feedback and makes any necessary changes to the song.

[1056] Through the above series of steps, the present invention provides a system that allows anyone to easily enjoy composing music.

[1057] Example 1

[1058] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1059] Existing music production systems require advanced musical knowledge and complex operations, making it difficult for average users to easily create music. Other issues include a lack of functionality to instantly turn a user's original melody into digital data, and a lack of systems that can easily convert it into multiple instrument sounds. Furthermore, the lack of support for users when considering the next phrase makes it difficult to smoothly compose music.

[1060] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1061] In this invention, the server includes means for recording a user's voice and saving it as digital data, means for transmitting the digital data to the server, means for analyzing the digital data on the server and generating a scale, rhythm, and melody, means for converting the music data into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for automatically using AI to suggest next phrase candidates for the edited music data, and means for uploading music data to a community and receiving feedback. This enables general users to easily record, generate, and edit music, convert it into multiple instrument sounds, and receive feedback from other users to smoothly compose music.

[1062] A "user" is an entity that uses the system to record humming and create, edit, and share music.

[1063] A "terminal" is a device used by a user, and has functions such as recording, playback, editing, sending, and receiving.

[1064] "Digital data" refers to a collection of bits that have been converted to store the user's humming or voice electronically.

[1065] A "server" is a computing device that receives and analyzes digital data sent from a terminal, generates music data, and provides it to the user.

[1066] A "generative AI model" is an artificial intelligence algorithm that runs on a server and extracts scales, rhythms, and melodies from audio data.

[1067] "Music data" refers to digital music data generated based on scale, rhythm, and melody information analyzed by a generative AI model.

[1068] "Instrument sound" refers to the tone or sound produced when music data is converted into the sound of a specific instrument.

[1069] "Editing" refers to the modification or addition of music data that the user performs on the generated music data, and includes changing the tempo, modifying notes, adding harmonies, and the like.

[1070] "Phrase candidates" are options for the next melody line or harmony that the generative AI model automatically suggests for the music data being edited.

[1071] "Community" is an online platform for publishing user-generated compositions and receiving feedback from other users.

[1072] "Feedback" is the opinions and ratings provided by other users of the community, and is information that helps improve and refine your songs.

[1073] The present invention is a system that allows users to record their humming and then create, edit, and share music based on that recording.

[1074] This system consists of a terminal used by the user, a server, and software for linking them.

[1075] Recording preparation and humming

[1076] The user launches the application installed on the device. The device provides the user with a record button through the interface. When the user presses the record button, the device uses the built-in microphone to record the humming sound and saves it as digital data. This digital data is temporarily stored in the device.

[1077] Sending voice data and analyzing it on the server

[1078] Once the recording is complete, the device sends the digital data to a server, for example, using the HTTPS protocol. The server then analyzes the received digital data using an AI model (generative AI model). This model extracts scale, rhythm, and melody information and structures it as music data. Specifically, the pitch and timing of each note are detected from the audio data.

[1079] Music data generation and transmission

[1080] Next, the server converts the extracted music data into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sounds. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sounds. The converted music data is then sent back to the terminal. The terminal receives this data and notifies the user.

[1081] Playing and editing songs

[1082] Users can play the generated music by pressing the play button on their device. Furthermore, they can use the editing tools to change the tempo, modify the notes, or add harmonies as needed. This editing function allows users to create more refined music.

[1083] Phrase suggestion feature

[1084] If a user is unsure of the next phrase during the editing process, they can press the suggest button. The device then sends a request to the server, which uses an AI model to generate melody line and harmony suggestions. These suggestions are then presented to the user via their device, allowing them to select their favorite phrase and continue composing.

[1085] Community Uploads and Feedback

[1086] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song, and the server collects this feedback and notifies the song poster via their device. This allows the user to improve their song based on the opinions and ratings of other users.

[1087] Specific examples

[1088] For example, consider the case where a user hums "Happy Birthday." The user records the song on their device and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the user uploads the completed song to the community and receives feedback.

[1089] Prompt Sentence Examples

[1090] An example of a prompt to be input to the generative AI model is, "Please convert the recorded humming into instrument sounds and generate the song 'Happy Birthday.' Please also accurately determine the rhythm and scale and play it as a piano." Based on this prompt, the AI ​​model analyzes the humming audio data and generates a song using the specified instrument sounds.

[1091] As described above, by using the system of the present invention, users can easily create, edit, and share music even if they do not have advanced musical knowledge.

[1092] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1093] Step 1: Prepare to record

[1094] Input: User actions

[1095] Output: Recording ready state

[1096] Specific operation: The user launches the application on the device. The device displays the application's main screen and prepares for recording by displaying record, play, and edit buttons on the interface.

[1097] Step 2: Record your humming

[1098] Input: User's humming (audio data)

[1099] Output: Recorded digital data

[1100] Specific operation: When the user presses the record button, the device will use the built-in microphone to capture audio data in real time and save it as digital data. Once the recording is complete, the digital data will be temporarily stored on the device.

[1101] Step 3: Sending audio data

[1102] Input: Recorded digital data

[1103] Output: Notification of completion of transmission to the server

[1104] Specific operation: When the user finishes recording, the device compresses the saved audio data and sends it to the server using the HTTPS protocol. If the data is successfully sent, the server returns a successful reception response to the device.

[1105] Step 4: Analyzing the audio data

[1106] Input: Transmitted digital data

[1107] Output: Scale, rhythm, melody information

[1108] How it works: The server passes the received audio data to a generative AI model, which then analyzes the data for pitch and timing of each note, extracting information about the scale, rhythm, and melody.

[1109] Step 5: Generate music data

[1110] Input: scale, rhythm, melody information

[1111] Output: Music data converted into specific instrument sounds

[1112] Specific operation: The server converts the analyzed music data into the instrument sound selected by the user. For example, if the user selects piano, the AI ​​model converts the note data into piano sound. The converted music data is saved in a file format.

[1113] Step 6: Send your music

[1114] Input: Generated music data

[1115] Output: Notification of completion of transmission to the terminal

[1116] Specific operation: The server sends the converted music data back to the device. The device receives the data and displays "Music created" in the notification bar. The music data is also saved in the application.

[1117] Step 7: Play and edit your song

[1118] Input: Generated music data

[1119] Output: Played songs and edited song data

[1120] Specific operation: The user presses the play button to play a song. The device plays the specified song data, and when the user switches to edit mode, they can change the tempo, modify notes, add harmonies, etc. The edited song data is updated in real time.

[1121] Step 8: Use the Phrase Suggestion Feature

[1122] Input: Request to server and current music data

[1123] Output: Suggested phrase candidates

[1124] Specific operation: When the user presses the suggest button, the device sends the current song data to the server and requests suggestions for the next phrase. The server-side AI model generates multiple melody and harmony suggestions and sends them to the device. The user then selects from the suggested phrase suggestions through the device interface.

[1125] Step 9: Upload to the Community

[1126] Input: Completed song data

[1127] Output: Upload completion notification to the community

[1128] Specific operation: When a user presses the button to upload a song to the community, the device sends the song data to the server and makes it available within the community. The server stores the song and makes it accessible to other users. When the upload is successful, a notification is displayed on the device.

[1129] Step 10: Get feedback

[1130] Input: Feedback from other users

[1131] Output: Feedback notification and feedback content

[1132] What it does: Other users in the community add comments and ratings to uploaded songs. The server collects this feedback and sends it to the uploader's device as a notification. By tapping the notification, the user can view the feedback details.

[1133] As described above, by using this system, users can easily create, edit, and share music based on their humming.

[1134] (Application example 1)

[1135] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1136] In conventional music production systems, it was difficult for users without musical knowledge or skills to create and distribute music and share it with other users. They also lacked the functionality to receive feedback on the music they created and incorporate improvements. Furthermore, there was no system that provided users with the next phrase or idea during the music generation process based on humming, which increased the time and effort required for music production.

[1137] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1138] In this invention, the server includes means for recording a user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for uploading the generated music data to a content distribution service and sharing it with other users, means for receiving feedback from other users, and means for AI to automatically suggest the next phrase or idea. This allows even users with no musical knowledge or skills to easily create music from their humming, share it with other users, and receive feedback. Furthermore, AI's suggestions of the next phrase or idea significantly reduce the effort and time required for music production.

[1139] "User" refers to an individual or organization that uses the system to record humming, create music, and share it.

[1140] "Humming" refers to an informal singing voice with a melody or rhythm that is hummed by a user.

[1141] "Audio data" refers to data that is a digital recording of the user's humming.

[1142] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[1143] A "scale" is an arrangement of pitches that indicate the pitch of a sound, and is an element that makes up the melody of a piece of music.

[1144] "Rhythm" refers to the pattern of the placement and spacing of musical notes in time, which forms the tempo and beat of a piece of music.

[1145] A "melody" is the main theme of a piece of music, which is formed by combining scales and rhythms.

[1146] "Instrument sounds" are sounds produced by a particular instrument and are used to reproduce the scale, rhythm, and melody generated from humming.

[1147] A "terminal" refers to an electronic device used by a user, such as a computer, smartphone, or tablet.

[1148] A "content distribution service" is a platform for sharing and distributing music and other media content between users over the Internet.

[1149] "Feedback" means comments and ratings provided by other users within the community.

[1150] The present invention is a system that allows users to record their humming and then create, edit, and share music based on the audio data. The system mainly includes means for data communication between a terminal and a server and for data analysis. Specific embodiments of the present invention are described below.

[1151] Recording and Data Transmission

[1152] The user records their humming using an application installed on the device. At this time, the audio data is captured through the device's microphone and saved as digital data. After recording is complete, the audio data is sent from the device to the server. The software used includes a requests library for making HTTP requests.

[1153] Analysis and music generation on the server

[1154] The server receives the audio data sent from the device and analyzes the scale, rhythm, and melody information using an AI model. This analysis includes detecting the pitch and timing of musical notes from the audio data and structuring it as music data. Next, based on the analyzed music data, it converts it into the sound of an instrument selected by the user (e.g., piano). This conversion process involves analyzing the audio using an AI model and generating acoustic data. Audio analysis on the server is performed using an audio processing library (e.g., librosa).

[1155] Sending and playing music data

[1156] The generated music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. An audio data manipulation library (e.g., pydub) is used for playback. While listening to the music, the user can edit it by changing the tempo, modifying notes, adding harmonies, and so on.

[1157] Editing features and phrase suggestions

[1158] When editing a song on a device, if a user is unsure of the next phrase, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and presents them to the user via the device. This significantly reduces the effort required for music production.

[1159] Community and Feedback

[1160] The completed song is uploaded by the user to a community on the content distribution service. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. The user can then receive the feedback and make corrections to the song.

[1161] Examples of specific examples and prompts

[1162] For example, if a user hums "Happy Birthday," the device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1163] An example of a prompt for a generative AI model is:

[1164] "User sang 'Happy Birthday' with a humming voice. Convert this recording into a piano music file."

[1165] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1166] Step 1:

[1167] The user uses the terminal to record a hum.

[1168] Input: User's humming

[1169] Data processing: Capture and store audio data digitally through the device's microphone

[1170] Output: Digital audio data

[1171] Step 2:

[1172] The terminal transmits the recorded voice data to the server.

[1173] Input: Digital audio data

[1174] Data processing: Upload audio data to the server using an HTTP request

[1175] Output: Audio data stored on the server

[1176] Step 3:

[1177] The server receives the audio data and begins analyzing it.

[1178] Input: Audio data stored on the server

[1179] Data processing: Using audio analysis algorithms to analyze and extract pitch, rhythm, and melodic information

[1180] Output: Scale information, rhythm information, melody information

[1181] Step 4:

[1182] The server converts the analyzed music data into the sound of an instrument selected by the user.

[1183] Input: Scale information, rhythm information, melody information, and the type of instrument sound selected by the user

[1184] Data processing: Using a generative AI model to convert note data into selected instrument sounds

[1185] Output: Music data converted into instrument sounds

[1186] Step 5:

[1187] The server transmits the generated music data to the terminal.

[1188] Input: Music data converted into instrument sounds

[1189] Data processing: Send music data to the device using HTTP responses

[1190] Output: Song data stored on the device

[1191] Step 6:

[1192] The terminal allows the user to play the music data.

[1193] Input: Song data stored on the device

[1194] Data processing: Playing audio data using the pydub library

[1195] Output: The song played by the user

[1196] Step 7:

[1197] The user edits the music as needed.

[1198] Input: Song data to edit

[1199] Data processing: Editing operations such as changing the tempo, correcting notes, and adding harmonies can be performed through the application.

[1200] Output: Edited song data

[1201] Step 8:

[1202] The user uploads the completed song to a content distribution service.

[1203] Input: Completed song data

[1204] Data processing: Upload music data using HTTP requests and make it available to the community

[1205] Output: Songs published on content distribution services

[1206] Step 9:

[1207] Other users in the community provide feedback on the published songs.

[1208] Input: Published songs

[1209] Data Processing: Post a comment or rating

[1210] Output: Feedback provided

[1211] Step 10:

[1212] The server collects feedback from other users and notifies the song poster via the terminal.

[1213] Input: Feedback from other users

[1214] Data processing: Collect feedback data and notify song submitters

[1215] Output: Feedback notification received by song submitter

[1216] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1217] The following describes in detail an embodiment of the present invention: The present invention is a system that allows a user to record a humming tune and then create, edit, and share music based on the humming tune. It also incorporates an emotion engine that recognizes the user's emotional state and applies it to the creation and editing of music.

[1218] Overall system overview

[1219] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server then analyzes the received audio data and generates scale, rhythm, and melody information. The server then uses an emotion engine to analyze the user's emotions and automatically adjusts the mood and tempo of the music based on the emotion recognition results. A song is then generated based on the sounds of the instruments selected by the user and sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[1220] Recording and Data Transmission

[1221] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[1222] Analysis and music generation on the server

[1223] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[1224] The server then uses an emotion engine to analyze the user's emotional state. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and other biometric signals. Based on the emotion recognition results, the server automatically adjusts the mood and tempo of the music.

[1225] Furthermore, the extracted music data is converted into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the scale information is converted into note data corresponding to the respective instrument sound.

[1226] Sending and playing music data

[1227] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[1228] Editing features and phrase suggestions

[1229] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[1230] Community and Feedback

[1231] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[1232] Specific examples

[1233] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion, and if it recognizes it as "joy," for example, it speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1234] In this way, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music by generating music that takes the user's emotions into consideration.

[1235] The processing flow will be explained below.

[1236] Step 1:

[1237] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[1238] Step 2:

[1239] The device records the user's humming and saves it as audio data.

[1240] Step 3:

[1241] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[1242] Step 4:

[1243] The terminal transmits the voice data to the server via the Internet.

[1244] Step 5:

[1245] The server receives the voice data sent from the device and converts it into a format for analysis, which makes it easier to analyze.

[1246] Step 6:

[1247] The server uses AI models to analyze the audio data and extract pitch, rhythm, and melody information.

[1248] Example: Detecting pitch and beat from a recorded humming and converting it into musical note data.

[1249] Step 7:

[1250] Users provide emotional data, such as facial expressions and tone of voice, through their device's camera and microphone.

[1251] The device acquires the emotion data and transmits it to the server.

[1252] Step 8:

[1253] The server uses an emotion engine to analyze the received emotion data, thereby recognizing the user's emotional state (e.g., joy, sadness).

[1254] Example: AI analyzes facial expressions and recognizes that if the user is smiling, it is expressing "joy."

[1255] Step 9:

[1256] The server automatically adjusts the mood and tempo of the music based on the emotion recognition results. For example, if it recognizes "joy," it will speed up the tempo or add a more upbeat arrangement.

[1257] Step 10:

[1258] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[1259] For example, if the user selects piano, the note data is converted into piano sounds.

[1260] Step 11:

[1261] The server transmits the generated music data to the terminal.

[1262] Step 12:

[1263] The terminal receives the music data sent from the server and displays music playback options to the user.

[1264] Step 13:

[1265] The user presses the play button to listen to the generated music.

[1266] Step 14:

[1267] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[1268] Step 15:

[1269] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[1270] Step 16:

[1271] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[1272] Step 17:

[1273] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[1274] Step 18:

[1275] The device displays suggested phrases and adds the user's selection to the song.

[1276] Step 19:

[1277] The user presses the upload button to upload the completed song to the community.

[1278] Step 20:

[1279] The terminal transmits the completed song data to a community server, and the song is made public.

[1280] Step 21:

[1281] Other users in the community provide feedback on the published songs.

[1282] Step 22:

[1283] The server notifies the song poster of the collected feedback.

[1284] Step 23:

[1285] The user reviews the feedback and makes any necessary changes to the song.

[1286] Through the above series of steps, the present invention allows anyone to easily enjoy composing music, and by taking the user's emotions into consideration when generating music, it provides a more personal music-making experience.

[1287] Example 2

[1288] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1289] Conventional music generation systems require users to manually edit music, often requiring technical knowledge. Furthermore, they are unable to generate music based on the user's emotional state, making it difficult to generate optimal music that matches individual emotions. Furthermore, there is no function to automatically suggest the next phrase or idea for a generated piece of music. There is a need for a system that can resolve these issues and allow users to generate music more easily and based on their emotions.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data using a generative AI model to generate a scale, rhythm, and melody, means for analyzing the user's emotional state and automatically adjusting the mood and tempo of the music, and means for automatically suggesting the user's next phrase or idea using the generative AI model via a suggestion button. This allows users to easily generate music without technical knowledge, and the music is optimized according to the user's emotional state, enabling more satisfying music production. Furthermore, the automatic suggestion of the next phrase or idea can support the user's creativity.

[1291] "User" means an individual who uses the System to record humming and create, edit, and share music.

[1292] "Terminal" refers to a communication device used by a user to record humming, including a smartphone, tablet, etc.

[1293] A "server" is a computer system that receives audio data sent by a user, analyzes it, and creates music.

[1294] "Audio data" refers to information stored in digital form of a user's humming.

[1295] A "generative AI model" is an artificial intelligence algorithm that analyzes audio data to generate scales, rhythms, and melodies.

[1296] A "scale" refers to the arrangement of pitches that make up the melody of a piece of music.

[1297] "Rhythm" refers to the pattern of timing and duration of notes in a piece of music.

[1298] "Melody" refers to the main theme of a piece of music that is formed by combining scales and rhythms.

[1299] The "emotion engine" is an algorithm that analyzes the user's emotional state and adjusts the mood and tempo of the music based on the results.

[1300] "Instrumental sounds" refers to sounds produced by different musical instruments such as piano, guitar, violin, drums, etc.

[1301] "Music Data" means music information in digital form that has been generated through an analysis and conversion process.

[1302] "Suggestion button" refers to an interface element that a user uses to request an automatic suggestion of the next phrase or idea.

[1303] "Community" means the online platform where users can upload their created Music and receive feedback from other users.

[1304] "Feedback" means ratings and opinions provided by other users within the Community regarding the Generated Song.

[1305] The system of the present invention allows users to record their humming and then use it to create, edit, and share music, and is characterized by using an emotion engine to apply the user's emotional state to the music creation. Below, we will explain in detail how this system is implemented.

[1306] Recording and Data Transmission

[1307] Users record their humming using an application installed on their smartphone, tablet, or other device. When they tap the record button, the device captures the audio data through the built-in microphone and saves it as digital data in PCM or AAC format. Once recording is complete, the device sends the audio data to a server via the Internet.

[1308] Hardware used: Smartphone, tablet

[1309] Software used: Recording application

[1310] Analysis and music generation on the server

[1311] The server receives the audio data sent from the device and uses a generative AI model to analyze the audio data and extract musical scale, rhythm, and melody information. This analysis involves using a speech recognition algorithm to detect the pitch and timing of musical notes.

[1312] The server then analyzes the user's emotional state using an emotion engine, which analyzes the tone of the recorded voice and facial expression images provided by the user, and automatically adjusts the mood and tempo of the music based on the emotion recognition results.

[1313] Finally, the server converts the extracted music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If another instrument (e.g., guitar or violin) is selected, the information is converted into note data corresponding to the instrument sound.

[1314] Hardware used: Server

[1315] Software used: Generative AI model, emotion engine

[1316] Sending and playing music data

[1317] The server then sends the generated music data back to the terminal, which receives it and displays it on its user interface. The user can listen to the music by pressing the play button.

[1318] Hardware used: Server, terminal

[1319] Software used: Music playback application

[1320] Editing features and phrase suggestions

[1321] Users can edit songs through the application, for example, by changing the tempo, modifying specific notes, or adding harmonies. If users are unsure of the next phrase during editing, they can press the suggestion button. This causes the server to use a generative AI model to generate several melody line and harmony candidates and suggest them to the user via their device.

[1322] Hardware used: Server, terminal

[1323] Software used: Music editing application, generative AI model

[1324] Community and Feedback

[1325] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device, allowing the user to receive opinions and ratings from other users and make corrections to the song.

[1326] Hardware used: Server, terminal

[1327] Software used: Community Platform

[1328] Specific examples

[1329] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server uses a generative AI model to analyze the scale and extract the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion and, if it recognizes it as "joy," for example, speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1330] Example prompt sentence:

[1331] "I'm humming Happy Birthday. Identify the user's emotion as joy, speed up the tempo, and emphasize the bright tone. Finally, convert it into a piano sound."

[1332] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1333] Step 1:

[1334] The user launches the device's recording application and taps the record button. The device uses the built-in microphone to capture the user's humming and saves it as digital audio data in, for example, PCM or AAC format. The input is the user's humming, and the output is the captured audio data. This recording is saved as a temporary file in the device's storage.

[1335] Specific behavior:

[1336] User taps the record button

[1337] The device captures audio through the microphone

[1338] The device generates and stores digital audio data

[1339] Step 2:

[1340] After recording is complete, the device sends the audio data to the server via the Internet. The input is the stored audio data, and the output is the audio data sent to the server. The data is transferred securely using the HTTP protocol or WebSocket.

[1341] Specific behavior:

[1342] The device is waiting for audio data

[1343] The device sends the voice data to the server

[1344] Step 3:

[1345] The server analyzes the received audio data. First, it uses a generative AI model to extract scale, rhythm, and melody from the audio data. The input is audio data, and the output is structured musical data. It uses a speech recognition algorithm (e.g., a deep learning model) to detect the pitch and timing of each note and stores this in a database format.

[1346] Specific behavior:

[1347] The server receives the audio data

[1348] The server applies generative AI models to extract scale, rhythm, and melody

[1349] The server structures and stores music data

[1350] Step 4:

[1351] The server then analyzes the user's emotional state using an emotion engine. The emotion engine recognizes emotions by analyzing voice tone and provided facial expression images. The input is voice data and facial expression data, and the output is the recognized emotional state. The emotion engine uses machine learning algorithms to determine emotions and incorporates this information into the music data.

[1352] Specific behavior:

[1353] The server applies the emotion engine

[1354] The server analyzes voice tone and facial expression data

[1355] The server recognizes and records the user's emotional state.

[1356] Step 5:

[1357] The server converts the generated music data into the instrument sounds selected by the user. If the user selects piano, it is converted into piano sounds. The input is structured music data and instrument selection information, and the output is music data converted into specific instrument sounds. The instrument sounds are synthesized using sound fonts and audio sampling.

[1358] Specific behavior:

[1359] User selects instrument

[1360] The server converts the music data into instrument sounds

[1361] Step 6:

[1362] The server sends the converted music data to the terminal. The input is the music data, and the output is the music data sent to the terminal. The data is transferred via a secure communication channel.

[1363] Specific behavior:

[1364] The server waits for music data

[1365] The server sends the music data to the device

[1366] Step 7:

[1367] The terminal displays the received music data on the user interface, and the user can listen to the music by pressing the play button. The input is the received music data, and the output is the played music.

[1368] Specific behavior:

[1369] The device receives the music data

[1370] User taps the play button

[1371] The device plays the music

[1372] Step 8:

[1373] The user edits the music through the application, changing the tempo, modifying specific notes, adding harmonies, etc. The input is the music data and editing operations, and the output is the edited music data.

[1374] Specific behavior:

[1375] User uses editing functions

[1376] The device reflects the edited content in the song data

[1377] Step 9:

[1378] When a user is unsure of the next phrase, they press the suggest button. The server uses a generative AI model to generate several melody line and harmony candidates and suggests them to the user via the device. The input is a suggestion request, and the output is the suggested melody or harmony.

[1379] Specific behavior:

[1380] User taps the suggest button

[1381] Server generates candidates

[1382] Your device will display suggestions

[1383] Step 10:

[1384] A completed song is uploaded to the community by the user. Other users in the community can provide feedback on this song. The input is the completed song data, and the output is the feedback from the user. The server collects the feedback and notifies the song uploader via their device.

[1385] Specific behavior:

[1386] Users upload songs to the community

[1387] Other users provide feedback

[1388] Server collects and notifies feedback

[1389] (Application example 2)

[1390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1391] Previous technologies have not provided sufficient concrete methods for improving work efficiency and worker morale in factories. Furthermore, systems that automatically generate music suited to the work environment have not been able to combine it with emotion recognition technology. Therefore, there is a need for technology that automatically generates music suited to the work environment while taking into account the user's emotions.

[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for converting music into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music suitable for the work environment, means for playing the background music on robots in the factory, and means for uploading music to the community and receiving feedback. This makes it possible to generate appropriate background music based on the user's humming and emotions, thereby improving work efficiency and morale in the factory.

[1393] definition statement

[1394] The "means for recording a user's humming and saving it as audio data" is a device or program that captures the audio of a user humming as digital data and saves it for later processing.

[1395] The "means for transmitting the audio data to the server" is a system that transfers the recorded audio data to a remote server via the Internet or a local network.

[1396] The "means for analyzing audio data on a server and generating scales, rhythms, and melodies" refers to a server system that performs processing to extract highly accurate musical elements based on audio data.

[1397] The "means for converting a piece of music into various instrument sounds based on the generated scale, rhythm, and melody" is a system that uses the analyzed musical elements to convert note data into different instrument sounds selected by the user.

[1398] "Means for transmitting generated music data to a terminal and enabling playback and editing" refers to a system for transferring music data generated on a server to a user's operating terminal, allowing the data to be played back and further edited.

[1399] "Means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music appropriate for the work environment" refers to a system that analyzes the user's emotional state and automatically adjusts the tempo and tone of the music to an appropriate level based on the results.

[1400] The "means for playing the background music by a robot in a factory" refers to a robot system used to play the generated background music in a physical space.

[1401] "Means for uploading music to a community and receiving feedback" is a system that allows users to share music they have created with an online community and collect ratings and comments from other users.

[1402] MODE FOR CARRYING OUT THE INVENTION

[1403] The embodiment of the present invention is a music generation system aimed at improving work efficiency and worker morale in a factory. This system allows a user to record a humming tune and then generates, edits, and plays appropriate background music based on that humming. The system also recognizes the user's emotions and reflects them in the music it generates, providing music that is optimal for the work environment.

[1404] The server processes and calculates data using the following hardware and software: The "librosa" library is used for analyzing audio data, the "EmotionEngine" is used for emotion recognition, and the "MusicGenerator" is used for music generation.

[1405] System operation explanation

[1406] Humming recording and data transmission

[1407] A user records their humming using a terminal equipped with a recording function. The recorded audio data is saved as digital data by the terminal. This digital data is then transmitted to a server via a network.

[1408] Analysis and music generation on the server

[1409] The server receives the transmitted audio data and uses the librosa library to analyze the pitch and timing of each note, extracting scale, rhythm, and melody information. It then uses the Emotion Engine to analyze the user's emotions. Based on this emotional analysis, the tempo and timbre of the music are automatically adjusted.

[1410] Based on the analyzed musical scale information, the "Music Generator" is used to convert it into an appropriate instrument sound. At this stage, the note data is converted based on the instrument sound selected by the user (e.g. piano, guitar, drums, etc.).

[1411] Sending and playing music data

[1412] The generated music data is then sent back to the terminal. The terminal displays the received music data, and the user can listen to the music by pressing the play button. The user can also edit the music by changing the tempo or correcting the notes. In particular, it is possible to use robots in factories to play background music.

[1413] Community Features and Feedback

[1414] Users can upload their created compositions to the community and receive feedback from other users, allowing for further improvements and suggestions for new ideas.

[1415] Examples and prompts

[1416] As a specific example, let us consider a case where a user working in a factory hums, saying, "I want to concentrate, so I want some calming music." The tempo and tone of the music generated from this humming are adjusted appropriately based on the tone of the user's voice and emotional analysis.

[1417] Prompt Sentence Examples

[1418] Prompt: "I need some calming music to help me concentrate."

[1419] In this way, the present invention provides a music generation system that aims to improve work efficiency and morale in factories. Furthermore, by generating music that takes emotions into consideration, it is possible to provide appropriate music that matches the user's psychological state.

[1420] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1421] System program processing flow

[1422] Step 1:

[1423] The user uses the device to record their humming. When they press the record button, the audio is captured through the device's microphone and saved as digital audio data.

[1424] Input: User's humming (audio)

[1425] Output: Digital audio data (.wav, .mp3, etc.)

[1426] Step 2:

[1427] The stored digital audio data is sent from the device to the server, where it is transferred securely and quickly using a network protocol for transmission (e.g., HTTP, FTP).

[1428] Input: Digital audio data

[1429] Output: Digital audio data received by the server

[1430] Step 3:

[1431] The server analyzes the received audio data and generates scale, rhythm, and melody information using the audio data analysis library "librosa," which detects the pitch and timing of each note.

[1432] Input: Digital audio data received by the server

[1433] Output: Scale, rhythm, melody information (structured data)

[1434] Step 4:

[1435] Based on the analyzed scale, rhythm, and melody information, the note data is converted into musical instrument sounds. To correspond to the instrument sounds selected by the user (e.g., piano, guitar), the note data is converted using "MusicGenerator."

[1436] Input: Scale, rhythm, melody information, and selected instrument information

[1437] Output: Musical note data converted into instrument sounds (music data)

[1438] Step 5:

[1439] The server analyzes the user's emotions and reflects them in the generated music data. Using the emotion recognition engine "EmotionEngine," it analyzes the user's emotional state from the audio data. Based on that emotional state, the tempo and tone of the music are automatically adjusted.

[1440] Input: Digital voice data and analyzed emotional information

[1441] Output: Music data with tempo and tone adjusted to match the emotion

[1442] Step 6:

[1443] The generated and adjusted music data is sent back to the device. The server sends the data to the device via a transfer protocol.

[1444] Input: Music data generated and adjusted on the server

[1445] Output: Music data received on the device

[1446] Step 7:

[1447] The device plays and edits the received music data. The user can listen to the created music by pressing the play button. The edit button can also be used to change the tempo or modify the notes.

[1448] Input: Music data received on the device

[1449] Output: Played and edited song

[1450] Step 8:

[1451] Users can then play the final edited song as background music on a robot in the factory, which has a built-in speaker and plays the music based on the song data.

[1452] Input: Song data edited on the device

[1453] Output: Background music played in your work environment

[1454] Step 9:

[1455] The completed songs are uploaded to the community by the users, and other users provide feedback on the uploaded songs, which is collected by the server and sent to the device.

[1456] Input: Completed song data

[1457] Output: Songs shared with the community and given feedback

[1458] Specific prompt examples:

[1459] Prompt: "I need some calming music to help me concentrate."

[1460] In this way, a system is provided that generates music that reflects the user's humming or emotional state and plays it as appropriate background music in the factory.

[1461] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1462] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1463] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1464] [Fourth embodiment]

[1465] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1466] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1467] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1468] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1469] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1470] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1471] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1472] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1473] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1474] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1475] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1476] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1477] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1478] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention provides a system that allows a user to record a humming tune and then create, edit, and share a song based on the humming tune.

[1479] Overall system overview

[1480] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server analyzes the received audio data and generates scale, rhythm, and melody information. A song is then generated based on the sounds of the instrument selected by the user, and the generated song is sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[1481] Recording and Data Transmission

[1482] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[1483] Analysis and music generation on the server

[1484] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[1485] Next, the server converts the generated music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sound.

[1486] Sending and playing music data

[1487] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[1488] Editing features and phrase suggestions

[1489] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[1490] Community and Feedback

[1491] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[1492] Specific examples

[1493] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on the device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1494] As a result, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music.

[1495] The processing flow will be explained below.

[1496] Step 1:

[1497] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[1498] Step 2:

[1499] The device records the user's humming and saves it as audio data.

[1500] Step 3:

[1501] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[1502] Step 4:

[1503] The terminal transmits the voice data to the server via the Internet.

[1504] Step 5:

[1505] The server receives the voice data sent from the terminal and converts the voice data into a format for analysis.

[1506] Step 6:

[1507] The server uses an AI model to analyze the audio data and extract pitch, rhythm, and melody information.

[1508] Specifically, the AI ​​model analyzes audio waveforms to identify the pitch and timing of each note.

[1509] Step 7:

[1510] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[1511] For example, if the user selects piano, the note data is converted into piano sounds.

[1512] Step 8:

[1513] The server transmits the generated music data to the terminal.

[1514] Step 9:

[1515] The terminal receives the music data sent from the server and displays music playback options to the user.

[1516] Step 10:

[1517] The user presses the play button to listen to the generated music.

[1518] Step 11:

[1519] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[1520] Step 12:

[1521] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[1522] Step 13:

[1523] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[1524] Step 14:

[1525] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[1526] Step 15:

[1527] The device displays suggested phrases and adds the user's selection to the song.

[1528] Step 16:

[1529] The user presses the upload button to upload the completed song to the community.

[1530] Step 17:

[1531] The terminal transmits the completed song data to a community server, and the song is made public.

[1532] Step 18:

[1533] Other users in the community provide feedback on the published songs.

[1534] Step 19:

[1535] The server notifies the song poster of the collected feedback.

[1536] Step 20:

[1537] The user reviews the feedback and makes any necessary changes to the song.

[1538] Through the above series of steps, the present invention provides a system that allows anyone to easily enjoy composing music.

[1539] Example 1

[1540] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1541] Existing music production systems require advanced musical knowledge and complex operations, making it difficult for average users to easily create music. Other issues include a lack of functionality to instantly turn a user's original melody into digital data, and a lack of systems that can easily convert it into multiple instrument sounds. Furthermore, the lack of support for users when considering the next phrase makes it difficult to smoothly compose music.

[1542] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1543] In this invention, the server includes means for recording a user's voice and saving it as digital data, means for transmitting the digital data to the server, means for analyzing the digital data on the server and generating a scale, rhythm, and melody, means for converting the music data into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for automatically using AI to suggest next phrase candidates for the edited music data, and means for uploading music data to a community and receiving feedback. This enables general users to easily record, generate, and edit music, convert it into multiple instrument sounds, and receive feedback from other users to smoothly compose music.

[1544] A "user" is an entity that uses the system to record humming and create, edit, and share music.

[1545] A "terminal" is a device used by a user, and has functions such as recording, playback, editing, sending, and receiving.

[1546] "Digital data" refers to a collection of bits that have been converted to store the user's humming or voice electronically.

[1547] A "server" is a computing device that receives and analyzes digital data sent from a terminal, generates music data, and provides it to the user.

[1548] A "generative AI model" is an artificial intelligence algorithm that runs on a server and extracts scales, rhythms, and melodies from audio data.

[1549] "Music data" refers to digital music data generated based on scale, rhythm, and melody information analyzed by a generative AI model.

[1550] "Instrument sound" refers to the tone or sound produced when music data is converted into the sound of a specific instrument.

[1551] "Editing" refers to the modification or addition of music data that the user performs on the generated music data, and includes changing the tempo, modifying notes, adding harmonies, and the like.

[1552] "Phrase candidates" are options for the next melody line or harmony that the generative AI model automatically suggests for the music data being edited.

[1553] "Community" is an online platform for publishing user-generated compositions and receiving feedback from other users.

[1554] "Feedback" is the opinions and ratings provided by other users of the community, and is information that helps improve and refine your songs.

[1555] The present invention is a system that allows users to record their humming and then create, edit, and share music based on that recording.

[1556] This system consists of a terminal used by the user, a server, and software for linking them.

[1557] Recording preparation and humming

[1558] The user launches the application installed on the device. The device provides the user with a record button through the interface. When the user presses the record button, the device uses the built-in microphone to record the humming sound and saves it as digital data. This digital data is temporarily stored in the device.

[1559] Sending voice data and analyzing it on the server

[1560] Once the recording is complete, the device sends the digital data to a server, for example, using the HTTPS protocol. The server then analyzes the received digital data using an AI model (generative AI model). This model extracts scale, rhythm, and melody information and structures it as music data. Specifically, the pitch and timing of each note are detected from the audio data.

[1561] Music data generation and transmission

[1562] Next, the server converts the extracted music data into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sounds. If the user selects another instrument (e.g., guitar or violin), the information is converted into note data corresponding to the respective instrument sounds. The converted music data is then sent back to the terminal. The terminal receives this data and notifies the user.

[1563] Playing and editing songs

[1564] Users can play the generated music by pressing the play button on their device. Furthermore, they can use the editing tools to change the tempo, modify the notes, or add harmonies as needed. This editing function allows users to create more refined music.

[1565] Phrase suggestion feature

[1566] If a user is unsure of the next phrase during the editing process, they can press the suggest button. The device then sends a request to the server, which uses an AI model to generate melody line and harmony suggestions. These suggestions are then presented to the user via their device, allowing them to select their favorite phrase and continue composing.

[1567] Community Uploads and Feedback

[1568] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song, and the server collects this feedback and notifies the song poster via their device. This allows the user to improve their song based on the opinions and ratings of other users.

[1569] Specific examples

[1570] For example, consider the case where a user hums "Happy Birthday." The user records the song on their device and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the user uploads the completed song to the community and receives feedback.

[1571] Prompt Sentence Examples

[1572] An example of a prompt to be input to the generative AI model is, "Please convert the recorded humming into instrument sounds and generate the song 'Happy Birthday.' Please also accurately determine the rhythm and scale and play it as a piano." Based on this prompt, the AI ​​model analyzes the humming audio data and generates a song using the specified instrument sounds.

[1573] As described above, by using the system of the present invention, users can easily create, edit, and share music even if they do not have advanced musical knowledge.

[1574] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1575] Step 1: Prepare to record

[1576] Input: User actions

[1577] Output: Recording ready state

[1578] Specific operation: The user launches the application on the device. The device displays the application's main screen and prepares for recording by displaying record, play, and edit buttons on the interface.

[1579] Step 2: Record your humming

[1580] Input: User's humming (audio data)

[1581] Output: Recorded digital data

[1582] Specific operation: When the user presses the record button, the device will use the built-in microphone to capture audio data in real time and save it as digital data. Once the recording is complete, the digital data will be temporarily stored on the device.

[1583] Step 3: Sending audio data

[1584] Input: Recorded digital data

[1585] Output: Notification of completion of transmission to the server

[1586] Specific operation: When the user finishes recording, the device compresses the saved audio data and sends it to the server using the HTTPS protocol. If the data is successfully sent, the server returns a successful reception response to the device.

[1587] Step 4: Analyzing the audio data

[1588] Input: Transmitted digital data

[1589] Output: Scale, rhythm, melody information

[1590] How it works: The server passes the received audio data to a generative AI model, which then analyzes the data for pitch and timing of each note, extracting information about the scale, rhythm, and melody.

[1591] Step 5: Generate music data

[1592] Input: scale, rhythm, melody information

[1593] Output: Music data converted into specific instrument sounds

[1594] Specific operation: The server converts the analyzed music data into the instrument sound selected by the user. For example, if the user selects piano, the AI ​​model converts the note data into piano sound. The converted music data is saved in a file format.

[1595] Step 6: Send your music

[1596] Input: Generated music data

[1597] Output: Notification of completion of transmission to the terminal

[1598] Specific operation: The server sends the converted music data back to the device. The device receives the data and displays "Music created" in the notification bar. The music data is also saved in the application.

[1599] Step 7: Play and edit your song

[1600] Input: Generated music data

[1601] Output: Played songs and edited song data

[1602] Specific operation: The user presses the play button to play a song. The device plays the specified song data, and when the user switches to edit mode, they can change the tempo, modify notes, add harmonies, etc. The edited song data is updated in real time.

[1603] Step 8: Use the Phrase Suggestion Feature

[1604] Input: Request to server and current music data

[1605] Output: Suggested phrase candidates

[1606] Specific operation: When the user presses the suggest button, the device sends the current song data to the server and requests suggestions for the next phrase. The server-side AI model generates multiple melody and harmony suggestions and sends them to the device. The user then selects from the suggested phrase suggestions through the device interface.

[1607] Step 9: Upload to the Community

[1608] Input: Completed song data

[1609] Output: Upload completion notification to the community

[1610] Specific operation: When a user presses the button to upload a song to the community, the device sends the song data to the server and makes it available within the community. The server stores the song and makes it accessible to other users. When the upload is successful, a notification is displayed on the device.

[1611] Step 10: Get feedback

[1612] Input: Feedback from other users

[1613] Output: Feedback notification and feedback content

[1614] What it does: Other users in the community add comments and ratings to uploaded songs. The server collects this feedback and sends it to the uploader's device as a notification. By tapping the notification, the user can view the feedback details.

[1615] As described above, by using this system, users can easily create, edit, and share music based on their humming.

[1616] (Application example 1)

[1617] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1618] In conventional music production systems, it was difficult for users without musical knowledge or skills to create and distribute music and share it with other users. They also lacked the functionality to receive feedback on the music they created and incorporate improvements. Furthermore, there was no system that provided users with the next phrase or idea during the music generation process based on humming, which increased the time and effort required for music production.

[1619] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1620] In this invention, the server includes means for recording a user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for uploading the generated music data to a content distribution service and sharing it with other users, means for receiving feedback from other users, and means for AI to automatically suggest the next phrase or idea. This allows even users with no musical knowledge or skills to easily create music from their humming, share it with other users, and receive feedback. Furthermore, AI's suggestions of the next phrase or idea significantly reduce the effort and time required for music production.

[1621] "User" refers to an individual or organization that uses the system to record humming, create music, and share it.

[1622] "Humming" refers to an informal singing voice with a melody or rhythm that is hummed by a user.

[1623] "Audio data" refers to data that is a digital recording of the user's humming.

[1624] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[1625] A "scale" is an arrangement of pitches that indicate the pitch of a sound, and is an element that makes up the melody of a piece of music.

[1626] "Rhythm" refers to the pattern of the placement and spacing of musical notes in time, which forms the tempo and beat of a piece of music.

[1627] A "melody" is the main theme of a piece of music, which is formed by combining scales and rhythms.

[1628] "Instrument sounds" are sounds produced by a particular instrument and are used to reproduce the scale, rhythm, and melody generated from humming.

[1629] A "terminal" refers to an electronic device used by a user, such as a computer, smartphone, or tablet.

[1630] A "content distribution service" is a platform for sharing and distributing music and other media content between users over the Internet.

[1631] "Feedback" means comments and ratings provided by other users within the community.

[1632] The present invention is a system that allows users to record their humming and then create, edit, and share music based on the audio data. The system mainly includes means for data communication between a terminal and a server and for data analysis. Specific embodiments of the present invention are described below.

[1633] Recording and Data Transmission

[1634] The user records their humming using an application installed on the device. At this time, the audio data is captured through the device's microphone and saved as digital data. After recording is complete, the audio data is sent from the device to the server. The software used includes a requests library for making HTTP requests.

[1635] Analysis and music generation on the server

[1636] The server receives the audio data sent from the device and analyzes the scale, rhythm, and melody information using an AI model. This analysis includes detecting the pitch and timing of musical notes from the audio data and structuring it as music data. Next, based on the analyzed music data, it converts it into the sound of an instrument selected by the user (e.g., piano). This conversion process involves analyzing the audio using an AI model and generating acoustic data. Audio analysis on the server is performed using an audio processing library (e.g., librosa).

[1637] Sending and playing music data

[1638] The generated music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. An audio data manipulation library (e.g., pydub) is used for playback. While listening to the music, the user can edit it by changing the tempo, modifying notes, adding harmonies, and so on.

[1639] Editing features and phrase suggestions

[1640] When editing a song on a device, if a user is unsure of the next phrase, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and presents them to the user via the device. This significantly reduces the effort required for music production.

[1641] Community and Feedback

[1642] The completed song is uploaded by the user to a community on the content distribution service. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. The user can then receive the feedback and make corrections to the song.

[1643] Examples of specific examples and prompts

[1644] For example, if a user hums "Happy Birthday," the device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1645] An example of a prompt for a generative AI model is:

[1646] "User sang 'Happy Birthday' with a humming voice. Convert this recording into a piano music file."

[1647] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1648] Step 1:

[1649] The user uses the terminal to record a hum.

[1650] Input: User's humming

[1651] Data processing: Capture and store audio data digitally through the device's microphone

[1652] Output: Digital audio data

[1653] Step 2:

[1654] The terminal transmits the recorded voice data to the server.

[1655] Input: Digital audio data

[1656] Data processing: Upload audio data to the server using an HTTP request

[1657] Output: Audio data stored on the server

[1658] Step 3:

[1659] The server receives the audio data and begins analyzing it.

[1660] Input: Audio data stored on the server

[1661] Data processing: Using audio analysis algorithms to analyze and extract pitch, rhythm, and melodic information

[1662] Output: Scale information, rhythm information, melody information

[1663] Step 4:

[1664] The server converts the analyzed music data into the sound of an instrument selected by the user.

[1665] Input: Scale information, rhythm information, melody information, and the type of instrument sound selected by the user

[1666] Data processing: Using a generative AI model to convert note data into selected instrument sounds

[1667] Output: Music data converted into instrument sounds

[1668] Step 5:

[1669] The server transmits the generated music data to the terminal.

[1670] Input: Music data converted into instrument sounds

[1671] Data processing: Send music data to the device using HTTP responses

[1672] Output: Song data stored on the device

[1673] Step 6:

[1674] The terminal allows the user to play the music data.

[1675] Input: Song data stored on the device

[1676] Data processing: Playing audio data using the pydub library

[1677] Output: The song played by the user

[1678] Step 7:

[1679] The user edits the music as needed.

[1680] Input: Song data to edit

[1681] Data processing: Editing operations such as changing the tempo, correcting notes, and adding harmonies can be performed through the application.

[1682] Output: Edited song data

[1683] Step 8:

[1684] The user uploads the completed song to a content distribution service.

[1685] Input: Completed song data

[1686] Data processing: Upload music data using HTTP requests and make it available to the community

[1687] Output: Songs published on content distribution services

[1688] Step 9:

[1689] Other users in the community provide feedback on the published songs.

[1690] Input: Published songs

[1691] Data Processing: Post a comment or rating

[1692] Output: Feedback provided

[1693] Step 10:

[1694] The server collects feedback from other users and notifies the song poster via the terminal.

[1695] Input: Feedback from other users

[1696] Data processing: Collect feedback data and notify song submitters

[1697] Output: Feedback notification received by song submitter

[1698] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1699] The following describes in detail an embodiment of the present invention: The present invention is a system that allows a user to record a humming tune and then create, edit, and share music based on the humming tune. It also incorporates an emotion engine that recognizes the user's emotional state and applies it to the creation and editing of music.

[1700] Overall system overview

[1701] The system begins when a user uses a device to record themselves humming and sends the audio data to a server. The server then analyzes the received audio data and generates scale, rhythm, and melody information. The server then uses an emotion engine to analyze the user's emotions and automatically adjusts the mood and tempo of the music based on the emotion recognition results. A song is then generated based on the sounds of the instruments selected by the user and sent back to the device. The user can then play and edit the generated song on their device, and by uploading the completed song to the community, they can receive feedback from other users.

[1702] Recording and Data Transmission

[1703] The user records their humming using the application installed on the device. When the user presses the record button, the device captures the audio data through the microphone and stores it as digital data. After the recording is complete, the device sends the audio data to the server.

[1704] Analysis and music generation on the server

[1705] The server receives the audio data sent from the device and analyzes it using an AI model. This analysis extracts information about the scale, rhythm, and melody. Specifically, the pitch and timing of each note are detected from the audio data, and this information is structured as music data.

[1706] The server then uses an emotion engine to analyze the user's emotional state. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and other biometric signals. Based on the emotion recognition results, the server automatically adjusts the mood and tempo of the music.

[1707] Furthermore, the extracted music data is converted into the sound of the instrument selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If the user selects another instrument (e.g., guitar or violin), the scale information is converted into note data corresponding to the respective instrument sound.

[1708] Sending and playing music data

[1709] The converted music data is then sent back to the device from the server. The device displays the received music data to the user, who can listen to the music by clicking the play button. The user can then edit the music as needed while listening to it.

[1710] Editing features and phrase suggestions

[1711] Users can edit songs on their devices. For example, they can change the tempo, modify notes, and add additional harmonies. If they are unsure of the next phrase during the editing process, they can press the suggestion button. The server then uses AI to generate several melody line and harmony candidates and present them to the user via their device.

[1712] Community and Feedback

[1713] The completed song is uploaded to the community by the user. Other users of the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device. This allows the user to receive opinions and ratings from other users and, in some cases, make corrections to the song.

[1714] Specific examples

[1715] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server analyzes the scale and extracts the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion, and if it recognizes it as "joy," for example, it speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1716] In this way, the present invention significantly reduces the technical barriers to music production and provides a system that allows anyone to easily enjoy composing music by generating music that takes the user's emotions into consideration.

[1717] The processing flow will be explained below.

[1718] Step 1:

[1719] The user launches the application on their device and presses the record button, which starts recording and captures their humming through the microphone.

[1720] Step 2:

[1721] The device records the user's humming and saves it as audio data.

[1722] Step 3:

[1723] When the user is done recording, they press the stop button, which causes the device to stop recording and make a request to send the saved audio data to the server.

[1724] Step 4:

[1725] The terminal transmits the voice data to the server via the Internet.

[1726] Step 5:

[1727] The server receives the voice data sent from the device and converts it into a format for analysis, which makes it easier to analyze.

[1728] Step 6:

[1729] The server uses AI models to analyze the audio data and extract pitch, rhythm, and melody information.

[1730] Example: Detecting pitch and beat from a recorded humming and converting it into musical note data.

[1731] Step 7:

[1732] Users provide emotional data, such as facial expressions and tone of voice, through their device's camera and microphone.

[1733] The device acquires the emotion data and transmits it to the server.

[1734] Step 8:

[1735] The server uses an emotion engine to analyze the received emotion data, thereby recognizing the user's emotional state (e.g., joy, sadness).

[1736] Example: AI analyzes facial expressions and recognizes that if the user is smiling, it is expressing "joy."

[1737] Step 9:

[1738] The server automatically adjusts the mood and tempo of the music based on the emotion recognition results. For example, if it recognizes "joy," it will speed up the tempo or add a more upbeat arrangement.

[1739] Step 10:

[1740] The server converts the extracted scale, rhythm, and melody into the sound of an instrument selected by the user.

[1741] For example, if the user selects piano, the note data is converted into piano sounds.

[1742] Step 11:

[1743] The server transmits the generated music data to the terminal.

[1744] Step 12:

[1745] The terminal receives the music data sent from the server and displays music playback options to the user.

[1746] Step 13:

[1747] The user presses the play button to listen to the generated music.

[1748] Step 14:

[1749] The terminal activates the playback function and plays the music with the instrument sounds selected by the user.

[1750] Step 15:

[1751] The user can edit the music while listening to it, for example, by changing the tempo or correcting the notes.

[1752] Step 16:

[1753] When the user is unsure of the next phrase in the song, he or she presses the suggestion button.

[1754] Step 17:

[1755] The server uses AI to generate candidates for the next phrase or melody line and suggest them to the user.

[1756] Step 18:

[1757] The device displays suggested phrases and adds the user's selection to the song.

[1758] Step 19:

[1759] The user presses the upload button to upload the completed song to the community.

[1760] Step 20:

[1761] The terminal transmits the completed song data to a community server, and the song is made public.

[1762] Step 21:

[1763] Other users in the community provide feedback on the published songs.

[1764] Step 22:

[1765] The server notifies the song poster of the collected feedback.

[1766] Step 23:

[1767] The user reviews the feedback and makes any necessary changes to the song.

[1768] Through the above series of steps, the present invention allows anyone to easily enjoy composing music, and by taking the user's emotions into consideration when generating music, it provides a more personal music-making experience.

[1769] Example 2

[1770] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1771] Conventional music generation systems require users to manually edit music, often requiring technical knowledge. Furthermore, they are unable to generate music based on the user's emotional state, making it difficult to generate optimal music that matches individual emotions. Furthermore, there is no function to automatically suggest the next phrase or idea for a generated piece of music. There is a need for a system that can resolve these issues and allow users to generate music more easily and based on their emotions.

[1772] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data using a generative AI model to generate a scale, rhythm, and melody, means for analyzing the user's emotional state and automatically adjusting the mood and tempo of the music, and means for automatically suggesting the user's next phrase or idea using the generative AI model via a suggestion button. This allows users to easily generate music without technical knowledge, and the music is optimized according to the user's emotional state, enabling more satisfying music production. Furthermore, the automatic suggestion of the next phrase or idea can support the user's creativity.

[1773] "User" means an individual who uses the System to record humming and create, edit, and share music.

[1774] "Terminal" refers to a communication device used by a user to record humming, including a smartphone, tablet, etc.

[1775] A "server" is a computer system that receives audio data sent by a user, analyzes it, and creates music.

[1776] "Audio data" refers to information stored in digital form of a user's humming.

[1777] A "generative AI model" is an artificial intelligence algorithm that analyzes audio data to generate scales, rhythms, and melodies.

[1778] A "scale" refers to the arrangement of pitches that make up the melody of a piece of music.

[1779] "Rhythm" refers to the pattern of timing and duration of notes in a piece of music.

[1780] "Melody" refers to the main theme of a piece of music that is formed by combining scales and rhythms.

[1781] The "emotion engine" is an algorithm that analyzes the user's emotional state and adjusts the mood and tempo of the music based on the results.

[1782] "Instrumental sounds" refers to sounds produced by different musical instruments such as piano, guitar, violin, drums, etc.

[1783] "Music Data" means music information in digital form that has been generated through an analysis and conversion process.

[1784] "Suggestion button" refers to an interface element that a user uses to request an automatic suggestion of the next phrase or idea.

[1785] "Community" means the online platform where users can upload their created Music and receive feedback from other users.

[1786] "Feedback" means ratings and opinions provided by other users within the Community regarding the Generated Song.

[1787] The system of the present invention allows users to record their humming and then use it to create, edit, and share music, and is characterized by using an emotion engine to apply the user's emotional state to the music creation. Below, we will explain in detail how this system is implemented.

[1788] Recording and Data Transmission

[1789] Users record their humming using an application installed on their smartphone, tablet, or other device. When they tap the record button, the device captures the audio data through the built-in microphone and saves it as digital data in PCM or AAC format. Once recording is complete, the device sends the audio data to a server via the Internet.

[1790] Hardware used: Smartphone, tablet

[1791] Software used: Recording application

[1792] Analysis and music generation on the server

[1793] The server receives the audio data sent from the device and uses a generative AI model to analyze the audio data and extract musical scale, rhythm, and melody information. This analysis involves using a speech recognition algorithm to detect the pitch and timing of musical notes.

[1794] The server then analyzes the user's emotional state using an emotion engine, which analyzes the tone of the recorded voice and facial expression images provided by the user, and automatically adjusts the mood and tempo of the music based on the emotion recognition results.

[1795] Finally, the server converts the extracted music data into the instrument sound selected by the user. For example, if the user selects piano, the scale information is converted into piano sound. If another instrument (e.g., guitar or violin) is selected, the information is converted into note data corresponding to the instrument sound.

[1796] Hardware used: Server

[1797] Software used: Generative AI model, emotion engine

[1798] Sending and playing music data

[1799] The server then sends the generated music data back to the terminal, which receives it and displays it on its user interface. The user can listen to the music by pressing the play button.

[1800] Hardware used: Server, terminal

[1801] Software used: Music playback application

[1802] Editing features and phrase suggestions

[1803] Users can edit songs through the application, for example, by changing the tempo, modifying specific notes, or adding harmonies. If users are unsure of the next phrase during editing, they can press the suggestion button. This causes the server to use a generative AI model to generate several melody line and harmony candidates and suggest them to the user via their device.

[1804] Hardware used: Server, terminal

[1805] Software used: Music editing application, generative AI model

[1806] Community and Feedback

[1807] The completed song is uploaded to the community by the user. Other users in the community can provide feedback on the song. The server collects this feedback and notifies the song poster via their device, allowing the user to receive opinions and ratings from other users and make corrections to the song.

[1808] Hardware used: Server, terminal

[1809] Software used: Community Platform

[1810] Specific examples

[1811] As an example, let's consider the case where a user hums "Happy Birthday." The device records the song and sends the audio data to the server. The server uses a generative AI model to analyze the scale and extract the melody and rhythm of "Happy Birthday." The emotion engine then analyzes the user's emotion and, if it recognizes it as "joy," for example, speeds up the tempo and adds arrangements that emphasize brighter tones. If the user selects piano, the server converts the note data into piano sounds and sends them to the device as song data. The user plays the song on their device and edits the tempo and notes as needed. Finally, the completed song is uploaded to the community and feedback is received.

[1812] Example prompt sentence:

[1813] "I'm humming Happy Birthday. Identify the user's emotion as joy, speed up the tempo, and emphasize the bright tone. Finally, convert it into a piano sound."

[1814] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1815] Step 1:

[1816] The user launches the device's recording application and taps the record button. The device uses the built-in microphone to capture the user's humming and saves it as digital audio data in, for example, PCM or AAC format. The input is the user's humming, and the output is the captured audio data. This recording is saved as a temporary file in the device's storage.

[1817] Specific behavior:

[1818] User taps the record button

[1819] The device captures audio through the microphone

[1820] The device generates and stores digital audio data

[1821] Step 2:

[1822] After recording is complete, the device sends the audio data to the server via the Internet. The input is the stored audio data, and the output is the audio data sent to the server. The data is transferred securely using the HTTP protocol or WebSocket.

[1823] Specific behavior:

[1824] The device is waiting for audio data

[1825] The device sends the voice data to the server

[1826] Step 3:

[1827] The server analyzes the received audio data. First, it uses a generative AI model to extract scale, rhythm, and melody from the audio data. The input is audio data, and the output is structured musical data. It uses a speech recognition algorithm (e.g., a deep learning model) to detect the pitch and timing of each note and stores this in a database format.

[1828] Specific behavior:

[1829] The server receives the audio data

[1830] The server applies generative AI models to extract scale, rhythm, and melody

[1831] The server structures and stores music data

[1832] Step 4:

[1833] The server then analyzes the user's emotional state using an emotion engine. The emotion engine recognizes emotions by analyzing voice tone and provided facial expression images. The input is voice data and facial expression data, and the output is the recognized emotional state. The emotion engine uses machine learning algorithms to determine emotions and incorporates this information into the music data.

[1834] Specific behavior:

[1835] The server applies the emotion engine

[1836] The server analyzes voice tone and facial expression data

[1837] The server recognizes and records the user's emotional state.

[1838] Step 5:

[1839] The server converts the generated music data into the instrument sounds selected by the user. If the user selects piano, it is converted into piano sounds. The input is structured music data and instrument selection information, and the output is music data converted into specific instrument sounds. The instrument sounds are synthesized using sound fonts and audio sampling.

[1840] Specific behavior:

[1841] User selects instrument

[1842] The server converts the music data into instrument sounds

[1843] Step 6:

[1844] The server sends the converted music data to the terminal. The input is the music data, and the output is the music data sent to the terminal. The data is transferred via a secure communication channel.

[1845] Specific behavior:

[1846] The server waits for music data

[1847] The server sends the music data to the device

[1848] Step 7:

[1849] The terminal displays the received music data on the user interface, and the user can listen to the music by pressing the play button. The input is the received music data, and the output is the played music.

[1850] Specific behavior:

[1851] The device receives the music data

[1852] User taps the play button

[1853] The device plays the music

[1854] Step 8:

[1855] The user edits the music through the application, changing the tempo, modifying specific notes, adding harmonies, etc. The input is the music data and editing operations, and the output is the edited music data.

[1856] Specific behavior:

[1857] User uses editing functions

[1858] The device reflects the edited content in the song data

[1859] Step 9:

[1860] When a user is unsure of the next phrase, they press the suggest button. The server uses a generative AI model to generate several melody line and harmony candidates and suggests them to the user via the device. The input is a suggestion request, and the output is the suggested melody or harmony.

[1861] Specific behavior:

[1862] User taps the suggest button

[1863] Server generates candidates

[1864] Your device will display suggestions

[1865] Step 10:

[1866] A completed song is uploaded to the community by the user. Other users in the community can provide feedback on this song. The input is the completed song data, and the output is the feedback from the user. The server collects the feedback and notifies the song uploader via their device.

[1867] Specific behavior:

[1868] Users upload songs to the community

[1869] Other users provide feedback

[1870] Server collects and notifies feedback

[1871] (Application example 2)

[1872] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1873] Previous technologies have not provided sufficient concrete methods for improving work efficiency and worker morale in factories. Furthermore, systems that automatically generate music suited to the work environment have not been able to combine it with emotion recognition technology. Therefore, there is a need for technology that automatically generates music suited to the work environment while taking into account the user's emotions.

[1874] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's humming and saving it as audio data, means for transmitting the audio data to the server, means for analyzing the audio data on the server and generating a scale, rhythm, and melody, means for converting music into various instrument sounds based on the generated scale, rhythm, and melody, means for transmitting the generated music data to a terminal and enabling playback and editing, means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music suitable for the work environment, means for playing the background music on robots in the factory, and means for uploading music to the community and receiving feedback. This makes it possible to generate appropriate background music based on the user's humming and emotions, thereby improving work efficiency and morale in the factory.

[1875] definition statement

[1876] The "means for recording a user's humming and saving it as audio data" is a device or program that captures the audio of a user humming as digital data and saves it for later processing.

[1877] The "means for transmitting the audio data to the server" is a system that transfers the recorded audio data to a remote server via the Internet or a local network.

[1878] The "means for analyzing audio data on a server and generating scales, rhythms, and melodies" refers to a server system that performs processing to extract highly accurate musical elements based on audio data.

[1879] The "means for converting a piece of music into various instrument sounds based on the generated scale, rhythm, and melody" is a system that uses the analyzed musical elements to convert note data into different instrument sounds selected by the user.

[1880] "Means for transmitting generated music data to a terminal and enabling playback and editing" refers to a system for transferring music data generated on a server to a user's operating terminal, allowing the data to be played back and further edited.

[1881] "Means for recognizing the user's emotions and automatically adjusting the tempo and tone to generate background music appropriate for the work environment" refers to a system that analyzes the user's emotional state and automatically adjusts the tempo and tone of the music to an appropriate level based on the results.

[1882] The "means for playing the background music by a robot in a factory" refers to a robot system used to play the generated background music in a physical space.

[1883] "Means for uploading music to a community and receiving feedback" is a system that allows users to share music they have created with an online community and collect ratings and comments from other users.

[1884] MODE FOR CARRYING OUT THE INVENTION

[1885] The embodiment of the present invention is a music generation system aimed at improving work efficiency and worker morale in a factory. This system allows a user to record a humming tune and then generates, edits, and plays appropriate background music based on that humming. The system also recognizes the user's emotions and reflects them in the music it generates, providing music that is optimal for the work environment.

[1886] The server processes and calculates data using the following hardware and software: The "librosa" library is used for analyzing audio data, the "EmotionEngine" is used for emotion recognition, and the "MusicGenerator" is used for music generation.

[1887] System operation explanation

[1888] Humming recording and data transmission

[1889] A user records their humming using a terminal equipped with a recording function. The recorded audio data is saved as digital data by the terminal. This digital data is then transmitted to a server via a network.

[1890] Analysis and music generation on the server

[1891] The server receives the transmitted audio data and uses the librosa library to analyze the pitch and timing of each note, extracting scale, rhythm, and melody information. It then uses the Emotion Engine to analyze the user's emotions. Based on this emotional analysis, the tempo and timbre of the music are automatically adjusted.

[1892] Based on the analyzed musical scale information, the "Music Generator" is used to convert it into an appropriate instrument sound. At this stage, the note data is converted based on the instrument sound selected by the user (e.g. piano, guitar, drums, etc.).

[1893] Sending and playing music data

[1894] The generated music data is then sent back to the terminal. The terminal displays the received music data, and the user can listen to the music by pressing the play button. The user can also edit the music by changing the tempo or correcting the notes. In particular, it is possible to use robots in factories to play background music.

[1895] Community Features and Feedback

[1896] Users can upload their created compositions to the community and receive feedback from other users, allowing for further improvements and suggestions for new ideas.

[1897] Examples and prompts

[1898] As a specific example, let us consider a case where a user working in a factory hums, saying, "I want to concentrate, so I want some calming music." The tempo and tone of the music generated from this humming are adjusted appropriately based on the tone of the user's voice and emotional analysis.

[1899] Prompt Sentence Examples

[1900] Prompt: "I need some calming music to help me concentrate."

[1901] In this way, the present invention provides a music generation system that aims to improve work efficiency and morale in factories. Furthermore, by generating music that takes emotions into consideration, it is possible to provide appropriate music that matches the user's psychological state.

[1902] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1903] System program processing flow

[1904] Step 1:

[1905] The user uses the device to record their humming. When they press the record button, the audio is captured through the device's microphone and saved as digital audio data.

[1906] Input: User's humming (audio)

[1907] Output: Digital audio data (.wav, .mp3, etc.)

[1908] Step 2:

[1909] The stored digital audio data is sent from the device to the server, where it is transferred securely and quickly using a network protocol for transmission (e.g., HTTP, FTP).

[1910] Input: Digital audio data

[1911] Output: Digital audio data received by the server

[1912] Step 3:

[1913] The server analyzes the received audio data and generates scale, rhythm, and melody information using the audio data analysis library "librosa," which detects the pitch and timing of each note.

[1914] Input: Digital audio data received by the server

[1915] Output: Scale, rhythm, melody information (structured data)

[1916] Step 4:

[1917] Based on the analyzed scale, rhythm, and melody information, the note data is converted into musical instrument sounds. To correspond to the instrument sounds selected by the user (e.g., piano, guitar), the note data is converted using "MusicGenerator."

[1918] Input: Scale, rhythm, melody information, and selected instrument information

[1919] Output: Musical note data converted into instrument sounds (music data)

[1920] Step 5:

[1921] The server analyzes the user's emotions and reflects them in the generated music data. Using the emotion recognition engine "EmotionEngine," it analyzes the user's emotional state from the audio data. Based on that emotional state, the tempo and tone of the music are automatically adjusted.

[1922] Input: Digital voice data and analyzed emotional information

[1923] Output: Music data with tempo and tone adjusted to match the emotion

[1924] Step 6:

[1925] The generated and adjusted music data is sent back to the device. The server sends the data to the device via a transfer protocol.

[1926] Input: Music data generated and adjusted on the server

[1927] Output: Music data received on the device

[1928] Step 7:

[1929] The device plays and edits the received music data. The user can listen to the created music by pressing the play button. The edit button can also be used to change the tempo or modify the notes.

[1930] Input: Music data received on the device

[1931] Output: Played and edited song

[1932] Step 8:

[1933] Users can then play the final edited song as background music on a robot in the factory, which has a built-in speaker and plays the music based on the song data.

[1934] Input: Song data edited on the device

[1935] Output: Background music played in your work environment

[1936] Step 9:

[1937] The completed songs are uploaded to the community by the users, and other users provide feedback on the uploaded songs, which is collected by the server and sent to the device.

[1938] Input: Completed song data

[1939] Output: Songs shared with the community and given feedback

[1940] Specific prompt examples:

[1941] Prompt: "I need some calming music to help me concentrate."

[1942] In this way, a system is provided that generates music that reflects the user's humming or emotional state and plays it as appropriate background music in the factory.

[1943] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1944] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1945] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1946] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1947] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1948] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1949] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1950] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1951] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1952] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1953] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1954] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1955] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1956] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1957] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1958] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1959] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1960] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1961] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1962] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1963] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1964] The following is further disclosed regarding the above embodiment.

[1965] (Claim 1)

[1966] A means for recording a user's humming and saving it as audio data;

[1967] means for transmitting the voice data to a server;

[1968] means for analyzing the audio data on the server and generating a scale, rhythm, and melody;

[1969] A means for converting the musical piece into various musical instrument sounds based on the generated scale, rhythm, and melody;

[1970] means for transmitting the generated music data to a terminal and enabling playback and editing;

[1971] A system that includes a means to upload songs to the community and receive feedback.

[1972] (Claim 2)

[1973] The system according to claim 1, further comprising a means for the AI ​​to automatically suggest the next phrase or idea for the song generated from the user's humming.

[1974] (Claim 3)

[1975] 10. The system of claim 1, which allows a user to convert a generated musical composition into multiple instrument sounds, such as piano, guitar, violin, and drums.

[1976] "Example 1"

[1977] (Claim 1)

[1978] means for recording and storing the user's voice as digital data;

[1979] means for transmitting the digital data to a server;

[1980] means for analyzing digital data on a server and generating scales, rhythms, and melodies;

[1981] a means for converting the music data into various musical instrument sounds based on the generated scale, rhythm, and melody;

[1982] means for transmitting the generated music data to a terminal and enabling playback and editing;

[1983] A method for AI to automatically suggest the next phrase candidate for edited music data,

[1984] A system that includes a means for uploading music data to a community and receiving feedback.

[1985] (Claim 2)

[1986] The system according to claim 1, further comprising a means for AI to automatically suggest the next phrase or idea for the music data generated from the user's voice.

[1987] (Claim 3)

[1988] 10. The system of claim 1, further comprising means for enabling a user to convert generated music data into multiple instrument sounds.

[1989] "Application Example 1"

[1990] (Claim 1)

[1991] A means for recording a user's humming and saving it as audio data;

[1992] means for transmitting the voice data to a server;

[1993] means for analyzing the audio data on the server and generating a scale, rhythm, and melody;

[1994] A means for converting the musical piece into various musical instrument sounds based on the generated scale, rhythm, and melody;

[1995] means for transmitting the generated music data to a terminal and enabling playback and editing;

[1996] A means for uploading the generated music data to a content distribution service and sharing it with other users;

[1997] A system that includes a means for receiving feedback from other users.

[1998] (Claim 2)

[1999] The system according to claim 1, further comprising a means for the AI ​​to automatically suggest the next phrase or idea for the song generated from the user's humming.

[2000] (Claim 3)

[2001] 10. The system of claim 1, which allows a user to convert a generated musical piece into multiple instrument sounds.

[2002] "Example 2: Combining Emotion Engines"

[2003] (Claim 1)

[2004] A means for recording a user's humming and saving it as audio data;

[2005] means for transmitting the voice data to a server;

[2006] A means using a generative AI model that analyzes audio data on a server and generates scales, rhythms, and melodies;

[2007] A means for converting the musical piece into various musical instrument sounds based on the generated scale, rhythm, and melody;

[2008] means for analyzing the emotional state of the user based on the generated music data and automatically adjusting the mood and tempo of the music;

[2009] means for transmitting the generated music data to a terminal and enabling playback and editing;

[2010] A way to upload your music to the community and receive feedback,

[2011] A suggestion button uses generative AI models to automatically suggest the next phrase or idea to the user.

[2012] A system including:

[2013] (Claim 2)

[2014] 2. The system according to claim 1, further comprising means for using an emotion engine on a server to analyze a user's emotions and apply the results to music generation.

[2015] (Claim 3)

[2016] 10. The system of claim 1, which allows a user to convert a generated musical piece into multiple instrument sounds.

[2017] "Application example 2 when combining emotion engines"

[2018] Rewriting of claims

[2019] (Claim 1)

[2020] A means for recording a user's humming and saving it as audio data;

[2021] means for transmitting the voice data to a server;

[2022] means for analyzing the audio data on the server and generating a scale, rhythm, and melody;

[2023] A means for converting the musical piece into various musical instrument sounds based on the generated scale, rhythm, and melody;

[2024] means for transmitting the generated music data to a terminal and enabling playback and editing;

[2025] A means for recognizing a user's emotions and automatically adjusting the tempo and tone to generate background music suitable for the work environment;

[2026] means for playing the background music by a robot in a factory;

[2027] A system that includes a means to upload songs to the community and receive feedback.

[2028] (Claim 2)

[2029] The system according to claim 1, further comprising a means for the AI ​​to automatically suggest the next phrase or idea for the song generated from the user's humming.

[2030] (Claim 3)

[2031] 10. The system of claim 1, which allows a user to convert a generated musical composition into multiple instrument sounds, such as piano, guitar, violin, and drums. [Explanation of symbols]

[2032] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for recording a user's humming and saving it as audio data; means for transmitting the voice data to a server; means for analyzing the audio data on the server and generating a scale, rhythm, and melody; A means for converting the musical piece into various musical instrument sounds based on the generated scale, rhythm, and melody; means for transmitting the generated music data to a terminal and enabling playback and editing; A system that includes a means to upload songs to the community and receive feedback.

2. The system according to claim 1, further comprising a means for the AI ​​to automatically suggest the next phrase or idea for the music generated from the user's humming.

3. 10. The system of claim 1, which allows a user to convert a generated musical composition into multiple instrument sounds, such as piano, guitar, violin, drums, etc.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A