System

The system facilitates the conversion of self-recorded songs into professional-quality music by using AI to generate and adjust melodies, addressing the barrier of technical knowledge in music creation.

JP2026019764APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121512
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Creating music requires advanced technical knowledge and specialized musical skills, making it difficult for average users and the elderly to produce professional-quality music, and converting self-recorded songs into such quality is challenging.

Method used

A system that allows users to record their voice, transmit the data to a server, perform voice analysis to extract melody lines and lyrics, generate a professional-quality melody line using an AI model, adjust the recorded voice to fit the melody, and provide the completed song, all without specialized knowledge or skills.

Benefits of technology

Enables users to easily convert their recorded songs into professional-quality music, allowing anyone to enjoy music production without needing specialized knowledge or skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019764000001_ABST
    Figure 2026019764000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system including a means for a user to record voice, a means for transmitting the recorded voice data to a server, a means for performing voice analysis in the server and extracting a melody line and lyrics, a means for generating a professional-quality melody line using an AI model based on the extracted melody line, a means for adjusting the recorded voice using the generated melody line and generating a completed musical piece, and a means for providing the generated completed musical piece to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventionally, creating music on one's own requires advanced technical knowledge and specialized musical skills, posing a significant hurdle for average users and the elderly. Furthermore, it has been difficult to turn a self-recorded song into a professional-quality piece of music. As a result, many people have hesitated to create their own music and have missed out on the joy of music production. To address the above issues, the present invention aims to provide a system that allows users to easily convert recorded songs into professional-quality music. [Means for solving the problem]

[0005] The present invention proposes a system including a means for a user to record voice, a means for transmitting the recorded voice data to a server, a means for performing voice analysis on the server and extracting a melody line and lyrics, a means for generating a professional-quality melody line using an AI model based on the extracted melody line, a means for adjusting the recorded voice using the generated melody line to generate a completed song, and a means for providing the generated completed song to a user. This system allows users to convert their own recorded songs into professional-quality songs without specialized knowledge or skills, and easily enjoy the fun of music production.

[0006] "User" means any individual or entity that uses the System.

[0007] "Audio" refers to recorded data of the user's voice.

[0008] A "recording means" refers to a device or software that records audio as digital data.

[0009] "Transmission means" refers to the device or protocol that sends the recorded audio data to the server over the network.

[0010] "Server" means a remote computer system that analyzes, processes, and stores audio data.

[0011] "Audio analysis" is a technology that extracts specific musical elements from recorded audio data.

[0012] A "melody line" is a pitch or rhythm pattern extracted through audio analysis.

[0013] "Lyrics" refers to the spoken part of the song sung by the user.

[0014] An "AI model" is an algorithm or system that uses artificial intelligence techniques to generate a specific output from data.

[0015] A "professional quality melody line" is a sophisticated melody pattern that sounds like it was created by a professional.

[0016] "Adjusting means" refers to methods or techniques for modifying the user's recorded voice to fit the generated melody line.

[0017] A "finished song" is a music file that is finally generated based on the adjusted audio data.

[0018] "Means of providing" refers to the procedures and techniques used to make the completed music available to users.

[0019] "Pitch" refers to the pitch of each note.

[0020] A "rhythm pattern" refers to the structure of the timing and intervals of sounds.

[0021] "Sound quality optimization" refers to processing to improve the quality of audio data.

[0022] A "filter" is an audio signal processing technique that emphasizes or suppresses specific frequency components.

[0023] "Effect" refers to acoustic processing that adds additional effects to audio data. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0025] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0026] First, the terms used in the following description will be explained.

[0027] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0028] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0029] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0030] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0032] [First embodiment]

[0033] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0034] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0036] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0037] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0039] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0040] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0041] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0043] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0045] As an embodiment of the present invention, a system is provided that allows users to easily convert their own recorded songs into professional quality musical compositions.

[0046] System Overview

[0047] First, the user uses a device to record their own voice. A recording application is installed on the device, and the user starts the application and presses the "Start Recording" button to begin recording. Once recording is complete, the device sends the recorded data to the server.

[0048] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. It then uses an AI model to generate a professional-quality melody line based on the extracted melody line. It then uses the generated melody line to adjust the user's recorded voice to create a finished song. The finished song is then stored by the server and can be downloaded by the user via their device.

[0049] Explanation of program processing (with concrete examples)

[0050] 1. Start recording and send data

[0051] The user starts a recording app on their smartphone and sings the lyrics of the song they want to sing, for example, "One afternoon, relaxing in the garden," for one minute.

[0052] Once the recording is complete, the device sends the audio data to the server.

[0053] 2. Audio analysis and melody extraction

[0054] The server receives the audio data sent from the device and uses a speech analysis module to analyze the recording, for example, to identify the pitch and rhythmic patterns of the recorded lyrics, "One afternoon, relaxing in the garden."

[0055] The server extracts the lyrics and melody line and stores in a database the pitch and rhythm in which each word of the lyrics is sung.

[0056] 3. Melody adjustment and generation

[0057] The server inputs the extracted melody line and lyrics into the AI ​​model. For example, the lyric "One afternoon, relaxing in the garden" can be input into the AI ​​model to generate a more refined melody line.

[0058] Based on the input data, the AI ​​model generates professional-quality melody lines and enhances the original melody lines.

[0059] 4. Music Generation and Enhancement

[0060] The server uses the generated professional-quality melody line to adjust the recording of the user's voice, so that the user's voice is modified to fit the newly generated melody line.

[0061] The server uses an enhancement module to optimize the sound quality and apply filters and effects to improve the overall sound quality of the song, for example using a noise reduction filter or a reverb effect.

[0062] 5. Submission of completed music

[0063] The server generates the finished song and stores it in a database, which the user can then download from the server using their device.

[0064] Users can check the completed song and share it via social media or email if they wish.

[0065] In this way, the present invention provides a concrete means for users to easily create professional-quality music, allowing anyone to easily enjoy the joy of music production, even without specialized knowledge or skills.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording.

[0069] Step 2:

[0070] The device uses a built-in microphone to record the user's singing voice and saves the recorded data as a temporary file.

[0071] Step 3:

[0072] When the user presses the button to stop recording, the recording ends.

[0073] Step 4:

[0074] The terminal transmits the recorded voice data to the server.

[0075] Step 5:

[0076] The server receives and stores the voice data sent from the terminal.

[0077] Step 6:

[0078] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythm, and lyric timing.

[0079] Step 7:

[0080] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0081] Step 8:

[0082] The server inputs the extracted melody line and lyrics into the AI ​​model.

[0083] Step 9:

[0084] The AI ​​model generates a professional-quality melody line based on the input data and returns the generated melody line to the server.

[0085] Step 10:

[0086] The server adjusts the user's recorded voice based on the generated melody line and modifies it to fit the new melody line.

[0087] Step 11:

[0088] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song.

[0089] Step 12:

[0090] The server generates professional quality finished music and stores it in a database.

[0091] Step 13:

[0092] The user uses a terminal to access the server and download the completed song.

[0093] Step 14:

[0094] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0095] Through this series of steps, users can transform their recorded songs into professional-quality songs and easily enjoy the music production process.

[0096] Example 1

[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0098] In the past, users needed advanced musical knowledge and specialized equipment to create professional-quality music, making it extremely difficult for the average user. Furthermore, adjusting recorded audio to match a high-quality melody line required specialized skills. Therefore, a method for easily generating high-quality music was needed.

[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0100] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed piece of music, and means for providing the generated completed piece of music to the user, thereby enabling users to easily generate high-quality music without having specialized knowledge or skills.

[0101] "User" means any person or entity that intends to use the System to record audio and generate musical compositions.

[0102] "Audio recording means" refers to a device or application that a user uses to record their own audio.

[0103] "Means for transmitting the recorded voice data to the server" refers to the communications protocols and procedures for uploading the recorded voice data to the server via the Internet.

[0104] "Server" refers to a computer system for receiving, analyzing, and processing recorded audio data.

[0105] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms and software for automatically analyzing and extracting pitch, rhythm patterns, and lyrics from recorded audio data.

[0106] "Generative AI Model" refers to an artificial intelligence program or algorithm used to generate professional-quality melody lines based on extracted melody lines and lyrics.

[0107] "Professional quality melody line" refers to a melody line of professional quality.

[0108] "Means for adjusting the audio and generating a finished piece of music" refers to the methods and techniques for modifying the recorded audio to match the generated melody line and generating the final piece of music.

[0109] "Means for providing the completed song to the user" refers to the method or procedure for providing the final song to the user in a downloadable format.

[0110] "Means for extracting pitch and rhythmic patterns" refers to algorithms or software that identify and extract pitch and rhythmic patterns from recorded audio data.

[0111] "Means for sound quality optimization and applying filters and effects" refers to techniques for applying audio filters and effects to improve the sound quality of the generated music.

[0112] "System" refers to the overall setup that includes the elements mentioned above and enables users to create professional quality music.

[0113] The present invention relates to a system that allows users to easily create professional quality music. Specific means for carrying out the invention will be described in detail below.

[0114] A user first launches a recording application on a device such as a smartphone or tablet. This application provides a recording function, allowing the user to record their own voice. Once recording is complete, the device sends the recorded data (voice data) to a server. To do this, the device establishes a connection to the server via the Internet and sends the voice data using a protocol such as an HTTP POST request.

[0115] The server uses a receiving module to receive audio data sent from the device. This audio data is stored in a database on the server. The server then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. The audio analysis module uses libraries such as FFT (Fast Fourier Transform) and LibROSA to analyze the frequency components of the audio data. This allows detailed analysis data of the recorded audio to be obtained.

[0116] The analyzed data is stored in a database on the server, and the next step is to input it into a generative AI model. The generative AI model generates a professional-quality melody line based on the extracted melody line and lyrics. This model can use, for example, a generative AI model or other similar artificial intelligence algorithms. As a concrete example, the prompt sentence is as follows:

[0117] Prompt statement:

[0118] "Generate a professional-quality melody line for "One Afternoon Relaxing in the Garden" based on your voice data. The voice data is in the attached file."

[0119] The generative AI model takes this prompt and the attached file and generates a professional-quality melody line. The generated melody line is sent back to the server and stored in a database. The server then uses this generated melody line to adjust the user's recorded voice, using techniques such as pitch shifting and time stretching to modify the user's voice to fit the new melody line.

[0120] The server then uses an enhancement module to optimize the sound quality, applying noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in the final finished song.

[0121] Finally, the server stores the completed song in a database and provides it to the user, who can download it from the server using their device. Users can also share the completed song via social media or email.

[0122] In this way, the present invention provides a concrete means for users to easily create professional-quality music, and is a system that allows anyone to easily enjoy the fun of music production, even without specialized knowledge or skills.

[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0124] Step 1:

[0125] The user launches a recording application on their smartphone or tablet. They press the "Start Recording" button and sing the lyrics of the song they want to sing. They sing specific lyrics, such as "One afternoon, relaxing in the garden," for one minute. Once the recording is complete, the device sends the recording to the server. The input is the audio data sung by the user, and the output is the audio data sent to the server.

[0126] Step 2:

[0127] The server receives the audio data sent from the device. The received audio data is stored in a database on the server. It then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. Specifically, it analyzes the frequency components of the audio data using libraries such as FFT (Fast Fourier Transform) and LibROSA. The input is the audio data sent to the server, and the output is the analyzed pitch, rhythm patterns, and timing information for the lyrics.

[0128] Step 3:

[0129] The server inputs the melody line and lyrics extracted from the analyzed audio data into the generative AI model. A prompt is sent to the generative AI model, which generates a professional-quality melody line according to the prompt. For example, the prompt could be, "Based on the user's audio data, please generate a professional-quality melody line for 'Relaxing in the Garden One Afternoon.' The audio data is in the attached file." The input is the extracted melody line and lyrics, and the output is the professional-quality melody line returned by the generative AI model.

[0130] Step 4:

[0131] The server adjusts the user's recorded voice based on the generated professional-quality melody line. Techniques used include pitch shifting and time stretching, which modify the user's voice to fit the new melody line. The input is the generated melody line and the user's recording, and the output is the adjusted voice data.

[0132] Step 5:

[0133] The server uses an enhancement module to optimize the sound quality. It applies noise reduction filters and reverb effects to improve the overall sound quality of the song. Specifically, it removes recorded background noise and adds reverb to add depth to the audio. The input is the adjusted audio data, and the output is an enhanced, sound-optimized finished song.

[0134] Step 6:

[0135] The server stores the generated completed music in a database and provides it to the user. The user can download and check the generated music on their own device. They can also share it via social media or email. The input is the enhanced completed music, and the output is music data that can be accessed by providing the user with a download link.

[0136] (Application example 1)

[0137] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0138] The traditional music production process requires specialized knowledge and skills, making it difficult for many users to easily create professional-quality music. Furthermore, there are limited ways to easily share completed music, and in particular, there is a lack of an environment in place that allows ordinary users to easily upload their own music to streaming platforms.

[0139] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0140] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed song, means for providing the generated completed song to the user, and means for uploading the completed song to a streaming platform. This enables users to create high-quality songs without requiring specialized knowledge or skills and easily upload them to a streaming platform for sharing.

[0141] "User" means an individual or entity that uses the system to record audio and generate and upload music.

[0142] An "audio recording means" is a device or application that a user uses to record a song.

[0143] "Means for transmitting recorded voice data to a server" refers to the process or function that transmits a user's recorded voice data to a server.

[0144] A "server" is a central computer system for analyzing and processing audio data.

[0145] "Means for performing audio analysis" refers to software or algorithms for extracting melody lines and lyrics from recorded audio data.

[0146] The "means for extracting a melody line" is a function for identifying pitch and rhythm patterns from audio data.

[0147] A "generative AI model" is an artificial intelligence (AI) model that generates professional-quality melody lines based on extracted melody lines.

[0148] "Means for generating professional-quality melody lines" refers to the process of using a generative AI model to create high-quality melodies based on a user's recordings.

[0149] The "means for adjusting the recorded voice and generating a completed piece of music" refers to the process of adjusting the user's recorded voice to match the generated melody line, thereby generating a completed piece of music as a whole.

[0150] The "means for providing the generated completed music to the user" is a process or function that provides the completed music in a form that the user can access and download.

[0151] "Means for uploading the completed song to a streaming platform" refers to a function for uploading the completed song to a streaming service on the Internet.

[0152] The present invention provides a system that allows users to convert their recorded songs into professional quality music and easily upload the music to a streaming platform. Specific embodiments of the present invention are described below.

[0153] System Overview

[0154] 1. Recording and data transmission

[0155] The user starts a recording application on a device such as a smartphone and records a song. When the recording is finished, the device sends the audio data to the server.

[0156] For example, a user can sing a song such as "One Afternoon, Relaxing in the Garden" using a recording app on their smartphone and send the recording data to a server.

[0157] 2. Audio analysis and melody extraction

[0158] The server receives the transmitted audio data and analyzes it using audio analysis software (e.g., librosa), which extracts pitch, rhythm patterns, and the timing of lyrics.

[0159] For example, the pitch and rhythmic patterns of the recorded lyrics "One afternoon, relaxing in the garden" are analyzed.

[0160] 3. Melody Generation

[0161] The server inputs the extracted melody line and lyrics into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line.

[0162] For example, a sophisticated melody line can be generated based on the lyrics "One afternoon relaxing in the garden."

[0163] 4. Audio adjustment and music generation

[0164] The server then uses the generated professional-quality melody line to adjust the recorded audio and generate the finished song, applying filters and effects (e.g., noise reduction filters and reverb effects) to optimize the sound quality.

[0165] For example, the user's recording is adjusted to fit the newly generated melody line, and noise reduction filters and reverb effects are applied.

[0166] 5. Providing music and uploading to streaming services

[0167] The server generates the finished song and provides it to the user, who can then download it or upload it to a streaming platform.

[0168] For example, users can review the finished song on their smartphone and upload it directly to a streaming platform if they wish.

[0169] Examples of specific examples and prompts

[0170] As a concrete example, consider a user relaxing in their garden and recording a song:

[0171] Example prompt sentence:

[0172] Users can simply record a song on their smartphone while relaxing in the garden one afternoon. The app then sends the recording to a server, where an AI model generates a professional-quality melody line. Users can then upload the finished song directly to streaming platforms.

[0173] This system allows users to create professional-quality music without any specialized knowledge or skills, and then easily upload and share the music on streaming platforms.

[0174] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0175] Step 1:

[0176] A user records a song using a recording app on their smartphone. The user launches the app, presses the "Start Recording" button, and sings. When the recording is finished, the application saves the audio data in a file format. The input of this step is the user's singing voice, and the output is the recorded audio file.

[0177] Step 2:

[0178] The device sends the recorded audio data to the server. After the recording is complete, the device uses the data transmission function to upload the audio file to the server. The input of this step is the saved audio file, and the output is the audio data uploaded to the server.

[0179] Step 3:

[0180] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. Audio analysis software (e.g., librosa) is used for the analysis. Specifically, the audio data is loaded using the librosa.load function, and the pitch and rhythmic patterns are extracted using the librosa.piptrack function. The input of this step is the audio data uploaded to the server, and the output is the analyzed melody line and rhythmic patterns.

[0181] Step 4:

[0182] The server inputs the extracted melody line and rhythm pattern into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line. Specifically, the analyzed melody line and rhythm pattern are fed into the AI ​​model, and a new melody line is generated using the model's inference function. The input of this step is the extracted melody line and rhythm pattern, and the output is the generated professional-quality melody line.

[0183] Step 5:

[0184] The server then uses the generated, professional-quality melody line to adjust the recorded audio and generate a finished song. Specifically, it uses audio editing software to adjust the user's recording to fit the new melody line and apply filters and effects (e.g., noise reduction filters and reverb effects). The input to this step is the generated, professional-quality melody line and the recorded audio data, and the output is the finished song.

[0185] Step 6:

[0186] The server provides the generated finished song to the user. Specifically, the server stores the song file on the server so that the user can download it, and provides the user with a download link. The input of this step is the finished song, and the output is a song file that the user can download.

[0187] Step 7:

[0188] The user uploads the completed song to a streaming platform. Specifically, the user uploads the completed song to the platform using the streaming platform's upload function provided by the smartphone application. The input of this step is the song file downloaded by the user, and the output is the song uploaded to the streaming platform.

[0189] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0190] As an embodiment of the present invention, a system is provided that allows a user to easily convert their own recorded songs into professional quality music, and furthermore, reflects emotional elements in the music using an emotion engine.

[0191] System Overview

[0192] The present invention consists of the following main components:

[0193] 1. A means for users to record audio (a device with a recording application)

[0194] 2. A means of sending recorded audio data to the server

[0195] 3. A method for analyzing audio on the server and extracting melody lines and lyrics

[0196] 4. A means to recognize emotions from the user's voice using an emotion engine

[0197] 5. A method for generating professional-quality melody lines using an AI model based on the extracted melody lines and the recognized emotions.

[0198] 6. A method for generating a complete song using the generated melody line and the adjusted user's voice

[0199] 7. Means of sound quality optimization and application of filters and effects

[0200] 8. Means for providing the generated finished music to the user

[0201] Explanation of program processing (with concrete examples)

[0202] Start recording and send data

[0203] 1. The user starts the recording app on their device (such as a smartphone) and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0204] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0205] 3. When the user presses the button to stop recording, the recording ends.

[0206] 4. The device sends the recorded audio data to the server.

[0207] Voice analysis, emotion recognition, and melody generation

[0208] 5. The server receives the voice data sent from the terminal.

[0209] 6. The server uses an audio analysis module to analyze the recording and detect pitch, rhythmic patterns, and lyric timing.

[0210] 7. The server extracts the melody line and lyrics based on the analysis results.

[0211] 8. The emotion engine on the server recognizes the user's emotion from the voice data. For example, if the user sings "happily," the emotion engine recognizes a positive emotion.

[0212] 9. The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0213] 10. The AI ​​model uses input data to generate professional-quality melody lines that match emotions. For example, if a positive emotion is recognized, an upbeat, rhythmic melody line will be generated.

[0214] Music Generation and Enhancement

[0215] 11. The server then adjusts the user's recorded voice based on the generated professional-quality melody line, modifying it to fit the new melody line.

[0216] 12. The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, for example, using a noise reduction filter or a reverb effect.

[0217] 13. The server generates a professional quality finished song and stores it in a database.

[0218] Submit and share your music

[0219] 14. The user uses the terminal to access the server and download the completed song.

[0220] 15. Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0221] Through this series of steps, users can transform their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0222] The processing flow will be explained below.

[0223] Step 1:

[0224] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording. For example, they sing lyrics such as "One afternoon, relaxing in the garden."

[0225] Step 2:

[0226] The device uses a built-in microphone to record the user's singing voice and saves the recording data as a temporary file.

[0227] Step 3:

[0228] When the user presses the button to stop recording, the recording ends.

[0229] Step 4:

[0230] The terminal transmits the recorded voice data to the server.

[0231] Step 5:

[0232] The server receives and stores the voice data sent from the terminal.

[0233] Step 6:

[0234] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythmic patterns, and lyric timing.

[0235] Step 7:

[0236] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0237] Step 8:

[0238] The server's emotion engine recognizes the user's emotions from the voice data. For example, it recognizes the positive emotion of "feeling happy" from the user's tone of voice and intonation.

[0239] Step 9:

[0240] The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0241] Step 10:

[0242] The AI ​​model uses the input data to generate professional-quality melody lines that match the emotion. For example, if a positive emotion is recognized, it will generate an upbeat, rhythmic melody line.

[0243] Step 11:

[0244] The server then adjusts the user's recorded voice based on the generated melody line, correcting it to fit the new melody line, including adjusting the pitch and timing.

[0245] Step 12:

[0246] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, such as noise reduction filters and reverb effects.

[0247] Step 13:

[0248] The server generates professional quality finished music and stores it in a database.

[0249] Step 14:

[0250] The user uses a terminal to access the server and download the completed song.

[0251] Step 15:

[0252] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0253] Through this series of processes, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0254] Example 2

[0255] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0256] Previous music production systems required specialized knowledge and skills, as well as a great deal of time and effort, to convert users' recorded audio into professional-quality music. It was also difficult for average users to incorporate emotional elements into music. This made music production a challenging process, making it difficult for anyone to easily create high-quality music.

[0257] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting melody lines and lyrics, means for recognizing emotions from voice data, and means for generating professional-quality melody lines using a generative AI model based on the extracted melody lines and the recognized emotions. This allows anyone to easily generate high-quality music and enjoy music that reflects their own emotions.

[0258] A "voice recording means" is an application or device that allows a user to record their own voice.

[0259] "Means for transmitting audio data to a server" refers to a communication means for uploading recorded audio data to a server, which includes an internet connection and a corresponding application.

[0260] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms or software for analyzing recorded audio data to identify and extract melody lines and lyrics therefrom.

[0261] The "means for recognizing emotions" is an algorithm or software for analyzing voice data and recognizing the user's emotions.

[0262] "Means for generating professional-quality melody lines using generative AI models" refers to a system or software that uses AI technology to automatically generate high-quality melody lines.

[0263] "Means for adjusting recorded voice and generating a completed piece of music" refers to a system or software for adjusting a user's recorded voice to match the generated melody line and combining them into a completed piece of music.

[0264] "Means for applying filters and effects to optimize sound quality" refers to a system or software for applying various audio filters and effects to improve the sound quality of the generated music.

[0265] "Means for providing the completed music to the user" refers to means for providing the created music in a form that the user can access, including providing a download link or using cloud storage.

[0266] The present invention provides a system that allows users to easily convert their own recorded songs into professional quality music and further reflects emotional elements in the music using an emotion engine. Specific embodiments will be described below.

[0267] Hardware and Software Configuration

[0268] User terminal

[0269] The user terminal is a device such as a smartphone or tablet that has a recording application installed. This terminal has a built-in microphone that is used to record the user's voice. It also has a communication function to send the recorded data to the server.

[0270] server

[0271] The server contains multiple modules for processing the received voice data, including a voice analysis module, an emotion recognition engine, an AI model, and a sound quality optimization module.

[0272] Voice recording and data transmission

[0273] 1. The user launches the recording app on the device and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0274] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0275] 3. When the user presses the stop recording button, the recording ends.

[0276] 4. The device sends the recorded audio data to the server via an HTTP POST request, using the HTTPS protocol.

[0277] Voice analysis and emotion recognition

[0278] 5. The server receives the audio data sent from the device and analyzes the recording using the Python LibROSA library.

[0279] 6. The server detects pitch, rhythmic patterns, and lyric timing and stores them in an internal database (e.g., MongoDB) in JSON format.

[0280] 7. The server's emotion engine uses an emotion recognition algorithm (e.g., Microsoft Azure Cognitive Services) to recognize the user's emotion from the voice data.

[0281] Melody line generation

[0282] 8. The server inputs the extracted melody line and the recognized emotion into the AI ​​model. For example, it inputs the following prompt to the AI ​​model: "Create a fun melody line in the style of James Brown."

[0283] 9. The AI ​​model uses the input data to generate a professional-quality melody line tailored to the emotion, using OpenAI GPT-3 or other music generation models.

[0284] Music Generation and Enhancement

[0285] 10. The server uses Auto-tune to adjust the user's recorded voice based on the generated professional-quality melody line.

[0286] 11. The server uses a sound quality optimization module (e.g., iZotope Neutron) to apply noise reduction filters and reverb effects to improve the overall sound quality of the song.

[0287] Providing completed music

[0288] 12. The server stores the generated professional quality finished music in a database.

[0289] 13. The user can use their device to access the server and download the completed song from the "My Songs" section.

[0290] Through this process, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0291] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0292] Step 1:

[0293] The user launches the device's recording app and taps the "Start Recording" button to begin recording. The input is the user's button operation, and the output is the trigger to start recording. Specifically, the recording app accesses the built-in microphone and begins capturing audio in real time.

[0294] Step 2:

[0295] The device uses a built-in microphone to record the user's singing voice and temporarily saves the recorded data. The input is the user's voice, and the output is the recorded voice data. Specifically, the recorded data is temporarily saved in the device's local storage in wav format.

[0296] Step 3:

[0297] Recording ends when the user presses the stop recording button. The input is the user's button operation, and the output is the end of the recording session. Specifically, the recording data capture stops and the final data file is committed.

[0298] Step 4:

[0299] The device sends the recorded audio data to the server. The input is the recorded data, and the output is the data sent to the server. Specifically, the audio data is sent to the server's API endpoint via an HTTP POST request, and the HTTPS protocol is used for communication.

[0300] Step 5:

[0301] The server receives the voice data sent from the device and stores it in storage. The input is the recorded data, and the output is the stored data. Specifically, the received data is stored in file storage and placed in a waiting state for processing.

[0302] Step 6:

[0303] The server uses an audio analysis module to analyze the recorded data and detect pitch, rhythmic patterns, and lyric timing. The input is the stored audio data, and the output is the analysis results. Specifically, it uses the Python LibROSA library to perform spectral analysis and pitch detection of the audio data.

[0304] Step 7:

[0305] The server extracts the melody line and lyrics based on the analysis results. The input is the audio analysis results, and the output is the extracted melody line and lyrics. Specifically, the analysis results are saved in JSON format in an internal database (e.g., MongoDB).

[0306] Step 8:

[0307] The server's emotion engine recognizes the user's emotions from the voice data. The input is the recorded data, and the output is an emotion score. Specifically, it uses Microsoft's Azure Cognitive Services to extract emotions from the voice data in real time and calculate the emotion score.

[0308] Step 9:

[0309] The server inputs the extracted melody line and the recognized emotion into the AI ​​model. The input is the melody line and emotion score, and the output is the generated melody line. Specifically, the server inputs the following prompt to the AI ​​model: "Please create a fun melody line in the style of James Brown."

[0310] Step 10:

[0311] The AI ​​model uses input data to generate a professional-quality melody line that matches the emotion. The input is a prompt, and the output is the generated melody line. Specifically, the generative AI model (e.g., OpenAI GPT-3 or other music generation model) generates a professional-quality melody line and returns the data to the server.

[0312] Step 11:

[0313] The server adjusts the user's recorded voice based on the generated professional-quality melody line. The input is the user's voice and the generated melody line, and the output is the adjusted song. Specifically, it uses Auto-tune to ensure consistency between the voice and the melody line.

[0314] Step 12:

[0315] The server uses a sound quality optimization module to improve the overall sound quality of the track. The input is the adjusted track, and the output is the sound-optimized track. Specifically, it uses iZotope's Neutron to apply noise reduction filters and reverb effects.

[0316] Step 13:

[0317] The server generates professional-quality finished songs and stores them in a database. The input is a quality-optimized song, and the output is a saved song file. Specifically, the final audio file is saved in the database in MP3 or WAV format.

[0318] Step 14:

[0319] The user uses their device to access the server and download the completed song. The input is the user's request, and the output is the downloaded song file. Specifically, the user selects a song from the "My Songs" section within the recording app and downloads it securely via HTTPS.

[0320] Step 15:

[0321] Users can play the completed song to check its contents, and if they wish, share it on social media or by email. The input is the user's operation, and the output is the shared song. Specifically, an API is used to share songs directly to social media from within the app, providing a seamless sharing experience.

[0322] (Application example 2)

[0323] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0324] Previously, users needed advanced music production skills and specialized equipment to transform their own songs into professional-quality music. Furthermore, there was limited technology available to capture the emotions conveyed in the song. This made it difficult for many users to easily create high-quality music.

[0325] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting a melody line and lyrics, means for recognizing a user's emotion from the analyzed voice data, and means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion data. This allows users to easily convert recorded songs into professional-quality music and enjoy music that reflects their emotions.

[0326] A "terminal" is an electronic device that a user uses to record audio, and includes a smartphone, tablet, etc.

[0327] "Audio Data" means audio information recorded by a user and stored in digital form.

[0328] A "server" is a computer system that receives recorded voice data and performs analysis and generation processing.

[0329] "Audio analysis" is the process of extracting elements such as melody lines, lyrics, pitch, and rhythmic patterns from recorded audio data.

[0330] A "melody line" refers to a sequence of notes that form a musical melody and constitutes the basic structure of a piece of music.

[0331] "Emotion recognition" is the process of identifying a user's emotional state from their voice data, determining whether their emotions are positive or negative.

[0332] A "generative AI model" is an algorithm that uses AI (artificial intelligence) technology to generate professional-quality melody lines based on input data.

[0333] "Sound quality optimization" is the process of applying audio filters such as noise reduction and reverb effects to improve the quality of your recordings.

[0334] A "noise reduction filter" is a filtering technique used to remove unwanted background noise from recorded audio data.

[0335] The "reverb effect" is an effect technique that adds artificial reverberation to recorded audio to increase the breadth and depth of the sound.

[0336] A "finished song" is the final musical work after the user's recorded vocals have been adjusted, a generated melody line has been generated, and sound quality optimization has been applied.

[0337] To implement this invention, the user, terminal, and server must fulfill their respective roles and work in cooperation with each other. The specific processing of each component and the hardware and software used will be described below.

[0338] Recording and Data Transmission

[0339] Users use devices such as smartphones to record audio. They start a recording app and tap the "Start Recording" button to begin recording, and audio data is collected using the smartphone's built-in microphone. When recording ends, the audio data is temporarily saved on the device. When the user presses the "Stop Recording" button, recording ends and the recorded audio data is sent to the server.

[0340] Voice analysis and emotion recognition

[0341] When the server receives the audio data, it uses a voice analysis module to analyze the recording. Specifically, it detects pitch, rhythm patterns, and lyric timing, and extracts the melody line and lyrics. At the same time, the emotion engine recognizes the user's emotions from the audio data. For example, if the user is singing "happily," the emotion engine will recognize a positive emotion.

[0342] Melody Generation

[0343] The server inputs the extracted melody line and the recognized emotion into a generative AI model, which then generates a professional-quality melody line based on the input data. For example, if a positive emotion is recognized, a bright and rhythmic melody line will be generated.

[0344] Music generation and sound quality optimization

[0345] The server then adjusts the user's recorded voice based on the generated professional-quality melody line, correcting it to fit the new melody line, and then uses a sound quality optimization module to apply noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in a professional-quality song.

[0346] Specific examples

[0347] For example, if a user records a song about "happy days," the server will recognize the positive emotions and generate an upbeat, rhythmic melody line, then add noise reduction and reverb effects to complete the final song.

[0348] Here are some example prompts to input to the generative AI model:

[0349] Generate a melody line that reflects positive emotions based on a happy singing voice recorded by the user.

[0350] Recording data: {Recording data}

[0351] Emotion: Positive, fun

[0352] Desired melody style: Bright and rhythmic

[0353] Submit and share your music

[0354] Users can check the completed song on their smartphone or other device and upload it to social media or content distribution platforms if they wish, allowing them to easily convert their recorded songs into professional-quality songs and share them.

[0355] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0356] Step 1:

[0357] The user launches the device's recording app and taps the "Start Recording" button. This causes the device to begin recording the user's voice using the built-in microphone. The input is the user's singing voice, and the output is the temporarily saved audio data.

[0358] Step 2:

[0359] When the user presses the "Stop Recording" button, recording ends. The device temporarily stores the recorded audio data. This audio data is used as input for subsequent processing.

[0360] Step 3:

[0361] The device sends the recorded audio data to the server. The input of this process is the temporarily stored audio data, and the output is the audio data sent to the server.

[0362] Step 4:

[0363] The server receives audio data from the terminal. It uses an audio analysis module to extract melody lines, pitches, rhythmic patterns, and lyric timing from the audio data. The input is the audio data, and the output is the extracted pitches, rhythmic patterns, melody lines, and lyric timing.

[0364] Step 5:

[0365] The server uses an emotion engine to recognize the user's emotion from the voice data. For example, if the user sings "happily," it recognizes a positive emotion. The input is the voice data, and the output is the recognized emotion data.

[0366] Step 6:

[0367] The server inputs the extracted melody line and the recognized emotion data into a generative AI model, which generates a new, professional-quality melody line based on the melody line and emotion data. The input is the melody line and emotion data, and the output is a new, professional-quality melody line.

[0368] Step 7:

[0369] The server then adjusts the user's recorded voice to fit the new melody line based on the generated professional-quality melody line. The input is the recorded voice data and the generated melody line, and the output is the adjusted voice data.

[0370] Step 8:

[0371] The server uses the sound quality optimization module to optimize the recorded audio with noise reduction filters and reverb effects. The input is the adjusted audio data, and the output is the optimized audio data.

[0372] Step 9:

[0373] The server generates a complete piece of music based on the optimized audio data and stores it in a database. The input is the optimized audio data, and the output is the stored complete piece of music.

[0374] Step 10:

[0375] The user accesses the server using a device and downloads the completed song. The user checks the song and, if desired, shares it on social media or a content distribution platform. The input is the completed song from the server, and the output is the downloaded song and the shared song.

[0376] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0377] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0378] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0379] [Second embodiment]

[0380] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0381] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0382] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0383] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0384] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0385] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0386] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0387] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0388] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0389] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0390] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0391] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0392] As an embodiment of the present invention, a system is provided that allows users to easily convert their own recorded songs into professional quality musical compositions.

[0393] System Overview

[0394] First, the user uses a device to record their own voice. A recording application is installed on the device, and the user starts the application and presses the "Start Recording" button to begin recording. Once recording is complete, the device sends the recorded data to the server.

[0395] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. It then uses an AI model to generate a professional-quality melody line based on the extracted melody line. It then uses the generated melody line to adjust the user's recorded voice to create a finished song. The finished song is then stored by the server and can be downloaded by the user via their device.

[0396] Explanation of program processing (with concrete examples)

[0397] 1. Start recording and send data

[0398] The user starts a recording app on their smartphone and sings the lyrics of the song they want to sing, for example, "One afternoon, relaxing in the garden," for one minute.

[0399] Once the recording is complete, the device sends the audio data to the server.

[0400] 2. Audio analysis and melody extraction

[0401] The server receives the audio data sent from the device and uses a speech analysis module to analyze the recording, for example, to identify the pitch and rhythmic patterns of the recorded lyrics, "One afternoon, relaxing in the garden."

[0402] The server extracts the lyrics and melody line and stores in a database the pitch and rhythm in which each word of the lyrics is sung.

[0403] 3. Melody adjustment and generation

[0404] The server inputs the extracted melody line and lyrics into the AI ​​model. For example, the lyric "One afternoon, relaxing in the garden" can be input into the AI ​​model to generate a more refined melody line.

[0405] Based on the input data, the AI ​​model generates professional-quality melody lines and enhances the original melody lines.

[0406] 4. Music Generation and Enhancement

[0407] The server uses the generated professional-quality melody line to adjust the recording of the user's voice, so that the user's voice is modified to fit the newly generated melody line.

[0408] The server uses an enhancement module to optimize the sound quality and apply filters and effects to improve the overall sound quality of the song, for example using a noise reduction filter or a reverb effect.

[0409] 5. Submission of completed music

[0410] The server generates the finished song and stores it in a database, which the user can then download from the server using their device.

[0411] Users can check the completed song and share it via social media or email if they wish.

[0412] In this way, the present invention provides a concrete means for users to easily create professional-quality music, allowing anyone to easily enjoy the joy of music production, even without specialized knowledge or skills.

[0413] The processing flow will be explained below.

[0414] Step 1:

[0415] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording.

[0416] Step 2:

[0417] The device uses a built-in microphone to record the user's singing voice and saves the recorded data as a temporary file.

[0418] Step 3:

[0419] When the user presses the button to stop recording, the recording ends.

[0420] Step 4:

[0421] The terminal transmits the recorded voice data to the server.

[0422] Step 5:

[0423] The server receives and stores the voice data sent from the terminal.

[0424] Step 6:

[0425] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythm, and lyric timing.

[0426] Step 7:

[0427] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0428] Step 8:

[0429] The server inputs the extracted melody line and lyrics into the AI ​​model.

[0430] Step 9:

[0431] The AI ​​model generates a professional-quality melody line based on the input data and returns the generated melody line to the server.

[0432] Step 10:

[0433] The server adjusts the user's recorded voice based on the generated melody line and modifies it to fit the new melody line.

[0434] Step 11:

[0435] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song.

[0436] Step 12:

[0437] The server generates professional quality finished music and stores it in a database.

[0438] Step 13:

[0439] The user uses a terminal to access the server and download the completed song.

[0440] Step 14:

[0441] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0442] Through this series of steps, users can transform their recorded songs into professional-quality songs and easily enjoy the music production process.

[0443] Example 1

[0444] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0445] In the past, users needed advanced musical knowledge and specialized equipment to create professional-quality music, making it extremely difficult for the average user. Furthermore, adjusting recorded audio to match a high-quality melody line required specialized skills. Therefore, a method for easily generating high-quality music was needed.

[0446] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0447] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed piece of music, and means for providing the generated completed piece of music to the user, thereby enabling users to easily generate high-quality music without having specialized knowledge or skills.

[0448] "User" means any person or entity that intends to use the System to record audio and generate musical compositions.

[0449] "Audio recording means" refers to a device or application that a user uses to record their own audio.

[0450] "Means for transmitting the recorded voice data to the server" refers to the communications protocols and procedures for uploading the recorded voice data to the server via the Internet.

[0451] "Server" refers to a computer system for receiving, analyzing, and processing recorded audio data.

[0452] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms and software for automatically analyzing and extracting pitch, rhythm patterns, and lyrics from recorded audio data.

[0453] "Generative AI Model" refers to an artificial intelligence program or algorithm used to generate professional-quality melody lines based on extracted melody lines and lyrics.

[0454] "Professional quality melody line" refers to a melody line of professional quality.

[0455] "Means for adjusting the audio and generating a finished piece of music" refers to the methods and techniques for modifying the recorded audio to match the generated melody line and generating the final piece of music.

[0456] "Means for providing the completed song to the user" refers to the method or procedure for providing the final song to the user in a downloadable format.

[0457] "Means for extracting pitch and rhythmic patterns" refers to algorithms or software that identify and extract pitch and rhythmic patterns from recorded audio data.

[0458] "Means for sound quality optimization and applying filters and effects" refers to techniques for applying audio filters and effects to improve the sound quality of the generated music.

[0459] "System" refers to the overall setup that includes the elements mentioned above and enables users to create professional quality music.

[0460] The present invention relates to a system that allows users to easily create professional quality music. Specific means for carrying out the invention will be described in detail below.

[0461] A user first launches a recording application on a device such as a smartphone or tablet. This application provides a recording function, allowing the user to record their own voice. Once recording is complete, the device sends the recorded data (voice data) to a server. To do this, the device establishes a connection to the server via the Internet and sends the voice data using a protocol such as an HTTP POST request.

[0462] The server uses a receiving module to receive audio data sent from the device. This audio data is stored in a database on the server. The server then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. The audio analysis module uses libraries such as FFT (Fast Fourier Transform) and LibROSA to analyze the frequency components of the audio data. This allows detailed analysis data of the recorded audio to be obtained.

[0463] The analyzed data is stored in a database on the server, and the next step is to input it into a generative AI model. The generative AI model generates a professional-quality melody line based on the extracted melody line and lyrics. This model can use, for example, a generative AI model or other similar artificial intelligence algorithms. As a concrete example, the prompt sentence is as follows:

[0464] Prompt statement:

[0465] "Generate a professional-quality melody line for "One Afternoon Relaxing in the Garden" based on your voice data. The voice data is in the attached file."

[0466] The generative AI model takes this prompt and the attached file and generates a professional-quality melody line. The generated melody line is sent back to the server and stored in a database. The server then uses this generated melody line to adjust the user's recorded voice, using techniques such as pitch shifting and time stretching to modify the user's voice to fit the new melody line.

[0467] The server then uses an enhancement module to optimize the sound quality, applying noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in the final finished song.

[0468] Finally, the server stores the completed song in a database and provides it to the user, who can download it from the server using their device. Users can also share the completed song via social media or email.

[0469] In this way, the present invention provides a concrete means for users to easily create professional-quality music, and is a system that allows anyone to easily enjoy the fun of music production, even without specialized knowledge or skills.

[0470] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0471] Step 1:

[0472] The user launches a recording application on their smartphone or tablet. They press the "Start Recording" button and sing the lyrics of the song they want to sing. They sing specific lyrics, such as "One afternoon, relaxing in the garden," for one minute. Once the recording is complete, the device sends the recording to the server. The input is the audio data sung by the user, and the output is the audio data sent to the server.

[0473] Step 2:

[0474] The server receives the audio data sent from the device. The received audio data is stored in a database on the server. It then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. Specifically, it analyzes the frequency components of the audio data using libraries such as FFT (Fast Fourier Transform) and LibROSA. The input is the audio data sent to the server, and the output is the analyzed pitch, rhythm patterns, and timing information for the lyrics.

[0475] Step 3:

[0476] The server inputs the melody line and lyrics extracted from the analyzed audio data into the generative AI model. A prompt is sent to the generative AI model, which generates a professional-quality melody line according to the prompt. For example, the prompt could be, "Based on the user's audio data, please generate a professional-quality melody line for 'Relaxing in the Garden One Afternoon.' The audio data is in the attached file." The input is the extracted melody line and lyrics, and the output is the professional-quality melody line returned by the generative AI model.

[0477] Step 4:

[0478] The server adjusts the user's recorded voice based on the generated professional-quality melody line. Techniques used include pitch shifting and time stretching, which modify the user's voice to fit the new melody line. The input is the generated melody line and the user's recording, and the output is the adjusted voice data.

[0479] Step 5:

[0480] The server uses an enhancement module to optimize the sound quality. It applies noise reduction filters and reverb effects to improve the overall sound quality of the song. Specifically, it removes recorded background noise and adds reverb to add depth to the audio. The input is the adjusted audio data, and the output is an enhanced, sound-optimized finished song.

[0481] Step 6:

[0482] The server stores the generated completed music in a database and provides it to the user. The user can download and check the generated music on their own device. They can also share it via social media or email. The input is the enhanced completed music, and the output is music data that can be accessed by providing the user with a download link.

[0483] (Application example 1)

[0484] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0485] The traditional music production process requires specialized knowledge and skills, making it difficult for many users to easily create professional-quality music. Furthermore, there are limited ways to easily share completed music, and in particular, there is a lack of an environment in place that allows ordinary users to easily upload their own music to streaming platforms.

[0486] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0487] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed song, means for providing the generated completed song to the user, and means for uploading the completed song to a streaming platform. This enables users to create high-quality songs without requiring specialized knowledge or skills and easily upload them to a streaming platform for sharing.

[0488] "User" means an individual or entity that uses the system to record audio and generate and upload music.

[0489] An "audio recording means" is a device or application that a user uses to record a song.

[0490] "Means for transmitting recorded voice data to a server" refers to the process or function that transmits a user's recorded voice data to a server.

[0491] A "server" is a central computer system for analyzing and processing audio data.

[0492] "Means for performing audio analysis" refers to software or algorithms for extracting melody lines and lyrics from recorded audio data.

[0493] The "means for extracting a melody line" is a function for identifying pitch and rhythm patterns from audio data.

[0494] A "generative AI model" is an artificial intelligence (AI) model that generates professional-quality melody lines based on extracted melody lines.

[0495] "Means for generating professional-quality melody lines" refers to the process of using a generative AI model to create high-quality melodies based on a user's recordings.

[0496] The "means for adjusting the recorded voice and generating a completed piece of music" refers to the process of adjusting the user's recorded voice to match the generated melody line, thereby generating a completed piece of music as a whole.

[0497] The "means for providing the generated completed music to the user" is a process or function that provides the completed music in a form that the user can access and download.

[0498] "Means for uploading the completed song to a streaming platform" refers to a function for uploading the completed song to a streaming service on the Internet.

[0499] The present invention provides a system that allows users to convert their recorded songs into professional quality music and easily upload the music to a streaming platform. Specific embodiments of the present invention are described below.

[0500] System Overview

[0501] 1. Recording and data transmission

[0502] The user starts a recording application on a device such as a smartphone and records a song. When the recording is finished, the device sends the audio data to the server.

[0503] For example, a user can sing a song such as "One Afternoon, Relaxing in the Garden" using a recording app on their smartphone and send the recording data to a server.

[0504] 2. Audio analysis and melody extraction

[0505] The server receives the transmitted audio data and analyzes it using audio analysis software (e.g., librosa), which extracts pitch, rhythm patterns, and the timing of lyrics.

[0506] For example, the pitch and rhythmic patterns of the recorded lyrics "One afternoon, relaxing in the garden" are analyzed.

[0507] 3. Melody Generation

[0508] The server inputs the extracted melody line and lyrics into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line.

[0509] For example, a sophisticated melody line can be generated based on the lyrics "One afternoon relaxing in the garden."

[0510] 4. Audio adjustment and music generation

[0511] The server then uses the generated professional-quality melody line to adjust the recorded audio and generate the finished song, applying filters and effects (e.g., noise reduction filters and reverb effects) to optimize the sound quality.

[0512] For example, the user's recording is adjusted to fit the newly generated melody line, and noise reduction filters and reverb effects are applied.

[0513] 5. Providing music and uploading to streaming services

[0514] The server generates the finished song and provides it to the user, who can then download it or upload it to a streaming platform.

[0515] For example, users can review the finished song on their smartphone and upload it directly to a streaming platform if they wish.

[0516] Examples of specific examples and prompts

[0517] As a concrete example, consider a user relaxing in their garden and recording a song:

[0518] Example prompt sentence:

[0519] Users can simply record a song on their smartphone while relaxing in the garden one afternoon. The app then sends the recording to a server, where an AI model generates a professional-quality melody line. Users can then upload the finished song directly to streaming platforms.

[0520] This system allows users to create professional-quality music without any specialized knowledge or skills, and then easily upload and share the music on streaming platforms.

[0521] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0522] Step 1:

[0523] A user records a song using a recording app on their smartphone. The user launches the app, presses the "Start Recording" button, and sings. When the recording is finished, the application saves the audio data in a file format. The input of this step is the user's singing voice, and the output is the recorded audio file.

[0524] Step 2:

[0525] The device sends the recorded audio data to the server. After the recording is complete, the device uses the data transmission function to upload the audio file to the server. The input of this step is the saved audio file, and the output is the audio data uploaded to the server.

[0526] Step 3:

[0527] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. Audio analysis software (e.g., librosa) is used for the analysis. Specifically, the audio data is loaded using the librosa.load function, and the pitch and rhythmic patterns are extracted using the librosa.piptrack function. The input of this step is the audio data uploaded to the server, and the output is the analyzed melody line and rhythmic patterns.

[0528] Step 4:

[0529] The server inputs the extracted melody line and rhythm pattern into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line. Specifically, the analyzed melody line and rhythm pattern are fed into the AI ​​model, and a new melody line is generated using the model's inference function. The input of this step is the extracted melody line and rhythm pattern, and the output is the generated professional-quality melody line.

[0530] Step 5:

[0531] The server then uses the generated, professional-quality melody line to adjust the recorded audio and generate a finished song. Specifically, it uses audio editing software to adjust the user's recording to fit the new melody line and apply filters and effects (e.g., noise reduction filters and reverb effects). The input to this step is the generated, professional-quality melody line and the recorded audio data, and the output is the finished song.

[0532] Step 6:

[0533] The server provides the generated finished song to the user. Specifically, the server stores the song file on the server so that the user can download it, and provides the user with a download link. The input of this step is the finished song, and the output is a song file that the user can download.

[0534] Step 7:

[0535] The user uploads the completed song to a streaming platform. Specifically, the user uploads the completed song to the platform using the streaming platform's upload function provided by the smartphone application. The input of this step is the song file downloaded by the user, and the output is the song uploaded to the streaming platform.

[0536] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0537] As an embodiment of the present invention, a system is provided that allows a user to easily convert their own recorded songs into professional quality music, and furthermore, reflects emotional elements in the music using an emotion engine.

[0538] System Overview

[0539] The present invention consists of the following main components:

[0540] 1. A means for users to record audio (a device with a recording application)

[0541] 2. A means of sending recorded audio data to the server

[0542] 3. A method for analyzing audio on the server and extracting melody lines and lyrics

[0543] 4. A means to recognize emotions from the user's voice using an emotion engine

[0544] 5. A method for generating professional-quality melody lines using an AI model based on the extracted melody lines and the recognized emotions.

[0545] 6. A method for generating a complete song using the generated melody line and the adjusted user's voice

[0546] 7. Means of sound quality optimization and application of filters and effects

[0547] 8. Means for providing the generated finished music to the user

[0548] Explanation of program processing (with concrete examples)

[0549] Start recording and send data

[0550] 1. The user starts the recording app on their device (such as a smartphone) and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0551] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0552] 3. When the user presses the button to stop recording, the recording ends.

[0553] 4. The device sends the recorded audio data to the server.

[0554] Voice analysis, emotion recognition, and melody generation

[0555] 5. The server receives the voice data sent from the terminal.

[0556] 6. The server uses an audio analysis module to analyze the recording and detect pitch, rhythmic patterns, and lyric timing.

[0557] 7. The server extracts the melody line and lyrics based on the analysis results.

[0558] 8. The emotion engine on the server recognizes the user's emotion from the voice data. For example, if the user sings "happily," the emotion engine recognizes a positive emotion.

[0559] 9. The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0560] 10. The AI ​​model uses input data to generate professional-quality melody lines that match emotions. For example, if a positive emotion is recognized, an upbeat, rhythmic melody line will be generated.

[0561] Music Generation and Enhancement

[0562] 11. The server then adjusts the user's recorded voice based on the generated professional-quality melody line, modifying it to fit the new melody line.

[0563] 12. The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, for example, using a noise reduction filter or a reverb effect.

[0564] 13. The server generates a professional quality finished song and stores it in a database.

[0565] Submit and share your music

[0566] 14. The user uses the terminal to access the server and download the completed song.

[0567] 15. Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0568] Through this series of steps, users can transform their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0569] The processing flow will be explained below.

[0570] Step 1:

[0571] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording. For example, they sing lyrics such as "One afternoon, relaxing in the garden."

[0572] Step 2:

[0573] The device uses a built-in microphone to record the user's singing voice and saves the recording data as a temporary file.

[0574] Step 3:

[0575] When the user presses the button to stop recording, the recording ends.

[0576] Step 4:

[0577] The terminal transmits the recorded voice data to the server.

[0578] Step 5:

[0579] The server receives and stores the voice data sent from the terminal.

[0580] Step 6:

[0581] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythmic patterns, and lyric timing.

[0582] Step 7:

[0583] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0584] Step 8:

[0585] The server's emotion engine recognizes the user's emotions from the voice data. For example, it recognizes the positive emotion of "feeling happy" from the user's tone of voice and intonation.

[0586] Step 9:

[0587] The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0588] Step 10:

[0589] The AI ​​model uses the input data to generate professional-quality melody lines that match the emotion. For example, if a positive emotion is recognized, it will generate an upbeat, rhythmic melody line.

[0590] Step 11:

[0591] The server then adjusts the user's recorded voice based on the generated melody line, correcting it to fit the new melody line, including adjusting the pitch and timing.

[0592] Step 12:

[0593] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, such as noise reduction filters and reverb effects.

[0594] Step 13:

[0595] The server generates professional quality finished music and stores it in a database.

[0596] Step 14:

[0597] The user uses a terminal to access the server and download the completed song.

[0598] Step 15:

[0599] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0600] Through this series of processes, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0601] Example 2

[0602] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0603] Previous music production systems required specialized knowledge and skills, as well as a great deal of time and effort, to convert users' recorded audio into professional-quality music. It was also difficult for average users to incorporate emotional elements into music. This made music production a challenging process, making it difficult for anyone to easily create high-quality music.

[0604] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting melody lines and lyrics, means for recognizing emotions from voice data, and means for generating professional-quality melody lines using a generative AI model based on the extracted melody lines and the recognized emotions. This allows anyone to easily generate high-quality music and enjoy music that reflects their own emotions.

[0605] A "voice recording means" is an application or device that allows a user to record their own voice.

[0606] "Means for transmitting audio data to a server" refers to a communication means for uploading recorded audio data to a server, which includes an internet connection and a corresponding application.

[0607] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms or software for analyzing recorded audio data to identify and extract melody lines and lyrics therefrom.

[0608] The "means for recognizing emotions" is an algorithm or software for analyzing voice data and recognizing the user's emotions.

[0609] "Means for generating professional-quality melody lines using generative AI models" refers to a system or software that uses AI technology to automatically generate high-quality melody lines.

[0610] "Means for adjusting recorded voice and generating a completed piece of music" refers to a system or software for adjusting a user's recorded voice to match the generated melody line and combining them into a completed piece of music.

[0611] "Means for applying filters and effects to optimize sound quality" refers to a system or software for applying various audio filters and effects to improve the sound quality of the generated music.

[0612] "Means for providing the completed music to the user" refers to means for providing the created music in a form that the user can access, including providing a download link or using cloud storage.

[0613] The present invention provides a system that allows users to easily convert their own recorded songs into professional quality music and further reflects emotional elements in the music using an emotion engine. Specific embodiments will be described below.

[0614] Hardware and Software Configuration

[0615] User terminal

[0616] The user terminal is a device such as a smartphone or tablet that has a recording application installed. This terminal has a built-in microphone that is used to record the user's voice. It also has a communication function to send the recorded data to the server.

[0617] server

[0618] The server contains multiple modules for processing the received voice data, including a voice analysis module, an emotion recognition engine, an AI model, and a sound quality optimization module.

[0619] Voice recording and data transmission

[0620] 1. The user launches the recording app on the device and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0621] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0622] 3. When the user presses the stop recording button, the recording ends.

[0623] 4. The device sends the recorded audio data to the server via an HTTP POST request, using the HTTPS protocol.

[0624] Voice analysis and emotion recognition

[0625] 5. The server receives the audio data sent from the device and analyzes the recording using the Python LibROSA library.

[0626] 6. The server detects pitch, rhythmic patterns, and lyric timing and stores them in an internal database (e.g., MongoDB) in JSON format.

[0627] 7. The server's emotion engine uses an emotion recognition algorithm (e.g., Microsoft Azure Cognitive Services) to recognize the user's emotion from the voice data.

[0628] Melody line generation

[0629] 8. The server inputs the extracted melody line and the recognized emotion into the AI ​​model. For example, it inputs the following prompt to the AI ​​model: "Create a fun melody line in the style of James Brown."

[0630] 9. The AI ​​model uses the input data to generate a professional-quality melody line tailored to the emotion, using OpenAI GPT-3 or other music generation models.

[0631] Music Generation and Enhancement

[0632] 10. The server uses Auto-tune to adjust the user's recorded voice based on the generated professional-quality melody line.

[0633] 11. The server uses a sound quality optimization module (e.g., iZotope Neutron) to apply noise reduction filters and reverb effects to improve the overall sound quality of the song.

[0634] Providing completed music

[0635] 12. The server stores the generated professional quality finished music in a database.

[0636] 13. The user can use their device to access the server and download the completed song from the "My Songs" section.

[0637] Through this process, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0638] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0639] Step 1:

[0640] The user launches the device's recording app and taps the "Start Recording" button to begin recording. The input is the user's button operation, and the output is the trigger to start recording. Specifically, the recording app accesses the built-in microphone and begins capturing audio in real time.

[0641] Step 2:

[0642] The device uses a built-in microphone to record the user's singing voice and temporarily saves the recorded data. The input is the user's voice, and the output is the recorded voice data. Specifically, the recorded data is temporarily saved in the device's local storage in wav format.

[0643] Step 3:

[0644] Recording ends when the user presses the stop recording button. The input is the user's button operation, and the output is the end of the recording session. Specifically, the recording data capture stops and the final data file is committed.

[0645] Step 4:

[0646] The device sends the recorded audio data to the server. The input is the recorded data, and the output is the data sent to the server. Specifically, the audio data is sent to the server's API endpoint via an HTTP POST request, and the HTTPS protocol is used for communication.

[0647] Step 5:

[0648] The server receives the voice data sent from the device and stores it in storage. The input is the recorded data, and the output is the stored data. Specifically, the received data is stored in file storage and placed in a waiting state for processing.

[0649] Step 6:

[0650] The server uses an audio analysis module to analyze the recorded data and detect pitch, rhythmic patterns, and lyric timing. The input is the stored audio data, and the output is the analysis results. Specifically, it uses the Python LibROSA library to perform spectral analysis and pitch detection of the audio data.

[0651] Step 7:

[0652] The server extracts the melody line and lyrics based on the analysis results. The input is the audio analysis results, and the output is the extracted melody line and lyrics. Specifically, the analysis results are saved in JSON format in an internal database (e.g., MongoDB).

[0653] Step 8:

[0654] The server's emotion engine recognizes the user's emotions from the voice data. The input is the recorded data, and the output is an emotion score. Specifically, it uses Microsoft's Azure Cognitive Services to extract emotions from the voice data in real time and calculate the emotion score.

[0655] Step 9:

[0656] The server inputs the extracted melody line and the recognized emotion into the AI ​​model. The input is the melody line and emotion score, and the output is the generated melody line. Specifically, the server inputs the following prompt to the AI ​​model: "Please create a fun melody line in the style of James Brown."

[0657] Step 10:

[0658] The AI ​​model uses input data to generate a professional-quality melody line that matches the emotion. The input is a prompt, and the output is the generated melody line. Specifically, the generative AI model (e.g., OpenAI GPT-3 or other music generation model) generates a professional-quality melody line and returns the data to the server.

[0659] Step 11:

[0660] The server adjusts the user's recorded voice based on the generated professional-quality melody line. The input is the user's voice and the generated melody line, and the output is the adjusted song. Specifically, it uses Auto-tune to ensure consistency between the voice and the melody line.

[0661] Step 12:

[0662] The server uses a sound quality optimization module to improve the overall sound quality of the track. The input is the adjusted track, and the output is the sound-optimized track. Specifically, it uses iZotope's Neutron to apply noise reduction filters and reverb effects.

[0663] Step 13:

[0664] The server generates professional-quality finished songs and stores them in a database. The input is a quality-optimized song, and the output is a saved song file. Specifically, the final audio file is saved in the database in MP3 or WAV format.

[0665] Step 14:

[0666] The user uses their device to access the server and download the completed song. The input is the user's request, and the output is the downloaded song file. Specifically, the user selects a song from the "My Songs" section within the recording app and downloads it securely via HTTPS.

[0667] Step 15:

[0668] Users can play the completed song to check its contents, and if they wish, share it on social media or by email. The input is the user's operation, and the output is the shared song. Specifically, an API is used to share songs directly to social media from within the app, providing a seamless sharing experience.

[0669] (Application example 2)

[0670] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0671] Previously, users needed advanced music production skills and specialized equipment to transform their own songs into professional-quality music. Furthermore, there was limited technology available to capture the emotions conveyed in the song. This made it difficult for many users to easily create high-quality music.

[0672] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting a melody line and lyrics, means for recognizing a user's emotion from the analyzed voice data, and means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion data. This allows users to easily convert recorded songs into professional-quality music and enjoy music that reflects their emotions.

[0673] A "terminal" is an electronic device that a user uses to record audio, and includes a smartphone, tablet, etc.

[0674] "Audio Data" means audio information recorded by a user and stored in digital form.

[0675] A "server" is a computer system that receives recorded voice data and performs analysis and generation processing.

[0676] "Audio analysis" is the process of extracting elements such as melody lines, lyrics, pitch, and rhythmic patterns from recorded audio data.

[0677] A "melody line" refers to a sequence of notes that form a musical melody and constitutes the basic structure of a piece of music.

[0678] "Emotion recognition" is the process of identifying a user's emotional state from their voice data, determining whether their emotions are positive or negative.

[0679] A "generative AI model" is an algorithm that uses AI (artificial intelligence) technology to generate professional-quality melody lines based on input data.

[0680] "Sound quality optimization" is the process of applying audio filters such as noise reduction and reverb effects to improve the quality of your recordings.

[0681] A "noise reduction filter" is a filtering technique used to remove unwanted background noise from recorded audio data.

[0682] The "reverb effect" is an effect technique that adds artificial reverberation to recorded audio to increase the breadth and depth of the sound.

[0683] A "finished song" is the final musical work after the user's recorded vocals have been adjusted, a generated melody line has been generated, and sound quality optimization has been applied.

[0684] To implement this invention, the user, terminal, and server must fulfill their respective roles and work in cooperation with each other. The specific processing of each component and the hardware and software used will be described below.

[0685] Recording and Data Transmission

[0686] Users use devices such as smartphones to record audio. They start a recording app and tap the "Start Recording" button to begin recording, and audio data is collected using the smartphone's built-in microphone. When recording ends, the audio data is temporarily saved on the device. When the user presses the "Stop Recording" button, recording ends and the recorded audio data is sent to the server.

[0687] Voice analysis and emotion recognition

[0688] When the server receives the audio data, it uses a voice analysis module to analyze the recording. Specifically, it detects pitch, rhythm patterns, and lyric timing, and extracts the melody line and lyrics. At the same time, the emotion engine recognizes the user's emotions from the audio data. For example, if the user is singing "happily," the emotion engine will recognize a positive emotion.

[0689] Melody Generation

[0690] The server inputs the extracted melody line and the recognized emotion into a generative AI model, which then generates a professional-quality melody line based on the input data. For example, if a positive emotion is recognized, a bright and rhythmic melody line will be generated.

[0691] Music generation and sound quality optimization

[0692] The server then adjusts the user's recorded voice based on the generated professional-quality melody line, correcting it to fit the new melody line, and then uses a sound quality optimization module to apply noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in a professional-quality song.

[0693] Specific examples

[0694] For example, if a user records a song about "happy days," the server will recognize the positive emotions and generate an upbeat, rhythmic melody line, then add noise reduction and reverb effects to complete the final song.

[0695] Here are some example prompts to input to the generative AI model:

[0696] Generate a melody line that reflects positive emotions based on a happy singing voice recorded by the user.

[0697] Recording data: {Recording data}

[0698] Emotion: Positive, fun

[0699] Desired melody style: Bright and rhythmic

[0700] Submit and share your music

[0701] Users can check the completed song on their smartphone or other device and upload it to social media or content distribution platforms if they wish, allowing them to easily convert their recorded songs into professional-quality songs and share them.

[0702] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0703] Step 1:

[0704] The user launches the device's recording app and taps the "Start Recording" button. This causes the device to begin recording the user's voice using the built-in microphone. The input is the user's singing voice, and the output is the temporarily saved audio data.

[0705] Step 2:

[0706] When the user presses the "Stop Recording" button, recording ends. The device temporarily stores the recorded audio data. This audio data is used as input for subsequent processing.

[0707] Step 3:

[0708] The device sends the recorded audio data to the server. The input of this process is the temporarily stored audio data, and the output is the audio data sent to the server.

[0709] Step 4:

[0710] The server receives audio data from the terminal. It uses an audio analysis module to extract melody lines, pitches, rhythmic patterns, and lyric timing from the audio data. The input is the audio data, and the output is the extracted pitches, rhythmic patterns, melody lines, and lyric timing.

[0711] Step 5:

[0712] The server uses an emotion engine to recognize the user's emotion from the voice data. For example, if the user sings "happily," it recognizes a positive emotion. The input is the voice data, and the output is the recognized emotion data.

[0713] Step 6:

[0714] The server inputs the extracted melody line and the recognized emotion data into a generative AI model, which generates a new, professional-quality melody line based on the melody line and emotion data. The input is the melody line and emotion data, and the output is a new, professional-quality melody line.

[0715] Step 7:

[0716] The server then adjusts the user's recorded voice to fit the new melody line based on the generated professional-quality melody line. The input is the recorded voice data and the generated melody line, and the output is the adjusted voice data.

[0717] Step 8:

[0718] The server uses the sound quality optimization module to optimize the recorded audio with noise reduction filters and reverb effects. The input is the adjusted audio data, and the output is the optimized audio data.

[0719] Step 9:

[0720] The server generates a complete piece of music based on the optimized audio data and stores it in a database. The input is the optimized audio data, and the output is the stored complete piece of music.

[0721] Step 10:

[0722] The user accesses the server using a device and downloads the completed song. The user checks the song and, if desired, shares it on social media or a content distribution platform. The input is the completed song from the server, and the output is the downloaded song and the shared song.

[0723] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0724] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0725] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0726] [Third embodiment]

[0727] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0728] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0729] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0730] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0731] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0732] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0733] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0734] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0735] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0736] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0737] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0738] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0739] As an embodiment of the present invention, a system is provided that allows users to easily convert their own recorded songs into professional quality musical compositions.

[0740] System Overview

[0741] First, the user uses a device to record their own voice. A recording application is installed on the device, and the user starts the application and presses the "Start Recording" button to begin recording. Once recording is complete, the device sends the recorded data to the server.

[0742] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. It then uses an AI model to generate a professional-quality melody line based on the extracted melody line. It then uses the generated melody line to adjust the user's recorded voice to create a finished song. The finished song is then stored by the server and can be downloaded by the user via their device.

[0743] Explanation of program processing (with concrete examples)

[0744] 1. Start recording and send data

[0745] The user starts a recording app on their smartphone and sings the lyrics of the song they want to sing, for example, "One afternoon, relaxing in the garden," for one minute.

[0746] Once the recording is complete, the device sends the audio data to the server.

[0747] 2. Audio analysis and melody extraction

[0748] The server receives the audio data sent from the device and uses a speech analysis module to analyze the recording, for example, to identify the pitch and rhythmic patterns of the recorded lyrics, "One afternoon, relaxing in the garden."

[0749] The server extracts the lyrics and melody line and stores in a database the pitch and rhythm in which each word of the lyrics is sung.

[0750] 3. Melody adjustment and generation

[0751] The server inputs the extracted melody line and lyrics into the AI ​​model. For example, the lyric "One afternoon, relaxing in the garden" can be input into the AI ​​model to generate a more refined melody line.

[0752] Based on the input data, the AI ​​model generates professional-quality melody lines and enhances the original melody lines.

[0753] 4. Music Generation and Enhancement

[0754] The server uses the generated professional-quality melody line to adjust the recording of the user's voice, so that the user's voice is modified to fit the newly generated melody line.

[0755] The server uses an enhancement module to optimize the sound quality and apply filters and effects to improve the overall sound quality of the song, for example using a noise reduction filter or a reverb effect.

[0756] 5. Submission of completed music

[0757] The server generates the finished song and stores it in a database, which the user can then download from the server using their device.

[0758] Users can check the completed song and share it via social media or email if they wish.

[0759] In this way, the present invention provides a concrete means for users to easily create professional-quality music, allowing anyone to easily enjoy the joy of music production, even without specialized knowledge or skills.

[0760] The processing flow will be explained below.

[0761] Step 1:

[0762] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording.

[0763] Step 2:

[0764] The device uses a built-in microphone to record the user's singing voice and saves the recorded data as a temporary file.

[0765] Step 3:

[0766] When the user presses the button to stop recording, the recording ends.

[0767] Step 4:

[0768] The terminal transmits the recorded voice data to the server.

[0769] Step 5:

[0770] The server receives and stores the voice data sent from the terminal.

[0771] Step 6:

[0772] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythm, and lyric timing.

[0773] Step 7:

[0774] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0775] Step 8:

[0776] The server inputs the extracted melody line and lyrics into the AI ​​model.

[0777] Step 9:

[0778] The AI ​​model generates a professional-quality melody line based on the input data and returns the generated melody line to the server.

[0779] Step 10:

[0780] The server adjusts the user's recorded voice based on the generated melody line and modifies it to fit the new melody line.

[0781] Step 11:

[0782] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song.

[0783] Step 12:

[0784] The server generates professional quality finished music and stores it in a database.

[0785] Step 13:

[0786] The user uses a terminal to access the server and download the completed song.

[0787] Step 14:

[0788] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0789] Through this series of steps, users can transform their recorded songs into professional-quality songs and easily enjoy the music production process.

[0790] Example 1

[0791] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0792] In the past, users needed advanced musical knowledge and specialized equipment to create professional-quality music, making it extremely difficult for the average user. Furthermore, adjusting recorded audio to match a high-quality melody line required specialized skills. Therefore, a method for easily generating high-quality music was needed.

[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0794] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed piece of music, and means for providing the generated completed piece of music to the user, thereby enabling users to easily generate high-quality music without having specialized knowledge or skills.

[0795] "User" means any person or entity that intends to use the System to record audio and generate musical compositions.

[0796] "Audio recording means" refers to a device or application that a user uses to record their own audio.

[0797] "Means for transmitting the recorded voice data to the server" refers to the communications protocols and procedures for uploading the recorded voice data to the server via the Internet.

[0798] "Server" refers to a computer system for receiving, analyzing, and processing recorded audio data.

[0799] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms and software for automatically analyzing and extracting pitch, rhythm patterns, and lyrics from recorded audio data.

[0800] "Generative AI Model" refers to an artificial intelligence program or algorithm used to generate professional-quality melody lines based on extracted melody lines and lyrics.

[0801] "Professional quality melody line" refers to a melody line of professional quality.

[0802] "Means for adjusting the audio and generating a finished piece of music" refers to the methods and techniques for modifying the recorded audio to match the generated melody line and generating the final piece of music.

[0803] "Means for providing the completed song to the user" refers to the method or procedure for providing the final song to the user in a downloadable format.

[0804] "Means for extracting pitch and rhythmic patterns" refers to algorithms or software that identify and extract pitch and rhythmic patterns from recorded audio data.

[0805] "Means for sound quality optimization and applying filters and effects" refers to techniques for applying audio filters and effects to improve the sound quality of the generated music.

[0806] "System" refers to the overall setup that includes the elements mentioned above and enables users to create professional quality music.

[0807] The present invention relates to a system that allows users to easily create professional quality music. Specific means for carrying out the invention will be described in detail below.

[0808] A user first launches a recording application on a device such as a smartphone or tablet. This application provides a recording function, allowing the user to record their own voice. Once recording is complete, the device sends the recorded data (voice data) to a server. To do this, the device establishes a connection to the server via the Internet and sends the voice data using a protocol such as an HTTP POST request.

[0809] The server uses a receiving module to receive audio data sent from the device. This audio data is stored in a database on the server. The server then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. The audio analysis module uses libraries such as FFT (Fast Fourier Transform) and LibROSA to analyze the frequency components of the audio data. This allows detailed analysis data of the recorded audio to be obtained.

[0810] The analyzed data is stored in a database on the server, and the next step is to input it into a generative AI model. The generative AI model generates a professional-quality melody line based on the extracted melody line and lyrics. This model can use, for example, a generative AI model or other similar artificial intelligence algorithms. As a concrete example, the prompt sentence is as follows:

[0811] Prompt statement:

[0812] "Generate a professional-quality melody line for "One Afternoon Relaxing in the Garden" based on your voice data. The voice data is in the attached file."

[0813] The generative AI model takes this prompt and the attached file and generates a professional-quality melody line. The generated melody line is sent back to the server and stored in a database. The server then uses this generated melody line to adjust the user's recorded voice, using techniques such as pitch shifting and time stretching to modify the user's voice to fit the new melody line.

[0814] The server then uses an enhancement module to optimize the sound quality, applying noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in the final finished song.

[0815] Finally, the server stores the completed song in a database and provides it to the user, who can download it from the server using their device. Users can also share the completed song via social media or email.

[0816] In this way, the present invention provides a concrete means for users to easily create professional-quality music, and is a system that allows anyone to easily enjoy the fun of music production, even without specialized knowledge or skills.

[0817] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0818] Step 1:

[0819] The user launches a recording application on their smartphone or tablet. They press the "Start Recording" button and sing the lyrics of the song they want to sing. They sing specific lyrics, such as "One afternoon, relaxing in the garden," for one minute. Once the recording is complete, the device sends the recording to the server. The input is the audio data sung by the user, and the output is the audio data sent to the server.

[0820] Step 2:

[0821] The server receives the audio data sent from the device. The received audio data is stored in a database on the server. It then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. Specifically, it analyzes the frequency components of the audio data using libraries such as FFT (Fast Fourier Transform) and LibROSA. The input is the audio data sent to the server, and the output is the analyzed pitch, rhythm patterns, and timing information for the lyrics.

[0822] Step 3:

[0823] The server inputs the melody line and lyrics extracted from the analyzed audio data into the generative AI model. A prompt is sent to the generative AI model, which generates a professional-quality melody line according to the prompt. For example, the prompt could be, "Based on the user's audio data, please generate a professional-quality melody line for 'Relaxing in the Garden One Afternoon.' The audio data is in the attached file." The input is the extracted melody line and lyrics, and the output is the professional-quality melody line returned by the generative AI model.

[0824] Step 4:

[0825] The server adjusts the user's recorded voice based on the generated professional-quality melody line. Techniques used include pitch shifting and time stretching, which modify the user's voice to fit the new melody line. The input is the generated melody line and the user's recording, and the output is the adjusted voice data.

[0826] Step 5:

[0827] The server uses an enhancement module to optimize the sound quality. It applies noise reduction filters and reverb effects to improve the overall sound quality of the song. Specifically, it removes recorded background noise and adds reverb to add depth to the audio. The input is the adjusted audio data, and the output is an enhanced, sound-optimized finished song.

[0828] Step 6:

[0829] The server stores the generated completed music in a database and provides it to the user. The user can download and check the generated music on their own device. They can also share it via social media or email. The input is the enhanced completed music, and the output is music data that can be accessed by providing the user with a download link.

[0830] (Application example 1)

[0831] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0832] The traditional music production process requires specialized knowledge and skills, making it difficult for many users to easily create professional-quality music. Furthermore, there are limited ways to easily share completed music, and in particular, there is a lack of an environment in place that allows ordinary users to easily upload their own music to streaming platforms.

[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0834] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed song, means for providing the generated completed song to the user, and means for uploading the completed song to a streaming platform. This enables users to create high-quality songs without requiring specialized knowledge or skills and easily upload them to a streaming platform for sharing.

[0835] "User" means an individual or entity that uses the system to record audio and generate and upload music.

[0836] An "audio recording means" is a device or application that a user uses to record a song.

[0837] "Means for transmitting recorded voice data to a server" refers to the process or function that transmits a user's recorded voice data to a server.

[0838] A "server" is a central computer system for analyzing and processing audio data.

[0839] "Means for performing audio analysis" refers to software or algorithms for extracting melody lines and lyrics from recorded audio data.

[0840] The "means for extracting a melody line" is a function for identifying pitch and rhythm patterns from audio data.

[0841] A "generative AI model" is an artificial intelligence (AI) model that generates professional-quality melody lines based on extracted melody lines.

[0842] "Means for generating professional-quality melody lines" refers to the process of using a generative AI model to create high-quality melodies based on a user's recordings.

[0843] The "means for adjusting the recorded voice and generating a completed piece of music" refers to the process of adjusting the user's recorded voice to match the generated melody line, thereby generating a completed piece of music as a whole.

[0844] The "means for providing the generated completed music to the user" is a process or function that provides the completed music in a form that the user can access and download.

[0845] "Means for uploading the completed song to a streaming platform" refers to a function for uploading the completed song to a streaming service on the Internet.

[0846] The present invention provides a system that allows users to convert their recorded songs into professional quality music and easily upload the music to a streaming platform. Specific embodiments of the present invention are described below.

[0847] System Overview

[0848] 1. Recording and data transmission

[0849] The user starts a recording application on a device such as a smartphone and records a song. When the recording is finished, the device sends the audio data to the server.

[0850] For example, a user can sing a song such as "One Afternoon, Relaxing in the Garden" using a recording app on their smartphone and send the recording data to a server.

[0851] 2. Audio analysis and melody extraction

[0852] The server receives the transmitted audio data and analyzes it using audio analysis software (e.g., librosa), which extracts pitch, rhythm patterns, and the timing of lyrics.

[0853] For example, the pitch and rhythmic patterns of the recorded lyrics "One afternoon, relaxing in the garden" are analyzed.

[0854] 3. Melody Generation

[0855] The server inputs the extracted melody line and lyrics into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line.

[0856] For example, a sophisticated melody line can be generated based on the lyrics "One afternoon relaxing in the garden."

[0857] 4. Audio adjustment and music generation

[0858] The server then uses the generated professional-quality melody line to adjust the recorded audio and generate the finished song, applying filters and effects (e.g., noise reduction filters and reverb effects) to optimize the sound quality.

[0859] For example, the user's recording is adjusted to fit the newly generated melody line, and noise reduction filters and reverb effects are applied.

[0860] 5. Providing music and uploading to streaming services

[0861] The server generates the finished song and provides it to the user, who can then download it or upload it to a streaming platform.

[0862] For example, users can review the finished song on their smartphone and upload it directly to a streaming platform if they wish.

[0863] Examples of specific examples and prompts

[0864] As a concrete example, consider a user relaxing in their garden and recording a song:

[0865] Example prompt sentence:

[0866] Users can simply record a song on their smartphone while relaxing in the garden one afternoon. The app then sends the recording to a server, where an AI model generates a professional-quality melody line. Users can then upload the finished song directly to streaming platforms.

[0867] This system allows users to create professional-quality music without any specialized knowledge or skills, and then easily upload and share the music on streaming platforms.

[0868] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0869] Step 1:

[0870] A user records a song using a recording app on their smartphone. The user launches the app, presses the "Start Recording" button, and sings. When the recording is finished, the application saves the audio data in a file format. The input of this step is the user's singing voice, and the output is the recorded audio file.

[0871] Step 2:

[0872] The device sends the recorded audio data to the server. After the recording is complete, the device uses the data transmission function to upload the audio file to the server. The input of this step is the saved audio file, and the output is the audio data uploaded to the server.

[0873] Step 3:

[0874] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. Audio analysis software (e.g., librosa) is used for the analysis. Specifically, the audio data is loaded using the librosa.load function, and the pitch and rhythmic patterns are extracted using the librosa.piptrack function. The input of this step is the audio data uploaded to the server, and the output is the analyzed melody line and rhythmic patterns.

[0875] Step 4:

[0876] The server inputs the extracted melody line and rhythm pattern into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line. Specifically, the analyzed melody line and rhythm pattern are fed into the AI ​​model, and a new melody line is generated using the model's inference function. The input of this step is the extracted melody line and rhythm pattern, and the output is the generated professional-quality melody line.

[0877] Step 5:

[0878] The server then uses the generated, professional-quality melody line to adjust the recorded audio and generate a finished song. Specifically, it uses audio editing software to adjust the user's recording to fit the new melody line and apply filters and effects (e.g., noise reduction filters and reverb effects). The input to this step is the generated, professional-quality melody line and the recorded audio data, and the output is the finished song.

[0879] Step 6:

[0880] The server provides the generated finished song to the user. Specifically, the server stores the song file on the server so that the user can download it, and provides the user with a download link. The input of this step is the finished song, and the output is a song file that the user can download.

[0881] Step 7:

[0882] The user uploads the completed song to a streaming platform. Specifically, the user uploads the completed song to the platform using the streaming platform's upload function provided by the smartphone application. The input of this step is the song file downloaded by the user, and the output is the song uploaded to the streaming platform.

[0883] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0884] As an embodiment of the present invention, a system is provided that allows a user to easily convert their own recorded songs into professional quality music, and furthermore, reflects emotional elements in the music using an emotion engine.

[0885] System Overview

[0886] The present invention consists of the following main components:

[0887] 1. A means for users to record audio (a device with a recording application)

[0888] 2. A means of sending recorded audio data to the server

[0889] 3. A method for analyzing audio on the server and extracting melody lines and lyrics

[0890] 4. A means to recognize emotions from the user's voice using an emotion engine

[0891] 5. A method for generating professional-quality melody lines using an AI model based on the extracted melody lines and the recognized emotions.

[0892] 6. A method for generating a complete song using the generated melody line and the adjusted user's voice

[0893] 7. Means of sound quality optimization and application of filters and effects

[0894] 8. Means for providing the generated finished music to the user

[0895] Explanation of program processing (with concrete examples)

[0896] Start recording and send data

[0897] 1. The user starts the recording app on their device (such as a smartphone) and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0898] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0899] 3. When the user presses the button to stop recording, the recording ends.

[0900] 4. The device sends the recorded audio data to the server.

[0901] Voice analysis, emotion recognition, and melody generation

[0902] 5. The server receives the voice data sent from the terminal.

[0903] 6. The server uses an audio analysis module to analyze the recording and detect pitch, rhythmic patterns, and lyric timing.

[0904] 7. The server extracts the melody line and lyrics based on the analysis results.

[0905] 8. The emotion engine on the server recognizes the user's emotion from the voice data. For example, if the user sings "happily," the emotion engine recognizes a positive emotion.

[0906] 9. The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0907] 10. The AI ​​model uses input data to generate professional-quality melody lines that match emotions. For example, if a positive emotion is recognized, an upbeat, rhythmic melody line will be generated.

[0908] Music Generation and Enhancement

[0909] 11. The server then adjusts the user's recorded voice based on the generated professional-quality melody line, modifying it to fit the new melody line.

[0910] 12. The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, for example, using a noise reduction filter or a reverb effect.

[0911] 13. The server generates a professional quality finished song and stores it in a database.

[0912] Submit and share your music

[0913] 14. The user uses the terminal to access the server and download the completed song.

[0914] 15. Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0915] Through this series of steps, users can transform their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0916] The processing flow will be explained below.

[0917] Step 1:

[0918] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording. For example, they sing lyrics such as "One afternoon, relaxing in the garden."

[0919] Step 2:

[0920] The device uses a built-in microphone to record the user's singing voice and saves the recording data as a temporary file.

[0921] Step 3:

[0922] When the user presses the button to stop recording, the recording ends.

[0923] Step 4:

[0924] The terminal transmits the recorded voice data to the server.

[0925] Step 5:

[0926] The server receives and stores the voice data sent from the terminal.

[0927] Step 6:

[0928] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythmic patterns, and lyric timing.

[0929] Step 7:

[0930] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[0931] Step 8:

[0932] The server's emotion engine recognizes the user's emotions from the voice data. For example, it recognizes the positive emotion of "feeling happy" from the user's tone of voice and intonation.

[0933] Step 9:

[0934] The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[0935] Step 10:

[0936] The AI ​​model uses the input data to generate professional-quality melody lines that match the emotion. For example, if a positive emotion is recognized, it will generate an upbeat, rhythmic melody line.

[0937] Step 11:

[0938] The server then adjusts the user's recorded voice based on the generated melody line, correcting it to fit the new melody line, including adjusting the pitch and timing.

[0939] Step 12:

[0940] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, such as noise reduction filters and reverb effects.

[0941] Step 13:

[0942] The server generates professional quality finished music and stores it in a database.

[0943] Step 14:

[0944] The user uses a terminal to access the server and download the completed song.

[0945] Step 15:

[0946] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[0947] Through this series of processes, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0948] Example 2

[0949] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0950] Previous music production systems required specialized knowledge and skills, as well as a great deal of time and effort, to convert users' recorded audio into professional-quality music. It was also difficult for average users to incorporate emotional elements into music. This made music production a challenging process, making it difficult for anyone to easily create high-quality music.

[0951] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting melody lines and lyrics, means for recognizing emotions from voice data, and means for generating professional-quality melody lines using a generative AI model based on the extracted melody lines and the recognized emotions. This allows anyone to easily generate high-quality music and enjoy music that reflects their own emotions.

[0952] A "voice recording means" is an application or device that allows a user to record their own voice.

[0953] "Means for transmitting audio data to a server" refers to a communication means for uploading recorded audio data to a server, which includes an internet connection and a corresponding application.

[0954] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms or software for analyzing recorded audio data to identify and extract melody lines and lyrics therefrom.

[0955] The "means for recognizing emotions" is an algorithm or software for analyzing voice data and recognizing the user's emotions.

[0956] "Means for generating professional-quality melody lines using generative AI models" refers to a system or software that uses AI technology to automatically generate high-quality melody lines.

[0957] "Means for adjusting recorded voice and generating a completed piece of music" refers to a system or software for adjusting a user's recorded voice to match the generated melody line and combining them into a completed piece of music.

[0958] "Means for applying filters and effects to optimize sound quality" refers to a system or software for applying various audio filters and effects to improve the sound quality of the generated music.

[0959] "Means for providing the completed music to the user" refers to means for providing the created music in a form that the user can access, including providing a download link or using cloud storage.

[0960] The present invention provides a system that allows users to easily convert their own recorded songs into professional quality music and further reflects emotional elements in the music using an emotion engine. Specific embodiments will be described below.

[0961] Hardware and Software Configuration

[0962] User terminal

[0963] The user terminal is a device such as a smartphone or tablet that has a recording application installed. This terminal has a built-in microphone that is used to record the user's voice. It also has a communication function to send the recorded data to the server.

[0964] server

[0965] The server contains multiple modules for processing the received voice data, including a voice analysis module, an emotion recognition engine, an AI model, and a sound quality optimization module.

[0966] Voice recording and data transmission

[0967] 1. The user launches the recording app on the device and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[0968] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[0969] 3. When the user presses the stop recording button, the recording ends.

[0970] 4. The device sends the recorded audio data to the server via an HTTP POST request, using the HTTPS protocol.

[0971] Voice analysis and emotion recognition

[0972] 5. The server receives the audio data sent from the device and analyzes the recording using the Python LibROSA library.

[0973] 6. The server detects pitch, rhythmic patterns, and lyric timing and stores them in an internal database (e.g., MongoDB) in JSON format.

[0974] 7. The server's emotion engine uses an emotion recognition algorithm (e.g., Microsoft Azure Cognitive Services) to recognize the user's emotion from the voice data.

[0975] Melody line generation

[0976] 8. The server inputs the extracted melody line and the recognized emotion into the AI ​​model. For example, it inputs the following prompt to the AI ​​model: "Create a fun melody line in the style of James Brown."

[0977] 9. The AI ​​model uses the input data to generate a professional-quality melody line tailored to the emotion, using OpenAI GPT-3 or other music generation models.

[0978] Music Generation and Enhancement

[0979] 10. The server uses Auto-tune to adjust the user's recorded voice based on the generated professional-quality melody line.

[0980] 11. The server uses a sound quality optimization module (e.g., iZotope Neutron) to apply noise reduction filters and reverb effects to improve the overall sound quality of the song.

[0981] Providing completed music

[0982] 12. The server stores the generated professional quality finished music in a database.

[0983] 13. The user can use their device to access the server and download the completed song from the "My Songs" section.

[0984] Through this process, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[0985] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0986] Step 1:

[0987] The user launches the device's recording app and taps the "Start Recording" button to begin recording. The input is the user's button operation, and the output is the trigger to start recording. Specifically, the recording app accesses the built-in microphone and begins capturing audio in real time.

[0988] Step 2:

[0989] The device uses a built-in microphone to record the user's singing voice and temporarily saves the recorded data. The input is the user's voice, and the output is the recorded voice data. Specifically, the recorded data is temporarily saved in the device's local storage in wav format.

[0990] Step 3:

[0991] Recording ends when the user presses the stop recording button. The input is the user's button operation, and the output is the end of the recording session. Specifically, the recording data capture stops and the final data file is committed.

[0992] Step 4:

[0993] The device sends the recorded audio data to the server. The input is the recorded data, and the output is the data sent to the server. Specifically, the audio data is sent to the server's API endpoint via an HTTP POST request, and the HTTPS protocol is used for communication.

[0994] Step 5:

[0995] The server receives the voice data sent from the device and stores it in storage. The input is the recorded data, and the output is the stored data. Specifically, the received data is stored in file storage and placed in a waiting state for processing.

[0996] Step 6:

[0997] The server uses an audio analysis module to analyze the recorded data and detect pitch, rhythmic patterns, and lyric timing. The input is the stored audio data, and the output is the analysis results. Specifically, it uses the Python LibROSA library to perform spectral analysis and pitch detection of the audio data.

[0998] Step 7:

[0999] The server extracts the melody line and lyrics based on the analysis results. The input is the audio analysis results, and the output is the extracted melody line and lyrics. Specifically, the analysis results are saved in JSON format in an internal database (e.g., MongoDB).

[1000] Step 8:

[1001] The server's emotion engine recognizes the user's emotions from the voice data. The input is the recorded data, and the output is an emotion score. Specifically, it uses Microsoft's Azure Cognitive Services to extract emotions from the voice data in real time and calculate the emotion score.

[1002] Step 9:

[1003] The server inputs the extracted melody line and the recognized emotion into the AI ​​model. The input is the melody line and emotion score, and the output is the generated melody line. Specifically, the server inputs the following prompt to the AI ​​model: "Please create a fun melody line in the style of James Brown."

[1004] Step 10:

[1005] The AI ​​model uses input data to generate a professional-quality melody line that matches the emotion. The input is a prompt, and the output is the generated melody line. Specifically, the generative AI model (e.g., OpenAI GPT-3 or other music generation model) generates a professional-quality melody line and returns the data to the server.

[1006] Step 11:

[1007] The server adjusts the user's recorded voice based on the generated professional-quality melody line. The input is the user's voice and the generated melody line, and the output is the adjusted song. Specifically, it uses Auto-tune to ensure consistency between the voice and the melody line.

[1008] Step 12:

[1009] The server uses a sound quality optimization module to improve the overall sound quality of the track. The input is the adjusted track, and the output is the sound-optimized track. Specifically, it uses iZotope's Neutron to apply noise reduction filters and reverb effects.

[1010] Step 13:

[1011] The server generates professional-quality finished songs and stores them in a database. The input is a quality-optimized song, and the output is a saved song file. Specifically, the final audio file is saved in the database in MP3 or WAV format.

[1012] Step 14:

[1013] The user uses their device to access the server and download the completed song. The input is the user's request, and the output is the downloaded song file. Specifically, the user selects a song from the "My Songs" section within the recording app and downloads it securely via HTTPS.

[1014] Step 15:

[1015] Users can play the completed song to check its contents, and if they wish, share it on social media or by email. The input is the user's operation, and the output is the shared song. Specifically, an API is used to share songs directly to social media from within the app, providing a seamless sharing experience.

[1016] (Application example 2)

[1017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1018] Previously, users needed advanced music production skills and specialized equipment to transform their own songs into professional-quality music. Furthermore, there was limited technology available to capture the emotions conveyed in the song. This made it difficult for many users to easily create high-quality music.

[1019] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting a melody line and lyrics, means for recognizing a user's emotion from the analyzed voice data, and means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion data. This allows users to easily convert recorded songs into professional-quality music and enjoy music that reflects their emotions.

[1020] A "terminal" is an electronic device that a user uses to record audio, and includes a smartphone, tablet, etc.

[1021] "Audio Data" means audio information recorded by a user and stored in digital form.

[1022] A "server" is a computer system that receives recorded voice data and performs analysis and generation processing.

[1023] "Audio analysis" is the process of extracting elements such as melody lines, lyrics, pitch, and rhythmic patterns from recorded audio data.

[1024] A "melody line" refers to a sequence of notes that form a musical melody and constitutes the basic structure of a piece of music.

[1025] "Emotion recognition" is the process of identifying a user's emotional state from their voice data, determining whether their emotions are positive or negative.

[1026] A "generative AI model" is an algorithm that uses AI (artificial intelligence) technology to generate professional-quality melody lines based on input data.

[1027] "Sound quality optimization" is the process of applying audio filters such as noise reduction and reverb effects to improve the quality of your recordings.

[1028] A "noise reduction filter" is a filtering technique used to remove unwanted background noise from recorded audio data.

[1029] The "reverb effect" is an effect technique that adds artificial reverberation to recorded audio to increase the breadth and depth of the sound.

[1030] A "finished song" is the final musical work after the user's recorded vocals have been adjusted, a generated melody line has been generated, and sound quality optimization has been applied.

[1031] To implement this invention, the user, terminal, and server must fulfill their respective roles and work in cooperation with each other. The specific processing of each component and the hardware and software used will be described below.

[1032] Recording and Data Transmission

[1033] Users use devices such as smartphones to record audio. They start a recording app and tap the "Start Recording" button to begin recording, and audio data is collected using the smartphone's built-in microphone. When recording ends, the audio data is temporarily saved on the device. When the user presses the "Stop Recording" button, recording ends and the recorded audio data is sent to the server.

[1034] Voice analysis and emotion recognition

[1035] When the server receives the audio data, it uses a voice analysis module to analyze the recording. Specifically, it detects pitch, rhythm patterns, and lyric timing, and extracts the melody line and lyrics. At the same time, the emotion engine recognizes the user's emotions from the audio data. For example, if the user is singing "happily," the emotion engine will recognize a positive emotion.

[1036] Melody Generation

[1037] The server inputs the extracted melody line and the recognized emotion into a generative AI model, which then generates a professional-quality melody line based on the input data. For example, if a positive emotion is recognized, a bright and rhythmic melody line will be generated.

[1038] Music generation and sound quality optimization

[1039] The server then adjusts the user's recorded voice based on the generated professional-quality melody line, correcting it to fit the new melody line, and then uses a sound quality optimization module to apply noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in a professional-quality song.

[1040] Specific examples

[1041] For example, if a user records a song about "happy days," the server will recognize the positive emotions and generate an upbeat, rhythmic melody line, then add noise reduction and reverb effects to complete the final song.

[1042] Here are some example prompts to input to the generative AI model:

[1043] Generate a melody line that reflects positive emotions based on a happy singing voice recorded by the user.

[1044] Recording data: {Recording data}

[1045] Emotion: Positive, fun

[1046] Desired melody style: Bright and rhythmic

[1047] Submit and share your music

[1048] Users can check the completed song on their smartphone or other device and upload it to social media or content distribution platforms if they wish, allowing them to easily convert their recorded songs into professional-quality songs and share them.

[1049] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1050] Step 1:

[1051] The user launches the device's recording app and taps the "Start Recording" button. This causes the device to begin recording the user's voice using the built-in microphone. The input is the user's singing voice, and the output is the temporarily saved audio data.

[1052] Step 2:

[1053] When the user presses the "Stop Recording" button, recording ends. The device temporarily stores the recorded audio data. This audio data is used as input for subsequent processing.

[1054] Step 3:

[1055] The device sends the recorded audio data to the server. The input of this process is the temporarily stored audio data, and the output is the audio data sent to the server.

[1056] Step 4:

[1057] The server receives audio data from the terminal. It uses an audio analysis module to extract melody lines, pitches, rhythmic patterns, and lyric timing from the audio data. The input is the audio data, and the output is the extracted pitches, rhythmic patterns, melody lines, and lyric timing.

[1058] Step 5:

[1059] The server uses an emotion engine to recognize the user's emotion from the voice data. For example, if the user sings "happily," it recognizes a positive emotion. The input is the voice data, and the output is the recognized emotion data.

[1060] Step 6:

[1061] The server inputs the extracted melody line and the recognized emotion data into a generative AI model, which generates a new, professional-quality melody line based on the melody line and emotion data. The input is the melody line and emotion data, and the output is a new, professional-quality melody line.

[1062] Step 7:

[1063] The server then adjusts the user's recorded voice to fit the new melody line based on the generated professional-quality melody line. The input is the recorded voice data and the generated melody line, and the output is the adjusted voice data.

[1064] Step 8:

[1065] The server uses the sound quality optimization module to optimize the recorded audio with noise reduction filters and reverb effects. The input is the adjusted audio data, and the output is the optimized audio data.

[1066] Step 9:

[1067] The server generates a complete piece of music based on the optimized audio data and stores it in a database. The input is the optimized audio data, and the output is the stored complete piece of music.

[1068] Step 10:

[1069] The user accesses the server using a device and downloads the completed song. The user checks the song and, if desired, shares it on social media or a content distribution platform. The input is the completed song from the server, and the output is the downloaded song and the shared song.

[1070] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1071] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1072] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1073] [Fourth embodiment]

[1074] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1075] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1076] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1077] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1078] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1079] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1080] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1081] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1082] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1083] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1084] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1085] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1086] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1087] As an embodiment of the present invention, a system is provided that allows users to easily convert their own recorded songs into professional quality musical compositions.

[1088] System Overview

[1089] First, the user uses a device to record their own voice. A recording application is installed on the device, and the user starts the application and presses the "Start Recording" button to begin recording. Once recording is complete, the device sends the recorded data to the server.

[1090] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. It then uses an AI model to generate a professional-quality melody line based on the extracted melody line. It then uses the generated melody line to adjust the user's recorded voice to create a finished song. The finished song is then stored by the server and can be downloaded by the user via their device.

[1091] Explanation of program processing (with concrete examples)

[1092] 1. Start recording and send data

[1093] The user starts a recording app on their smartphone and sings the lyrics of the song they want to sing, for example, "One afternoon, relaxing in the garden," for one minute.

[1094] Once the recording is complete, the device sends the audio data to the server.

[1095] 2. Audio analysis and melody extraction

[1096] The server receives the audio data sent from the device and uses a speech analysis module to analyze the recording, for example, to identify the pitch and rhythmic patterns of the recorded lyrics, "One afternoon, relaxing in the garden."

[1097] The server extracts the lyrics and melody line and stores in a database the pitch and rhythm in which each word of the lyrics is sung.

[1098] 3. Melody adjustment and generation

[1099] The server inputs the extracted melody line and lyrics into the AI ​​model. For example, the lyric "One afternoon, relaxing in the garden" can be input into the AI ​​model to generate a more refined melody line.

[1100] Based on the input data, the AI ​​model generates professional-quality melody lines and enhances the original melody lines.

[1101] 4. Music Generation and Enhancement

[1102] The server uses the generated professional-quality melody line to adjust the recording of the user's voice, so that the user's voice is modified to fit the newly generated melody line.

[1103] The server uses an enhancement module to optimize the sound quality and apply filters and effects to improve the overall sound quality of the song, for example using a noise reduction filter or a reverb effect.

[1104] 5. Submission of completed music

[1105] The server generates the finished song and stores it in a database, which the user can then download from the server using their device.

[1106] Users can check the completed song and share it via social media or email if they wish.

[1107] In this way, the present invention provides a concrete means for users to easily create professional-quality music, allowing anyone to easily enjoy the joy of music production, even without specialized knowledge or skills.

[1108] The processing flow will be explained below.

[1109] Step 1:

[1110] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording.

[1111] Step 2:

[1112] The device uses a built-in microphone to record the user's singing voice and saves the recorded data as a temporary file.

[1113] Step 3:

[1114] When the user presses the button to stop recording, the recording ends.

[1115] Step 4:

[1116] The terminal transmits the recorded voice data to the server.

[1117] Step 5:

[1118] The server receives and stores the voice data sent from the terminal.

[1119] Step 6:

[1120] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythm, and lyric timing.

[1121] Step 7:

[1122] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[1123] Step 8:

[1124] The server inputs the extracted melody line and lyrics into the AI ​​model.

[1125] Step 9:

[1126] The AI ​​model generates a professional-quality melody line based on the input data and returns the generated melody line to the server.

[1127] Step 10:

[1128] The server adjusts the user's recorded voice based on the generated melody line and modifies it to fit the new melody line.

[1129] Step 11:

[1130] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song.

[1131] Step 12:

[1132] The server generates professional quality finished music and stores it in a database.

[1133] Step 13:

[1134] The user uses a terminal to access the server and download the completed song.

[1135] Step 14:

[1136] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[1137] Through this series of steps, users can transform their recorded songs into professional-quality songs and easily enjoy the music production process.

[1138] Example 1

[1139] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1140] In the past, users needed advanced musical knowledge and specialized equipment to create professional-quality music, making it extremely difficult for the average user. Furthermore, adjusting recorded audio to match a high-quality melody line required specialized skills. Therefore, a method for easily generating high-quality music was needed.

[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1142] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed piece of music, and means for providing the generated completed piece of music to the user, thereby enabling users to easily generate high-quality music without having specialized knowledge or skills.

[1143] "User" means any person or entity that intends to use the System to record audio and generate musical compositions.

[1144] "Audio recording means" refers to a device or application that a user uses to record their own audio.

[1145] "Means for transmitting the recorded voice data to the server" refers to the communications protocols and procedures for uploading the recorded voice data to the server via the Internet.

[1146] "Server" refers to a computer system for receiving, analyzing, and processing recorded audio data.

[1147] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms and software for automatically analyzing and extracting pitch, rhythm patterns, and lyrics from recorded audio data.

[1148] "Generative AI Model" refers to an artificial intelligence program or algorithm used to generate professional-quality melody lines based on extracted melody lines and lyrics.

[1149] "Professional quality melody line" refers to a melody line of professional quality.

[1150] "Means for adjusting the audio and generating a finished piece of music" refers to the methods and techniques for modifying the recorded audio to match the generated melody line and generating the final piece of music.

[1151] "Means for providing the completed song to the user" refers to the method or procedure for providing the final song to the user in a downloadable format.

[1152] "Means for extracting pitch and rhythmic patterns" refers to algorithms or software that identify and extract pitch and rhythmic patterns from recorded audio data.

[1153] "Means for sound quality optimization and applying filters and effects" refers to techniques for applying audio filters and effects to improve the sound quality of the generated music.

[1154] "System" refers to the overall setup that includes the elements mentioned above and enables users to create professional quality music.

[1155] The present invention relates to a system that allows users to easily create professional quality music. Specific means for carrying out the invention will be described in detail below.

[1156] A user first launches a recording application on a device such as a smartphone or tablet. This application provides a recording function, allowing the user to record their own voice. Once recording is complete, the device sends the recorded data (voice data) to a server. To do this, the device establishes a connection to the server via the Internet and sends the voice data using a protocol such as an HTTP POST request.

[1157] The server uses a receiving module to receive audio data sent from the device. This audio data is stored in a database on the server. The server then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. The audio analysis module uses libraries such as FFT (Fast Fourier Transform) and LibROSA to analyze the frequency components of the audio data. This allows detailed analysis data of the recorded audio to be obtained.

[1158] The analyzed data is stored in a database on the server, and the next step is to input it into a generative AI model. The generative AI model generates a professional-quality melody line based on the extracted melody line and lyrics. This model can use, for example, a generative AI model or other similar artificial intelligence algorithms. As a concrete example, the prompt sentence is as follows:

[1159] Prompt statement:

[1160] "Generate a professional-quality melody line for "One Afternoon Relaxing in the Garden" based on your voice data. The voice data is in the attached file."

[1161] The generative AI model takes this prompt and the attached file and generates a professional-quality melody line. The generated melody line is sent back to the server and stored in a database. The server then uses this generated melody line to adjust the user's recorded voice, using techniques such as pitch shifting and time stretching to modify the user's voice to fit the new melody line.

[1162] The server then uses an enhancement module to optimize the sound quality, applying noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in the final finished song.

[1163] Finally, the server stores the completed song in a database and provides it to the user, who can download it from the server using their device. Users can also share the completed song via social media or email.

[1164] In this way, the present invention provides a concrete means for users to easily create professional-quality music, and is a system that allows anyone to easily enjoy the fun of music production, even without specialized knowledge or skills.

[1165] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1166] Step 1:

[1167] The user launches a recording application on their smartphone or tablet. They press the "Start Recording" button and sing the lyrics of the song they want to sing. They sing specific lyrics, such as "One afternoon, relaxing in the garden," for one minute. Once the recording is complete, the device sends the recording to the server. The input is the audio data sung by the user, and the output is the audio data sent to the server.

[1168] Step 2:

[1169] The server receives the audio data sent from the device. The received audio data is stored in a database on the server. It then launches an audio analysis module to analyze pitch, rhythm patterns, and the timing of lyrics. Specifically, it analyzes the frequency components of the audio data using libraries such as FFT (Fast Fourier Transform) and LibROSA. The input is the audio data sent to the server, and the output is the analyzed pitch, rhythm patterns, and timing information for the lyrics.

[1170] Step 3:

[1171] The server inputs the melody line and lyrics extracted from the analyzed audio data into the generative AI model. A prompt is sent to the generative AI model, which generates a professional-quality melody line according to the prompt. For example, the prompt could be, "Based on the user's audio data, please generate a professional-quality melody line for 'Relaxing in the Garden One Afternoon.' The audio data is in the attached file." The input is the extracted melody line and lyrics, and the output is the professional-quality melody line returned by the generative AI model.

[1172] Step 4:

[1173] The server adjusts the user's recorded voice based on the generated professional-quality melody line. Techniques used include pitch shifting and time stretching, which modify the user's voice to fit the new melody line. The input is the generated melody line and the user's recording, and the output is the adjusted voice data.

[1174] Step 5:

[1175] The server uses an enhancement module to optimize the sound quality. It applies noise reduction filters and reverb effects to improve the overall sound quality of the song. Specifically, it removes recorded background noise and adds reverb to add depth to the audio. The input is the adjusted audio data, and the output is an enhanced, sound-optimized finished song.

[1176] Step 6:

[1177] The server stores the generated completed music in a database and provides it to the user. The user can download and check the generated music on their own device. They can also share it via social media or email. The input is the enhanced completed music, and the output is music data that can be accessed by providing the user with a download link.

[1178] (Application example 1)

[1179] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1180] The traditional music production process requires specialized knowledge and skills, making it difficult for many users to easily create professional-quality music. Furthermore, there are limited ways to easily share completed music, and in particular, there is a lack of an environment in place that allows ordinary users to easily upload their own music to streaming platforms.

[1181] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1182] In this invention, the server includes means for a user to record voice, means for transmitting the recorded voice data to the server, means for performing voice analysis in the server and extracting a melody line and lyrics, means for generating a professional-quality melody line using a generative AI model based on the extracted melody line, means for adjusting the recorded voice using the generated melody line to generate a completed song, means for providing the generated completed song to the user, and means for uploading the completed song to a streaming platform. This enables users to create high-quality songs without requiring specialized knowledge or skills and easily upload them to a streaming platform for sharing.

[1183] "User" means an individual or entity that uses the system to record audio and generate and upload music.

[1184] An "audio recording means" is a device or application that a user uses to record a song.

[1185] "Means for transmitting recorded voice data to a server" refers to the process or function that transmits a user's recorded voice data to a server.

[1186] A "server" is a central computer system for analyzing and processing audio data.

[1187] "Means for performing audio analysis" refers to software or algorithms for extracting melody lines and lyrics from recorded audio data.

[1188] The "means for extracting a melody line" is a function for identifying pitch and rhythm patterns from audio data.

[1189] A "generative AI model" is an artificial intelligence (AI) model that generates professional-quality melody lines based on extracted melody lines.

[1190] "Means for generating professional-quality melody lines" refers to the process of using a generative AI model to create high-quality melodies based on a user's recordings.

[1191] The "means for adjusting the recorded voice and generating a completed piece of music" refers to the process of adjusting the user's recorded voice to match the generated melody line, thereby generating a completed piece of music as a whole.

[1192] The "means for providing the generated completed music to the user" is a process or function that provides the completed music in a form that the user can access and download.

[1193] "Means for uploading the completed song to a streaming platform" refers to a function for uploading the completed song to a streaming service on the Internet.

[1194] The present invention provides a system that allows users to convert their recorded songs into professional quality music and easily upload the music to a streaming platform. Specific embodiments of the present invention are described below.

[1195] System Overview

[1196] 1. Recording and data transmission

[1197] The user starts a recording application on a device such as a smartphone and records a song. When the recording is finished, the device sends the audio data to the server.

[1198] For example, a user can sing a song such as "One Afternoon, Relaxing in the Garden" using a recording app on their smartphone and send the recording data to a server.

[1199] 2. Audio analysis and melody extraction

[1200] The server receives the transmitted audio data and analyzes it using audio analysis software (e.g., librosa), which extracts pitch, rhythm patterns, and the timing of lyrics.

[1201] For example, the pitch and rhythmic patterns of the recorded lyrics "One afternoon, relaxing in the garden" are analyzed.

[1202] 3. Melody Generation

[1203] The server inputs the extracted melody line and lyrics into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line.

[1204] For example, a sophisticated melody line can be generated based on the lyrics "One afternoon relaxing in the garden."

[1205] 4. Audio adjustment and music generation

[1206] The server then uses the generated professional-quality melody line to adjust the recorded audio and generate the finished song, applying filters and effects (e.g., noise reduction filters and reverb effects) to optimize the sound quality.

[1207] For example, the user's recording is adjusted to fit the newly generated melody line, and noise reduction filters and reverb effects are applied.

[1208] 5. Providing music and uploading to streaming services

[1209] The server generates the finished song and provides it to the user, who can then download it or upload it to a streaming platform.

[1210] For example, users can review the finished song on their smartphone and upload it directly to a streaming platform if they wish.

[1211] Examples of specific examples and prompts

[1212] As a concrete example, consider a user relaxing in their garden and recording a song:

[1213] Example prompt sentence:

[1214] Users can simply record a song on their smartphone while relaxing in the garden one afternoon. The app then sends the recording to a server, where an AI model generates a professional-quality melody line. Users can then upload the finished song directly to streaming platforms.

[1215] This system allows users to create professional-quality music without any specialized knowledge or skills, and then easily upload and share the music on streaming platforms.

[1216] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1217] Step 1:

[1218] A user records a song using a recording app on their smartphone. The user launches the app, presses the "Start Recording" button, and sings. When the recording is finished, the application saves the audio data in a file format. The input of this step is the user's singing voice, and the output is the recorded audio file.

[1219] Step 2:

[1220] The device sends the recorded audio data to the server. After the recording is complete, the device uses the data transmission function to upload the audio file to the server. The input of this step is the saved audio file, and the output is the audio data uploaded to the server.

[1221] Step 3:

[1222] The server analyzes the received audio data and extracts pitch, rhythmic patterns, and lyric timing. Audio analysis software (e.g., librosa) is used for the analysis. Specifically, the audio data is loaded using the librosa.load function, and the pitch and rhythmic patterns are extracted using the librosa.piptrack function. The input of this step is the audio data uploaded to the server, and the output is the analyzed melody line and rhythmic patterns.

[1223] Step 4:

[1224] The server inputs the extracted melody line and rhythm pattern into a generative AI model (e.g., using TensorFlow) to generate a professional-quality melody line. Specifically, the analyzed melody line and rhythm pattern are fed into the AI ​​model, and a new melody line is generated using the model's inference function. The input of this step is the extracted melody line and rhythm pattern, and the output is the generated professional-quality melody line.

[1225] Step 5:

[1226] The server then uses the generated, professional-quality melody line to adjust the recorded audio and generate a finished song. Specifically, it uses audio editing software to adjust the user's recording to fit the new melody line and apply filters and effects (e.g., noise reduction filters and reverb effects). The input to this step is the generated, professional-quality melody line and the recorded audio data, and the output is the finished song.

[1227] Step 6:

[1228] The server provides the generated finished song to the user. Specifically, the server stores the song file on the server so that the user can download it, and provides the user with a download link. The input of this step is the finished song, and the output is a song file that the user can download.

[1229] Step 7:

[1230] The user uploads the completed song to a streaming platform. Specifically, the user uploads the completed song to the platform using the streaming platform's upload function provided by the smartphone application. The input of this step is the song file downloaded by the user, and the output is the song uploaded to the streaming platform.

[1231] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1232] As an embodiment of the present invention, a system is provided that allows a user to easily convert their own recorded songs into professional quality music, and furthermore, reflects emotional elements in the music using an emotion engine.

[1233] System Overview

[1234] The present invention consists of the following main components:

[1235] 1. A means for users to record audio (a device with a recording application)

[1236] 2. A means of sending recorded audio data to the server

[1237] 3. A method for analyzing audio on the server and extracting melody lines and lyrics

[1238] 4. A means to recognize emotions from the user's voice using an emotion engine

[1239] 5. A method for generating professional-quality melody lines using an AI model based on the extracted melody lines and the recognized emotions.

[1240] 6. A method for generating a complete song using the generated melody line and the adjusted user's voice

[1241] 7. Means of sound quality optimization and application of filters and effects

[1242] 8. Means for providing the generated finished music to the user

[1243] Explanation of program processing (with concrete examples)

[1244] Start recording and send data

[1245] 1. The user starts the recording app on their device (such as a smartphone) and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[1246] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[1247] 3. When the user presses the button to stop recording, the recording ends.

[1248] 4. The device sends the recorded audio data to the server.

[1249] Voice analysis, emotion recognition, and melody generation

[1250] 5. The server receives the voice data sent from the terminal.

[1251] 6. The server uses an audio analysis module to analyze the recording and detect pitch, rhythmic patterns, and lyric timing.

[1252] 7. The server extracts the melody line and lyrics based on the analysis results.

[1253] 8. The emotion engine on the server recognizes the user's emotion from the voice data. For example, if the user sings "happily," the emotion engine recognizes a positive emotion.

[1254] 9. The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[1255] 10. The AI ​​model uses input data to generate professional-quality melody lines that match emotions. For example, if a positive emotion is recognized, an upbeat, rhythmic melody line will be generated.

[1256] Music Generation and Enhancement

[1257] 11. The server then adjusts the user's recorded voice based on the generated professional-quality melody line, modifying it to fit the new melody line.

[1258] 12. The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, for example, using a noise reduction filter or a reverb effect.

[1259] 13. The server generates a professional quality finished song and stores it in a database.

[1260] Submit and share your music

[1261] 14. The user uses the terminal to access the server and download the completed song.

[1262] 15. Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[1263] Through this series of steps, users can transform their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[1264] The processing flow will be explained below.

[1265] Step 1:

[1266] The user starts the recording app on the device (such as a smartphone) and taps the record button to start recording. For example, they sing lyrics such as "One afternoon, relaxing in the garden."

[1267] Step 2:

[1268] The device uses a built-in microphone to record the user's singing voice and saves the recording data as a temporary file.

[1269] Step 3:

[1270] When the user presses the button to stop recording, the recording ends.

[1271] Step 4:

[1272] The terminal transmits the recorded voice data to the server.

[1273] Step 5:

[1274] The server receives and stores the voice data sent from the terminal.

[1275] Step 6:

[1276] The server uses an audio analysis module to analyze the received audio data and detect pitch, rhythmic patterns, and lyric timing.

[1277] Step 7:

[1278] The server extracts the melody line and lyrics based on the analysis results and stores them in a database.

[1279] Step 8:

[1280] The server's emotion engine recognizes the user's emotions from the voice data. For example, it recognizes the positive emotion of "feeling happy" from the user's tone of voice and intonation.

[1281] Step 9:

[1282] The server inputs the extracted melody line and the recognized emotion into the AI ​​model.

[1283] Step 10:

[1284] The AI ​​model uses the input data to generate professional-quality melody lines that match the emotion. For example, if a positive emotion is recognized, it will generate an upbeat, rhythmic melody line.

[1285] Step 11:

[1286] The server then adjusts the user's recorded voice based on the generated melody line, correcting it to fit the new melody line, including adjusting the pitch and timing.

[1287] Step 12:

[1288] The server uses a sound quality optimization module to apply filters and effects to improve the overall sound quality of the song, such as noise reduction filters and reverb effects.

[1289] Step 13:

[1290] The server generates professional quality finished music and stores it in a database.

[1291] Step 14:

[1292] The user uses a terminal to access the server and download the completed song.

[1293] Step 15:

[1294] Users can play the completed song to check its content, and if they wish, they can share the song via social media or email.

[1295] Through this series of processes, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[1296] Example 2

[1297] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1298] Previous music production systems required specialized knowledge and skills, as well as a great deal of time and effort, to convert users' recorded audio into professional-quality music. It was also difficult for average users to incorporate emotional elements into music. This made music production a challenging process, making it difficult for anyone to easily create high-quality music.

[1299] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting melody lines and lyrics, means for recognizing emotions from voice data, and means for generating professional-quality melody lines using a generative AI model based on the extracted melody lines and the recognized emotions. This allows anyone to easily generate high-quality music and enjoy music that reflects their own emotions.

[1300] A "voice recording means" is an application or device that allows a user to record their own voice.

[1301] "Means for transmitting audio data to a server" refers to a communication means for uploading recorded audio data to a server, which includes an internet connection and a corresponding application.

[1302] "Means for performing audio analysis and extracting melody lines and lyrics" refers to algorithms or software for analyzing recorded audio data to identify and extract melody lines and lyrics therefrom.

[1303] The "means for recognizing emotions" is an algorithm or software for analyzing voice data and recognizing the user's emotions.

[1304] "Means for generating professional-quality melody lines using generative AI models" refers to a system or software that uses AI technology to automatically generate high-quality melody lines.

[1305] "Means for adjusting recorded voice and generating a completed piece of music" refers to a system or software for adjusting a user's recorded voice to match the generated melody line and combining them into a completed piece of music.

[1306] "Means for applying filters and effects to optimize sound quality" refers to a system or software for applying various audio filters and effects to improve the sound quality of the generated music.

[1307] "Means for providing the completed music to the user" refers to means for providing the created music in a form that the user can access, including providing a download link or using cloud storage.

[1308] The present invention provides a system that allows users to easily convert their own recorded songs into professional quality music and further reflects emotional elements in the music using an emotion engine. Specific embodiments will be described below.

[1309] Hardware and Software Configuration

[1310] User terminal

[1311] The user terminal is a device such as a smartphone or tablet that has a recording application installed. This terminal has a built-in microphone that is used to record the user's voice. It also has a communication function to send the recorded data to the server.

[1312] server

[1313] The server contains multiple modules for processing the received voice data, including a voice analysis module, an emotion recognition engine, an AI model, and a sound quality optimization module.

[1314] Voice recording and data transmission

[1315] 1. The user launches the recording app on the device and taps the "Start Recording" button to begin recording. For example, the user sings lyrics such as "One afternoon, relaxing in the garden."

[1316] 2. The device uses the built-in microphone to record the user's singing voice and temporarily saves the recording data.

[1317] 3. When the user presses the stop recording button, the recording ends.

[1318] 4. The device sends the recorded audio data to the server via an HTTP POST request, using the HTTPS protocol.

[1319] Voice analysis and emotion recognition

[1320] 5. The server receives the audio data sent from the device and analyzes the recording using the Python LibROSA library.

[1321] 6. The server detects pitch, rhythmic patterns, and lyric timing and stores them in an internal database (e.g., MongoDB) in JSON format.

[1322] 7. The server's emotion engine uses an emotion recognition algorithm (e.g., Microsoft Azure Cognitive Services) to recognize the user's emotion from the voice data.

[1323] Melody line generation

[1324] 8. The server inputs the extracted melody line and the recognized emotion into the AI ​​model. For example, it inputs the following prompt to the AI ​​model: "Create a fun melody line in the style of James Brown."

[1325] 9. The AI ​​model uses the input data to generate a professional-quality melody line tailored to the emotion, using OpenAI GPT-3 or other music generation models.

[1326] Music Generation and Enhancement

[1327] 10. The server uses Auto-tune to adjust the user's recorded voice based on the generated professional-quality melody line.

[1328] 11. The server uses a sound quality optimization module (e.g., iZotope Neutron) to apply noise reduction filters and reverb effects to improve the overall sound quality of the song.

[1329] Providing completed music

[1330] 12. The server stores the generated professional quality finished music in a database.

[1331] 13. The user can use their device to access the server and download the completed song from the "My Songs" section.

[1332] Through this process, users can convert their recorded songs into professional quality music and enjoy music that reflects their own emotions.

[1333] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1334] Step 1:

[1335] The user launches the device's recording app and taps the "Start Recording" button to begin recording. The input is the user's button operation, and the output is the trigger to start recording. Specifically, the recording app accesses the built-in microphone and begins capturing audio in real time.

[1336] Step 2:

[1337] The device uses a built-in microphone to record the user's singing voice and temporarily saves the recorded data. The input is the user's voice, and the output is the recorded voice data. Specifically, the recorded data is temporarily saved in the device's local storage in wav format.

[1338] Step 3:

[1339] Recording ends when the user presses the stop recording button. The input is the user's button operation, and the output is the end of the recording session. Specifically, the recording data capture stops and the final data file is committed.

[1340] Step 4:

[1341] The device sends the recorded audio data to the server. The input is the recorded data, and the output is the data sent to the server. Specifically, the audio data is sent to the server's API endpoint via an HTTP POST request, and the HTTPS protocol is used for communication.

[1342] Step 5:

[1343] The server receives the voice data sent from the device and stores it in storage. The input is the recorded data, and the output is the stored data. Specifically, the received data is stored in file storage and placed in a waiting state for processing.

[1344] Step 6:

[1345] The server uses an audio analysis module to analyze the recorded data and detect pitch, rhythmic patterns, and lyric timing. The input is the stored audio data, and the output is the analysis results. Specifically, it uses the Python LibROSA library to perform spectral analysis and pitch detection of the audio data.

[1346] Step 7:

[1347] The server extracts the melody line and lyrics based on the analysis results. The input is the audio analysis results, and the output is the extracted melody line and lyrics. Specifically, the analysis results are saved in JSON format in an internal database (e.g., MongoDB).

[1348] Step 8:

[1349] The server's emotion engine recognizes the user's emotions from the voice data. The input is the recorded data, and the output is an emotion score. Specifically, it uses Microsoft's Azure Cognitive Services to extract emotions from the voice data in real time and calculate the emotion score.

[1350] Step 9:

[1351] The server inputs the extracted melody line and the recognized emotion into the AI ​​model. The input is the melody line and emotion score, and the output is the generated melody line. Specifically, the server inputs the following prompt to the AI ​​model: "Please create a fun melody line in the style of James Brown."

[1352] Step 10:

[1353] The AI ​​model uses input data to generate a professional-quality melody line that matches the emotion. The input is a prompt, and the output is the generated melody line. Specifically, the generative AI model (e.g., OpenAI GPT-3 or other music generation model) generates a professional-quality melody line and returns the data to the server.

[1354] Step 11:

[1355] The server adjusts the user's recorded voice based on the generated professional-quality melody line. The input is the user's voice and the generated melody line, and the output is the adjusted song. Specifically, it uses Auto-tune to ensure consistency between the voice and the melody line.

[1356] Step 12:

[1357] The server uses a sound quality optimization module to improve the overall sound quality of the track. The input is the adjusted track, and the output is the sound-optimized track. Specifically, it uses iZotope's Neutron to apply noise reduction filters and reverb effects.

[1358] Step 13:

[1359] The server generates professional-quality finished songs and stores them in a database. The input is a quality-optimized song, and the output is a saved song file. Specifically, the final audio file is saved in the database in MP3 or WAV format.

[1360] Step 14:

[1361] The user uses their device to access the server and download the completed song. The input is the user's request, and the output is the downloaded song file. Specifically, the user selects a song from the "My Songs" section within the recording app and downloads it securely via HTTPS.

[1362] Step 15:

[1363] Users can play the completed song to check its contents, and if they wish, share it on social media or by email. The input is the user's operation, and the output is the shared song. Specifically, an API is used to share songs directly to social media from within the app, providing a seamless sharing experience.

[1364] (Application example 2)

[1365] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1366] Previously, users needed advanced music production skills and specialized equipment to transform their own songs into professional-quality music. Furthermore, there was limited technology available to capture the emotions conveyed in the song. This made it difficult for many users to easily create high-quality music.

[1367] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for performing voice analysis and extracting a melody line and lyrics, means for recognizing a user's emotion from the analyzed voice data, and means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion data. This allows users to easily convert recorded songs into professional-quality music and enjoy music that reflects their emotions.

[1368] A "terminal" is an electronic device that a user uses to record audio, and includes a smartphone, tablet, etc.

[1369] "Audio Data" means audio information recorded by a user and stored in digital form.

[1370] A "server" is a computer system that receives recorded voice data and performs analysis and generation processing.

[1371] "Audio analysis" is the process of extracting elements such as melody lines, lyrics, pitch, and rhythmic patterns from recorded audio data.

[1372] A "melody line" refers to a sequence of notes that form a musical melody and constitutes the basic structure of a piece of music.

[1373] "Emotion recognition" is the process of identifying a user's emotional state from their voice data, determining whether their emotions are positive or negative.

[1374] A "generative AI model" is an algorithm that uses AI (artificial intelligence) technology to generate professional-quality melody lines based on input data.

[1375] "Sound quality optimization" is the process of applying audio filters such as noise reduction and reverb effects to improve the quality of your recordings.

[1376] A "noise reduction filter" is a filtering technique used to remove unwanted background noise from recorded audio data.

[1377] The "reverb effect" is an effect technique that adds artificial reverberation to recorded audio to increase the breadth and depth of the sound.

[1378] A "finished song" is the final musical work after the user's recorded vocals have been adjusted, a generated melody line has been generated, and sound quality optimization has been applied.

[1379] To implement this invention, the user, terminal, and server must fulfill their respective roles and work in cooperation with each other. The specific processing of each component and the hardware and software used will be described below.

[1380] Recording and Data Transmission

[1381] Users use devices such as smartphones to record audio. They start a recording app and tap the "Start Recording" button to begin recording, and audio data is collected using the smartphone's built-in microphone. When recording ends, the audio data is temporarily saved on the device. When the user presses the "Stop Recording" button, recording ends and the recorded audio data is sent to the server.

[1382] Voice analysis and emotion recognition

[1383] When the server receives the audio data, it uses a voice analysis module to analyze the recording. Specifically, it detects pitch, rhythm patterns, and lyric timing, and extracts the melody line and lyrics. At the same time, the emotion engine recognizes the user's emotions from the audio data. For example, if the user is singing "happily," the emotion engine will recognize a positive emotion.

[1384] Melody Generation

[1385] The server inputs the extracted melody line and the recognized emotion into a generative AI model, which then generates a professional-quality melody line based on the input data. For example, if a positive emotion is recognized, a bright and rhythmic melody line will be generated.

[1386] Music generation and sound quality optimization

[1387] The server then adjusts the user's recorded voice based on the generated professional-quality melody line, correcting it to fit the new melody line, and then uses a sound quality optimization module to apply noise reduction filters and reverb effects to improve the overall sound quality of the song, resulting in a professional-quality song.

[1388] Specific examples

[1389] For example, if a user records a song about "happy days," the server will recognize the positive emotions and generate an upbeat, rhythmic melody line, then add noise reduction and reverb effects to complete the final song.

[1390] Here are some example prompts to input to the generative AI model:

[1391] Generate a melody line that reflects positive emotions based on a happy singing voice recorded by the user.

[1392] Recording data: {Recording data}

[1393] Emotion: Positive, fun

[1394] Desired melody style: Bright and rhythmic

[1395] Submit and share your music

[1396] Users can check the completed song on their smartphone or other device and upload it to social media or content distribution platforms if they wish, allowing them to easily convert their recorded songs into professional-quality songs and share them.

[1397] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1398] Step 1:

[1399] The user launches the device's recording app and taps the "Start Recording" button. This causes the device to begin recording the user's voice using the built-in microphone. The input is the user's singing voice, and the output is the temporarily saved audio data.

[1400] Step 2:

[1401] When the user presses the "Stop Recording" button, recording ends. The device temporarily stores the recorded audio data. This audio data is used as input for subsequent processing.

[1402] Step 3:

[1403] The device sends the recorded audio data to the server. The input of this process is the temporarily stored audio data, and the output is the audio data sent to the server.

[1404] Step 4:

[1405] The server receives audio data from the terminal. It uses an audio analysis module to extract melody lines, pitches, rhythmic patterns, and lyric timing from the audio data. The input is the audio data, and the output is the extracted pitches, rhythmic patterns, melody lines, and lyric timing.

[1406] Step 5:

[1407] The server uses an emotion engine to recognize the user's emotion from the voice data. For example, if the user sings "happily," it recognizes a positive emotion. The input is the voice data, and the output is the recognized emotion data.

[1408] Step 6:

[1409] The server inputs the extracted melody line and the recognized emotion data into a generative AI model, which generates a new, professional-quality melody line based on the melody line and emotion data. The input is the melody line and emotion data, and the output is a new, professional-quality melody line.

[1410] Step 7:

[1411] The server then adjusts the user's recorded voice to fit the new melody line based on the generated professional-quality melody line. The input is the recorded voice data and the generated melody line, and the output is the adjusted voice data.

[1412] Step 8:

[1413] The server uses the sound quality optimization module to optimize the recorded audio with noise reduction filters and reverb effects. The input is the adjusted audio data, and the output is the optimized audio data.

[1414] Step 9:

[1415] The server generates a complete piece of music based on the optimized audio data and stores it in a database. The input is the optimized audio data, and the output is the stored complete piece of music.

[1416] Step 10:

[1417] The user accesses the server using a device and downloads the completed song. The user checks the song and, if desired, shares it on social media or a content distribution platform. The input is the completed song from the server, and the output is the downloaded song and the shared song.

[1418] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1419] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1420] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1421] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1422] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1423] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1424] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1425] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1426] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1427] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1428] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1429] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1430] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1431] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1432] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1433] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1434] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1435] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1436] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1437] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1438] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1439] The following is further disclosed regarding the above embodiment.

[1440] (Claim 1)

[1441] a means for a user to record audio;

[1442] means for transmitting the recorded voice data to a server;

[1443] A means for performing audio analysis in a server and extracting a melody line and lyrics;

[1444] A means for generating a professional-quality melody line using an AI model based on the extracted melody line;

[1445] a means for adjusting the recorded voice using the generated melody line to generate a completed piece of music;

[1446] a means for providing the generated completed musical piece to a user;

[1447] A system including:

[1448] (Claim 2)

[1449] 10. The system of claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the audio analysis.

[1450] (Claim 3)

[1451] 10. The system of claim 1, further comprising means for performing sound quality optimization based on the generated melody line and applying filters and effects.

[1452] "Example 1"

[1453] (Claim 1)

[1454] a means for a user to record audio;

[1455] means for transmitting the recorded voice data to a server;

[1456] A means for performing audio analysis in a server and extracting a melody line and lyrics;

[1457] A means for generating a professional-quality melody line using a generative AI model based on the extracted melody line;

[1458] a means for adjusting the recorded voice using the generated melody line to generate a completed piece of music;

[1459] a means for providing the generated completed musical piece to a user;

[1460] A system including:

[1461] (Claim 2)

[1462] 10. The system of claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the audio analysis.

[1463] (Claim 3)

[1464] 10. The system of claim 1, further comprising means for performing sound quality optimization based on the generated melody line and applying filters and effects.

[1465] "Application Example 1"

[1466] (Claim 1)

[1467] a means for a user to record audio;

[1468] means for transmitting the recorded voice data to a server;

[1469] A means for performing audio analysis in a server and extracting a melody line and lyrics;

[1470] A means for generating a professional-quality melody line using a generative AI model based on the extracted melody line;

[1471] a means for adjusting the recorded voice using the generated melody line to generate a completed piece of music;

[1472] a means for providing the generated completed musical piece to a user;

[1473] A way to upload the finished song to a streaming platform,

[1474] A system including:

[1475] (Claim 2)

[1476] 10. The system of claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the audio analysis.

[1477] (Claim 3)

[1478] 10. The system of claim 1, further comprising means for performing sound quality optimization based on the generated melody line and applying filters and effects.

[1479] "Example 2: Combining Emotion Engines"

[1480] (Claim 1)

[1481] a means for a user to record audio;

[1482] means for transmitting the recorded voice data to a server;

[1483] A means for performing audio analysis in a server and extracting a melody line and lyrics;

[1484] means for recognizing emotions from speech data;

[1485] A means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion;

[1486] a means for adjusting the recorded voice using the generated melody line to generate a completed piece of music;

[1487] means for applying filters and effects for sound quality optimization;

[1488] a means for providing the generated completed musical piece to a user;

[1489] A system including:

[1490] (Claim 2)

[1491] 10. The system of claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the audio analysis.

[1492] (Claim 3)

[1493] 10. The system of claim 1, further comprising means for performing sound quality optimization based on the generated melody line and applying filters and effects.

[1494] "Application example 2 when combining emotion engines"

[1495] (Claim 1)

[1496] means for providing a terminal for a user to record audio;

[1497] means for transmitting the recorded voice data to a server;

[1498] A means for performing audio analysis in a server and extracting a melody line and lyrics;

[1499] means for recognizing a user's emotion from the analyzed voice data;

[1500] a means for generating a professional-quality melody line using a generative AI model based on the extracted melody line and the recognized emotion data;

[1501] a means for adjusting the recorded voice to match the generated melody line and generating a completed song;

[1502] A means for providing and sharing the generated completed music to users;

[1503] A system including:

[1504] (Claim 2)

[1505] 10. The system of claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the audio analysis.

[1506] (Claim 3)

[1507] 10. The system of claim 1, further comprising means for performing sound quality optimization based on the generated melody line and applying a noise reduction filter and a reverb effect. [Explanation of symbols]

[1508] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to record audio; means for transmitting the recorded voice data to a server; A means for performing audio analysis in a server and extracting a melody line and lyrics; A means for generating a professional-quality melody line using an AI model based on the extracted melody line; a means for adjusting the recorded voice using the generated melody line to generate a completed piece of music; a means for providing the generated completed musical piece to a user; A system including:

2. 2. The system according to claim 1, further comprising means for extracting pitch and rhythm patterns from the results of the voice analysis.

3. The system of claim 1 further comprising means for performing tone optimization based on the generated melody line and applying filters and effects.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A