system
The system addresses the challenge of specialized knowledge and techniques in the field of music production by enabling real-time accompaniment generation and promotional video creation, allowing users to easily produce and share high-quality music, thereby enhancing creative expression and collaboration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Existing music production systems require specialized knowledge and techniques, making it difficult for beginners and users with immature skills to create high-quality music, especially in real-time jam sessions or improvisations, limiting their creative expression and collaboration opportunities.
A system that analyzes a user's voice signal in real-time using AI to generate accompaniments and backing tracks, accompanied by automatic promotional video creation and sharing features, allowing users to easily produce and disseminate music without advanced technical skills.
Enables users to enjoy high-quality music production experiences, facilitates real-time accompaniment generation, and promotes collaboration by lowering the technical barriers for music creation and sharing.
Smart Images

Figure 2026068428000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to the description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] [In music production, since specialized knowledge and techniques have been conventionally required, it has been difficult for beginners and users with immature techniques to create high-quality music. Also, it has been difficult to generate accompaniments in real-time jam sessions or improvisations, and as a result, many potential music creators have lost the opportunity to demonstrate their creativity. Thus, it is required to lower the technical barriers in music production and provide a music production environment in which many users can easily participate.]
Means for Solving the Problems
[0005] This invention provides a system that analyzes a user's voice signal and generates additional music in real time based on that analysis. This allows users to enjoy a rich musical experience by utilizing AI-powered accompaniment, even without advanced technical skills. Furthermore, by incorporating features for automatic generation of promotional videos based on the generated music and sharing them with other users, users can widely disseminate their work and create new opportunities for collaboration. Such a system makes it easy for beginners and users with limited technical skills to enjoy high-quality music production, significantly lowering the barrier to entry for music creation.
[0006] "Audio signals" are electrical signals that are converted into sounds, such as musical instrument performances or vocals, that a user inputs into a device.
[0007] "Analysis" is the process of extracting features from a received audio signal and identifying musical elements such as tempo, key, and rhythm.
[0008] "Generative AI" is an artificial intelligence technology that automatically creates additional music or accompaniment based on the analysis results of audio signals.
[0009] "Additional music" refers to musical elements, including accompaniments and backing tracks, that are generated in response to the user's voice signal.
[0010] A "promotional video" is [visual content automatically created to accompany generated music, intended for the promotion and sharing of music].
[0011] "Sharing" refers to a feature that allows users to publish music and videos they have created and interact with other users through viewing and rating. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, a communication I / F (Interface) with a reference numeral is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] The system according to the present invention supports improvisational performances and jam sessions by allowing users to participate through playing musical instruments or singing, and by having a generative AI generate additional music in real time. Users can participate in music sessions using their devices and enjoy an advanced musical experience through communication with the server.
[0034] First, the user launches an application on their device and inputs an audio signal to the system using an instrument or microphone. The device then transmits the input audio signal to the server in real time. After receiving the audio signal, the server analyzes it using technologies such as deep learning. Based on the analysis results, a generative AI generates an appropriate accompaniment or backing track.
[0035] The generated music is sent to the user's device, allowing them to enjoy new songs in real time, combined with their own performance or singing. Users can adjust the tempo and volume according to the generated music, enabling them to enjoy an even more creative musical experience.
[0036] Furthermore, the system also has a function to automatically generate promotional videos. This function provides users with rich video content that matches the music they create. The server selects effects and video scenes based on the characteristics of the generated music and delivers the video data to the terminal.
[0037] Finally, users can share their completed music and videos with other users on the platform. The server manages the shared content and provides a mechanism to facilitate communication through ratings and comments from other users. This makes it possible to create new spaces for collaboration and expand opportunities to participate in music production.
[0038] As a concrete example, let's say a user uses this system to sing while playing the guitar. The user launches the application on their device and begins recording their performance and singing. The server analyzes the input data, and the AI generates drum and bass accompaniment that matches the tempo of the performance. The user can continue playing along with the accompaniment, or add riffs or melodies. In addition, a promotional video matching the generated performance is created, which can be shared with other users to receive feedback and explore new musical ideas.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The user launches an application on their device and inputs an audio signal using a musical instrument or microphone. The device digitizes the audio signal and transmits the signal data to the server in real time.
[0042] Step 2:
[0043] The server buffers the received audio signal and analyzes it using digital signal processing algorithms. This analysis extracts musical characteristics such as tempo, key, and rhythm.
[0044] Step 3:
[0045] The server invokes a generation AI based on the analysis results. The AI generates additional music, namely accompaniments and backing tracks, in real time. Deep learning technology is used for generation to create a natural flow of music.
[0046] Step 4:
[0047] The server sends the generated additional music to the terminal. The terminal plays this music in sync with the user's original performance or singing. The user can then add further performances to accompany the generated music.
[0048] Step 5:
[0049] Users can use music editing tools on their devices to adjust the tempo, volume, and harmony of the music. They can also add and edit effects to further refine the final track.
[0050] Step 6:
[0051] The server will then start a function to automatically generate a promotional video based on the completed music. It will select visual effects that match the characteristics of the music, build the video using a video template, and deliver it to the device.
[0052] Step 7:
[0053] Users share their completed music and videos on the platform via the server. The server notifies other users of the published content and facilitates communication through comment and rating functions.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] In recent years, creating and sharing music individually has become easier, but there remains a challenge in generating accompaniments and sharing music in real time during improvisational performances and jam sessions. In particular, there is a need for technology that allows for the easy creation and sharing of appropriate accompaniments and visual content when playing instruments or singing.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes means for acquiring a user's acoustic signal and transmitting the acoustic signal to a data processing device; means for analyzing the acquired acoustic signal using information processing technology in the data processing device; and means for generating artificial intelligence to instantly create music based on the analyzed acoustic data. This enables the user to generate appropriate accompaniment in real time during improvisational performances and to easily create and share visual content based on it.
[0059] An "acoustic signal" is an electrical signal that represents the sound waves generated by a user through playing a musical instrument or singing.
[0060] A "data processing device" is a computer system used to analyze acoustic signals and perform calculations for music generation.
[0061] "Information processing technology" refers to techniques used to analyze acoustic signals and extract features, and primarily includes methods such as deep learning.
[0062] "Generative artificial intelligence" refers to algorithms and models that automatically generate new music based on analysis results.
[0063] "Accompaniment" refers to additional sounds that add harmony and rhythm to a song played by a user, providing musical richness.
[0064] "Visual content" refers to videos, animations, and other materials synchronized with the music, provided to complement the musical experience.
[0065] "Shareable" refers to a state where music and visual content created by users can be shared with other users via the internet or other means.
[0066] This document describes embodiments for carrying out the invention. This invention is a system that analyzes a user's acoustic signal in real time and instantly generates a corresponding accompaniment using a generative AI model. Furthermore, this system automatically creates visual content based on the generated music and enables sharing with other users.
[0067] The user first launches a dedicated application on the terminal. The user inputs an audio signal using an instrument or microphone. The terminal captures this audio signal as a digital signal and transmits it to the server via the internet. The software on the terminal utilizes microphones and audio interfaces as hardware to efficiently collect the audio signal.
[0068] The server processes the received acoustic signals using a deep learning-based acoustic analysis method. This extracts the characteristics of the acoustic signals, and based on the analysis results, a generative AI model generates appropriate music. This generative AI model runs on a computer system with high computing power and utilizes large datasets to improve the accuracy of music generation.
[0069] The generated music is streamed to the user's device, allowing the user to enjoy the music experience in real time. The user can adjust the volume and tempo on their device and continue playing along with the music.
[0070] Furthermore, the server automatically generates visual content based on the generated music. During this process, effects and scenes that match the characteristics of the music are selected and provided to the user in the form of a promotional video. Users can share this visual content with other users, leading to the formation of new music communities.
[0071] As a concrete example, when a user inputs an audio signal to the system using an acoustic guitar, the server analyzes the signal and generates an accompaniment suitable for a folk style. An example of a prompt in this case might be, "Generate a folk-style accompaniment that suits this acoustic guitar performance." This allows the user to obtain an accompaniment that matches their performance in real time.
[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0073] Step 1:
[0074] The user launches a dedicated application on their device. The user inputs an audio signal into the system using an instrument or microphone. Specifically, the system converts the analog audio signal into a digital signal via the microphone or audio interface, generating data that includes volume and frequency information.
[0075] Step 2:
[0076] The terminal transmits the acquired digital audio signal to the server. Specifically, it sends digitized audio data at a constant sampling rate to the server in real time via internet communication. Data compression technology may be used to improve the efficiency of the transfer.
[0077] Step 3:
[0078] The server analyzes the received acoustic signal using information processing technology. Specifically, it uses a deep learning model to extract musical features such as tempo, key, and volume from the acoustic signal. Based on the input acoustic data, it calculates feature quantities and transfers the results to the next processing step.
[0079] Step 4:
[0080] The generative AI model instantly generates accompaniment based on the analysis results. The server executes an algorithm to generate music data that matches the musical style and tempo. Based on the input features, it obtains output that generates accompaniment music using virtual instruments.
[0081] Step 5:
[0082] The server streams the generated music to the user's device. The music data is encoded using a codec and output in a format that allows for real-time playback. The user can then decode and play the music on their device.
[0083] Step 6:
[0084] The server automatically generates visual content based on the generated music. The server selects effects and scenes according to the music's tempo and mood, creating data in a promotional video format. The input for the visual content is the generated music data, and the output is the completed video data.
[0085] Step 7:
[0086] Users share their completed music and visual content with other users. The server uploads the data to the platform used for sharing, facilitating content sharing among users. This allows users to receive feedback from others and explore new possibilities in musical expression.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Traditional music production systems lacked the flexibility to allow users to generate accompaniments in real time while performing themselves, and they also had difficulty sharing the generated music as visual content. Furthermore, providing a platform for collaborative creation among users was also challenging.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes means for receiving and analyzing user voice information, means for generating additional sound in real time based on the analyzed voice information, and means for generating visual content and automatically generating video based on the generated sound. This makes it possible for users to obtain accompaniment in real time while playing, easily share the generated music and video with other users, and enjoy collaborative creation.
[0092] A "user" is an individual or group that uses a music production system to input audio information, enjoys the generated audio and visual content, and shares or co-creates it with other users.
[0093] "Audio information" refers to audio signals, such as instrument performances or vocals, that users input into the system.
[0094] "Analysis" is the process of extracting features from received audio information using digital signal processing and artificial intelligence technology to obtain the information necessary for generating accompaniment.
[0095] "Additional sound effects" refer to accompaniments and backing tracks generated in real time by a generating AI based on the voice information input by the user.
[0096] "Visual content" refers to content that provides visual expressions such as videos and effects synchronized with generated music.
[0097] An "information processing device" is a device that allows a user to input voice information and receive generated audio or visual content, and includes smartphones, tablets, and personal computers.
[0098] "Sharing" refers to the activity of providing generated music and visual content to other users via a network, and engaging in communication and collaboration.
[0099] "Collaborative creation" is a process in which multiple users exchange ideas and make revisions based on the music and visual content they have generated, in order to create a single work.
[0100] This invention is a system for improving the user's music production experience. The user can input audio information using an information processing device and generate and share music and visual content in real time.
[0101] This system mainly consists of the following components:
[0102] First, there is the user's information processing device for inputting voice information. This includes smartphones, tablets, and personal computers, and users can record musical instrument performances or vocals. Voice input uses the device's microphone, and the audio signal is captured using a specific API (e.g., CoreAudio on iOS or AudioRecord on Android®).
[0103] Next, the server receives and analyzes the audio information sent by the user. This analysis uses deep learning technology and extracts audio features using libraries such as TENSORFLOW®. Based on this information, a generative AI model generates additional sounds in real time. This allows the user to instantly obtain accompaniment or backing tracks that match their performance.
[0104] Furthermore, the server uses the FFmpeg library to generate visual content based on the music data. The generated video is synchronized with the music, and this is designed to allow users to easily share it with others.
[0105] As a concrete example, suppose a user records themselves singing while playing the guitar at home. The user provides the AI with a prompt such as, "Please provide an upbeat rock-style accompaniment," which then generates an accompaniment perfectly suited to their performance. The generated music, along with corresponding visual content, plays simultaneously within the app and can be shared with other users via the network. This process promotes collaborative creation and feedback exchange among users, leading to new and exciting musical experiences.
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The user inputs voice information through their device.
[0109] The user records instrument or vocal sounds using the device's microphone, and the input audio data is acquired using an audio capture API. The input audio data is then ready to be sent to the server in real time.
[0110] Step 2:
[0111] The device sends voice information to the server.
[0112] The terminal uses the WebSocket protocol to send the input audio data to the server. The transmitted audio data is immediately received by the server.
[0113] Step 3:
[0114] The server analyzes the audio information and extracts features.
[0115] The server analyzes the received audio data using deep learning technology. Specifically, it uses a TensorFlow model to extract features from the audio signal and identifies musical elements based on those features. This provides the basic data that the generative AI uses to generate accompaniment.
[0116] Step 4:
[0117] The server generates additional sounds using a generated AI model.
[0118] Based on the analysis results, the server uses a generative AI model to create accompaniments and backing tracks in real time. The model adjusts the musical style according to the prompt "Please provide accompaniment in a light rock style" and generates appropriate sounds. This results in the generated music data as output.
[0119] Step 5:
[0120] The server generates visual content based on music data.
[0121] Based on the generated music data, the server automatically creates synchronized visual content using the FFmpeg library. Specifically, effects and scene transitions are set according to the music's tempo and rhythm, generating visually rich content. The generated visual content is output in a playable format.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] The system according to the present invention is an advanced music generation platform equipped with an emotion engine that utilizes the user's voice signal and extracts emotional information from it. This makes it possible to provide a music experience synchronized with the user's emotions.
[0124] First, the user records their performance or vocals through a device and inputs the audio signal into the system. In this process, the device sends the audio data to a server. The server performs digital signal processing to analyze the received audio signal and extract musical features. This analysis identifies the tempo, key, and rhythm.
[0125] A key feature here is the emotion engine, which is implemented on the server. Based on the analysis of audio signals, the emotion engine recognizes emotions from the user's voice and the tone of their performance. This emotional information becomes a crucial element in subsequent additional music generation processes.
[0126] The generation AI combines analyzed musical characteristics with recognized emotional information to generate additional music in real time. If the user's emotions, such as joy, sadness, or excitement, are detected, the accompaniment style and musical atmosphere are automatically adjusted to match each emotion. The generated music is then sent to the device to seamlessly integrate with the user's performance.
[0127] Furthermore, the system creates a promotional video based on the generated music. The video generation function takes extracted emotional information into account and dynamically adjusts the video's colors and effects. As a result, it is possible to provide a visual experience that perfectly matches the user's emotions.
[0128] As a concrete example, suppose a user plays the guitar in an excited state. In this case, the device captures the audio signal and sends it to the server. The server uses an emotion engine to detect strong energy and excitement. The generating AI takes this emotion into account and generates an upbeat, cheerful accompaniment, which is then delivered to the user in real time. The server also generates vivid and dynamic promotional videos, enhancing the user experience. With such a system, users can enjoy creating nuanced and personalized music tailored to their individual emotional state.
[0129] The following describes the processing flow.
[0130] Step 1:
[0131] The user launches an application on their device and inputs an audio signal through a microphone or musical instrument. The device digitizes the audio signal and transmits the data to the server in real time.
[0132] Step 2:
[0133] The server analyzes the received audio signal. Digital signal processing algorithms are used for the analysis to extract musical characteristics such as tempo, key, and rhythm.
[0134] Step 3:
[0135] The server uses an emotion engine to recognize the user's emotions from the analysis of the audio signal. This process analyzes changes in tone and volume of the voice to identify emotions such as joy, sadness, and excitement.
[0136] Step 4:
[0137] The server invokes a generation AI to generate additional music in real time based on analyzed musical characteristics and recognized emotions. The generated music reflects a specific accompaniment style and musical expression that aligns with the user's emotions.
[0138] Step 5:
[0139] The server sends the generated music to the terminal. The terminal synchronizes the user's performance or singing with the generated music and plays a unified track. The user can continue playing or making adjustments while monitoring this track in real time.
[0140] Step 6:
[0141] The server activates the promotional video generation function and generates a video with visual effects adjusted based on the generated music and recognized emotions. The video is constructed so that its colors and dynamics match the user's emotions.
[0142] Step 7:
[0143] Users share their completed music and promotional videos on the platform via their devices. The server performs management functions to provide user-uploaded content to other users and facilitates communication through comments and feedback.
[0144] (Example 2)
[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0146] There is a growing demand to provide a richer user experience than traditional music generation platforms can offer by making the music experience more personal and emotionally synchronized. However, current technology faces the challenge of effectively integrating user emotions with music generation and visual content creation.
[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0148] In this invention, the server includes means for acquiring audio data via a communication device and transmitting the audio data to an information processing device; means for a generative model to generate music content in real time using extracted musical characteristics and emotional information as input; and means for generating visual content based on the emotional information and adjusting it by adding dynamic effects. This makes it possible to generate and provide music and visual experiences optimized for the user's emotions in real time.
[0149] "Audio data" refers to information recorded in digital format of a user's performance or voice.
[0150] A "communication device" is a device used to acquire voice data and transmit it to an information processing device.
[0151] An "information processing device" is a device that analyzes audio data received from a communication device and extracts musical characteristics and emotional information.
[0152] "Signal processing" is a computational process used to extract musical characteristics such as tempo and key from audio data.
[0153] "Musical characteristics" refer to attributes related to music, such as tempo, key, and rhythm, that are analyzed from audio data.
[0154] "Emotional information" refers to information that indicates the user's emotional state, derived from the tone and analysis results of voice data.
[0155] A "generative model" is an artificial intelligence model that generates musical content in real time based on musical characteristics and emotional information.
[0156] "Musical content" refers to the musical elements and accompaniment generated by the generative model.
[0157] "Visual content" refers to media that includes images and effects generated in conjunction with music.
[0158] "Dynamic effects" refer to the use of colors and effects in visual content that change in response to emotional information.
[0159] The system according to the present invention acquires the user's voice as digital data, analyzes it to extract emotional information, and generates music and visual content based on that information.
[0160] Users record their performances or voices as audio data through a device. This device functions as a communication device and transmits that audio data to a server, which is an information processing device. The server is equipped with advanced digital signal processing capabilities, making it possible to extract musical characteristics such as tempo, key, and rhythm from the received audio data.
[0161] The server also features an emotion engine that analyzes musical characteristics and voice tone to recognize the user's emotional information. This emotion engine identifies emotions such as joy and sadness based on specific patterns and intonations contained in the voice data.
[0162] Next, the generative AI model is activated and generates a prompt sentence that combines musical characteristics and emotional information. For example, a possible prompt sentence would be, "The user has an audio recording of themselves playing while excited. Please generate up-tempo music that is appropriate for this." Based on this prompt sentence, the generative AI model generates music that matches the user's emotions in real time and seamlessly integrates it with the user's audio.
[0163] Furthermore, the server generates visual content based on emotional information, dynamically adjusting the video's colors and effects. This process reconstructs the user's provided audio data into music and visual experiences that resonate with their emotions.
[0164] For example, if a user performs an energetic piece, the device records it and sends it to a server. The server, using an emotion engine, detects high energy and excitement in the performance, and a generative AI model generates bright and lively music to match. Visual content is also dynamically adjusted to match this atmosphere, resulting in a new musical experience for the user.
[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0166] Step 1:
[0167] The user records their own performance or vocals using a terminal. The recorded audio data is sent directly from the communication device to the server, which is an information processing device. Here, the terminal converts the user's voice into a digital format and securely transmits that data to the server over the network. The input is the user's raw voice, and the output is digital audio data.
[0168] Step 2:
[0169] The server first passes the received audio data to a digital signal processing unit to extract musical features. This process uses algorithms such as FFT (Fast Fourier Transform) to analyze musical attributes such as tempo, key, and rhythm. The input is the digital audio data transmitted from the terminal, and the output is the extracted musical features.
[0170] Step 3:
[0171] The server analyzes the tone of the voice using an emotion engine based on musical characteristics and extracts the user's emotional information. The emotion engine employs algorithms that infer specific emotional states (e.g., joy, sadness, excitement, etc.) from the voice waveform and tone. The input is musical characteristics and the original voice data, and the output is the user's emotional information.
[0172] Step 4:
[0173] The server generates prompt statements for the generative AI model that integrate musical characteristics and emotional information, and uses these as input to generate musical content. These prompt statements specifically instruct the user's emotional state and the appropriate musical style for it. The generative AI model then regenerates the musical content in real time based on this information. The input is the prompt statement, and the output is the generated musical content.
[0174] Step 5:
[0175] The generated music is seamlessly integrated with the user's original audio data by the server. The integrated music is then sent back to the user and played on their device. At this stage, digital mixing technology is applied to harmonize the different sound sources. The input consists of the generated music and the original audio data, and the output is the integrated music data.
[0176] Step 6:
[0177] The server generates visual content using emotional information and optimizes it by combining it with dynamic effects. This creates visuals synchronized with the music. The visual content is presented to the user with its colors and effects adjusted according to the generated music. The input is emotional information and the generated music content, and the output is visual content.
[0178] Step 7:
[0179] Ultimately, the generated music and visual content are sent back from the server to the terminal, where the user can view it and even share it with other users. This process utilizes a distribution protocol to send and receive data, enabling interaction between users. The input consists of integrated music data and visual content, while the output is ready for distribution and sharing with end users.
[0180] (Application Example 2)
[0181] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0182] Conventional music generation systems provide music and videos without considering the user's emotional state, making it difficult to provide an experience tailored to individual emotions. As a result, users cannot enjoy a musical experience as an expression of their inner selves, and the generation of personalized, interactive content has been limited.
[0183] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0184] In this invention, the server includes means for analyzing the user's voice information, means for generating music and visual content based on the analyzed voice information and emotional information, and means for providing the generated content to the user and dynamically adjusting the visual effects. This makes it possible to provide music and video content synchronized with the user's emotions.
[0185] "User voice information" refers to the voice signals emitted by users through speech, singing, playing musical instruments, etc., and includes elements that express their emotional state.
[0186] "Means for analysis" refers to devices and programs that analyze collected audio information using digital signal processing technology and other methods to extract its characteristics and emotional information.
[0187] "Emotional information" refers to data indicating the emotional state extracted from the user's voice information, and it is a factor that determines the style of music and video generated from it.
[0188] "Additional music" refers to music that is dynamically generated based on the user's emotional information and seamlessly integrated into the user's expressive activities.
[0189] "Visual content" refers to video expressions that are generated based on the user's emotional information and whose colors and effects are dynamically adjusted, adding a visual aspect to the music experience.
[0190] "Dynamic adjustment" refers to the process of changing the style of music and video generated by the system in real time in response to changes in the user's emotions and voice information.
[0191] "Generated AI" refers to algorithms and programs that utilize voice and emotional information to automatically generate music and visual content tailored to the user.
[0192] "Means of sharing" refers to providing functions that enable users to share generated music and visual content with other users via a communication network, allowing them to view and play the content.
[0193] Modes for carrying out the invention
[0194] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. In this system, the user's voice information is acquired by the terminal and transmitted to the server. The terminal can be a communication device such as a smartphone or smart glasses. Digital signal processing technology is applied to analyze the voice information, extracting musical characteristics such as the tempo and rhythm of the voice. Furthermore, the server is equipped with an emotion engine that extracts the user's emotional information from the voice information.
[0195] The server generates music in real time using generated AI based on analyzed audio and emotional information. This generating AI operates with an algorithm that inputs the analyzed data into a learning model and generates melodies and accompaniments that correspond to emotions. Visual content tailored to the user's emotions is also generated, with colors and video effects dynamically adjusted. Machine learning frameworks such as TensorFlow and PyTorch are used to execute the generated AI.
[0196] This system sends generated music and visual content to a device, allowing users to enjoy a personalized music experience tailored to their emotions. For example, a user can play a musical instrument while viewing background music and dynamic visuals generated according to their emotions through smart glasses.
[0197] Examples of prompt messages are as follows:
[0198] "Analyze the user's emotions from their voice and generate an appropriate music style based on those emotions. If the emotion is [joy], prepare a bright and cheerful music style. Also, generate a video style that matches the emotion."
[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0200] Step 1:
[0201] The user inputs voice information using a device. The device's microphone captures the voice signal and sends the data to the server. The input here is the user's speech or musical performance, and the output is digitized voice data. The voice signal is converted into a digital format through sampling and quantization.
[0202] Step 2:
[0203] The server analyzes the received audio data. Using digital signal processing (DSP), it extracts musical characteristics such as tempo, rhythm, and key. The input is the audio data transmitted from the terminal, and the output is the musical characteristic data resulting from the analysis. Specifically, frequency analysis and spectral analysis of time-domain signals are performed.
[0204] Step 3:
[0205] The server uses an emotion engine to extract emotional information from audio data. Based on the tone and intonation of the voice, the AI model recognizes emotions such as joy, sadness, and excitement. The input is musical feature data, and the output is emotional information. A pre-trained generative AI model is used to classify specific emotional patterns from the audio waveform.
[0206] Step 4:
[0207] The server uses analyzed musical feature data and emotional information to generate music in real time using a generative AI. It dynamically constructs melodies, rhythms, and accompaniments according to the emotions. The input is musical feature data and emotional information, and the output is generated music data. The generative AI model executes algorithms to generate and select melody candidates based on emotional prompts.
[0208] Step 5:
[0209] The server generates visual content based on the user's emotional information. It dynamically adjusts colors and applies visual effects to visually enhance the emotional experience. The input is emotional information, and the output is generated visual data. Visual effects are dynamically determined by a programmed rule engine.
[0210] Step 6:
[0211] The generated music and visual content are sent to the device, where the user can experience it. In this process, the device displays the music and visual data in sync, enriching the user's experience. The input is the generated music and visual data, and the output is the presentation of content that resonates with the user's emotions.
[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0215] [Second Embodiment]
[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0228] The system according to the present invention supports improvisational performances and jam sessions by allowing users to participate through playing musical instruments or singing, and by having a generative AI generate additional music in real time. Users can participate in music sessions using their devices and enjoy an advanced musical experience through communication with the server.
[0229] First, the user launches an application on their device and inputs an audio signal to the system using an instrument or microphone. The device then transmits the input audio signal to the server in real time. After receiving the audio signal, the server analyzes it using technologies such as deep learning. Based on the analysis results, a generative AI generates an appropriate accompaniment or backing track.
[0230] The generated music is sent to the user's device, allowing them to enjoy new songs in real time, combined with their own performance or singing. Users can adjust the tempo and volume according to the generated music, enabling them to enjoy an even more creative musical experience.
[0231] Furthermore, the system also has a function to automatically generate promotional videos. This function provides users with rich video content that matches the music they create. The server selects effects and video scenes based on the characteristics of the generated music and delivers the video data to the terminal.
[0232] Finally, users can share their completed music and videos with other users on the platform. The server manages the shared content and provides a mechanism to facilitate communication through ratings and comments from other users. This makes it possible to create new spaces for collaboration and expand opportunities to participate in music production.
[0233] As a concrete example, let's say a user uses this system to sing while playing the guitar. The user launches the application on their device and begins recording their performance and singing. The server analyzes the input data, and the AI generates drum and bass accompaniment that matches the tempo of the performance. The user can continue playing along with the accompaniment, or add riffs or melodies. In addition, a promotional video matching the generated performance is created, which can be shared with other users to receive feedback and explore new musical ideas.
[0234] The following describes the processing flow.
[0235] Step 1:
[0236] The user launches an application on their device and inputs an audio signal using a musical instrument or microphone. The device digitizes the audio signal and transmits the signal data to the server in real time.
[0237] Step 2:
[0238] The server buffers the received audio signal and analyzes it using digital signal processing algorithms. This analysis extracts musical characteristics such as tempo, key, and rhythm.
[0239] Step 3:
[0240] The server invokes a generation AI based on the analysis results. The AI generates additional music, namely accompaniments and backing tracks, in real time. Deep learning technology is used for generation to create a natural flow of music.
[0241] Step 4:
[0242] The server sends the generated additional music to the terminal. The terminal plays this music in sync with the user's original performance or singing. The user can then add further performances to accompany the generated music.
[0243] Step 5:
[0244] Users can use music editing tools on their devices to adjust the tempo, volume, and harmony of the music. They can also add and edit effects to further refine the final track.
[0245] Step 6:
[0246] The server will then start a function to automatically generate a promotional video based on the completed music. It will select visual effects that match the characteristics of the music, build the video using a video template, and deliver it to the device.
[0247] Step 7:
[0248] Users share their completed music and videos on the platform via the server. The server notifies other users of the published content and facilitates communication through comment and rating functions.
[0249] (Example 1)
[0250] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0251] In recent years, creating and sharing music individually has become easier, but there remains a challenge in generating accompaniments and sharing music in real time during improvisational performances and jam sessions. In particular, there is a need for technology that allows for the easy creation and sharing of appropriate accompaniments and visual content when playing instruments or singing.
[0252] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0253] In this invention, the server includes means for acquiring a user's acoustic signal and transmitting the acoustic signal to a data processing device; means for analyzing the acquired acoustic signal using information processing technology in the data processing device; and means for generating artificial intelligence to instantly create music based on the analyzed acoustic data. This enables the user to generate appropriate accompaniment in real time during improvisational performances and to easily create and share visual content based on it.
[0254] An "acoustic signal" is an electrical signal that represents the sound waves generated by a user through playing a musical instrument or singing.
[0255] A "data processing device" is a computer system used to analyze acoustic signals and perform calculations for music generation.
[0256] "Information processing technology" refers to techniques used to analyze acoustic signals and extract features, and primarily includes methods such as deep learning.
[0257] "Generative artificial intelligence" refers to algorithms and models that automatically generate new music based on analysis results.
[0258] "Accompaniment" refers to additional sounds that add harmony and rhythm to a song played by a user, providing musical richness.
[0259] "Visual content" refers to videos, animations, and other materials synchronized with the music, provided to complement the musical experience.
[0260] "Shareable" refers to a state where music and visual content created by users can be shared with other users via the internet or other means.
[0261] This document describes embodiments for carrying out the invention. This invention is a system that analyzes a user's acoustic signal in real time and instantly generates a corresponding accompaniment using a generative AI model. Furthermore, this system automatically creates visual content based on the generated music and enables sharing with other users.
[0262] The user first launches a dedicated application on the terminal. The user inputs an audio signal using an instrument or microphone. The terminal captures this audio signal as a digital signal and transmits it to the server via the internet. The software on the terminal utilizes microphones and audio interfaces as hardware to efficiently collect the audio signal.
[0263] The server processes the received acoustic signals using a deep learning-based acoustic analysis method. This extracts the characteristics of the acoustic signals, and based on the analysis results, a generative AI model generates appropriate music. This generative AI model runs on a computer system with high computing power and utilizes large datasets to improve the accuracy of music generation.
[0264] The generated music is streamed to the user's device, allowing the user to enjoy the music experience in real time. The user can adjust the volume and tempo on their device and continue playing along with the music.
[0265] Furthermore, the server automatically generates visual content based on the generated music. During this process, effects and scenes that match the characteristics of the music are selected and provided to the user in the form of a promotional video. Users can share this visual content with other users, leading to the formation of new music communities.
[0266] As a concrete example, when a user inputs an audio signal to the system using an acoustic guitar, the server analyzes the signal and generates an accompaniment suitable for a folk style. An example of a prompt in this case might be, "Generate a folk-style accompaniment that suits this acoustic guitar performance." This allows the user to obtain an accompaniment that matches their performance in real time.
[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0268] Step 1:
[0269] The user launches a dedicated application on their device. The user inputs an audio signal into the system using an instrument or microphone. Specifically, the system converts the analog audio signal into a digital signal via the microphone or audio interface, generating data that includes volume and frequency information.
[0270] Step 2:
[0271] The terminal transmits the acquired digital audio signal to the server. Specifically, it sends digitized audio data at a constant sampling rate to the server in real time via internet communication. Data compression technology may be used to improve the efficiency of the transfer.
[0272] Step 3:
[0273] The server analyzes the received acoustic signal using information processing technology. Specifically, it uses a deep learning model to extract musical features such as tempo, key, and volume from the acoustic signal. Based on the input acoustic data, it calculates feature quantities and transfers the results to the next processing step.
[0274] Step 4:
[0275] The generative AI model instantly generates accompaniment based on the analysis results. The server executes an algorithm to generate music data that matches the musical style and tempo. Based on the input features, it obtains output that generates accompaniment music using virtual instruments.
[0276] Step 5:
[0277] The server streams the generated music to the user's device. The music data is encoded using a codec and output in a format that allows for real-time playback. The user can then decode and play the music on their device.
[0278] Step 6:
[0279] The server automatically generates visual content based on the generated music. The server selects effects and scenes according to the tempo and mood of the music and creates data in the form of a promotional video. The input of the visual content is the generated music data, and the output is the completed video data.
[0280] Step 7:
[0281] The user shares the completed music and visual content with other users. The server uploads the data to the platform used for sharing and promotes content sharing among users. As a result, the user can receive feedback from others and explore new possibilities for musical expression.
[0282] (Application Example 1)
[0283] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0284] In a conventional music production system, there are problems in that the user lacks flexibility to generate accompaniment in real time while performing by themselves, and it is difficult to easily share the generated music as visual content. Furthermore, it was also difficult to provide a place for collaborative production among users.
[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0286] In this invention, the server includes: means for receiving the user's voice information and analyzing the voice information; means for generating additional sounds in real time based on the analyzed voice information; and means for using means for generating visual content to automatically generate a video based on the generated sounds. As a result, the user can obtain accompaniment in real time while performing, and it becomes possible to easily share the generated music and video with other users and enjoy co-creation.
[0287] The "user" refers to an individual or group who inputs voice information using a music production system, enjoys the generated sounds and visual content, and shares or co-creates with other users.
[0288] The "voice information" refers to voice signals such as instrument performances and vocals input by the user into the system.
[0289] "Analysis" is a process of extracting the characteristics of the received voice information using digital signal processing and artificial intelligence technology based on the received voice information, and obtaining the information necessary for accompaniment generation.
[0290] The "additional sounds" refer to accompaniments and backing tracks generated in real time by the generation AI based on the voice information input by the user.
[0291] The "visual content" refers to content that provides visual expressions such as videos and effects synchronized with the generated music.
[0292] The "information processing device" is a device for the user to input voice information and receive the generated sounds and visual content, and includes smartphones, tablets, and personal computers.
[0293] "Sharing" is an activity of providing the generated music and visual content to other users via a network for communication and collaboration.
[0294] "Collaborative creation" is a process in which multiple users exchange ideas and make revisions based on the music and visual content they have generated, in order to create a single work.
[0295] This invention is a system for improving the user's music production experience. The user can input audio information using an information processing device and generate and share music and visual content in real time.
[0296] This system mainly consists of the following components:
[0297] First, there is the user's information processing device for inputting audio information. This includes smartphones, tablets, and personal computers, and users can record musical instrument performances or vocals. For audio input, the device's microphone is used, and the audio signal is captured using a specific API (e.g., CoreAudio on iOS or AudioRecord on Android).
[0298] Next, the server receives and analyzes the audio information sent by the user. This analysis uses deep learning techniques, extracting audio features using libraries such as TensorFlow. Based on this information, a generative AI model generates additional sounds in real time. This allows the user to instantly obtain accompaniment or backing tracks that match their performance.
[0299] Furthermore, the server uses the FFmpeg library to generate visual content based on the music data. The generated video is synchronized with the music, and this is designed to allow users to easily share it with others.
[0300] As a specific example, assume that a user records a song while playing the guitar at home. At this time, the user provides a prompt sentence such as "Please provide an accompaniment in a lively rock style" to the AI, and an accompaniment suitable for the user's performance is generated. The generated music is played in parallel within the app together with the corresponding visual content and can be directly shared with other users through the network. Through this process, co-creation and feedback exchange among users are promoted, and a new music experience can be obtained.
[0301] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0302] Step 1:
[0303] The user inputs voice information through the terminal.
[0304] The user uses the microphone of the terminal to record the sound of the instrument or vocals, and obtains the input voice data using the voice capture API. The input voice data is ready to be sent to the server in real time.
[0305] Step 2:
[0306] The terminal sends the voice information to the server.
[0307] The terminal uses the WebSocket protocol to send the input voice data to the server. The sent voice data is immediately received by the server.
[0308] Step 3:
[0309] The server analyzes the voice information and extracts features.
[0310] The server analyzes the received audio data using deep learning technology. Specifically, it uses a TensorFlow model to extract features from the audio signal and identifies musical elements based on those features. This provides the basic data that the generative AI uses to generate accompaniment.
[0311] Step 4:
[0312] The server generates additional sounds using a generated AI model.
[0313] Based on the analysis results, the server uses a generative AI model to create accompaniments and backing tracks in real time. The model adjusts the musical style according to the prompt "Please provide accompaniment in a light rock style" and generates appropriate sounds. This results in the generated music data as output.
[0314] Step 5:
[0315] The server generates visual content based on music data.
[0316] Based on the generated music data, the server automatically creates synchronized visual content using the FFmpeg library. Specifically, effects and scene transitions are set according to the music's tempo and rhythm, generating visually rich content. The generated visual content is output in a playable format.
[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0318] The system according to the present invention is an advanced music generation platform equipped with an emotion engine that utilizes the user's voice signal and extracts emotional information from it. This makes it possible to provide a music experience synchronized with the user's emotions.
[0319] First, the user records their performance or vocals through a device and inputs the audio signal into the system. In this process, the device sends the audio data to a server. The server performs digital signal processing to analyze the received audio signal and extract musical features. This analysis identifies the tempo, key, and rhythm.
[0320] A key feature here is the emotion engine, which is implemented on the server. Based on the analysis of audio signals, the emotion engine recognizes emotions from the user's voice and the tone of their performance. This emotional information becomes a crucial element in subsequent additional music generation processes.
[0321] The generation AI combines analyzed musical characteristics with recognized emotional information to generate additional music in real time. If the user's emotions, such as joy, sadness, or excitement, are detected, the accompaniment style and musical atmosphere are automatically adjusted to match each emotion. The generated music is then sent to the device to seamlessly integrate with the user's performance.
[0322] Furthermore, the system creates a promotional video based on the generated music. The video generation function takes extracted emotional information into account and dynamically adjusts the video's colors and effects. As a result, it is possible to provide a visual experience that perfectly matches the user's emotions.
[0323] As a concrete example, suppose a user plays the guitar in an excited state. In this case, the device captures the audio signal and sends it to the server. The server uses an emotion engine to detect strong energy and excitement. The generating AI takes this emotion into account and generates an upbeat, cheerful accompaniment, which is then delivered to the user in real time. The server also generates vivid and dynamic promotional videos, enhancing the user experience. With such a system, users can enjoy creating nuanced and personalized music tailored to their individual emotional state.
[0324] The following describes the processing flow.
[0325] Step 1:
[0326] The user launches an application on their device and inputs an audio signal through a microphone or musical instrument. The device digitizes the audio signal and transmits the data to the server in real time.
[0327] Step 2:
[0328] The server analyzes the received audio signal. Digital signal processing algorithms are used for the analysis to extract musical characteristics such as tempo, key, and rhythm.
[0329] Step 3:
[0330] The server uses an emotion engine to recognize the user's emotions from the analysis of the audio signal. This process analyzes changes in tone and volume of the voice to identify emotions such as joy, sadness, and excitement.
[0331] Step 4:
[0332] The server invokes a generation AI to generate additional music in real time based on analyzed musical characteristics and recognized emotions. The generated music reflects a specific accompaniment style and musical expression that aligns with the user's emotions.
[0333] Step 5:
[0334] The server sends the generated music to the terminal. The terminal synchronizes the user's performance or singing with the generated music and plays a unified track. The user can continue playing or making adjustments while monitoring this track in real time.
[0335] Step 6:
[0336] The server activates the promotional video generation function and generates a video with visual effects adjusted based on the generated music and recognized emotions. The video is constructed so that its colors and dynamics match the user's emotions.
[0337] Step 7:
[0338] Users share their completed music and promotional videos on the platform via their devices. The server performs management functions to provide user-uploaded content to other users and facilitates communication through comments and feedback.
[0339] (Example 2)
[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0341] There is a growing demand to provide a richer user experience than traditional music generation platforms can offer by making the music experience more personal and emotionally synchronized. However, current technology faces the challenge of effectively integrating user emotions with music generation and visual content creation.
[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0343] In this invention, the server includes means for acquiring audio data via a communication device and transmitting the audio data to an information processing device; means for a generative model to generate music content in real time using extracted musical characteristics and emotional information as input; and means for generating visual content based on the emotional information and adjusting it by adding dynamic effects. This makes it possible to generate and provide music and visual experiences optimized for the user's emotions in real time.
[0344] "Audio data" refers to information recorded in digital format of a user's performance or voice.
[0345] A "communication device" is a device used to acquire voice data and transmit it to an information processing device.
[0346] An "information processing device" is a device that analyzes audio data received from a communication device and extracts musical characteristics and emotional information.
[0347] "Signal processing" is a computational process used to extract musical characteristics such as tempo and key from audio data.
[0348] "Musical characteristics" refer to attributes related to music, such as tempo, key, and rhythm, that are analyzed from audio data.
[0349] "Emotional information" refers to information that indicates the user's emotional state, derived from the tone and analysis results of voice data.
[0350] A "generative model" is an artificial intelligence model that generates musical content in real time based on musical characteristics and emotional information.
[0351] "Musical content" refers to the musical elements and accompaniment generated by the generative model.
[0352] "Visual content" refers to media that includes images and effects generated in conjunction with music.
[0353] "Dynamic effects" refer to the use of colors and effects in visual content that change in response to emotional information.
[0354] The system according to the present invention acquires the user's voice as digital data, analyzes it to extract emotional information, and generates music and visual content based on that information.
[0355] Users record their performances or voices as audio data through a device. This device functions as a communication device and transmits that audio data to a server, which is an information processing device. The server is equipped with advanced digital signal processing capabilities, making it possible to extract musical characteristics such as tempo, key, and rhythm from the received audio data.
[0356] The server also features an emotion engine that analyzes musical characteristics and voice tone to recognize the user's emotional information. This emotion engine identifies emotions such as joy and sadness based on specific patterns and intonations contained in the voice data.
[0357] Next, the generative AI model is activated and generates a prompt sentence that combines musical characteristics and emotional information. For example, a possible prompt sentence would be, "The user has an audio recording of themselves playing while excited. Please generate up-tempo music that is appropriate for this." Based on this prompt sentence, the generative AI model generates music that matches the user's emotions in real time and seamlessly integrates it with the user's audio.
[0358] Furthermore, the server generates visual content based on emotional information, dynamically adjusting the video's colors and effects. This process reconstructs the user's provided audio data into music and visual experiences that resonate with their emotions.
[0359] For example, if a user performs an energetic piece, the device records it and sends it to a server. The server, using an emotion engine, detects high energy and excitement in the performance, and a generative AI model generates bright and lively music to match. Visual content is also dynamically adjusted to match this atmosphere, resulting in a new musical experience for the user.
[0360] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0361] Step 1:
[0362] The user records their own performance or vocals using a terminal. The recorded audio data is sent directly from the communication device to the server, which is an information processing device. Here, the terminal converts the user's voice into a digital format and securely transmits that data to the server over the network. The input is the user's raw voice, and the output is digital audio data.
[0363] Step 2:
[0364] The server first passes the received audio data to a digital signal processing unit to extract musical features. This process uses algorithms such as FFT (Fast Fourier Transform) to analyze musical attributes such as tempo, key, and rhythm. The input is the digital audio data transmitted from the terminal, and the output is the extracted musical features.
[0365] Step 3:
[0366] The server analyzes the tone of the voice using an emotion engine based on musical characteristics and extracts the user's emotional information. The emotion engine employs algorithms that infer specific emotional states (e.g., joy, sadness, excitement, etc.) from the voice waveform and tone. The input is musical characteristics and the original voice data, and the output is the user's emotional information.
[0367] Step 4:
[0368] The server generates prompt statements for the generative AI model that integrate musical characteristics and emotional information, and uses these as input to generate musical content. These prompt statements specifically instruct the user's emotional state and the appropriate musical style for it. The generative AI model then regenerates the musical content in real time based on this information. The input is the prompt statement, and the output is the generated musical content.
[0369] Step 5:
[0370] The generated music is seamlessly integrated with the user's original audio data by the server. The integrated music is then sent back to the user and played on their device. At this stage, digital mixing technology is applied to harmonize the different sound sources. The input consists of the generated music and the original audio data, and the output is the integrated music data.
[0371] Step 6:
[0372] The server generates visual content using emotional information and optimizes it by combining it with dynamic effects. This creates visuals synchronized with the music. The visual content is presented to the user with its colors and effects adjusted according to the generated music. The input is emotional information and the generated music content, and the output is visual content.
[0373] Step 7:
[0374] Ultimately, the generated music and visual content are sent back from the server to the terminal, where the user can view it and even share it with other users. This process utilizes a distribution protocol to send and receive data, enabling interaction between users. The input consists of integrated music data and visual content, while the output is ready for distribution and sharing with end users.
[0375] (Application Example 2)
[0376] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0377] Conventional music generation systems provide music and videos without considering the user's emotional state, making it difficult to provide an experience tailored to individual emotions. As a result, users cannot enjoy a musical experience as an expression of their inner selves, and the generation of personalized, interactive content has been limited.
[0378] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0379] In this invention, the server includes means for analyzing the user's voice information, means for generating music and visual content based on the analyzed voice information and emotional information, and means for providing the generated content to the user and dynamically adjusting the visual effects. This makes it possible to provide music and video content synchronized with the user's emotions.
[0380] "User voice information" refers to the voice signals emitted by users through speech, singing, playing musical instruments, etc., and includes elements that express their emotional state.
[0381] "Means for analysis" refers to devices and programs that analyze collected audio information using digital signal processing technology and other methods to extract its characteristics and emotional information.
[0382] "Emotional information" refers to data indicating the emotional state extracted from the user's voice information, and it is a factor that determines the style of music and video generated from it.
[0383] "Additional music" refers to music that is dynamically generated based on the user's emotional information and seamlessly integrated into the user's expressive activities.
[0384] "Visual content" refers to video expressions that are generated based on the user's emotional information and whose colors and effects are dynamically adjusted, adding a visual aspect to the music experience.
[0385] "Dynamic adjustment" refers to the process of changing the style of music and video generated by the system in real time in response to changes in the user's emotions and voice information.
[0386] "Generated AI" refers to algorithms and programs that utilize voice and emotional information to automatically generate music and visual content tailored to the user.
[0387] "Means of sharing" refers to providing functions that enable users to share generated music and visual content with other users via a communication network, allowing them to view and play the content.
[0388] Modes for carrying out the invention
[0389] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. In this system, the user's voice information is acquired by the terminal and transmitted to the server. The terminal can be a communication device such as a smartphone or smart glasses. Digital signal processing technology is applied to analyze the voice information, extracting musical characteristics such as the tempo and rhythm of the voice. Furthermore, the server is equipped with an emotion engine that extracts the user's emotional information from the voice information.
[0390] The server generates music in real time using generated AI based on analyzed audio and emotional information. This generating AI operates with an algorithm that inputs the analyzed data into a learning model and generates melodies and accompaniments that correspond to emotions. Visual content tailored to the user's emotions is also generated, with colors and video effects dynamically adjusted. Machine learning frameworks such as TensorFlow and PyTorch are used to execute the generated AI.
[0391] This system sends generated music and visual content to a device, allowing users to enjoy a personalized music experience tailored to their emotions. For example, a user can play a musical instrument while viewing background music and dynamic visuals generated according to their emotions through smart glasses.
[0392] Examples of prompt messages are as follows:
[0393] "Analyze the user's emotions from their voice and generate an appropriate music style based on those emotions. If the emotion is [joy], prepare a bright and cheerful music style. Also, generate a video style that matches the emotion."
[0394] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0395] Step 1:
[0396] The user inputs voice information using a device. The device's microphone captures the voice signal and sends the data to the server. The input here is the user's speech or musical performance, and the output is digitized voice data. The voice signal is converted into a digital format through sampling and quantization.
[0397] Step 2:
[0398] The server analyzes the received audio data. Using digital signal processing (DSP), it extracts musical characteristics such as tempo, rhythm, and key. The input is the audio data transmitted from the terminal, and the output is the musical characteristic data resulting from the analysis. Specifically, frequency analysis and spectral analysis of time-domain signals are performed.
[0399] Step 3:
[0400] The server uses an emotion engine to extract emotional information from audio data. Based on the tone and intonation of the voice, the AI model recognizes emotions such as joy, sadness, and excitement. The input is musical feature data, and the output is emotional information. A pre-trained generative AI model is used to classify specific emotional patterns from the audio waveform.
[0401] Step 4:
[0402] The server uses analyzed musical feature data and emotional information to generate music in real time using a generative AI. It dynamically constructs melodies, rhythms, and accompaniments according to the emotions. The input is musical feature data and emotional information, and the output is generated music data. The generative AI model executes algorithms to generate and select melody candidates based on emotional prompts.
[0403] Step 5:
[0404] The server generates visual content based on the user's emotional information. It dynamically adjusts colors and applies visual effects to visually enhance the emotional experience. The input is emotional information, and the output is generated visual data. Visual effects are dynamically determined by a programmed rule engine.
[0405] Step 6:
[0406] The generated music and visual content are sent to the device, where the user can experience it. In this process, the device displays the music and visual data in sync, enriching the user's experience. The input is the generated music and visual data, and the output is the presentation of content that resonates with the user's emotions.
[0407] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0410] [Third Embodiment]
[0411] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0412] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0413] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0414] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0415] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0417] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0418] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0419] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0420] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0421] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0422] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0423] The system according to the present invention supports improvisational performances and jam sessions by allowing users to participate through playing musical instruments or singing, and by having a generative AI generate additional music in real time. Users can participate in music sessions using their devices and enjoy an advanced musical experience through communication with the server.
[0424] First, the user launches an application on their device and inputs an audio signal to the system using an instrument or microphone. The device then transmits the input audio signal to the server in real time. After receiving the audio signal, the server analyzes it using technologies such as deep learning. Based on the analysis results, a generative AI generates an appropriate accompaniment or backing track.
[0425] The generated music is sent to the user's device, allowing them to enjoy new songs in real time, combined with their own performance or singing. Users can adjust the tempo and volume according to the generated music, enabling them to enjoy an even more creative musical experience.
[0426] Furthermore, the system also has a function to automatically generate promotional videos. This function provides users with rich video content that matches the music they create. The server selects effects and video scenes based on the characteristics of the generated music and delivers the video data to the terminal.
[0427] Finally, users can share their completed music and videos with other users on the platform. The server manages the shared content and provides a mechanism to facilitate communication through ratings and comments from other users. This makes it possible to create new spaces for collaboration and expand opportunities to participate in music production.
[0428] As a concrete example, let's say a user uses this system to sing while playing the guitar. The user launches the application on their device and begins recording their performance and singing. The server analyzes the input data, and the AI generates drum and bass accompaniment that matches the tempo of the performance. The user can continue playing along with the accompaniment, or add riffs or melodies. In addition, a promotional video matching the generated performance is created, which can be shared with other users to receive feedback and explore new musical ideas.
[0429] The following describes the processing flow.
[0430] Step 1:
[0431] The user launches an application on their device and inputs an audio signal using a musical instrument or microphone. The device digitizes the audio signal and transmits the signal data to the server in real time.
[0432] Step 2:
[0433] The server buffers the received audio signal and analyzes it using digital signal processing algorithms. This analysis extracts musical characteristics such as tempo, key, and rhythm.
[0434] Step 3:
[0435] The server invokes a generation AI based on the analysis results. The AI generates additional music, namely accompaniments and backing tracks, in real time. Deep learning technology is used for generation to create a natural flow of music.
[0436] Step 4:
[0437] The server sends the generated additional music to the terminal. The terminal plays this music in sync with the user's original performance or singing. The user can then add further performances to accompany the generated music.
[0438] Step 5:
[0439] Users can use music editing tools on their devices to adjust the tempo, volume, and harmony of the music. They can also add and edit effects to further refine the final track.
[0440] Step 6:
[0441] The server will then start a function to automatically generate a promotional video based on the completed music. It will select visual effects that match the characteristics of the music, build the video using a video template, and deliver it to the device.
[0442] Step 7:
[0443] Users share their completed music and videos on the platform via the server. The server notifies other users of the published content and facilitates communication through comment and rating functions.
[0444] (Example 1)
[0445] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0446] In recent years, creating and sharing music individually has become easier, but there remains a challenge in generating accompaniments and sharing music in real time during improvisational performances and jam sessions. In particular, there is a need for technology that allows for the easy creation and sharing of appropriate accompaniments and visual content when playing instruments or singing.
[0447] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0448] In this invention, the server includes means for acquiring a user's acoustic signal and transmitting the acoustic signal to a data processing device; means for analyzing the acquired acoustic signal using information processing technology in the data processing device; and means for generating artificial intelligence to instantly create music based on the analyzed acoustic data. This enables the user to generate appropriate accompaniment in real time during improvisational performances and to easily create and share visual content based on it.
[0449] An "acoustic signal" is an electrical signal that represents the sound waves generated by a user through playing a musical instrument or singing.
[0450] A "data processing device" is a computer system used to analyze acoustic signals and perform calculations for music generation.
[0451] "Information processing technology" refers to techniques used to analyze acoustic signals and extract features, and primarily includes methods such as deep learning.
[0452] "Generative artificial intelligence" refers to algorithms and models that automatically generate new music based on analysis results.
[0453] "Accompaniment" refers to additional sounds that add harmony and rhythm to a song played by a user, providing musical richness.
[0454] "Visual content" refers to videos, animations, and other materials synchronized with the music, provided to complement the musical experience.
[0455] "Shareable" refers to a state where music and visual content created by users can be shared with other users via the internet or other means.
[0456] This document describes embodiments for carrying out the invention. This invention is a system that analyzes a user's acoustic signal in real time and instantly generates a corresponding accompaniment using a generative AI model. Furthermore, this system automatically creates visual content based on the generated music and enables sharing with other users.
[0457] The user first launches a dedicated application on the terminal. The user inputs an audio signal using an instrument or microphone. The terminal captures this audio signal as a digital signal and transmits it to the server via the internet. The software on the terminal utilizes microphones and audio interfaces as hardware to efficiently collect the audio signal.
[0458] The server processes the received acoustic signals using a deep learning-based acoustic analysis method. This extracts the characteristics of the acoustic signals, and based on the analysis results, a generative AI model generates appropriate music. This generative AI model runs on a computer system with high computing power and utilizes large datasets to improve the accuracy of music generation.
[0459] The generated music is streamed to the user's device, allowing the user to enjoy the music experience in real time. The user can adjust the volume and tempo on their device and continue playing along with the music.
[0460] Furthermore, the server automatically generates visual content based on the generated music. During this process, effects and scenes that match the characteristics of the music are selected and provided to the user in the form of a promotional video. Users can share this visual content with other users, leading to the formation of new music communities.
[0461] As a concrete example, when a user inputs an audio signal to the system using an acoustic guitar, the server analyzes the signal and generates an accompaniment suitable for a folk style. An example of a prompt in this case might be, "Generate a folk-style accompaniment that suits this acoustic guitar performance." This allows the user to obtain an accompaniment that matches their performance in real time.
[0462] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0463] Step 1:
[0464] The user launches a dedicated application on their device. The user inputs an audio signal into the system using an instrument or microphone. Specifically, the system converts the analog audio signal into a digital signal via the microphone or audio interface, generating data that includes volume and frequency information.
[0465] Step 2:
[0466] The terminal transmits the acquired digital audio signal to the server. Specifically, it sends digitized audio data at a constant sampling rate to the server in real time via internet communication. Data compression technology may be used to improve the efficiency of the transfer.
[0467] Step 3:
[0468] The server analyzes the received acoustic signal using information processing technology. Specifically, it uses a deep learning model to extract musical features such as tempo, key, and volume from the acoustic signal. Based on the input acoustic data, it calculates feature quantities and transfers the results to the next processing step.
[0469] Step 4:
[0470] The generative AI model instantly generates accompaniment based on the analysis results. The server executes an algorithm to generate music data that matches the musical style and tempo. Based on the input features, it obtains output that generates accompaniment music using virtual instruments.
[0471] Step 5:
[0472] The server streams the generated music to the user's device. The music data is encoded using a codec and output in a format that allows for real-time playback. The user can then decode and play the music on their device.
[0473] Step 6:
[0474] The server automatically generates visual content based on the generated music. The server selects effects and scenes according to the music's tempo and mood, creating data in a promotional video format. The input for the visual content is the generated music data, and the output is the completed video data.
[0475] Step 7:
[0476] Users share their completed music and visual content with other users. The server uploads the data to the platform used for sharing, facilitating content sharing among users. This allows users to receive feedback from others and explore new possibilities in musical expression.
[0477] (Application Example 1)
[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] Traditional music production systems lacked the flexibility to allow users to generate accompaniments in real time while performing themselves, and they also had difficulty sharing the generated music as visual content. Furthermore, providing a platform for collaborative creation among users was also challenging.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0481] In this invention, the server includes means for receiving and analyzing user voice information, means for generating additional sound in real time based on the analyzed voice information, and means for generating visual content and automatically generating video based on the generated sound. This makes it possible for users to obtain accompaniment in real time while playing, easily share the generated music and video with other users, and enjoy collaborative creation.
[0482] A "user" is an individual or group that uses a music production system to input audio information, enjoys the generated audio and visual content, and shares or co-creates it with other users.
[0483] "Audio information" refers to audio signals, such as instrument performances or vocals, that users input into the system.
[0484] "Analysis" is the process of extracting features from received audio information using digital signal processing and artificial intelligence technology to obtain the information necessary for generating accompaniment.
[0485] "Additional sound effects" refer to accompaniments and backing tracks generated in real time by a generating AI based on the voice information input by the user.
[0486] "Visual content" refers to content that provides visual expressions such as videos and effects synchronized with generated music.
[0487] An "information processing device" is a device that allows a user to input voice information and receive generated audio or visual content, and includes smartphones, tablets, and personal computers.
[0488] "Sharing" refers to the activity of providing generated music and visual content to other users via a network, and engaging in communication and collaboration.
[0489] "Collaborative creation" is a process in which multiple users exchange ideas and make revisions based on the music and visual content they have generated, in order to create a single work.
[0490] This invention is a system for improving the user's music production experience. The user can input audio information using an information processing device and generate and share music and visual content in real time.
[0491] This system mainly consists of the following components:
[0492] First, there is the user's information processing device for inputting audio information. This includes smartphones, tablets, and personal computers, and users can record musical instrument performances or vocals. For audio input, the device's microphone is used, and the audio signal is captured using a specific API (e.g., CoreAudio on iOS or AudioRecord on Android).
[0493] Next, the server receives and analyzes the audio information sent by the user. This analysis uses deep learning techniques, extracting audio features using libraries such as TensorFlow. Based on this information, a generative AI model generates additional sounds in real time. This allows the user to instantly obtain accompaniment or backing tracks that match their performance.
[0494] Furthermore, the server uses the FFmpeg library to generate visual content based on the music data. The generated video is synchronized with the music, and this is designed to allow users to easily share it with others.
[0495] As a concrete example, suppose a user records themselves singing while playing the guitar at home. The user provides the AI with a prompt such as, "Please provide an upbeat rock-style accompaniment," which then generates an accompaniment perfectly suited to their performance. The generated music, along with corresponding visual content, plays simultaneously within the app and can be shared with other users via the network. This process promotes collaborative creation and feedback exchange among users, leading to new and exciting musical experiences.
[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0497] Step 1:
[0498] The user inputs voice information through their device.
[0499] The user records instrument or vocal sounds using the device's microphone, and the input audio data is acquired using an audio capture API. The input audio data is then ready to be sent to the server in real time.
[0500] Step 2:
[0501] The device sends voice information to the server.
[0502] The terminal uses the WebSocket protocol to send the input audio data to the server. The transmitted audio data is immediately received by the server.
[0503] Step 3:
[0504] The server analyzes the audio information and extracts features.
[0505] The server analyzes the received audio data using deep learning technology. Specifically, it uses a TensorFlow model to extract features from the audio signal and identifies musical elements based on those features. This provides the basic data that the generative AI uses to generate accompaniment.
[0506] Step 4:
[0507] The server generates additional sounds using a generated AI model.
[0508] Based on the analysis results, the server uses a generative AI model to create accompaniments and backing tracks in real time. The model adjusts the musical style according to the prompt "Please provide accompaniment in a light rock style" and generates appropriate sounds. This results in the generated music data as output.
[0509] Step 5:
[0510] The server generates visual content based on music data.
[0511] Based on the generated music data, the server automatically creates synchronized visual content using the FFmpeg library. Specifically, effects and scene transitions are set according to the music's tempo and rhythm, generating visually rich content. The generated visual content is output in a playable format.
[0512] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0513] The system according to the present invention is an advanced music generation platform equipped with an emotion engine that utilizes the user's voice signal and extracts emotional information from it. This makes it possible to provide a music experience synchronized with the user's emotions.
[0514] First, the user records their performance or vocals through a device and inputs the audio signal into the system. In this process, the device sends the audio data to a server. The server performs digital signal processing to analyze the received audio signal and extract musical features. This analysis identifies the tempo, key, and rhythm.
[0515] A key feature here is the emotion engine, which is implemented on the server. Based on the analysis of audio signals, the emotion engine recognizes emotions from the user's voice and the tone of their performance. This emotional information becomes a crucial element in subsequent additional music generation processes.
[0516] The generation AI combines analyzed musical characteristics with recognized emotional information to generate additional music in real time. If the user's emotions, such as joy, sadness, or excitement, are detected, the accompaniment style and musical atmosphere are automatically adjusted to match each emotion. The generated music is then sent to the device to seamlessly integrate with the user's performance.
[0517] Furthermore, the system creates a promotional video based on the generated music. The video generation function takes extracted emotional information into account and dynamically adjusts the video's colors and effects. As a result, it is possible to provide a visual experience that perfectly matches the user's emotions.
[0518] As a concrete example, suppose a user plays the guitar in an excited state. In this case, the device captures the audio signal and sends it to the server. The server uses an emotion engine to detect strong energy and excitement. The generating AI takes this emotion into account and generates an upbeat, cheerful accompaniment, which is then delivered to the user in real time. The server also generates vivid and dynamic promotional videos, enhancing the user experience. With such a system, users can enjoy creating nuanced and personalized music tailored to their individual emotional state.
[0519] The following describes the processing flow.
[0520] Step 1:
[0521] The user launches an application on their device and inputs an audio signal through a microphone or musical instrument. The device digitizes the audio signal and transmits the data to the server in real time.
[0522] Step 2:
[0523] The server analyzes the received audio signal. Digital signal processing algorithms are used for the analysis to extract musical characteristics such as tempo, key, and rhythm.
[0524] Step 3:
[0525] The server uses an emotion engine to recognize the user's emotions from the analysis of the audio signal. This process analyzes changes in tone and volume of the voice to identify emotions such as joy, sadness, and excitement.
[0526] Step 4:
[0527] The server invokes a generation AI to generate additional music in real time based on analyzed musical characteristics and recognized emotions. The generated music reflects a specific accompaniment style and musical expression that aligns with the user's emotions.
[0528] Step 5:
[0529] The server sends the generated music to the terminal. The terminal synchronizes the user's performance or singing with the generated music and plays a unified track. The user can continue playing or making adjustments while monitoring this track in real time.
[0530] Step 6:
[0531] The server activates the promotional video generation function and generates a video with visual effects adjusted based on the generated music and recognized emotions. The video is constructed so that its colors and dynamics match the user's emotions.
[0532] Step 7:
[0533] Users share their completed music and promotional videos on the platform via their devices. The server performs management functions to provide user-uploaded content to other users and facilitates communication through comments and feedback.
[0534] (Example 2)
[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] There is a growing demand to provide a richer user experience than traditional music generation platforms can offer by making the music experience more personal and emotionally synchronized. However, current technology faces the challenge of effectively integrating user emotions with music generation and visual content creation.
[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0538] In this invention, the server includes means for acquiring audio data via a communication device and transmitting the audio data to an information processing device; means for a generative model to generate music content in real time using extracted musical characteristics and emotional information as input; and means for generating visual content based on the emotional information and adjusting it by adding dynamic effects. This makes it possible to generate and provide music and visual experiences optimized for the user's emotions in real time.
[0539] "Audio data" refers to information recorded in digital format of a user's performance or voice.
[0540] A "communication device" is a device used to acquire voice data and transmit it to an information processing device.
[0541] An "information processing device" is a device that analyzes audio data received from a communication device and extracts musical characteristics and emotional information.
[0542] "Signal processing" is a computational process used to extract musical characteristics such as tempo and key from audio data.
[0543] "Musical characteristics" refer to attributes related to music, such as tempo, key, and rhythm, that are analyzed from audio data.
[0544] "Emotional information" refers to information that indicates the user's emotional state, derived from the tone and analysis results of voice data.
[0545] A "generative model" is an artificial intelligence model that generates musical content in real time based on musical characteristics and emotional information.
[0546] "Musical content" refers to the musical elements and accompaniment generated by the generative model.
[0547] "Visual content" refers to media that includes images and effects generated in conjunction with music.
[0548] "Dynamic effects" refer to the use of colors and effects in visual content that change in response to emotional information.
[0549] The system according to the present invention acquires the user's voice as digital data, analyzes it to extract emotional information, and generates music and visual content based on that information.
[0550] Users record their performances or voices as audio data through a device. This device functions as a communication device and transmits that audio data to a server, which is an information processing device. The server is equipped with advanced digital signal processing capabilities, making it possible to extract musical characteristics such as tempo, key, and rhythm from the received audio data.
[0551] The server also features an emotion engine that analyzes musical characteristics and voice tone to recognize the user's emotional information. This emotion engine identifies emotions such as joy and sadness based on specific patterns and intonations contained in the voice data.
[0552] Next, the generative AI model is activated and generates a prompt sentence that combines musical characteristics and emotional information. For example, a possible prompt sentence would be, "The user has an audio recording of themselves playing while excited. Please generate up-tempo music that is appropriate for this." Based on this prompt sentence, the generative AI model generates music that matches the user's emotions in real time and seamlessly integrates it with the user's audio.
[0553] Furthermore, the server generates visual content based on emotional information, dynamically adjusting the video's colors and effects. This process reconstructs the user's provided audio data into music and visual experiences that resonate with their emotions.
[0554] For example, if a user performs an energetic piece, the device records it and sends it to a server. The server, using an emotion engine, detects high energy and excitement in the performance, and a generative AI model generates bright and lively music to match. Visual content is also dynamically adjusted to match this atmosphere, resulting in a new musical experience for the user.
[0555] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0556] Step 1:
[0557] The user records their own performance or vocals using a terminal. The recorded audio data is sent directly from the communication device to the server, which is an information processing device. Here, the terminal converts the user's voice into a digital format and securely transmits that data to the server over the network. The input is the user's raw voice, and the output is digital audio data.
[0558] Step 2:
[0559] The server first passes the received audio data to a digital signal processing unit to extract musical features. This process uses algorithms such as FFT (Fast Fourier Transform) to analyze musical attributes such as tempo, key, and rhythm. The input is the digital audio data transmitted from the terminal, and the output is the extracted musical features.
[0560] Step 3:
[0561] The server analyzes the tone of the voice using an emotion engine based on musical characteristics and extracts the user's emotional information. The emotion engine employs algorithms that infer specific emotional states (e.g., joy, sadness, excitement, etc.) from the voice waveform and tone. The input is musical characteristics and the original voice data, and the output is the user's emotional information.
[0562] Step 4:
[0563] The server generates prompt statements for the generative AI model that integrate musical characteristics and emotional information, and uses these as input to generate musical content. These prompt statements specifically instruct the user's emotional state and the appropriate musical style for it. The generative AI model then regenerates the musical content in real time based on this information. The input is the prompt statement, and the output is the generated musical content.
[0564] Step 5:
[0565] The generated music is seamlessly integrated with the user's original audio data by the server. The integrated music is then sent back to the user and played on their device. At this stage, digital mixing technology is applied to harmonize the different sound sources. The input consists of the generated music and the original audio data, and the output is the integrated music data.
[0566] Step 6:
[0567] The server generates visual content using emotional information and optimizes it by combining it with dynamic effects. This creates visuals synchronized with the music. The visual content is presented to the user with its colors and effects adjusted according to the generated music. The input is emotional information and the generated music content, and the output is visual content.
[0568] Step 7:
[0569] Ultimately, the generated music and visual content are sent back from the server to the terminal, where the user can view it and even share it with other users. This process utilizes a distribution protocol to send and receive data, enabling interaction between users. The input consists of integrated music data and visual content, while the output is ready for distribution and sharing with end users.
[0570] (Application Example 2)
[0571] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0572] Conventional music generation systems provide music and videos without considering the user's emotional state, making it difficult to provide an experience tailored to individual emotions. As a result, users cannot enjoy a musical experience as an expression of their inner selves, and the generation of personalized, interactive content has been limited.
[0573] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0574] In this invention, the server includes means for analyzing the user's voice information, means for generating music and visual content based on the analyzed voice information and emotional information, and means for providing the generated content to the user and dynamically adjusting the visual effects. This makes it possible to provide music and video content synchronized with the user's emotions.
[0575] "User voice information" refers to the voice signals emitted by users through speech, singing, playing musical instruments, etc., and includes elements that express their emotional state.
[0576] "Means for analysis" refers to devices and programs that analyze collected audio information using digital signal processing technology and other methods to extract its characteristics and emotional information.
[0577] "Emotional information" refers to data indicating the emotional state extracted from the user's voice information, and it is a factor that determines the style of music and video generated from it.
[0578] "Additional music" refers to music that is dynamically generated based on the user's emotional information and seamlessly integrated into the user's expressive activities.
[0579] "Visual content" refers to video expressions that are generated based on the user's emotional information and whose colors and effects are dynamically adjusted, adding a visual aspect to the music experience.
[0580] "Dynamic adjustment" refers to the process of changing the style of music and video generated by the system in real time in response to changes in the user's emotions and voice information.
[0581] "Generated AI" refers to algorithms and programs that utilize voice and emotional information to automatically generate music and visual content tailored to the user.
[0582] "Means of sharing" refers to providing functions that enable users to share generated music and visual content with other users via a communication network, allowing them to view and play the content.
[0583] Modes for carrying out the invention
[0584] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. In this system, the user's voice information is acquired by the terminal and transmitted to the server. The terminal can be a communication device such as a smartphone or smart glasses. Digital signal processing technology is applied to analyze the voice information, extracting musical characteristics such as the tempo and rhythm of the voice. Furthermore, the server is equipped with an emotion engine that extracts the user's emotional information from the voice information.
[0585] The server generates music in real time using generated AI based on analyzed audio and emotional information. This generating AI operates with an algorithm that inputs the analyzed data into a learning model and generates melodies and accompaniments that correspond to emotions. Visual content tailored to the user's emotions is also generated, with colors and video effects dynamically adjusted. Machine learning frameworks such as TensorFlow and PyTorch are used to execute the generated AI.
[0586] This system sends generated music and visual content to a device, allowing users to enjoy a personalized music experience tailored to their emotions. For example, a user can play a musical instrument while viewing background music and dynamic visuals generated according to their emotions through smart glasses.
[0587] Examples of prompt messages are as follows:
[0588] "Analyze the user's emotions from their voice and generate an appropriate music style based on those emotions. If the emotion is [joy], prepare a bright and cheerful music style. Also, generate a video style that matches the emotion."
[0589] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0590] Step 1:
[0591] The user inputs voice information using a device. The device's microphone captures the voice signal and sends the data to the server. The input here is the user's speech or musical performance, and the output is digitized voice data. The voice signal is converted into a digital format through sampling and quantization.
[0592] Step 2:
[0593] The server analyzes the received audio data. Using digital signal processing (DSP), it extracts musical characteristics such as tempo, rhythm, and key. The input is the audio data transmitted from the terminal, and the output is the musical characteristic data resulting from the analysis. Specifically, frequency analysis and spectral analysis of time-domain signals are performed.
[0594] Step 3:
[0595] The server uses an emotion engine to extract emotional information from audio data. Based on the tone and intonation of the voice, the AI model recognizes emotions such as joy, sadness, and excitement. The input is musical feature data, and the output is emotional information. A pre-trained generative AI model is used to classify specific emotional patterns from the audio waveform.
[0596] Step 4:
[0597] The server uses analyzed musical feature data and emotional information to generate music in real time using a generative AI. It dynamically constructs melodies, rhythms, and accompaniments according to the emotions. The input is musical feature data and emotional information, and the output is generated music data. The generative AI model executes algorithms to generate and select melody candidates based on emotional prompts.
[0598] Step 5:
[0599] The server generates visual content based on the user's emotional information. It dynamically adjusts colors and applies visual effects to visually enhance the emotional experience. The input is emotional information, and the output is generated visual data. Visual effects are dynamically determined by a programmed rule engine.
[0600] Step 6:
[0601] The generated music and visual content are sent to the device, where the user can experience it. In this process, the device displays the music and visual data in sync, enriching the user's experience. The input is the generated music and visual data, and the output is the presentation of content that resonates with the user's emotions.
[0602] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0603] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0604] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0605] [Fourth Embodiment]
[0606] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0607] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0608] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0609] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0610] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0611] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0612] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0613] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0614] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0615] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0616] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0617] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0618] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0619] The system according to the present invention supports improvisational performances and jam sessions by allowing users to participate through playing musical instruments or singing, and by having a generative AI generate additional music in real time. Users can participate in music sessions using their devices and enjoy an advanced musical experience through communication with the server.
[0620] First, the user launches an application on their device and inputs an audio signal to the system using an instrument or microphone. The device then transmits the input audio signal to the server in real time. After receiving the audio signal, the server analyzes it using technologies such as deep learning. Based on the analysis results, a generative AI generates an appropriate accompaniment or backing track.
[0621] The generated music is sent to the user's device, allowing them to enjoy new songs in real time, combined with their own performance or singing. Users can adjust the tempo and volume according to the generated music, enabling them to enjoy an even more creative musical experience.
[0622] Furthermore, the system also has a function to automatically generate promotional videos. This function provides users with rich video content that matches the music they create. The server selects effects and video scenes based on the characteristics of the generated music and delivers the video data to the terminal.
[0623] Finally, users can share their completed music and videos with other users on the platform. The server manages the shared content and provides a mechanism to facilitate communication through ratings and comments from other users. This makes it possible to create new spaces for collaboration and expand opportunities to participate in music production.
[0624] As a concrete example, let's say a user uses this system to sing while playing the guitar. The user launches the application on their device and begins recording their performance and singing. The server analyzes the input data, and the AI generates drum and bass accompaniment that matches the tempo of the performance. The user can continue playing along with the accompaniment, or add riffs or melodies. In addition, a promotional video matching the generated performance is created, which can be shared with other users to receive feedback and explore new musical ideas.
[0625] The following describes the processing flow.
[0626] Step 1:
[0627] The user launches an application on their device and inputs an audio signal using a musical instrument or microphone. The device digitizes the audio signal and transmits the signal data to the server in real time.
[0628] Step 2:
[0629] The server buffers the received audio signal and analyzes it using digital signal processing algorithms. This analysis extracts musical characteristics such as tempo, key, and rhythm.
[0630] Step 3:
[0631] The server invokes a generation AI based on the analysis results. The AI generates additional music, namely accompaniments and backing tracks, in real time. Deep learning technology is used for generation to create a natural flow of music.
[0632] Step 4:
[0633] The server sends the generated additional music to the terminal. The terminal plays this music in sync with the user's original performance or singing. The user can then add further performances to accompany the generated music.
[0634] Step 5:
[0635] Users can use music editing tools on their devices to adjust the tempo, volume, and harmony of the music. They can also add and edit effects to further refine the final track.
[0636] Step 6:
[0637] The server will then start a function to automatically generate a promotional video based on the completed music. It will select visual effects that match the characteristics of the music, build the video using a video template, and deliver it to the device.
[0638] Step 7:
[0639] Users share their completed music and videos on the platform via the server. The server notifies other users of the published content and facilitates communication through comment and rating functions.
[0640] (Example 1)
[0641] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0642] In recent years, creating and sharing music individually has become easier, but there remains a challenge in generating accompaniments and sharing music in real time during improvisational performances and jam sessions. In particular, there is a need for technology that allows for the easy creation and sharing of appropriate accompaniments and visual content when playing instruments or singing.
[0643] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0644] In this invention, the server includes means for acquiring a user's acoustic signal and transmitting the acoustic signal to a data processing device; means for analyzing the acquired acoustic signal using information processing technology in the data processing device; and means for generating artificial intelligence to instantly create music based on the analyzed acoustic data. This enables the user to generate appropriate accompaniment in real time during improvisational performances and to easily create and share visual content based on it.
[0645] An "acoustic signal" is an electrical signal that represents the sound waves generated by a user through playing a musical instrument or singing.
[0646] A "data processing device" is a computer system used to analyze acoustic signals and perform calculations for music generation.
[0647] "Information processing technology" refers to techniques used to analyze acoustic signals and extract features, and primarily includes methods such as deep learning.
[0648] "Generative artificial intelligence" refers to algorithms and models that automatically generate new music based on analysis results.
[0649] "Accompaniment" refers to additional sounds that add harmony and rhythm to a song played by a user, providing musical richness.
[0650] "Visual content" refers to videos, animations, and other materials synchronized with the music, provided to complement the musical experience.
[0651] "Shareable" refers to a state where music and visual content created by users can be shared with other users via the internet or other means.
[0652] This document describes embodiments for carrying out the invention. This invention is a system that analyzes a user's acoustic signal in real time and instantly generates a corresponding accompaniment using a generative AI model. Furthermore, this system automatically creates visual content based on the generated music and enables sharing with other users.
[0653] The user first launches a dedicated application on the terminal. The user inputs an audio signal using an instrument or microphone. The terminal captures this audio signal as a digital signal and transmits it to the server via the internet. The software on the terminal utilizes microphones and audio interfaces as hardware to efficiently collect the audio signal.
[0654] The server processes the received acoustic signals using a deep learning-based acoustic analysis method. This extracts the characteristics of the acoustic signals, and based on the analysis results, a generative AI model generates appropriate music. This generative AI model runs on a computer system with high computing power and utilizes large datasets to improve the accuracy of music generation.
[0655] The generated music is streamed to the user's device, allowing the user to enjoy the music experience in real time. The user can adjust the volume and tempo on their device and continue playing along with the music.
[0656] Furthermore, the server automatically generates visual content based on the generated music. During this process, effects and scenes that match the characteristics of the music are selected and provided to the user in the form of a promotional video. Users can share this visual content with other users, leading to the formation of new music communities.
[0657] As a concrete example, when a user inputs an audio signal to the system using an acoustic guitar, the server analyzes the signal and generates an accompaniment suitable for a folk style. An example of a prompt in this case might be, "Generate a folk-style accompaniment that suits this acoustic guitar performance." This allows the user to obtain an accompaniment that matches their performance in real time.
[0658] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0659] Step 1:
[0660] The user launches a dedicated application on their device. The user inputs an audio signal into the system using an instrument or microphone. Specifically, the system converts the analog audio signal into a digital signal via the microphone or audio interface, generating data that includes volume and frequency information.
[0661] Step 2:
[0662] The terminal transmits the acquired digital audio signal to the server. Specifically, it sends digitized audio data at a constant sampling rate to the server in real time via internet communication. Data compression technology may be used to improve the efficiency of the transfer.
[0663] Step 3:
[0664] The server analyzes the received acoustic signal using information processing technology. Specifically, it uses a deep learning model to extract musical features such as tempo, key, and volume from the acoustic signal. Based on the input acoustic data, it calculates feature quantities and transfers the results to the next processing step.
[0665] Step 4:
[0666] The generative AI model instantly generates accompaniment based on the analysis results. The server executes an algorithm to generate music data that matches the musical style and tempo. Based on the input features, it obtains output that generates accompaniment music using virtual instruments.
[0667] Step 5:
[0668] The server streams the generated music to the user's device. The music data is encoded using a codec and output in a format that allows for real-time playback. The user can then decode and play the music on their device.
[0669] Step 6:
[0670] The server automatically generates visual content based on the generated music. The server selects effects and scenes according to the music's tempo and mood, creating data in a promotional video format. The input for the visual content is the generated music data, and the output is the completed video data.
[0671] Step 7:
[0672] Users share their completed music and visual content with other users. The server uploads the data to the platform used for sharing, facilitating content sharing among users. This allows users to receive feedback from others and explore new possibilities in musical expression.
[0673] (Application Example 1)
[0674] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0675] Traditional music production systems lacked the flexibility to allow users to generate accompaniments in real time while performing themselves, and they also had difficulty sharing the generated music as visual content. Furthermore, providing a platform for collaborative creation among users was also challenging.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0677] In this invention, the server includes means for receiving and analyzing user voice information, means for generating additional sound in real time based on the analyzed voice information, and means for generating visual content and automatically generating video based on the generated sound. This makes it possible for users to obtain accompaniment in real time while playing, easily share the generated music and video with other users, and enjoy collaborative creation.
[0678] A "user" is an individual or group that uses a music production system to input audio information, enjoys the generated audio and visual content, and shares or co-creates it with other users.
[0679] "Audio information" refers to audio signals, such as instrument performances or vocals, that users input into the system.
[0680] "Analysis" is the process of extracting features from received audio information using digital signal processing and artificial intelligence technology to obtain the information necessary for generating accompaniment.
[0681] "Additional sound effects" refer to accompaniments and backing tracks generated in real time by a generating AI based on the voice information input by the user.
[0682] "Visual content" refers to content that provides visual expressions such as videos and effects synchronized with generated music.
[0683] An "information processing device" is a device that allows a user to input voice information and receive generated audio or visual content, and includes smartphones, tablets, and personal computers.
[0684] "Sharing" refers to the activity of providing generated music and visual content to other users via a network, and engaging in communication and collaboration.
[0685] "Collaborative creation" is a process in which multiple users exchange ideas and make revisions based on the music and visual content they have generated, in order to create a single work.
[0686] This invention is a system for improving the user's music production experience. The user can input audio information using an information processing device and generate and share music and visual content in real time.
[0687] This system mainly consists of the following components:
[0688] First, there is the user's information processing device for inputting audio information. This includes smartphones, tablets, and personal computers, and users can record musical instrument performances or vocals. For audio input, the device's microphone is used, and the audio signal is captured using a specific API (e.g., CoreAudio on iOS or AudioRecord on Android).
[0689] Next, the server receives and analyzes the audio information sent by the user. This analysis uses deep learning techniques, extracting audio features using libraries such as TensorFlow. Based on this information, a generative AI model generates additional sounds in real time. This allows the user to instantly obtain accompaniment or backing tracks that match their performance.
[0690] Furthermore, the server uses the FFmpeg library to generate visual content based on the music data. The generated video is synchronized with the music, and this is designed to allow users to easily share it with others.
[0691] As a concrete example, suppose a user records themselves singing while playing the guitar at home. The user provides the AI with a prompt such as, "Please provide an upbeat rock-style accompaniment," which then generates an accompaniment perfectly suited to their performance. The generated music, along with corresponding visual content, plays simultaneously within the app and can be shared with other users via the network. This process promotes collaborative creation and feedback exchange among users, leading to new and exciting musical experiences.
[0692] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0693] Step 1:
[0694] The user inputs voice information through their device.
[0695] The user records instrument or vocal sounds using the device's microphone, and the input audio data is acquired using an audio capture API. The input audio data is then ready to be sent to the server in real time.
[0696] Step 2:
[0697] The device sends voice information to the server.
[0698] The terminal uses the WebSocket protocol to send the input audio data to the server. The transmitted audio data is immediately received by the server.
[0699] Step 3:
[0700] The server analyzes the audio information and extracts features.
[0701] The server analyzes the received audio data using deep learning technology. Specifically, it uses a TensorFlow model to extract features from the audio signal and identifies musical elements based on those features. This provides the basic data that the generative AI uses to generate accompaniment.
[0702] Step 4:
[0703] The server generates additional sounds using a generated AI model.
[0704] Based on the analysis results, the server uses a generative AI model to create accompaniments and backing tracks in real time. The model adjusts the musical style according to the prompt "Please provide accompaniment in a light rock style" and generates appropriate sounds. This results in the generated music data as output.
[0705] Step 5:
[0706] The server generates visual content based on music data.
[0707] Based on the generated music data, the server automatically creates synchronized visual content using the FFmpeg library. Specifically, effects and scene transitions are set according to the music's tempo and rhythm, generating visually rich content. The generated visual content is output in a playable format.
[0708] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0709] The system according to the present invention is an advanced music generation platform equipped with an emotion engine that utilizes the user's voice signal and extracts emotional information from it. This makes it possible to provide a music experience synchronized with the user's emotions.
[0710] First, the user records their performance or vocals through a device and inputs the audio signal into the system. In this process, the device sends the audio data to a server. The server performs digital signal processing to analyze the received audio signal and extract musical features. This analysis identifies the tempo, key, and rhythm.
[0711] A key feature here is the emotion engine, which is implemented on the server. Based on the analysis of audio signals, the emotion engine recognizes emotions from the user's voice and the tone of their performance. This emotional information becomes a crucial element in subsequent additional music generation processes.
[0712] The generation AI combines analyzed musical characteristics with recognized emotional information to generate additional music in real time. If the user's emotions, such as joy, sadness, or excitement, are detected, the accompaniment style and musical atmosphere are automatically adjusted to match each emotion. The generated music is then sent to the device to seamlessly integrate with the user's performance.
[0713] Furthermore, the system creates a promotional video based on the generated music. The video generation function takes extracted emotional information into account and dynamically adjusts the video's colors and effects. As a result, it is possible to provide a visual experience that perfectly matches the user's emotions.
[0714] As a concrete example, suppose a user plays the guitar in an excited state. In this case, the device captures the audio signal and sends it to the server. The server uses an emotion engine to detect strong energy and excitement. The generating AI takes this emotion into account and generates an upbeat, cheerful accompaniment, which is then delivered to the user in real time. The server also generates vivid and dynamic promotional videos, enhancing the user experience. With such a system, users can enjoy creating nuanced and personalized music tailored to their individual emotional state.
[0715] The following describes the processing flow.
[0716] Step 1:
[0717] The user launches an application on their device and inputs an audio signal through a microphone or musical instrument. The device digitizes the audio signal and transmits the data to the server in real time.
[0718] Step 2:
[0719] The server analyzes the received audio signal. Digital signal processing algorithms are used for the analysis to extract musical characteristics such as tempo, key, and rhythm.
[0720] Step 3:
[0721] The server uses an emotion engine to recognize the user's emotions from the analysis of the audio signal. This process analyzes changes in tone and volume of the voice to identify emotions such as joy, sadness, and excitement.
[0722] Step 4:
[0723] The server invokes a generation AI to generate additional music in real time based on analyzed musical characteristics and recognized emotions. The generated music reflects a specific accompaniment style and musical expression that aligns with the user's emotions.
[0724] Step 5:
[0725] The server sends the generated music to the terminal. The terminal synchronizes the user's performance or singing with the generated music and plays a unified track. The user can continue playing or making adjustments while monitoring this track in real time.
[0726] Step 6:
[0727] The server activates the promotional video generation function and generates a video with visual effects adjusted based on the generated music and recognized emotions. The video is constructed so that its colors and dynamics match the user's emotions.
[0728] Step 7:
[0729] Users share their completed music and promotional videos on the platform via their devices. The server performs management functions to provide user-uploaded content to other users and facilitates communication through comments and feedback.
[0730] (Example 2)
[0731] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0732] There is a growing demand to provide a richer user experience than traditional music generation platforms can offer by making the music experience more personal and emotionally synchronized. However, current technology faces the challenge of effectively integrating user emotions with music generation and visual content creation.
[0733] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0734] In this invention, the server includes means for acquiring audio data via a communication device and transmitting the audio data to an information processing device; means for a generative model to generate music content in real time using extracted musical characteristics and emotional information as input; and means for generating visual content based on the emotional information and adjusting it by adding dynamic effects. This makes it possible to generate and provide music and visual experiences optimized for the user's emotions in real time.
[0735] "Audio data" refers to information recorded in digital format of a user's performance or voice.
[0736] A "communication device" is a device used to acquire voice data and transmit it to an information processing device.
[0737] An "information processing device" is a device that analyzes audio data received from a communication device and extracts musical characteristics and emotional information.
[0738] "Signal processing" is a computational process used to extract musical characteristics such as tempo and key from audio data.
[0739] "Musical characteristics" refer to attributes related to music, such as tempo, key, and rhythm, that are analyzed from audio data.
[0740] "Emotional information" refers to information that indicates the user's emotional state, derived from the tone and analysis results of voice data.
[0741] A "generative model" is an artificial intelligence model that generates musical content in real time based on musical characteristics and emotional information.
[0742] "Musical content" refers to the musical elements and accompaniment generated by the generative model.
[0743] "Visual content" refers to media that includes images and effects generated in conjunction with music.
[0744] "Dynamic effects" refer to the use of colors and effects in visual content that change in response to emotional information.
[0745] The system according to the present invention acquires the user's voice as digital data, analyzes it to extract emotional information, and generates music and visual content based on that information.
[0746] Users record their performances or voices as audio data through a device. This device functions as a communication device and transmits that audio data to a server, which is an information processing device. The server is equipped with advanced digital signal processing capabilities, making it possible to extract musical characteristics such as tempo, key, and rhythm from the received audio data.
[0747] The server also features an emotion engine that analyzes musical characteristics and voice tone to recognize the user's emotional information. This emotion engine identifies emotions such as joy and sadness based on specific patterns and intonations contained in the voice data.
[0748] Next, the generative AI model is activated and generates a prompt sentence that combines musical characteristics and emotional information. For example, a possible prompt sentence would be, "The user has an audio recording of themselves playing while excited. Please generate up-tempo music that is appropriate for this." Based on this prompt sentence, the generative AI model generates music that matches the user's emotions in real time and seamlessly integrates it with the user's audio.
[0749] Furthermore, the server generates visual content based on emotional information, dynamically adjusting the video's colors and effects. This process reconstructs the user's provided audio data into music and visual experiences that resonate with their emotions.
[0750] For example, if a user performs an energetic piece, the device records it and sends it to a server. The server, using an emotion engine, detects high energy and excitement in the performance, and a generative AI model generates bright and lively music to match. Visual content is also dynamically adjusted to match this atmosphere, resulting in a new musical experience for the user.
[0751] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0752] Step 1:
[0753] The user records their own performance or vocals using a terminal. The recorded audio data is sent directly from the communication device to the server, which is an information processing device. Here, the terminal converts the user's voice into a digital format and securely transmits that data to the server over the network. The input is the user's raw voice, and the output is digital audio data.
[0754] Step 2:
[0755] The server first passes the received audio data to a digital signal processing unit to extract musical features. This process uses algorithms such as FFT (Fast Fourier Transform) to analyze musical attributes such as tempo, key, and rhythm. The input is the digital audio data transmitted from the terminal, and the output is the extracted musical features.
[0756] Step 3:
[0757] The server analyzes the tone of the voice using an emotion engine based on musical characteristics and extracts the user's emotional information. The emotion engine employs algorithms that infer specific emotional states (e.g., joy, sadness, excitement, etc.) from the voice waveform and tone. The input is musical characteristics and the original voice data, and the output is the user's emotional information.
[0758] Step 4:
[0759] The server generates prompt statements for the generative AI model that integrate musical characteristics and emotional information, and uses these as input to generate musical content. These prompt statements specifically instruct the user's emotional state and the appropriate musical style for it. The generative AI model then regenerates the musical content in real time based on this information. The input is the prompt statement, and the output is the generated musical content.
[0760] Step 5:
[0761] The generated music is seamlessly integrated with the user's original audio data by the server. The integrated music is then sent back to the user and played on their device. At this stage, digital mixing technology is applied to harmonize the different sound sources. The input consists of the generated music and the original audio data, and the output is the integrated music data.
[0762] Step 6:
[0763] The server generates visual content using emotional information and optimizes it by combining it with dynamic effects. This creates visuals synchronized with the music. The visual content is presented to the user with its colors and effects adjusted according to the generated music. The input is emotional information and the generated music content, and the output is visual content.
[0764] Step 7:
[0765] Ultimately, the generated music and visual content are sent back from the server to the terminal, where the user can view it and even share it with other users. This process utilizes a distribution protocol to send and receive data, enabling interaction between users. The input consists of integrated music data and visual content, while the output is ready for distribution and sharing with end users.
[0766] (Application Example 2)
[0767] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0768] Conventional music generation systems provide music and videos without considering the user's emotional state, making it difficult to provide an experience tailored to individual emotions. As a result, users cannot enjoy a musical experience as an expression of their inner selves, and the generation of personalized, interactive content has been limited.
[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0770] In this invention, the server includes means for analyzing the user's voice information, means for generating music and visual content based on the analyzed voice information and emotional information, and means for providing the generated content to the user and dynamically adjusting the visual effects. This makes it possible to provide music and video content synchronized with the user's emotions.
[0771] "User voice information" refers to the voice signals emitted by users through speech, singing, playing musical instruments, etc., and includes elements that express their emotional state.
[0772] "Means for analysis" refers to devices and programs that analyze collected audio information using digital signal processing technology and other methods to extract its characteristics and emotional information.
[0773] "Emotional information" refers to data indicating the emotional state extracted from the user's voice information, and it is a factor that determines the style of music and video generated from it.
[0774] "Additional music" refers to music that is dynamically generated based on the user's emotional information and seamlessly integrated into the user's expressive activities.
[0775] "Visual content" refers to video expressions that are generated based on the user's emotional information and whose colors and effects are dynamically adjusted, adding a visual aspect to the music experience.
[0776] "Dynamic adjustment" refers to the process of changing the style of music and video generated by the system in real time in response to changes in the user's emotions and voice information.
[0777] "Generated AI" refers to algorithms and programs that utilize voice and emotional information to automatically generate music and visual content tailored to the user.
[0778] "Means of sharing" refers to providing functions that enable users to share generated music and visual content with other users via a communication network, allowing them to view and play the content.
[0779] Modes for carrying out the invention
[0780] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. In this system, the user's voice information is acquired by the terminal and transmitted to the server. The terminal can be a communication device such as a smartphone or smart glasses. Digital signal processing technology is applied to analyze the voice information, extracting musical characteristics such as the tempo and rhythm of the voice. Furthermore, the server is equipped with an emotion engine that extracts the user's emotional information from the voice information.
[0781] The server generates music in real time using generated AI based on analyzed audio and emotional information. This generating AI operates with an algorithm that inputs the analyzed data into a learning model and generates melodies and accompaniments that correspond to emotions. Visual content tailored to the user's emotions is also generated, with colors and video effects dynamically adjusted. Machine learning frameworks such as TensorFlow and PyTorch are used to execute the generated AI.
[0782] This system sends generated music and visual content to a device, allowing users to enjoy a personalized music experience tailored to their emotions. For example, a user can play a musical instrument while viewing background music and dynamic visuals generated according to their emotions through smart glasses.
[0783] Examples of prompt messages are as follows:
[0784] "Analyze the user's emotions from their voice and generate an appropriate music style based on those emotions. If the emotion is [joy], prepare a bright and cheerful music style. Also, generate a video style that matches the emotion."
[0785] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0786] Step 1:
[0787] The user inputs voice information using a device. The device's microphone captures the voice signal and sends the data to the server. The input here is the user's speech or musical performance, and the output is digitized voice data. The voice signal is converted into a digital format through sampling and quantization.
[0788] Step 2:
[0789] The server analyzes the received audio data. Using digital signal processing (DSP), it extracts musical characteristics such as tempo, rhythm, and key. The input is the audio data transmitted from the terminal, and the output is the musical characteristic data resulting from the analysis. Specifically, frequency analysis and spectral analysis of time-domain signals are performed.
[0790] Step 3:
[0791] The server uses an emotion engine to extract emotional information from audio data. Based on the tone and intonation of the voice, the AI model recognizes emotions such as joy, sadness, and excitement. The input is musical feature data, and the output is emotional information. A pre-trained generative AI model is used to classify specific emotional patterns from the audio waveform.
[0792] Step 4:
[0793] The server uses analyzed musical feature data and emotional information to generate music in real time using a generative AI. It dynamically constructs melodies, rhythms, and accompaniments according to the emotions. The input is musical feature data and emotional information, and the output is generated music data. The generative AI model executes algorithms to generate and select melody candidates based on emotional prompts.
[0794] Step 5:
[0795] The server generates visual content based on the user's emotional information. It dynamically adjusts colors and applies visual effects to visually enhance the emotional experience. The input is emotional information, and the output is generated visual data. Visual effects are dynamically determined by a programmed rule engine.
[0796] Step 6:
[0797] The generated music and visual content are sent to the device, where the user can experience it. In this process, the device displays the music and visual data in sync, enriching the user's experience. The input is the generated music and visual data, and the output is the presentation of content that resonates with the user's emotions.
[0798] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0799] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0800] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0801] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0802] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0803] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0804] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0805] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0806] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0807] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0808] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0809] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0810] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0811] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0812] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0813] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0814] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0815] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0816] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0817] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0818] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0819] The following is further disclosed regarding the embodiments described above.
[0820] (Claim 1)
[0821] [Means for receiving and analyzing the user's voice signal,
[0822] [Means for generating additional music in real time based on the analyzed audio signal,
[0823] [Means for providing the generated music to the user,
[0824] [Means for generating a promotional video,
[0825] [Means that allow users to share music and videos they have created with other users,
[0826] A system that includes this.
[0827] (Claim 2)
[0828] [The system according to claim 1, in which a user inputs an audio signal via a terminal.]
[0829] (Claim 3)
[0830] [The system according to claim 1, wherein the generating AI generates accompaniment in real time based on the analyzed audio signal.
[0831] "Example 1"
[0832] (Claim 1)
[0833] [Means for acquiring a user's acoustic signal and transmitting the acoustic signal to a data processing device,
[0834] [In a data processing device, means for analyzing acquired acoustic signals using information processing technology,
[0835] [Means for generating artificial intelligence to instantly create music based on analyzed acoustic data,
[0836] [Means for outputting the created music to the user,
[0837] [Means for automatically creating visual content according to the characteristics of music,
[0838] [Means for enabling users to share music and visual content they have created with other users,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, in which a user transmits an acoustic signal via an information terminal.
[0842] (Claim 3)
[0843] [The system according to claim 1, wherein the generating artificial intelligence instantly creates an accompaniment based on the analyzed acoustic data.
[0844] "Application Example 1"
[0845] (Claim 1)
[0846] [Means for receiving user voice information and analyzing said voice information,
[0847] [Means for generating additional sounds in real time based on analyzed audio information,
[0848] [Means for providing the generated sound to the user,
[0849] [Means for generating visual content, and means for automatically generating video based on the generated sound,
[0850] [Means for users to share generated audio and visual content with other users,
[0851] [A means for users to input their own sound creations and collaborate with other users on production,
[0852] A system that includes this.
[0853] (Claim 2)
[0854] The system according to claim 1, in which a user inputs voice information via an information processing device.
[0855] (Claim 3)
[0856] [The system according to claim 1, wherein the generating artificial intelligence generates accompaniment in real time based on the analyzed audio information.
[0857] "Example 2 of combining an emotion engine"
[0858] (Claim 1)
[0859] [Means for acquiring audio data using a communication device and transmitting that audio data to an information processing device,
[0860] [In an information processing device, means for processing audio data to extract musical features,
[0861] [Means for extracting emotional information based on musical characteristics and vocal tone,
[0862] [Means for a generative model to generate musical content in real time using extracted musical features and emotional information as input,
[0863] [Means for integrating the generated music content with the user's voice and distributing it,
[0864] [Means for generating visual content based on emotional information and adjusting it by adding dynamic effects,
[0865] [Means for sharing the generated music content and visual content with other users,
[0866] A system that includes this.
[0867] (Claim 2)
[0868] [The system according to claim 1, wherein the user transmits voice data via an input device.]
[0869] (Claim 3)
[0870] [In an information processing device, the system according to claim 1, wherein a generation model instantly generates accompaniment based on signal-processed audio data.
[0871] "Application example 2 when combining with an emotional engine"
[0872] (Claim 1)
[0873] [Means for receiving user voice information and analyzing said voice information,
[0874] [Means for generating additional music in real time based on analyzed audio information and emotional information,
[0875] [Means for providing generated music and visual content to the user and for dynamically adjusting various visual effects,
[0876] [Means for generating visual content based on user sentiment information,
[0877] [Means that allow users to share generated music and visual content with other users,
[0878] A system that includes this.
[0879] (Claim 2)
[0880] [A system according to claim 1, wherein a user inputs voice information via a communication device and emotional information is extracted.]
[0881] (Claim 3)
[0882] [The system according to claim 1, wherein the generated AI generates accompaniment and visual content in real time based on analyzed voice information and emotion information. [Explanation of Symbols]
[0883] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving a user's voice signal and analyzing said voice signal, A means for generating additional music in real time based on the analyzed audio signal, A means of providing the generated music to the user, Means for generating a promotional video, A means for users to share music and videos they have created with other users, A system that includes this.
2. The system according to claim 1, in which a user inputs an audio signal via a terminal.
3. The system according to claim 1, wherein the generating AI generates accompaniment in real time based on the analyzed audio signal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A