System
The system addresses band member absences by using a terminal, server, and AI model to generate consistent performances, ensuring smooth band operations and maintaining musical style.
Patent Information
- Application Number
- JP2024117349
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
Band activities face challenges due to sudden member absences or difficulties in recruiting new members, leading to disruptions and inconsistencies in performance, as it is time-consuming to find replacements and difficult to replicate the playing style of professional musicians.
A system comprising a terminal for uploading performance audio, a server for storing and processing audio files, and an AI model trained to generate pinch performances based on extracted features, allowing seamless integration of AI as a substitute member.
Enables continuous band activities by providing a harmonious performance even with sudden vacancies, maintaining musical consistency and style through AI-generated substitute performances.
Smart Images

Figure 2026016259000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Create a document that describes the "problem the invention aims to solve" and the "means for solving the problem."
[0005] Band practices and live performances can be difficult due to sudden member absences or difficulties in recruiting new members. In such situations, it takes time and effort to find replacement members, making it difficult to continue band activities. Even if replacement members are quickly found, they often do not harmonize with the band's direction or playing style. Another issue is the difficulty of recreating the performance characteristics of professional musicians. [Means for solving the problem]
[0006] The present invention solves the above-mentioned problems with a system that includes a terminal for uploading performance audio files, a server for storing the performance audio files received from the terminal, processing means for extracting features from the performance audio files and creating a dataset, processing means for training an AI model using the dataset, and a terminal for generating and playing back a pinch performance using the trained AI model. This allows the AI to play in place of a member even in the event of a sudden vacancy, allowing band activities to proceed smoothly. Furthermore, because the AI learns the playing style of each member based on past performance audio files, it can function as a substitute member with a harmonious performance.
[0007] "Performance sound source" refers to data that records the sounds of instruments, vocals, etc.
[0008] A "terminal" is a device operated by a user, and refers to electronic devices such as smartphones, tablets, and computers.
[0009] A "server" is a computer system for storing, processing, and transmitting data over a network.
[0010] "Preprocessing" refers to a series of processes, including noise removal and data shaping, that are carried out before data analysis.
[0011] A "feature" is numerical information that represents useful patterns or attributes extracted from data.
[0012] A "dataset" is a collection of data used to train an AI model.
[0013] An "AI model" is a computational model of artificial intelligence that has been trained using machine learning algorithms.
[0014] "Training" refers to the process by which an AI model learns patterns and rules from data.
[0015] "Substitute performance" refers to performance data generated by AI to replace a missing band member. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention relates to a system in which an AI plays as a substitute when a band suddenly becomes vacant. The following describes an embodiment of this system.
[0038] System Configuration
[0039] This system consists of the following main components:
[0040] 1. A device for uploading performance audio
[0041] 2. A server for storing the performance sound source received from the terminal.
[0042] 3. Processing means for extracting features from the performance sound source and creating a data set
[0043] 4. Processing means for training an AI model using said dataset.
[0044] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[0045] System Operation
[0046] The program of this system operates in the following procedure.
[0047] 1. Uploading the audio
[0048] Users use a smartphone app to upload recordings of their past performances, which provides the system with the audio for each part of the song they want to use.
[0049] 2. Receiving and saving audio files
[0050] The server receives the audio files uploaded by the user and stores them in a database. The stored audio files are then processed as required for each part.
[0051] 3. Sound source preprocessing and feature extraction
[0052] The server performs preprocessing such as noise removal and clipping on the received audio files, then extracts features such as rhythm, tempo, and timbre to create a data set for each part.
[0053] 4. Training the AI model
[0054] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[0055] 5. Deploying the AI model
[0056] The server deploys the trained AI model as an executable endpoint, making it accessible from a smartphone app.
[0057] 6. Generation and playback of pinch-hitting performances
[0058] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch performance in real time in response to the request and sends it back to the device. The device then plays the received performance data and outputs it as sound through an audio output device such as an amplifier.
[0059] Specific examples
[0060] Example 1: Guitarist sub-hiring scenario during band practice
[0061] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0062] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0063] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0064] Example 2: Reproducing the style of a professional musician
[0065] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0066] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0067] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0068] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs, and also maintains musical consistency of substitute members according to the performance style and direction of the band.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0072] Step 2:
[0073] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0074] Step 3:
[0075] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[0076] Step 4:
[0077] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[0078] Step 5:
[0079] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[0080] Step 6:
[0081] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0082] Step 7:
[0083] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[0084] Step 8:
[0085] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[0086] Step 9:
[0087] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[0088] Step 10:
[0089] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[0090] Step 11:
[0091] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[0092] Step 12:
[0093] Users can provide feedback on the playback of their performance, and if necessary, make fine adjustments within the app to achieve even more precise performance.
[0094] Through these steps, the system can use AI to quickly and effectively provide substitute performances even when a sudden vacancy occurs in a band activity.
[0095] Example 1
[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0097] When a band member suddenly becomes unavailable, it is difficult to find a substitute musician immediately, resulting in interruptions to the performance. Effective methods are needed to solve this problem and maintain the consistency and quality of the performance. It is also difficult to reproduce the performance style of a specific musician.
[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0099] In this invention, the server includes a means for storing performance audio, a means for extracting features from the performance audio to create a dataset, a means for training an AI model using the dataset, and a means for generating a pinch performance using the trained AI model and transmitting it to a terminal. This allows the AI to generate a pinch performance in real time based on audio uploaded by a user and play it on the terminal. In particular, the server performs noise removal and extracts features, allowing the AI model to be trained using high-quality data, allowing the band to continue performing effectively even when a member is absent.
[0100] A "performance sound source" is a recording of music performed by a user, and is audio data stored in digital format.
[0101] A "terminal" is a device used by a user, including a smartphone, tablet, or PC, which is used to upload and play back performance audio.
[0102] A "server" is a computer system that processes information via a network and manages the storage of performance audio, feature extraction, training of AI models, and generation of pinch performances.
[0103] "Features" are quantitative attributes such as rhythm, tempo, and timbre extracted from audio data, and are used to train AI models.
[0104] A "dataset" is a collection of data that compiles features extracted from performance audio, and serves as basic information for training an AI model.
[0105] An "AI model" is an algorithmic system built using artificial intelligence technology, and is a program that learns specific performances and generates substitute performances.
[0106] "Training" is the process of building an AI model, where the AI learns specific patterns and styles using a dataset.
[0107] A "substitute performance" is a replacement performance for an absent or missing band member, generated by a trained AI model.
[0108] "Noise reduction" is a process that removes unnecessary background sounds and unpleasant acoustics from sound source data, and is a preprocessing step to improve data quality.
[0109] An "audio output device" is a device that is connected to a terminal and outputs the generated pinch-hit performance as sound, and includes a speaker, an amplifier, and the like.
[0110] The present invention relates to a system that uses AI to fill in for a band member when a sudden vacancy occurs during a performance. The configuration and operation of this system are described in detail below.
[0111] System configuration
[0112] This system consists of the following main components:
[0113] 1. A device for uploading performance audio
[0114] 2. Server for storing performance audio
[0115] 3. Processing method for extracting features from performance audio and creating a dataset
[0116] 4. A processing method for training an AI model using a dataset
[0117] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[0118] 6. Audio Output Devices
[0119] Hardware and software used
[0120] Devices: Smartphones, tablets, computers, etc.
[0121] Server: A computer system that processes information over a network (e.g., AWS EC2, Google Cloud)
[0122] Software: Python (audio source preprocessing and feature extraction), TensorFlow or PyTorch (AI model training), LibROSA (music signal processing), AWS S3 (audio source storage)
[0123] Audio output devices: speakers, amplifiers, etc.
[0124] System Operation
[0125] This system operates in the following procedure.
[0126] 1. Uploading the audio
[0127] Users use a smartphone app to upload recordings of past performances, such as recordings of band members performing or rehearsing, which provides the AI with data to learn from.
[0128] 2. Receiving and saving audio files
[0129] The server receives the audio files uploaded by the user and stores them using cloud storage such as AWS S3. Meta information (such as song title and part name) is added to the stored audio data.
[0130] 3. Sound source preprocessing and feature extraction
[0131] The server uses a music signal processing library such as LibROSA to perform preprocessing such as noise removal and clipping, then extracts features such as rhythm, tempo, and timbre to create a dataset corresponding to each part.
[0132] 4. Training the AI model
[0133] The server uses the preprocessed dataset to train an AI model in TensorFlow or PyTorch, which learns specific playing styles and phrasing patterns for each musical part.
[0134] 5. Deploying the AI model
[0135] The server deploys the trained AI model as an executable endpoint using AWS SageMaker or Google AI Platform, making it accessible from a smartphone app.
[0136] 6. Generation and playback of pinch-hitting performances
[0137] The device sends a request to the AI model endpoint based on the song and part specified by the user. The server generates a substitute performance in response to the request and returns it to the device in real time. The device then plays back the received performance data and outputs it as sound through an audio output device.
[0138] Examples of specific examples and prompts
[0139] Example 1: Guitarist sub-hiring scenario during band practice
[0140] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0141] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0142] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0143] Example 2: Reproducing the style of a professional musician
[0144] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0145] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0146] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0147] Prompt Sentence Examples
[0148] "Generate guitar parts based on this musician's playing style."
[0149] "Create a new performance based on the audio you uploaded."
[0150] This system allows the band to accommodate sudden absences of members and continue activities while maintaining consistency in performance.
[0151] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0152] Step 1: Upload your audio
[0153] Users use a smartphone app to upload audio recordings of past performances. Specifically, the user taps the "Upload Audio" button on the smartphone app and selects an audio file using a file browser. The selected file is also given meta information such as the song title and part name. The input is the audio file and meta information selected by the user, and the output is the audio file and meta information sent to the server.
[0154] Step 2: Receiving and saving the audio
[0155] The server receives audio files uploaded by users and stores them in a database. Specifically, the server receives an HTTP request and stores the audio files in cloud storage such as AWS S3. The input is the audio file and meta information sent by the user, and the output is the file path of the saved audio file and the meta information registered in the database.
[0156] Step 3: Preprocessing and feature extraction of the audio source
[0157] The server performs preprocessing such as noise removal and clipping on the received audio file. It then uses a music signal processing library such as LibROSA to extract features such as rhythm, tempo, and timbre and create a dataset. Specifically, the server loads the audio file and applies a noise removal algorithm (e.g., spectral subtraction). Next, it extracts features such as rhythm, tempo, and timbre using LibROSA. The input is the received audio file, and the output is the extracted features and a dataset.
[0158] Step 4: Training the AI model
[0159] The server trains the AI model using the preprocessed dataset. It uses TensorFlow or PyTorch to train a recurrent neural network (RNN) or a transformer model. Specifically, the server builds the RNN or transformer model, supplies it with the preprocessed dataset, and performs training. It monitors the training progress and accuracy. The input is the preprocessed dataset, and the output is the trained AI model.
[0160] Step 5: Deploying the AI model
[0161] The server deploys the trained AI model as an executable endpoint. The endpoint is built using AWS SageMaker or Google AI Platform and made accessible from a smartphone app. Specifically, the server saves the trained AI model and deploys it as an endpoint using AWS SageMaker. The input is the trained AI model, and the output is the endpoint URL.
[0162] Step 6: Generate and play back a pinch performance
[0163] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch-hit performance in real time in response to the request and sends it back to the device. The device plays back the received performance data and outputs it as sound through an audio output device. Specifically, the user taps the "Generate Pinch-hit Performance" button in the smartphone app and selects the song and part. The device then sends a request to the AI model endpoint based on the selected information, the server sends the generated pinch-hit performance data to the device, and the device plays back the received performance data. The input is the song and part specified by the user, and the output is the generated pinch-hit performance data.
[0164] In this way, each processing step works in conjunction with each other to create a system that can also handle sudden vacancies in the band.
[0165] (Application example 1)
[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0167] Currently, musical performances at music events, theme parks, and other venues rely heavily on human resources. This means that sudden staff shortages or unexpected problems can disrupt performances. Furthermore, reproducing the style of a professional performer requires advanced technology, which is difficult to replace immediately. Therefore, there is a need to solve these issues by introducing an automated performance system using robots.
[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0169] In this invention, the server includes a terminal for uploading performance sound sources, a server for storing the performance sound sources received from the terminal, processing means for extracting features from the performance sound sources and creating a data set, processing means for training an AI model using the data set, a terminal for generating and playing a pinch-hit performance using the trained AI model, and control means for giving instructions to an automatic musical performance robot used at music events and theme parks. This allows the robot to automatically fill in the gaps in performance even when a sudden vacancy occurs.
[0170] "Performance sound source" refers to audio data that records musical performances.
[0171] An "uploading device" is a device through which a user sends data over the Internet.
[0172] A "server" is a computer system that stores and processes received data.
[0173] "Feature extraction" means extracting important information such as rhythm, tempo, dynamics, and timbre from audio data.
[0174] "Creating a dataset" means organizing and aggregating the extracted features and constructing a set of data to be used for training an AI model.
[0175] "Training an AI model" means using machine learning algorithms to learn patterns from a dataset.
[0176] A "terminal for generating and playing pinch performances" is a device for playing music generated by AI.
[0177] An "automatic playing robot" is a robot designed to play music automatically.
[0178] "Control means" refers to a system or mechanism that gives specific instructions to the automatic musical robot.
[0179] "Playback" means outputting stored sound sources or generated music data as sound.
[0180] This invention is a system for robots to automatically perform music at music events and theme parks. The system consists of the following main components:
[0181] 1. A device for uploading performance audio
[0182] Users upload recordings of past performances to the system using smartphones or computers, which have the ability to transmit audio data to a server via the Internet.
[0183] 2. Server for storing performance audio
[0184] The server receives the audio files uploaded by the user and stores them in a database, which contains the audio data for each instrument part.
[0185] 3. Processing method for extracting features and creating a dataset
[0186] The server performs preprocessing such as noise reduction and clipping on the received audio files, then extracts features such as rhythm, tempo, dynamics, and timbre to create a data set for each part.
[0187] 4. Processing means for training AI models
[0188] The server uses the preprocessed dataset to train the AI model, which is done using machine learning libraries such as TensorFlow, so that the AI model learns the patterns and styles of musical performance.
[0189] 5. Device for generating and playing pinch-hitting performances
[0190] The trained AI model generates a pinch-hit performance based on the song and parts specified by the user. The generated performance data is sent to a terminal that provides instructions to the robot. The terminal then gives the robot specific instructions to play and plays the audio using speakers and amplifiers.
[0191] 6. Control means for giving instructions to automatic musical robots used in music events and theme parks
[0192] The control means connects the server and the robot, and is designed to allow the robot to receive accurate performance instructions. This control means sends the generated performance data to the robot in real time, ensuring smooth performance.
[0193] Specific examples
[0194] Music events at theme parks
[0195] Imagine an evening music event at a theme park where the drumming robot is suddenly absent. The user uploads past drum part audio using a dedicated smartphone app. The server receives this audio, extracts features, and creates a dataset. After the AI model is trained, a new drum part is generated and sent to the robot. The robot then automatically plays the drums based on the generated performance data.
[0196] Prompt Sentence Examples
[0197] The robot drummer for a nighttime music event at a theme park suddenly misses an opportunity, so we need to have an AI generate a drum part using past audio data. Please generate prompts to reproduce the drumming patterns and rhythms in real time.
[0198] Required files: Previous drum part sound files
[0199] Goal: Jazz-style drumming at a tempo of 120 BPM
[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0201] Step 1:
[0202] Users upload audio files of past performances using smartphones or computers. The audio files uploaded by users are sent to a server via the Internet. The input is the audio files of past performances, and the output is the audio data stored on the server.
[0203] Step 2:
[0204] The server stores the audio files received from the user in a database. This stored audio data is used in subsequent processing. The input is the audio file sent from the terminal, and the output is the audio data stored in the database.
[0205] Step 3:
[0206] The server performs preprocessing such as noise removal and clipping on the received audio files. It then extracts features such as rhythm, tempo, dynamics, and timbre from the audio data. This creates a dataset. The input is the saved audio data, and the output is the preprocessed feature data.
[0207] Step 4:
[0208] The server uses the preprocessed feature data to train an AI model. In this step, it runs machine learning algorithms using libraries such as TensorFlow to learn musical patterns and styles. The input is the feature dataset, and the output is a trained AI model.
[0209] Step 5:
[0210] The server uses the trained AI model to generate a pinch-hit performance. The AI model generates a pinch-hit performance based on the song and parts specified by the user. This generated performance data is sent to the device for playback. The input is the song and parts specified by the user and the AI model, and the output is the generated pinch-hit performance data.
[0211] Step 6:
[0212] The terminal transmits the received pinch-hit performance data to the automatic musical robot. Specific performance instructions are given to the robot via the control means, and the robot plays music. The input is the pinch-hit performance data, and the output is the actual performance by the robot.
[0213] Step 7:
[0214] The control means sends the generated performance data to the robot in real time, supporting accurate performance. The input is the performance data, and the output is the robot's smooth playing movements.
[0215] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0216] The present invention relates to a system in which AI performs as a substitute to solve problems that arise when a band suddenly becomes empty or when it is difficult to recruit new members. In addition, by combining it with an emotion engine that recognizes the user's emotions, a more human-like performance can be achieved. The following describes in detail the mode for implementing this system.
[0217] System Configuration
[0218] This system consists of the following main components:
[0219] 1. A device for uploading performance audio
[0220] 2. A server for storing the performance sound source received from the terminal.
[0221] 3. Processing means for extracting features from the performance sound source and creating a data set
[0222] 4. Processing means for training an AI model using said dataset.
[0223] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[0224] 6. Emotion engine that recognizes user emotions
[0225] System Operation
[0226] The program of this system operates in the following procedure.
[0227] 1. Uploading the audio
[0228] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0229] 2. Receiving and saving audio files
[0230] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0231] 3. Sound source preprocessing and feature extraction
[0232] The server performs preprocessing on the received audio files. This includes noise removal, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied. After that, feature quantities such as rhythm, tempo, and timbre are extracted, and a data set corresponding to each part is created.
[0233] 4. Training the AI model
[0234] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[0235] 5. Deploying the AI model
[0236] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0237] 6. Operation of the Emotion Engine
[0238] The device uses an emotion engine to recognize the user's emotions in real time, which analyzes the user's facial expressions and tone of voice to detect their emotional state.
[0239] 7. Generating emotional performance
[0240] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the emotional data, and the AI model adjusts its playing style based on that data.
[0241] 8. Generating and Playing Pinch-Hit Performances
[0242] The device receives and plays the generated pinch-hitting performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play together with other band members.
[0243] Specific examples
[0244] Example 1: Guitarist sub-hiring scenario during band practice
[0245] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0246] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0247] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0248] The device also detects the user's emotions and reflects them in the performance, creating a more realistic performance.
[0249] Example 2: Recreating the style of a professional musician in a stage performance
[0250] If a user wants to reproduce the performance characteristics of a famous professional musician on a live stage, they can upload the musician's performance recording to the app.
[0251] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0252] The device uses a trained AI model to generate and play back a pinch-hit performance in the style of a professional musician, and an emotion engine detects the reactions of the user and audience and reflects them in the performance.
[0253] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs. In addition, by using an emotion engine, it is possible to provide a more human-like performance that responds to the user's emotions.
[0254] The processing flow will be explained below.
[0255] Step 1:
[0256] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0257] Step 2:
[0258] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0259] Step 3:
[0260] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[0261] Step 4:
[0262] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[0263] Step 5:
[0264] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[0265] Step 6:
[0266] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0267] Step 7:
[0268] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[0269] Step 8:
[0270] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[0271] Step 9:
[0272] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[0273] Step 10:
[0274] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[0275] Step 11:
[0276] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[0277] Step 12:
[0278] The device uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's facial expressions and tone of voice to detect emotion data.
[0279] Step 13:
[0280] The device resends the request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the user's emotional data, and the AI model adjusts its playing style based on that data.
[0281] Step 14:
[0282] The server generates new performance data that reflects the emotional data and returns it to the terminal. This data is audio data that includes a performance style that reflects the user's emotional state.
[0283] Step 15:
[0284] The terminal then plays back the received emotion-reflecting performance data, resulting in a more human-like performance that reflects the user's emotions being output through an amplifier.
[0285] Through these steps, this system can quickly and effectively provide substitute performances using AI even when a band member suddenly becomes unavailable. It also uses an emotion engine to generate performances that reflect the user's emotional state, providing a realistic and immersive band experience.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] There is a need to solve the problem of bands suddenly becoming empty or finding new members is difficult, and to provide more human-like performances that reflect the user's emotions. However, conventional technologies have difficulty in processing performance sound sources in real time and in realizing performances that reflect the user's emotions. Furthermore, there is a lack of a function to dynamically generate performances that respond to emotions while faithfully reproducing the performance style. It is necessary to solve these issues.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0290] In this invention, the server includes a means for storing performance sound sources, a means for extracting features such as rhythm, tempo, and timbre from the performance sound sources to create a dataset, and a means for training an AI model using deep learning technology. This enables highly accurate pre-processing and feature extraction of performance sound sources, generation and playback of pinch hit performances in real time, and adjustment of performance style to reflect the user's emotions.
[0291] "Performance sound source" refers to performance data of instruments or vocals that have been recorded in the past.
[0292] "Terminal" refers to an electronic device such as a smartphone or computer used by a user.
[0293] "Server" refers to a networked hardware or software system for storing performance audio, processing, and training AI models.
[0294] "Features" refer to individual characteristic data such as rhythm, tempo, and timbre extracted from a performance sound source.
[0295] A "dataset" refers to a collection of data composed of feature quantities.
[0296] "Deep learning technology" is a method of training AI models using large datasets.
[0297] An "AI model" refers to an artificial intelligence system that has been trained to perform a specific task.
[0298] "Pinch performance" refers to performance data generated by an AI model to fill in for a missing player.
[0299] An "emotion engine" refers to software or hardware that recognizes a user's emotions in real time and generates that data.
[0300] "Processing means" refers to the technical mechanisms that receive, store, preprocess, extract features from performance audio, create datasets, train AI models, and analyze emotional data.
[0301] "Noise reduction" refers to the process of removing unwanted noise from a recorded performance sound source.
[0302] "Normalization" refers to an adjustment process for unifying the volume levels of the performance sound sources.
[0303] "Clipping" is a process of cutting off the peaks of an audio signal to adjust the volume.
[0304] "High-precision filtering processing" refers to a sound source processing procedure that uses advanced algorithms to maintain sound quality.
[0305] "Real time" means that processing occurs immediately without delay.
[0306] "Performance style" refers to a collection of specific rhythms, tempos, tones, phrases, etc. when performing.
[0307] "REST API endpoint" refers to an interface for accessing the functionality of an AI model over a network.
[0308] This system is designed to solve the problem of bands suddenly becoming vacant or finding new members is difficult, making it difficult to perform. In this system, AI substitutes and plays music, and by reflecting the user's emotions, it achieves a more human-like performance.
[0309] System configuration
[0310] This system consists of the following main components:
[0311] 1. A device for uploading performance audio
[0312] 2. A server for storing the performance sound source received from the terminal.
[0313] 3. A processing method for extracting features such as rhythm, tempo, and timbre from performance audio and creating a data set
[0314] 4. Processing means for training AI models using deep learning techniques
[0315] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[0316] 6. Emotion engine that recognizes user emotions in real time
[0317] 7. AI model integration to adjust playing style based on data from the emotion engine
[0318] System Operation Overview
[0319] The specific operation of this system is outlined below.
[0320] Uploading and analyzing audio files
[0321] Users use a smartphone app to upload recordings of their past performances. The uploaded audio files are divided into guitar parts, bass parts, drum parts, etc. Each audio file is sent to a server for analysis.
[0322] Saving and preprocessing audio sources
[0323] The server receives the uploaded audio files and stores them in a database. The received files are first preprocessed using noise removal, normalization, clipping, and other preprocessing techniques. This is done using a high-precision filtering algorithm (e.g., Butterworth filter). After preprocessing is complete, features such as rhythm, tempo, and timbre are extracted from the audio files, and a dataset is created.
[0324] Training an AI model
[0325] The server uses deep learning technologies (e.g., TensorFlow, PyTorch) to train an AI model using the dataset after preprocessing and feature extraction. This AI model learns the playing style and phrase patterns of each part, and once training is complete, it is deployed as a web service.
[0326] Emotion Engine Operation
[0327] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and uses emotion recognition algorithms to detect the user's emotional state, which is then sent to an AI model on the server to reflect a specific playing style.
[0328] Pinch-hitting performance generation and playback
[0329] The device sends a request from the server to the AI model based on the emotional data recognized by the emotion engine. The AI model takes the emotional data into account, adjusts the playing style, and generates appropriate performance data for the substitute. This performance data is played back on the device and connected to an amplifier via the audio output terminal, allowing the user to play along with other band members.
[0330] Specific examples
[0331] A specific example of the system is shown below.
[0332] Example 1: Guitarist sub-hiring scenario during band practice
[0333] 1. The user uploads a recording of their guitar part using a smartphone app.
[0334] 2. The server receives the audio files and stores and preprocesses them.
[0335] 3. The server trains the AI model and sets up the endpoint.
[0336] 4. The device detects the user's emotions using its emotion engine and sends the "joy" emotion data to the AI model.
[0337] 5. The terminal plays the received pinch-hit performance data and outputs the sound through the amplifier.
[0338] Example prompt sentence:
[0339] "Please upload a replacement guitar part for the absent guitarist. Please generate a performance that reflects the emotion of joy."
[0340] In this way, the system can quickly respond to sudden vacancies and provide human-like performances that reflect the user's emotions.
[0341] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0342] Step 1: Upload your audio
[0343] Users upload audio recordings of past performances using a smartphone app. Specifically, they open the app's "Audio Upload" function and select audio files from their PC or smartphone. Audio files are uploaded categorized into guitar parts, bass parts, drum parts, etc.
[0344] Input: Audio files on your PC or smartphone
[0345] Output: Uploaded audio file (sent to server)
[0346] Step 2: Receiving and saving the audio
[0347] The server receives audio files uploaded by users. The received files are stored in a database, and each part is managed as individual data. Specifically, the server verifies the file format (e.g., WAV, MP3), converts it to the appropriate data format, and saves it.
[0348] Input: Uploaded audio file
[0349] Output: Saved audio data (database)
[0350] Step 3: Preprocessing and feature extraction of the audio source
[0351] The server performs preprocessing on the received audio files, such as noise removal, normalization, and clipping. In particular, it applies a high-precision filtering algorithm (e.g., Butterworth filter) to remove noise. It then extracts features such as rhythm, tempo, and timbre, creating a dataset.
[0352] Input: Saved audio data
[0353] Output: A preprocessed dataset with features
[0354] Step 4: Training the AI model
[0355] The server uses the preprocessed dataset to train the AI model using deep learning techniques (e.g., TensorFlow, PyTorch). This involves learning the playing style and phrase patterns of each part. Once trained, the AI model is deployed as a web service.
[0356] Input: Preprocessed dataset
[0357] Output: A trained AI model
[0358] Step 5: Deploying the AI model
[0359] The server deploys the trained AI model as a REST API endpoint, which allows the AI model to be accessed from a smartphone app.
[0360] Input: A trained AI model
[0361] Output: REST API endpoint
[0362] Step 6: Emotion Engine in Action
[0363] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and detects their emotional state using an emotion recognition algorithm (e.g., Facial Action Coding System (FACS)). For example, if the user is smiling, emotion data of "joy" is generated.
[0364] Input: Real-time video and audio data of the user
[0365] Output: Emotion data
[0366] Step 7: Generating emotional performance
[0367] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes emotional data such as "joy," and the AI model adjusts its playing style based on this. The generated pinch-hit performance data is then sent to the device.
[0368] Input: Emotion data
[0369] Output: Pinch hit performance data
[0370] Step 8: Generate and play back a pinch performance
[0371] The device plays the performance data received from the server. Specifically, it outputs the sound using the built-in audio player and connects to an amplifier via the audio output terminal. This allows the user to play along with other band members.
[0372] Input: Pinch-hitting performance data sent from the AI model
[0373] Output: Playback audio
[0374] (Application example 2)
[0375] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0376] The problem that this invention aims to solve is the problem of bands being unable to perform due to sudden vacancies or the absence of a member. Another problem is the difficulty of realizing a more human-like performance that reflects the emotions of the audience and users. In particular, in physical venues that offer live music, it is difficult to quickly find a replacement member when a sudden vacancy occurs, making it difficult to maintain the quality of the performance.
[0377] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for uploading performance sound sources, a server for storing performance sound sources received from the terminal, means for extracting features from the performance sound sources and creating a dataset, means for training a generative AI model using the dataset, a terminal for generating and playing a pinch-hit performance using the trained generative AI model, an emotion engine that recognizes the emotions of the audience and users, and means for adjusting the generative AI model based on the emotion data recognized by the emotion engine. This makes it possible to smoothly carry out band activities and live music even when there is a sudden vacancy, and to provide a realistic performance that corresponds to the emotions of the audience and users.
[0378] A "performance sound source" is audio data generated by instruments, human voices, etc., and is a digital file of a recorded song or part.
[0379] A "terminal" is an electronic device used to upload performance audio and play back pinch-hitting performances using generative models, such as a smartphone or PC.
[0380] A "server" is a central management system for receiving, storing, and processing data via a network, and is a device that manages performance sound sources and feature data.
[0381] "Features" are characteristic values such as rhythm, tempo, and timbre extracted from audio data, and are important data used to train AI models.
[0382] A "dataset" is a collection of feature-rich data configured for use in training a generative AI model.
[0383] A "generative AI model" is a trained artificial intelligence model, a program that generates appropriate performances based on input data.
[0384] The "emotion engine" is a system that recognizes the emotional state of the user and audience and adjusts the performance style based on that data, and has the ability to analyze emotions from facial expressions and tone of voice.
[0385] "Real-time" refers to time processing that can instantly recognize the actions and emotions of users and spectators and generate instantaneous responses.
[0386] "Substitute performance" refers to performance data generated and played by AI in place of the original performing members, and is a means of solving problems when a member is suddenly absent or unavailable.
[0387] This invention relates to a system in which AI substitutes for a band member to solve the problem of sudden vacancies or the absence of a band member, making it difficult to perform. Furthermore, this invention combines an emotion engine that recognizes the emotions of the user and the audience to achieve a more human-like performance. This system consists of the following steps:
[0388] System Configuration
[0389] 1. Upload your performance audio
[0390] Users upload recordings of past performances using their smartphones, which are then saved in separate parts, such as guitar, bass, and drum parts.
[0391] 2. Saving the audio
[0392] The server receives the audio files uploaded by the user and stores them in a database, which allows data management for each part.
[0393] 3. Sound source preprocessing and feature extraction
[0394] The server performs preprocessing on the received audio source, such as noise reduction, normalization, and clipping, and then extracts features such as rhythm, tempo, and timbre. This process uses an audio processing library such as librosa.
[0395] 4. Training the AI model
[0396] The server uses the preprocessed dataset to train a generative AI model, which uses TensorFlow to learn the playing style and phrasing patterns of each part.
[0397] 5. Operation of the Emotion Engine
[0398] The device uses an emotion engine to recognize the emotions of users and spectators in real time. This engine analyzes the emotional state of users from their facial expressions and tone of voice, and uses a camera and microphone.
[0399] 6. Emotional Performance Adjustment
[0400] The device sends a request to the generative AI model on the server based on the emotional data recognized by the emotion engine. The generative AI model uses this emotional data to adjust the playing style.
[0401] 7. Generation and playback of pinch-hitting performances
[0402] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback, allowing the user to play in sync with other band members.
[0403] Specific examples
[0404] Example 1: Emotional responses during live performance
[0405] The user configures the emotion engine to detect facial and vocal reactions, such as surprise or cheers, from the audience during a live performance. For example, the user can configure the prompt as follows: "When facial and vocal reactions, such as surprise or cheers, are detected from the audience, please adjust your performance style to be more energetic to match that emotion." If this condition is met, a groove-like rhythm pattern is added, and an energetic performance is generated and played.
[0406] Example 2: Playing a ballad in a quiet scene
[0407] The user sets the goal to generate an intimate, moving ballad-style performance in a quiet scene. In this case, the prompt statement is "Please generate an intimate, moving ballad-style performance from a scene where the audience is quietly listening." If the emotion engine detects that the audience is sitting quietly, a slow, moving ballad-style performance is generated and played.
[0408] In this way, the present invention adjusts the performance style in real time based on the emotions of the audience or user, making it possible to smoothly carry out band activities or live music even when there is a sudden vacancy.
[0409] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0410] Step 1:
[0411] The device allows users to upload past performance recordings using a smartphone app. The input is a recorded music file, and the output is a music file sent to the server. This music file is divided into multiple parts (e.g., guitar part, bass part, drum part).
[0412] Step 2:
[0413] The server stores the performance sound files received from the terminal. The input is the music file sent from the terminal, and the output is the file stored in the database. This data is categorized and managed by part in the database.
[0414] Step 3:
[0415] The server performs preprocessing of received audio files by removing noise, normalizing, and clipping. The input is a saved music file, and the output is the preprocessed music data. Specifically, it uses the librosa library to remove noise and normalize the audio to improve the sound quality.
[0416] Step 4:
[0417] The server extracts features such as rhythm, tempo, and timbre from the preprocessed music data. The input is the preprocessed music data, and the output is a feature dataset. This extracts the necessary musical characteristics and constructs a dataset.
[0418] Step 5:
[0419] The server trains the generative AI model using the characteristic dataset. The input is the characteristic dataset, and the output is the trained generative AI model. Specifically, the AI model learns using TensorFlow.
[0420] Step 6:
[0421] The server deploys the trained generative AI model and configures it as an executable endpoint. The input is the trained generative AI model, and the output is the endpoint URL, which allows the AI model to be accessed from a terminal.
[0422] Step 7:
[0423] The device runs an emotion engine using a camera and microphone to recognize the emotions of users and spectators in real time. The input is real-time facial and voice data acquired from the camera and microphone, and the output is emotion data.
[0424] Step 8:
[0425] The device sends a request to the generative AI model based on the acquired emotion data. The input is emotion data, and the output is adjusted performance data. Specifically, the request including emotion data is sent to the endpoint, and the generative AI model adjusts the performance data.
[0426] Step 9:
[0427] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback. The input is the adjusted performance data, and the output is the music played in the physical store. This allows you to play in sync with other band members.
[0428] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0429] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0430] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0431] [Second embodiment]
[0432] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0433] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0434] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0435] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0436] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0437] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0438] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0439] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0440] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0441] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0442] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0443] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0444] The present invention relates to a system in which an AI plays as a substitute when a band suddenly becomes vacant. The following describes an embodiment of this system.
[0445] System Configuration
[0446] This system consists of the following main components:
[0447] 1. A device for uploading performance audio
[0448] 2. A server for storing the performance sound source received from the terminal.
[0449] 3. Processing means for extracting features from the performance sound source and creating a data set
[0450] 4. Processing means for training an AI model using said dataset.
[0451] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[0452] System Operation
[0453] The program of this system operates in the following procedure.
[0454] 1. Uploading the audio
[0455] Users use a smartphone app to upload recordings of their past performances, which provides the system with the audio for each part of the song they want to use.
[0456] 2. Receiving and saving audio files
[0457] The server receives the audio files uploaded by the user and stores them in a database. The stored audio files are then processed as required for each part.
[0458] 3. Sound source preprocessing and feature extraction
[0459] The server performs preprocessing such as noise removal and clipping on the received audio files, then extracts features such as rhythm, tempo, and timbre to create a data set for each part.
[0460] 4. Training the AI model
[0461] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[0462] 5. Deploying the AI model
[0463] The server deploys the trained AI model as an executable endpoint, making it accessible from a smartphone app.
[0464] 6. Generation and playback of pinch-hitting performances
[0465] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch performance in real time in response to the request and sends it back to the device. The device then plays the received performance data and outputs it as sound through an audio output device such as an amplifier.
[0466] Specific examples
[0467] Example 1: Guitarist sub-hiring scenario during band practice
[0468] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0469] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0470] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0471] Example 2: Reproducing the style of a professional musician
[0472] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0473] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0474] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0475] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs, and also maintains musical consistency of substitute members according to the performance style and direction of the band.
[0476] The processing flow will be explained below.
[0477] Step 1:
[0478] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0479] Step 2:
[0480] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0481] Step 3:
[0482] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[0483] Step 4:
[0484] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[0485] Step 5:
[0486] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[0487] Step 6:
[0488] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0489] Step 7:
[0490] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[0491] Step 8:
[0492] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[0493] Step 9:
[0494] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[0495] Step 10:
[0496] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[0497] Step 11:
[0498] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[0499] Step 12:
[0500] Users can provide feedback on the playback of their performance, and if necessary, make fine adjustments within the app to achieve even more precise performance.
[0501] Through these steps, the system can use AI to quickly and effectively provide substitute performances even when a sudden vacancy occurs in a band activity.
[0502] Example 1
[0503] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0504] When a band member suddenly becomes unavailable, it is difficult to find a substitute musician immediately, resulting in interruptions to the performance. Effective methods are needed to solve this problem and maintain the consistency and quality of the performance. It is also difficult to reproduce the performance style of a specific musician.
[0505] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0506] In this invention, the server includes a means for storing performance audio, a means for extracting features from the performance audio to create a dataset, a means for training an AI model using the dataset, and a means for generating a pinch performance using the trained AI model and transmitting it to a terminal. This allows the AI to generate a pinch performance in real time based on audio uploaded by a user and play it on the terminal. In particular, the server performs noise removal and extracts features, allowing the AI model to be trained using high-quality data, allowing the band to continue performing effectively even when a member is absent.
[0507] A "performance sound source" is a recording of music performed by a user, and is audio data stored in digital format.
[0508] A "terminal" is a device used by a user, including a smartphone, tablet, or PC, which is used to upload and play back performance audio.
[0509] A "server" is a computer system that processes information via a network and manages the storage of performance audio, feature extraction, training of AI models, and generation of pinch performances.
[0510] "Features" are quantitative attributes such as rhythm, tempo, and timbre extracted from audio data, and are used to train AI models.
[0511] A "dataset" is a collection of data that compiles features extracted from performance audio, and serves as basic information for training an AI model.
[0512] An "AI model" is an algorithmic system built using artificial intelligence technology, and is a program that learns specific performances and generates substitute performances.
[0513] "Training" is the process of building an AI model, where the AI learns specific patterns and styles using a dataset.
[0514] A "substitute performance" is a replacement performance for an absent or missing band member, generated by a trained AI model.
[0515] "Noise reduction" is a process that removes unnecessary background sounds and unpleasant acoustics from sound source data, and is a preprocessing step to improve data quality.
[0516] An "audio output device" is a device that is connected to a terminal and outputs the generated pinch-hit performance as sound, and includes a speaker, an amplifier, and the like.
[0517] The present invention relates to a system that uses AI to fill in for a band member when a sudden vacancy occurs during a performance. The configuration and operation of this system are described in detail below.
[0518] System configuration
[0519] This system consists of the following main components:
[0520] 1. A device for uploading performance audio
[0521] 2. Server for storing performance audio
[0522] 3. Processing method for extracting features from performance audio and creating a dataset
[0523] 4. A processing method for training an AI model using a dataset
[0524] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[0525] 6. Audio Output Devices
[0526] Hardware and software used
[0527] Devices: Smartphones, tablets, computers, etc.
[0528] Server: A computer system that processes information over a network (e.g., AWS EC2, Google Cloud)
[0529] Software: Python (audio source preprocessing and feature extraction), TensorFlow or PyTorch (AI model training), LibROSA (music signal processing), AWS S3 (audio source storage)
[0530] Audio output devices: speakers, amplifiers, etc.
[0531] System Operation
[0532] This system operates in the following procedure.
[0533] 1. Uploading the audio
[0534] Users use a smartphone app to upload recordings of past performances, such as recordings of band members performing or rehearsing, which provides the AI with data to learn from.
[0535] 2. Receiving and saving audio files
[0536] The server receives the audio files uploaded by the user and stores them using cloud storage such as AWS S3. Meta information (such as song title and part name) is added to the stored audio data.
[0537] 3. Sound source preprocessing and feature extraction
[0538] The server uses a music signal processing library such as LibROSA to perform preprocessing such as noise removal and clipping, then extracts features such as rhythm, tempo, and timbre to create a dataset corresponding to each part.
[0539] 4. Training the AI model
[0540] The server uses the preprocessed dataset to train an AI model in TensorFlow or PyTorch, which learns specific playing styles and phrasing patterns for each musical part.
[0541] 5. Deploying the AI model
[0542] The server deploys the trained AI model as an executable endpoint using AWS SageMaker or Google AI Platform, making it accessible from a smartphone app.
[0543] 6. Generation and playback of pinch-hitting performances
[0544] The device sends a request to the AI model endpoint based on the song and part specified by the user. The server generates a substitute performance in response to the request and returns it to the device in real time. The device then plays back the received performance data and outputs it as sound through an audio output device.
[0545] Examples of specific examples and prompts
[0546] Example 1: Guitarist sub-hiring scenario during band practice
[0547] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0548] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0549] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0550] Example 2: Reproducing the style of a professional musician
[0551] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0552] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0553] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0554] Prompt Sentence Examples
[0555] "Generate guitar parts based on this musician's playing style."
[0556] "Create a new performance based on the audio you uploaded."
[0557] This system allows the band to accommodate sudden absences of members and continue activities while maintaining consistency in performance.
[0558] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0559] Step 1: Upload your audio
[0560] Users use a smartphone app to upload audio recordings of past performances. Specifically, the user taps the "Upload Audio" button on the smartphone app and selects an audio file using a file browser. The selected file is also given meta information such as the song title and part name. The input is the audio file and meta information selected by the user, and the output is the audio file and meta information sent to the server.
[0561] Step 2: Receiving and saving the audio
[0562] The server receives audio files uploaded by users and stores them in a database. Specifically, the server receives an HTTP request and stores the audio files in cloud storage such as AWS S3. The input is the audio file and meta information sent by the user, and the output is the file path of the saved audio file and the meta information registered in the database.
[0563] Step 3: Preprocessing and feature extraction of the audio source
[0564] The server performs preprocessing such as noise removal and clipping on the received audio file. It then uses a music signal processing library such as LibROSA to extract features such as rhythm, tempo, and timbre and create a dataset. Specifically, the server loads the audio file and applies a noise removal algorithm (e.g., spectral subtraction). Next, it extracts features such as rhythm, tempo, and timbre using LibROSA. The input is the received audio file, and the output is the extracted features and a dataset.
[0565] Step 4: Training the AI model
[0566] The server trains the AI model using the preprocessed dataset. It uses TensorFlow or PyTorch to train a recurrent neural network (RNN) or a transformer model. Specifically, the server builds the RNN or transformer model, supplies it with the preprocessed dataset, and performs training. It monitors the training progress and accuracy. The input is the preprocessed dataset, and the output is the trained AI model.
[0567] Step 5: Deploying the AI model
[0568] The server deploys the trained AI model as an executable endpoint. The endpoint is built using AWS SageMaker or Google AI Platform and made accessible from a smartphone app. Specifically, the server saves the trained AI model and deploys it as an endpoint using AWS SageMaker. The input is the trained AI model, and the output is the endpoint URL.
[0569] Step 6: Generate and play back a pinch performance
[0570] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch-hit performance in real time in response to the request and sends it back to the device. The device plays back the received performance data and outputs it as sound through an audio output device. Specifically, the user taps the "Generate Pinch-hit Performance" button in the smartphone app and selects the song and part. The device then sends a request to the AI model endpoint based on the selected information, the server sends the generated pinch-hit performance data to the device, and the device plays back the received performance data. The input is the song and part specified by the user, and the output is the generated pinch-hit performance data.
[0571] In this way, each processing step works in conjunction with each other to create a system that can also handle sudden vacancies in the band.
[0572] (Application example 1)
[0573] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0574] Currently, musical performances at music events, theme parks, and other venues rely heavily on human resources. This means that sudden staff shortages or unexpected problems can disrupt performances. Furthermore, reproducing the style of a professional performer requires advanced technology, which is difficult to replace immediately. Therefore, there is a need to solve these issues by introducing an automated performance system using robots.
[0575] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0576] In this invention, the server includes a terminal for uploading performance sound sources, a server for storing the performance sound sources received from the terminal, processing means for extracting features from the performance sound sources and creating a data set, processing means for training an AI model using the data set, a terminal for generating and playing a pinch-hit performance using the trained AI model, and control means for giving instructions to an automatic musical performance robot used at music events and theme parks. This allows the robot to automatically fill in the gaps in performance even when a sudden vacancy occurs.
[0577] "Performance sound source" refers to audio data that records musical performances.
[0578] An "uploading device" is a device through which a user sends data over the Internet.
[0579] A "server" is a computer system that stores and processes received data.
[0580] "Feature extraction" means extracting important information such as rhythm, tempo, dynamics, and timbre from audio data.
[0581] "Creating a dataset" means organizing and aggregating the extracted features and constructing a set of data to be used for training an AI model.
[0582] "Training an AI model" means using machine learning algorithms to learn patterns from a dataset.
[0583] A "terminal for generating and playing pinch performances" is a device for playing music generated by AI.
[0584] An "automatic playing robot" is a robot designed to play music automatically.
[0585] "Control means" refers to a system or mechanism that gives specific instructions to the automatic musical robot.
[0586] "Playback" means outputting stored sound sources or generated music data as sound.
[0587] This invention is a system for robots to automatically perform music at music events and theme parks. The system consists of the following main components:
[0588] 1. A device for uploading performance audio
[0589] Users upload recordings of past performances to the system using smartphones or computers, which have the ability to transmit audio data to a server via the Internet.
[0590] 2. Server for storing performance audio
[0591] The server receives the audio files uploaded by the user and stores them in a database, which contains the audio data for each instrument part.
[0592] 3. Processing method for extracting features and creating a dataset
[0593] The server performs preprocessing such as noise reduction and clipping on the received audio files, then extracts features such as rhythm, tempo, dynamics, and timbre to create a data set for each part.
[0594] 4. Processing means for training AI models
[0595] The server uses the preprocessed dataset to train the AI model, which is done using machine learning libraries such as TensorFlow, so that the AI model learns the patterns and styles of musical performance.
[0596] 5. Device for generating and playing pinch-hitting performances
[0597] The trained AI model generates a pinch-hit performance based on the song and parts specified by the user. The generated performance data is sent to a terminal that provides instructions to the robot. The terminal then gives the robot specific instructions to play and plays the audio using speakers and amplifiers.
[0598] 6. Control means for giving instructions to automatic musical robots used in music events and theme parks
[0599] The control means connects the server and the robot, and is designed to allow the robot to receive accurate performance instructions. This control means sends the generated performance data to the robot in real time, ensuring smooth performance.
[0600] Specific examples
[0601] Music events at theme parks
[0602] Imagine an evening music event at a theme park where the drumming robot is suddenly absent. The user uploads past drum part audio using a dedicated smartphone app. The server receives this audio, extracts features, and creates a dataset. After the AI model is trained, a new drum part is generated and sent to the robot. The robot then automatically plays the drums based on the generated performance data.
[0603] Prompt Sentence Examples
[0604] The robot drummer for a nighttime music event at a theme park suddenly misses an opportunity, so we need to have an AI generate a drum part using past audio data. Please generate prompts to reproduce the drumming patterns and rhythms in real time.
[0605] Required files: Previous drum part sound files
[0606] Goal: Jazz-style drumming at a tempo of 120 BPM
[0607] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0608] Step 1:
[0609] Users upload audio files of past performances using smartphones or computers. The audio files uploaded by users are sent to a server via the Internet. The input is the audio files of past performances, and the output is the audio data stored on the server.
[0610] Step 2:
[0611] The server stores the audio files received from the user in a database. This stored audio data is used in subsequent processing. The input is the audio file sent from the terminal, and the output is the audio data stored in the database.
[0612] Step 3:
[0613] The server performs preprocessing such as noise removal and clipping on the received audio files. It then extracts features such as rhythm, tempo, dynamics, and timbre from the audio data. This creates a dataset. The input is the saved audio data, and the output is the preprocessed feature data.
[0614] Step 4:
[0615] The server uses the preprocessed feature data to train an AI model. In this step, it runs machine learning algorithms using libraries such as TensorFlow to learn musical patterns and styles. The input is the feature dataset, and the output is a trained AI model.
[0616] Step 5:
[0617] The server uses the trained AI model to generate a pinch-hit performance. The AI model generates a pinch-hit performance based on the song and parts specified by the user. This generated performance data is sent to the device for playback. The input is the song and parts specified by the user and the AI model, and the output is the generated pinch-hit performance data.
[0618] Step 6:
[0619] The terminal transmits the received pinch-hit performance data to the automatic musical robot. Specific performance instructions are given to the robot via the control means, and the robot plays music. The input is the pinch-hit performance data, and the output is the actual performance by the robot.
[0620] Step 7:
[0621] The control means sends the generated performance data to the robot in real time, supporting accurate performance. The input is the performance data, and the output is the robot's smooth playing movements.
[0622] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0623] The present invention relates to a system in which AI performs as a substitute to solve problems that arise when a band suddenly becomes empty or when it is difficult to recruit new members. In addition, by combining it with an emotion engine that recognizes the user's emotions, a more human-like performance can be achieved. The following describes in detail the mode for implementing this system.
[0624] System Configuration
[0625] This system consists of the following main components:
[0626] 1. A device for uploading performance audio
[0627] 2. A server for storing the performance sound source received from the terminal.
[0628] 3. Processing means for extracting features from the performance sound source and creating a data set
[0629] 4. Processing means for training an AI model using said dataset.
[0630] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[0631] 6. Emotion engine that recognizes user emotions
[0632] System Operation
[0633] The program of this system operates in the following procedure.
[0634] 1. Uploading the audio
[0635] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0636] 2. Receiving and saving audio files
[0637] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0638] 3. Sound source preprocessing and feature extraction
[0639] The server performs preprocessing on the received audio files. This includes noise removal, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied. After that, feature quantities such as rhythm, tempo, and timbre are extracted, and a data set corresponding to each part is created.
[0640] 4. Training the AI model
[0641] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[0642] 5. Deploying the AI model
[0643] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0644] 6. Operation of the Emotion Engine
[0645] The device uses an emotion engine to recognize the user's emotions in real time, which analyzes the user's facial expressions and tone of voice to detect their emotional state.
[0646] 7. Generating emotional performance
[0647] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the emotional data, and the AI model adjusts its playing style based on that data.
[0648] 8. Generating and Playing Pinch-Hit Performances
[0649] The device receives and plays the generated pinch-hitting performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play together with other band members.
[0650] Specific examples
[0651] Example 1: Guitarist sub-hiring scenario during band practice
[0652] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0653] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0654] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0655] The device also detects the user's emotions and reflects them in the performance, creating a more realistic performance.
[0656] Example 2: Recreating the style of a professional musician in a stage performance
[0657] If a user wants to reproduce the performance characteristics of a famous professional musician on a live stage, they can upload the musician's performance recording to the app.
[0658] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0659] The device uses a trained AI model to generate and play back a pinch-hit performance in the style of a professional musician, and an emotion engine detects the reactions of the user and audience and reflects them in the performance.
[0660] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs. In addition, by using an emotion engine, it is possible to provide a more human-like performance that responds to the user's emotions.
[0661] The processing flow will be explained below.
[0662] Step 1:
[0663] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0664] Step 2:
[0665] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0666] Step 3:
[0667] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[0668] Step 4:
[0669] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[0670] Step 5:
[0671] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[0672] Step 6:
[0673] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0674] Step 7:
[0675] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[0676] Step 8:
[0677] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[0678] Step 9:
[0679] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[0680] Step 10:
[0681] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[0682] Step 11:
[0683] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[0684] Step 12:
[0685] The device uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's facial expressions and tone of voice to detect emotion data.
[0686] Step 13:
[0687] The device resends the request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the user's emotional data, and the AI model adjusts its playing style based on that data.
[0688] Step 14:
[0689] The server generates new performance data that reflects the emotional data and returns it to the terminal. This data is audio data that includes a performance style that reflects the user's emotional state.
[0690] Step 15:
[0691] The terminal then plays back the received emotion-reflecting performance data, resulting in a more human-like performance that reflects the user's emotions being output through an amplifier.
[0692] Through these steps, this system can quickly and effectively provide substitute performances using AI even when a band member suddenly becomes unavailable. It also uses an emotion engine to generate performances that reflect the user's emotional state, providing a realistic and immersive band experience.
[0693] Example 2
[0694] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0695] There is a need to solve the problem of bands suddenly becoming empty or finding new members is difficult, and to provide more human-like performances that reflect the user's emotions. However, conventional technologies have difficulty in processing performance sound sources in real time and in realizing performances that reflect the user's emotions. Furthermore, there is a lack of a function to dynamically generate performances that respond to emotions while faithfully reproducing the performance style. It is necessary to solve these issues.
[0696] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0697] In this invention, the server includes a means for storing performance sound sources, a means for extracting features such as rhythm, tempo, and timbre from the performance sound sources to create a dataset, and a means for training an AI model using deep learning technology. This enables highly accurate pre-processing and feature extraction of performance sound sources, generation and playback of pinch hit performances in real time, and adjustment of performance style to reflect the user's emotions.
[0698] "Performance sound source" refers to performance data of instruments or vocals that have been recorded in the past.
[0699] "Terminal" refers to an electronic device such as a smartphone or computer used by a user.
[0700] "Server" refers to a networked hardware or software system for storing performance audio, processing, and training AI models.
[0701] "Features" refer to individual characteristic data such as rhythm, tempo, and timbre extracted from a performance sound source.
[0702] A "dataset" refers to a collection of data composed of feature quantities.
[0703] "Deep learning technology" is a method of training AI models using large datasets.
[0704] An "AI model" refers to an artificial intelligence system that has been trained to perform a specific task.
[0705] "Pinch performance" refers to performance data generated by an AI model to fill in for a missing player.
[0706] An "emotion engine" refers to software or hardware that recognizes a user's emotions in real time and generates that data.
[0707] "Processing means" refers to the technical mechanisms that receive, store, preprocess, extract features from performance audio, create datasets, train AI models, and analyze emotional data.
[0708] "Noise reduction" refers to the process of removing unwanted noise from a recorded performance sound source.
[0709] "Normalization" refers to an adjustment process for unifying the volume levels of the performance sound sources.
[0710] "Clipping" is a process of cutting off the peaks of an audio signal to adjust the volume.
[0711] "High-precision filtering processing" refers to a sound source processing procedure that uses advanced algorithms to maintain sound quality.
[0712] "Real time" means that processing occurs immediately without delay.
[0713] "Performance style" refers to a collection of specific rhythms, tempos, tones, phrases, etc. when performing.
[0714] "REST API endpoint" refers to an interface for accessing the functionality of an AI model over a network.
[0715] This system is designed to solve the problem of bands suddenly becoming vacant or finding new members is difficult, making it difficult to perform. In this system, AI substitutes and plays music, and by reflecting the user's emotions, it achieves a more human-like performance.
[0716] System configuration
[0717] This system consists of the following main components:
[0718] 1. A device for uploading performance audio
[0719] 2. A server for storing the performance sound source received from the terminal.
[0720] 3. A processing method for extracting features such as rhythm, tempo, and timbre from performance audio and creating a data set
[0721] 4. Processing means for training AI models using deep learning techniques
[0722] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[0723] 6. Emotion engine that recognizes user emotions in real time
[0724] 7. AI model integration to adjust playing style based on data from the emotion engine
[0725] System Operation Overview
[0726] The specific operation of this system is outlined below.
[0727] Uploading and analyzing audio files
[0728] Users use a smartphone app to upload recordings of their past performances. The uploaded audio files are divided into guitar parts, bass parts, drum parts, etc. Each audio file is sent to a server for analysis.
[0729] Saving and preprocessing audio sources
[0730] The server receives the uploaded audio files and stores them in a database. The received files are first preprocessed using noise removal, normalization, clipping, and other preprocessing techniques. This is done using a high-precision filtering algorithm (e.g., Butterworth filter). After preprocessing is complete, features such as rhythm, tempo, and timbre are extracted from the audio files, and a dataset is created.
[0731] Training an AI model
[0732] The server uses deep learning technologies (e.g., TensorFlow, PyTorch) to train an AI model using the dataset after preprocessing and feature extraction. This AI model learns the playing style and phrase patterns of each part, and once training is complete, it is deployed as a web service.
[0733] Emotion Engine Operation
[0734] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and uses emotion recognition algorithms to detect the user's emotional state, which is then sent to an AI model on the server to reflect a specific playing style.
[0735] Pinch-hitting performance generation and playback
[0736] The device sends a request from the server to the AI model based on the emotional data recognized by the emotion engine. The AI model takes the emotional data into account, adjusts the playing style, and generates appropriate performance data for the substitute. This performance data is played back on the device and connected to an amplifier via the audio output terminal, allowing the user to play along with other band members.
[0737] Specific examples
[0738] A specific example of the system is shown below.
[0739] Example 1: Guitarist sub-hiring scenario during band practice
[0740] 1. The user uploads a recording of their guitar part using a smartphone app.
[0741] 2. The server receives the audio files and stores and preprocesses them.
[0742] 3. The server trains the AI model and sets up the endpoint.
[0743] 4. The device detects the user's emotions using its emotion engine and sends the "joy" emotion data to the AI model.
[0744] 5. The terminal plays the received pinch-hit performance data and outputs the sound through the amplifier.
[0745] Example prompt sentence:
[0746] "Please upload a replacement guitar part for the absent guitarist. Please generate a performance that reflects the emotion of joy."
[0747] In this way, the system can quickly respond to sudden vacancies and provide human-like performances that reflect the user's emotions.
[0748] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0749] Step 1: Upload your audio
[0750] Users upload audio recordings of past performances using a smartphone app. Specifically, they open the app's "Audio Upload" function and select audio files from their PC or smartphone. Audio files are uploaded categorized into guitar parts, bass parts, drum parts, etc.
[0751] Input: Audio files on your PC or smartphone
[0752] Output: Uploaded audio file (sent to server)
[0753] Step 2: Receiving and saving the audio
[0754] The server receives audio files uploaded by users. The received files are stored in a database, and each part is managed as individual data. Specifically, the server verifies the file format (e.g., WAV, MP3), converts it to the appropriate data format, and saves it.
[0755] Input: Uploaded audio file
[0756] Output: Saved audio data (database)
[0757] Step 3: Preprocessing and feature extraction of the audio source
[0758] The server performs preprocessing on the received audio files, such as noise removal, normalization, and clipping. In particular, it applies a high-precision filtering algorithm (e.g., Butterworth filter) to remove noise. It then extracts features such as rhythm, tempo, and timbre, creating a dataset.
[0759] Input: Saved audio data
[0760] Output: A preprocessed dataset with features
[0761] Step 4: Training the AI model
[0762] The server uses the preprocessed dataset to train the AI model using deep learning techniques (e.g., TensorFlow, PyTorch). This involves learning the playing style and phrase patterns of each part. Once trained, the AI model is deployed as a web service.
[0763] Input: Preprocessed dataset
[0764] Output: A trained AI model
[0765] Step 5: Deploying the AI model
[0766] The server deploys the trained AI model as a REST API endpoint, which allows the AI model to be accessed from a smartphone app.
[0767] Input: A trained AI model
[0768] Output: REST API endpoint
[0769] Step 6: Emotion Engine in Action
[0770] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and detects their emotional state using an emotion recognition algorithm (e.g., Facial Action Coding System (FACS)). For example, if the user is smiling, emotion data of "joy" is generated.
[0771] Input: Real-time video and audio data of the user
[0772] Output: Emotion data
[0773] Step 7: Generating emotional performance
[0774] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes emotional data such as "joy," and the AI model adjusts its playing style based on this. The generated pinch-hit performance data is then sent to the device.
[0775] Input: Emotion data
[0776] Output: Pinch hit performance data
[0777] Step 8: Generate and play back a pinch performance
[0778] The device plays the performance data received from the server. Specifically, it outputs the sound using the built-in audio player and connects to an amplifier via the audio output terminal. This allows the user to play along with other band members.
[0779] Input: Pinch-hitting performance data sent from the AI model
[0780] Output: Playback audio
[0781] (Application example 2)
[0782] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0783] The problem that this invention aims to solve is the problem of bands being unable to perform due to sudden vacancies or the absence of a member. Another problem is the difficulty of realizing a more human-like performance that reflects the emotions of the audience and users. In particular, in physical venues that offer live music, it is difficult to quickly find a replacement member when a sudden vacancy occurs, making it difficult to maintain the quality of the performance.
[0784] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for uploading performance sound sources, a server for storing performance sound sources received from the terminal, means for extracting features from the performance sound sources and creating a dataset, means for training a generative AI model using the dataset, a terminal for generating and playing a pinch-hit performance using the trained generative AI model, an emotion engine that recognizes the emotions of the audience and users, and means for adjusting the generative AI model based on the emotion data recognized by the emotion engine. This makes it possible to smoothly carry out band activities and live music even when there is a sudden vacancy, and to provide a realistic performance that corresponds to the emotions of the audience and users.
[0785] A "performance sound source" is audio data generated by instruments, human voices, etc., and is a digital file of a recorded song or part.
[0786] A "terminal" is an electronic device used to upload performance audio and play back pinch-hitting performances using generative models, such as a smartphone or PC.
[0787] A "server" is a central management system for receiving, storing, and processing data via a network, and is a device that manages performance sound sources and feature data.
[0788] "Features" are characteristic values such as rhythm, tempo, and timbre extracted from audio data, and are important data used to train AI models.
[0789] A "dataset" is a collection of feature-rich data configured for use in training a generative AI model.
[0790] A "generative AI model" is a trained artificial intelligence model, a program that generates appropriate performances based on input data.
[0791] The "emotion engine" is a system that recognizes the emotional state of the user and audience and adjusts the performance style based on that data, and has the ability to analyze emotions from facial expressions and tone of voice.
[0792] "Real-time" refers to time processing that can instantly recognize the actions and emotions of users and spectators and generate instantaneous responses.
[0793] "Substitute performance" refers to performance data generated and played by AI in place of the original performing members, and is a means of solving problems when a member is suddenly absent or unavailable.
[0794] This invention relates to a system in which AI substitutes for a band member to solve the problem of sudden vacancies or the absence of a band member, making it difficult to perform. Furthermore, this invention combines an emotion engine that recognizes the emotions of the user and the audience to achieve a more human-like performance. This system consists of the following steps:
[0795] System Configuration
[0796] 1. Upload your performance audio
[0797] Users upload recordings of past performances using their smartphones, which are then saved in separate parts, such as guitar, bass, and drum parts.
[0798] 2. Saving the audio
[0799] The server receives the audio files uploaded by the user and stores them in a database, which allows data management for each part.
[0800] 3. Sound source preprocessing and feature extraction
[0801] The server performs preprocessing on the received audio source, such as noise reduction, normalization, and clipping, and then extracts features such as rhythm, tempo, and timbre. This process uses an audio processing library such as librosa.
[0802] 4. Training the AI model
[0803] The server uses the preprocessed dataset to train a generative AI model, which uses TensorFlow to learn the playing style and phrasing patterns of each part.
[0804] 5. Operation of the Emotion Engine
[0805] The device uses an emotion engine to recognize the emotions of users and spectators in real time. This engine analyzes the emotional state of users from their facial expressions and tone of voice, and uses a camera and microphone.
[0806] 6. Emotional Performance Adjustment
[0807] The device sends a request to the generative AI model on the server based on the emotional data recognized by the emotion engine. The generative AI model uses this emotional data to adjust the playing style.
[0808] 7. Generation and playback of pinch-hitting performances
[0809] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback, allowing the user to play in sync with other band members.
[0810] Specific examples
[0811] Example 1: Emotional responses during live performance
[0812] The user configures the emotion engine to detect facial and vocal reactions, such as surprise or cheers, from the audience during a live performance. For example, the user can configure the prompt as follows: "When facial and vocal reactions, such as surprise or cheers, are detected from the audience, please adjust your performance style to be more energetic to match that emotion." If this condition is met, a groove-like rhythm pattern is added, and an energetic performance is generated and played.
[0813] Example 2: Playing a ballad in a quiet scene
[0814] The user sets the goal to generate an intimate, moving ballad-style performance in a quiet scene. In this case, the prompt statement is "Please generate an intimate, moving ballad-style performance from a scene where the audience is quietly listening." If the emotion engine detects that the audience is sitting quietly, a slow, moving ballad-style performance is generated and played.
[0815] In this way, the present invention adjusts the performance style in real time based on the emotions of the audience or user, making it possible to smoothly carry out band activities or live music even when there is a sudden vacancy.
[0816] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0817] Step 1:
[0818] The device allows users to upload past performance recordings using a smartphone app. The input is a recorded music file, and the output is a music file sent to the server. This music file is divided into multiple parts (e.g., guitar part, bass part, drum part).
[0819] Step 2:
[0820] The server stores the performance sound files received from the terminal. The input is the music file sent from the terminal, and the output is the file stored in the database. This data is categorized and managed by part in the database.
[0821] Step 3:
[0822] The server performs preprocessing of received audio files by removing noise, normalizing, and clipping. The input is a saved music file, and the output is the preprocessed music data. Specifically, it uses the librosa library to remove noise and normalize the audio to improve the sound quality.
[0823] Step 4:
[0824] The server extracts features such as rhythm, tempo, and timbre from the preprocessed music data. The input is the preprocessed music data, and the output is a feature dataset. This extracts the necessary musical characteristics and constructs a dataset.
[0825] Step 5:
[0826] The server trains the generative AI model using the characteristic dataset. The input is the characteristic dataset, and the output is the trained generative AI model. Specifically, the AI model learns using TensorFlow.
[0827] Step 6:
[0828] The server deploys the trained generative AI model and configures it as an executable endpoint. The input is the trained generative AI model, and the output is the endpoint URL, which allows the AI model to be accessed from a terminal.
[0829] Step 7:
[0830] The device runs an emotion engine using a camera and microphone to recognize the emotions of users and spectators in real time. The input is real-time facial and voice data acquired from the camera and microphone, and the output is emotion data.
[0831] Step 8:
[0832] The device sends a request to the generative AI model based on the acquired emotion data. The input is emotion data, and the output is adjusted performance data. Specifically, the request including emotion data is sent to the endpoint, and the generative AI model adjusts the performance data.
[0833] Step 9:
[0834] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback. The input is the adjusted performance data, and the output is the music played in the physical store. This allows you to play in sync with other band members.
[0835] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0836] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0837] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0838] [Third embodiment]
[0839] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0840] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0841] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0842] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0843] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0844] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0845] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0846] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0847] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0848] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0849] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0850] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0851] The present invention relates to a system in which an AI plays as a substitute when a band suddenly becomes vacant. The following describes an embodiment of this system.
[0852] System Configuration
[0853] This system consists of the following main components:
[0854] 1. A device for uploading performance audio
[0855] 2. A server for storing the performance sound source received from the terminal.
[0856] 3. Processing means for extracting features from the performance sound source and creating a data set
[0857] 4. Processing means for training an AI model using said dataset.
[0858] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[0859] System Operation
[0860] The program of this system operates in the following procedure.
[0861] 1. Uploading the audio
[0862] Users use a smartphone app to upload recordings of their past performances, which provides the system with the audio for each part of the song they want to use.
[0863] 2. Receiving and saving audio files
[0864] The server receives the audio files uploaded by the user and stores them in a database. The stored audio files are then processed as required for each part.
[0865] 3. Sound source preprocessing and feature extraction
[0866] The server performs preprocessing such as noise removal and clipping on the received audio files, then extracts features such as rhythm, tempo, and timbre to create a data set for each part.
[0867] 4. Training the AI model
[0868] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[0869] 5. Deploying the AI model
[0870] The server deploys the trained AI model as an executable endpoint, making it accessible from a smartphone app.
[0871] 6. Generation and playback of pinch-hitting performances
[0872] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch performance in real time in response to the request and sends it back to the device. The device then plays the received performance data and outputs it as sound through an audio output device such as an amplifier.
[0873] Specific examples
[0874] Example 1: Guitarist sub-hiring scenario during band practice
[0875] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0876] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0877] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0878] Example 2: Reproducing the style of a professional musician
[0879] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0880] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0881] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0882] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs, and also maintains musical consistency of substitute members according to the performance style and direction of the band.
[0883] The processing flow will be explained below.
[0884] Step 1:
[0885] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[0886] Step 2:
[0887] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[0888] Step 3:
[0889] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[0890] Step 4:
[0891] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[0892] Step 5:
[0893] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[0894] Step 6:
[0895] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[0896] Step 7:
[0897] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[0898] Step 8:
[0899] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[0900] Step 9:
[0901] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[0902] Step 10:
[0903] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[0904] Step 11:
[0905] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[0906] Step 12:
[0907] Users can provide feedback on the playback of their performance, and if necessary, make fine adjustments within the app to achieve even more precise performance.
[0908] Through these steps, the system can use AI to quickly and effectively provide substitute performances even when a sudden vacancy occurs in a band activity.
[0909] Example 1
[0910] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0911] When a band member suddenly becomes unavailable, it is difficult to find a substitute musician to fill in, resulting in interruptions to the performance. Effective methods are needed to resolve this issue and maintain the consistency and quality of the performance. It can also be difficult to reproduce the performance style of a specific musician.
[0912] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0913] In this invention, the server includes a means for storing performance audio, a means for extracting features from the performance audio to create a dataset, a means for training an AI model using the dataset, and a means for generating a pinch performance using the trained AI model and transmitting it to a terminal. This allows the AI to generate a pinch performance in real time based on audio uploaded by a user and play it on the terminal. In particular, the server performs noise removal and extracts features, allowing the AI model to be trained using high-quality data, allowing the band to continue performing effectively even when a member is absent.
[0914] A "performance sound source" is a recording of music performed by a user, and is audio data stored in digital format.
[0915] A "terminal" is a device used by a user, including a smartphone, tablet, or PC, which is used to upload and play back performance audio.
[0916] A "server" is a computer system that processes information via a network and manages the storage of performance audio, feature extraction, training of AI models, and generation of pinch performances.
[0917] "Features" are quantitative attributes such as rhythm, tempo, and timbre extracted from audio data, and are used to train AI models.
[0918] A "dataset" is a collection of data that compiles features extracted from performance audio, and serves as basic information for training an AI model.
[0919] An "AI model" is an algorithmic system built using artificial intelligence technology, and is a program that learns specific performances and generates substitute performances.
[0920] "Training" is the process of building an AI model, where the AI learns specific patterns and styles using a dataset.
[0921] A "substitute performance" is a replacement performance for an absent or missing band member, generated by a trained AI model.
[0922] "Noise reduction" is a process that removes unnecessary background sounds and unpleasant acoustics from sound source data, and is a preprocessing step to improve data quality.
[0923] An "audio output device" is a device that is connected to a terminal and outputs the generated pinch-hit performance as sound, and includes a speaker, an amplifier, and the like.
[0924] The present invention relates to a system that uses AI to fill in for a band member when a sudden vacancy occurs during a performance. The configuration and operation of this system are described in detail below.
[0925] System configuration
[0926] This system consists of the following main components:
[0927] 1. A device for uploading performance audio
[0928] 2. Server for storing performance audio
[0929] 3. Processing method for extracting features from performance audio and creating a dataset
[0930] 4. A processing method for training an AI model using a dataset
[0931] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[0932] 6. Audio Output Devices
[0933] Hardware and software used
[0934] Devices: Smartphones, tablets, computers, etc.
[0935] Server: A computer system that processes information over a network (e.g., AWS EC2, Google Cloud)
[0936] Software: Python (audio source preprocessing and feature extraction), TensorFlow or PyTorch (AI model training), LibROSA (music signal processing), AWS S3 (audio source storage)
[0937] Audio output devices: speakers, amplifiers, etc.
[0938] System Operation
[0939] This system operates in the following procedure.
[0940] 1. Uploading the audio
[0941] Users use a smartphone app to upload recordings of past performances, such as recordings of band members performing or rehearsing, which provides the AI with data to learn from.
[0942] 2. Receiving and saving audio files
[0943] The server receives the audio files uploaded by the user and stores them using cloud storage such as AWS S3. Meta information (such as song title and part name) is added to the stored audio data.
[0944] 3. Sound source preprocessing and feature extraction
[0945] The server uses a music signal processing library such as LibROSA to perform preprocessing such as noise removal and clipping, then extracts features such as rhythm, tempo, and timbre to create a dataset corresponding to each part.
[0946] 4. Training the AI model
[0947] The server uses the preprocessed dataset to train an AI model in TensorFlow or PyTorch, which learns specific playing styles and phrasing patterns for each musical part.
[0948] 5. Deploying the AI model
[0949] The server deploys the trained AI model as an executable endpoint using AWS SageMaker or Google AI Platform, making it accessible from a smartphone app.
[0950] 6. Generation and playback of pinch-hitting performances
[0951] The device sends a request to the AI model endpoint based on the song and part specified by the user. The server generates a substitute performance in response to the request and returns it to the device in real time. The device then plays back the received performance data and outputs it as sound through an audio output device.
[0952] Examples of specific examples and prompts
[0953] Example 1: Guitarist sub-hiring scenario during band practice
[0954] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[0955] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[0956] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[0957] Example 2: Reproducing the style of a professional musician
[0958] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[0959] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[0960] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[0961] Prompt Sentence Examples
[0962] "Generate guitar parts based on this musician's playing style."
[0963] "Create a new performance based on the audio you uploaded."
[0964] This system allows the band to accommodate sudden absences of members and continue activities while maintaining consistency in performance.
[0965] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0966] Step 1: Upload your audio
[0967] Users use a smartphone app to upload audio recordings of past performances. Specifically, the user taps the "Upload Audio" button on the smartphone app and selects an audio file using a file browser. The selected file is also given meta information such as the song title and part name. The input is the audio file and meta information selected by the user, and the output is the audio file and meta information sent to the server.
[0968] Step 2: Receiving and saving the audio
[0969] The server receives audio files uploaded by users and stores them in a database. Specifically, the server receives an HTTP request and stores the audio files in cloud storage such as AWS S3. The input is the audio file and meta information sent by the user, and the output is the file path of the saved audio file and the meta information registered in the database.
[0970] Step 3: Preprocessing and feature extraction of the audio source
[0971] The server performs preprocessing such as noise removal and clipping on the received audio file. It then uses a music signal processing library such as LibROSA to extract features such as rhythm, tempo, and timbre and create a dataset. Specifically, the server loads the audio file and applies a noise removal algorithm (e.g., spectral subtraction). Next, it extracts features such as rhythm, tempo, and timbre using LibROSA. The input is the received audio file, and the output is the extracted features and a dataset.
[0972] Step 4: Training the AI model
[0973] The server trains the AI model using the preprocessed dataset. It uses TensorFlow or PyTorch to train a recurrent neural network (RNN) or a transformer model. Specifically, the server builds the RNN or transformer model, supplies it with the preprocessed dataset, and performs training. It monitors the training progress and accuracy. The input is the preprocessed dataset, and the output is the trained AI model.
[0974] Step 5: Deploying the AI model
[0975] The server deploys the trained AI model as an executable endpoint. The endpoint is built using AWS SageMaker or Google AI Platform and made accessible from a smartphone app. Specifically, the server saves the trained AI model and deploys it as an endpoint using AWS SageMaker. The input is the trained AI model, and the output is the endpoint URL.
[0976] Step 6: Generate and play back a pinch performance
[0977] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch-hit performance in real time in response to the request and sends it back to the device. The device plays back the received performance data and outputs it as sound through an audio output device. Specifically, the user taps the "Generate Pinch-hit Performance" button in the smartphone app and selects the song and part. The device then sends a request to the AI model endpoint based on the selected information, the server sends the generated pinch-hit performance data to the device, and the device plays back the received performance data. The input is the song and part specified by the user, and the output is the generated pinch-hit performance data.
[0978] In this way, each processing step works in conjunction with each other to create a system that can also handle sudden vacancies in the band.
[0979] (Application example 1)
[0980] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0981] Currently, musical performances at music events, theme parks, and other venues rely heavily on human resources. This means that sudden staff shortages or unexpected problems can disrupt performances. Furthermore, reproducing the style of a professional performer requires advanced technology, which is difficult to replace immediately. Therefore, there is a need to solve these issues by introducing an automated performance system using robots.
[0982] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0983] In this invention, the server includes a terminal for uploading performance sound sources, a server for storing the performance sound sources received from the terminal, processing means for extracting features from the performance sound sources and creating a data set, processing means for training an AI model using the data set, a terminal for generating and playing a pinch-hit performance using the trained AI model, and control means for giving instructions to an automatic musical performance robot used at music events and theme parks. This allows the robot to automatically fill in the gaps in performance even when a sudden vacancy occurs.
[0984] "Performance sound source" refers to audio data that records musical performances.
[0985] An "uploading device" is a device through which a user sends data over the Internet.
[0986] A "server" is a computer system that stores and processes received data.
[0987] "Feature extraction" means extracting important information such as rhythm, tempo, dynamics, and timbre from audio data.
[0988] "Creating a dataset" means organizing and aggregating the extracted features and constructing a set of data to be used for training an AI model.
[0989] "Training an AI model" means using machine learning algorithms to learn patterns from a dataset.
[0990] A "terminal for generating and playing pinch performances" is a device for playing music generated by AI.
[0991] An "automatic playing robot" is a robot designed to play music automatically.
[0992] "Control means" refers to a system or mechanism that gives specific instructions to the automatic musical robot.
[0993] "Playback" means outputting stored sound sources or generated music data as sound.
[0994] This invention is a system for robots to automatically perform music at music events and theme parks. The system consists of the following main components:
[0995] 1. A device for uploading performance audio
[0996] Users upload recordings of past performances to the system using smartphones or computers, which have the ability to transmit audio data to a server via the Internet.
[0997] 2. Server for storing performance audio
[0998] The server receives the audio files uploaded by the user and stores them in a database, which contains the audio data for each instrument part.
[0999] 3. Processing method for extracting features and creating a dataset
[1000] The server performs preprocessing such as noise reduction and clipping on the received audio files, then extracts features such as rhythm, tempo, dynamics, and timbre to create a data set for each part.
[1001] 4. Processing means for training AI models
[1002] The server uses the preprocessed dataset to train the AI model, which is done using machine learning libraries such as TensorFlow, so that the AI model learns the patterns and styles of musical performance.
[1003] 5. Device for generating and playing pinch-hitting performances
[1004] The trained AI model generates a pinch-hit performance based on the song and parts specified by the user. The generated performance data is sent to a terminal that provides instructions to the robot. The terminal then gives the robot specific instructions to play and plays the audio using speakers and amplifiers.
[1005] 6. Control means for giving instructions to automatic musical robots used in music events and theme parks
[1006] The control means connects the server and the robot, and is designed to allow the robot to receive accurate performance instructions. This control means sends the generated performance data to the robot in real time, ensuring smooth performance.
[1007] Specific examples
[1008] Music events at theme parks
[1009] Imagine an evening music event at a theme park where the drumming robot is suddenly absent. The user uploads past drum part audio using a dedicated smartphone app. The server receives this audio, extracts features, and creates a dataset. After the AI model is trained, a new drum part is generated and sent to the robot. The robot then automatically plays the drums based on the generated performance data.
[1010] Prompt Sentence Examples
[1011] The robot drummer for a nighttime music event at a theme park suddenly misses an opportunity, so we need to have an AI generate a drum part using past audio data. Please generate prompts to reproduce the drumming patterns and rhythms in real time.
[1012] Required files: Previous drum part sound files
[1013] Goal: Jazz-style drumming at a tempo of 120 BPM
[1014] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1015] Step 1:
[1016] Users upload audio files of past performances using smartphones or computers. The audio files uploaded by users are sent to a server via the Internet. The input is the audio files of past performances, and the output is the audio data stored on the server.
[1017] Step 2:
[1018] The server stores the audio files received from the user in a database. This stored audio data is used in subsequent processing. The input is the audio file sent from the terminal, and the output is the audio data stored in the database.
[1019] Step 3:
[1020] The server performs preprocessing such as noise removal and clipping on the received audio files. It then extracts features such as rhythm, tempo, dynamics, and timbre from the audio data. This creates a dataset. The input is the saved audio data, and the output is the preprocessed feature data.
[1021] Step 4:
[1022] The server uses the preprocessed feature data to train an AI model. In this step, it runs machine learning algorithms using libraries such as TensorFlow to learn musical patterns and styles. The input is the feature dataset, and the output is a trained AI model.
[1023] Step 5:
[1024] The server uses the trained AI model to generate a pinch-hit performance. The AI model generates a pinch-hit performance based on the song and parts specified by the user. This generated performance data is sent to the device for playback. The input is the song and parts specified by the user and the AI model, and the output is the generated pinch-hit performance data.
[1025] Step 6:
[1026] The terminal transmits the received pinch-hit performance data to the automatic musical robot. Specific performance instructions are given to the robot via the control means, and the robot plays music. The input is the pinch-hit performance data, and the output is the actual performance by the robot.
[1027] Step 7:
[1028] The control means sends the generated performance data to the robot in real time, supporting accurate performance. The input is the performance data, and the output is the robot's smooth playing movements.
[1029] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1030] The present invention relates to a system in which AI performs as a substitute to solve problems that arise when a band suddenly becomes empty or when it is difficult to recruit new members. In addition, by combining it with an emotion engine that recognizes the user's emotions, a more human-like performance can be achieved. The following describes in detail the mode for implementing this system.
[1031] System Configuration
[1032] This system consists of the following main components:
[1033] 1. A device for uploading performance audio
[1034] 2. A server for storing the performance sound source received from the terminal.
[1035] 3. Processing means for extracting features from the performance sound source and creating a data set
[1036] 4. Processing means for training an AI model using said dataset.
[1037] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[1038] 6. Emotion engine that recognizes user emotions
[1039] System Operation
[1040] The program of this system operates in the following procedure.
[1041] 1. Uploading the audio
[1042] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[1043] 2. Receiving and saving audio files
[1044] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[1045] 3. Sound source preprocessing and feature extraction
[1046] The server performs preprocessing on the received audio files. This includes noise removal, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied. After that, feature quantities such as rhythm, tempo, and timbre are extracted, and a data set corresponding to each part is created.
[1047] 4. Training the AI model
[1048] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[1049] 5. Deploying the AI model
[1050] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[1051] 6. Operation of the Emotion Engine
[1052] The device uses an emotion engine to recognize the user's emotions in real time, which analyzes the user's facial expressions and tone of voice to detect their emotional state.
[1053] 7. Generating emotional performance
[1054] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the emotional data, and the AI model adjusts its playing style based on that data.
[1055] 8. Generating and Playing Pinch-Hit Performances
[1056] The device receives and plays the generated pinch-hitting performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play together with other band members.
[1057] Specific examples
[1058] Example 1: Guitarist sub-hiring scenario during band practice
[1059] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[1060] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[1061] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[1062] The device also detects the user's emotions and reflects them in the performance, creating a more realistic performance.
[1063] Example 2: Recreating the style of a professional musician in a stage performance
[1064] If a user wants to reproduce the performance characteristics of a famous professional musician on a live stage, they can upload the musician's performance recording to the app.
[1065] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[1066] The device uses a trained AI model to generate and play back a pinch-hit performance in the style of a professional musician, and an emotion engine detects the reactions of the user and audience and reflects them in the performance.
[1067] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs. In addition, by using an emotion engine, it is possible to provide a more human-like performance that responds to the user's emotions.
[1068] The processing flow will be explained below.
[1069] Step 1:
[1070] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[1071] Step 2:
[1072] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[1073] Step 3:
[1074] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[1075] Step 4:
[1076] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[1077] Step 5:
[1078] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[1079] Step 6:
[1080] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[1081] Step 7:
[1082] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[1083] Step 8:
[1084] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[1085] Step 9:
[1086] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[1087] Step 10:
[1088] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[1089] Step 11:
[1090] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[1091] Step 12:
[1092] The device uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's facial expressions and tone of voice to detect emotion data.
[1093] Step 13:
[1094] The device resends the request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the user's emotional data, and the AI model adjusts its playing style based on that data.
[1095] Step 14:
[1096] The server generates new performance data that reflects the emotional data and returns it to the terminal. This data is audio data that includes a performance style that reflects the user's emotional state.
[1097] Step 15:
[1098] The terminal then plays back the received emotion-reflecting performance data, resulting in a more human-like performance that reflects the user's emotions being output through an amplifier.
[1099] Through these steps, this system can quickly and effectively provide substitute performances using AI even when a band member suddenly becomes unavailable. It also uses an emotion engine to generate performances that reflect the user's emotional state, providing a realistic and immersive band experience.
[1100] Example 2
[1101] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1102] There is a need to solve the problem of bands suddenly becoming empty or finding new members is difficult, and to provide more human-like performances that reflect the user's emotions. However, conventional technologies have difficulty in processing performance sound sources in real time and in realizing performances that reflect the user's emotions. Furthermore, there is a lack of a function to dynamically generate performances that respond to emotions while faithfully reproducing the performance style. It is necessary to solve these issues.
[1103] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1104] In this invention, the server includes a means for storing performance sound sources, a means for extracting features such as rhythm, tempo, and timbre from the performance sound sources to create a dataset, and a means for training an AI model using deep learning technology. This enables highly accurate pre-processing and feature extraction of performance sound sources, generation and playback of pinch hit performances in real time, and adjustment of performance style to reflect the user's emotions.
[1105] "Performance sound source" refers to performance data of instruments or vocals that have been recorded in the past.
[1106] "Terminal" refers to an electronic device such as a smartphone or computer used by a user.
[1107] "Server" refers to a networked hardware or software system for storing performance audio, processing, and training AI models.
[1108] "Features" refer to individual characteristic data such as rhythm, tempo, and timbre extracted from a performance sound source.
[1109] A "dataset" refers to a collection of data composed of feature quantities.
[1110] "Deep learning technology" is a method of training AI models using large datasets.
[1111] An "AI model" refers to an artificial intelligence system that has been trained to perform a specific task.
[1112] "Pinch performance" refers to performance data generated by an AI model to fill in for a missing player.
[1113] An "emotion engine" refers to software or hardware that recognizes a user's emotions in real time and generates that data.
[1114] "Processing means" refers to the technical mechanisms that receive, store, preprocess, extract features from performance audio, create datasets, train AI models, and analyze emotional data.
[1115] "Noise reduction" refers to the process of removing unwanted noise from a recorded performance sound source.
[1116] "Normalization" refers to an adjustment process for unifying the volume levels of the performance sound sources.
[1117] "Clipping" is a process of cutting off the peaks of an audio signal to adjust the volume.
[1118] "High-precision filtering processing" refers to a sound source processing procedure that uses advanced algorithms to maintain sound quality.
[1119] "Real time" means that processing occurs immediately without delay.
[1120] "Performance style" refers to a collection of specific rhythms, tempos, tones, phrases, etc. when performing.
[1121] "REST API endpoint" refers to an interface for accessing the functionality of an AI model over a network.
[1122] This system is designed to solve the problem of bands suddenly becoming vacant or finding new members is difficult, making it difficult to perform. In this system, AI substitutes and plays music, and by reflecting the user's emotions, it achieves a more human-like performance.
[1123] System configuration
[1124] This system consists of the following main components:
[1125] 1. A device for uploading performance audio
[1126] 2. A server for storing the performance sound source received from the terminal.
[1127] 3. A processing method for extracting features such as rhythm, tempo, and timbre from performance audio and creating a data set
[1128] 4. Processing means for training AI models using deep learning techniques
[1129] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[1130] 6. Emotion engine that recognizes user emotions in real time
[1131] 7. AI model integration to adjust playing style based on data from the emotion engine
[1132] System Operation Overview
[1133] The specific operation of this system is outlined below.
[1134] Uploading and analyzing audio files
[1135] Users use a smartphone app to upload recordings of their past performances. The uploaded audio files are divided into guitar parts, bass parts, drum parts, etc. Each audio file is sent to a server for analysis.
[1136] Saving and preprocessing audio sources
[1137] The server receives the uploaded audio files and stores them in a database. The received files are first preprocessed using noise removal, normalization, clipping, and other preprocessing techniques. This is done using a high-precision filtering algorithm (e.g., Butterworth filter). After preprocessing is complete, features such as rhythm, tempo, and timbre are extracted from the audio files, and a dataset is created.
[1138] Training an AI model
[1139] The server uses deep learning technologies (e.g., TensorFlow, PyTorch) to train an AI model using the dataset after preprocessing and feature extraction. This AI model learns the playing style and phrase patterns of each part, and once training is complete, it is deployed as a web service.
[1140] Emotion Engine Operation
[1141] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and uses emotion recognition algorithms to detect the user's emotional state, which is then sent to an AI model on the server to reflect a specific playing style.
[1142] Pinch-hitting performance generation and playback
[1143] The device sends a request from the server to the AI model based on the emotional data recognized by the emotion engine. The AI model takes the emotional data into account, adjusts the playing style, and generates appropriate performance data for the substitute. This performance data is played back on the device and connected to an amplifier via the audio output terminal, allowing the user to play along with other band members.
[1144] Specific examples
[1145] A specific example of the system is shown below.
[1146] Example 1: Guitarist sub-hiring scenario during band practice
[1147] 1. The user uploads a recording of their guitar part using a smartphone app.
[1148] 2. The server receives the audio files and stores and preprocesses them.
[1149] 3. The server trains the AI model and sets up the endpoint.
[1150] 4. The device detects the user's emotions using its emotion engine and sends the "joy" emotion data to the AI model.
[1151] 5. The terminal plays the received pinch-hit performance data and outputs the sound through the amplifier.
[1152] Example prompt sentence:
[1153] "Please upload a replacement guitar part for the absent guitarist. Please generate a performance that reflects the emotion of joy."
[1154] In this way, the system can quickly respond to sudden vacancies and provide human-like performances that reflect the user's emotions.
[1155] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1156] Step 1: Upload your audio
[1157] Users upload audio recordings of past performances using a smartphone app. Specifically, they open the app's "Audio Upload" function and select audio files from their PC or smartphone. Audio files are uploaded categorized into guitar parts, bass parts, drum parts, etc.
[1158] Input: Audio files on your PC or smartphone
[1159] Output: Uploaded audio file (sent to server)
[1160] Step 2: Receiving and saving the audio
[1161] The server receives audio files uploaded by users. The received files are stored in a database, and each part is managed as individual data. Specifically, the server verifies the file format (e.g., WAV, MP3), converts it to the appropriate data format, and saves it.
[1162] Input: Uploaded audio file
[1163] Output: Saved audio data (database)
[1164] Step 3: Preprocessing and feature extraction of the audio source
[1165] The server performs preprocessing on the received audio files, such as noise removal, normalization, and clipping. In particular, it applies a high-precision filtering algorithm (e.g., Butterworth filter) to remove noise. It then extracts features such as rhythm, tempo, and timbre, creating a dataset.
[1166] Input: Saved audio data
[1167] Output: A preprocessed dataset with features
[1168] Step 4: Training the AI model
[1169] The server uses the preprocessed dataset to train the AI model using deep learning techniques (e.g., TensorFlow, PyTorch). This involves learning the playing style and phrase patterns of each part. Once trained, the AI model is deployed as a web service.
[1170] Input: Preprocessed dataset
[1171] Output: A trained AI model
[1172] Step 5: Deploying the AI model
[1173] The server deploys the trained AI model as a REST API endpoint, which allows the AI model to be accessed from a smartphone app.
[1174] Input: A trained AI model
[1175] Output: REST API endpoint
[1176] Step 6: Emotion Engine in Action
[1177] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and detects their emotional state using an emotion recognition algorithm (e.g., Facial Action Coding System (FACS)). For example, if the user is smiling, emotion data of "joy" is generated.
[1178] Input: Real-time video and audio data of the user
[1179] Output: Emotion data
[1180] Step 7: Generating emotional performance
[1181] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes emotional data such as "joy," and the AI model adjusts its playing style based on this. The generated pinch-hit performance data is then sent to the device.
[1182] Input: Emotion data
[1183] Output: Pinch hit performance data
[1184] Step 8: Generate and play back a pinch performance
[1185] The device plays the performance data received from the server. Specifically, it outputs the sound using the built-in audio player and connects to an amplifier via the audio output terminal. This allows the user to play along with other band members.
[1186] Input: Pinch-hitting performance data sent from the AI model
[1187] Output: Playback audio
[1188] (Application example 2)
[1189] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1190] The problem that this invention aims to solve is the problem of bands being unable to perform due to sudden vacancies or the absence of a member. Another problem is the difficulty of realizing a more human-like performance that reflects the emotions of the audience and users. In particular, in physical venues that offer live music, it is difficult to quickly find a replacement member when a sudden vacancy occurs, making it difficult to maintain the quality of the performance.
[1191] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for uploading performance sound sources, a server for storing performance sound sources received from the terminal, means for extracting features from the performance sound sources and creating a dataset, means for training a generative AI model using the dataset, a terminal for generating and playing a pinch-hit performance using the trained generative AI model, an emotion engine that recognizes the emotions of the audience and users, and means for adjusting the generative AI model based on the emotion data recognized by the emotion engine. This makes it possible to smoothly carry out band activities and live music even when there is a sudden vacancy, and to provide a realistic performance that corresponds to the emotions of the audience and users.
[1192] A "performance sound source" is audio data generated by instruments, human voices, etc., and is a digital file of a recorded song or part.
[1193] A "terminal" is an electronic device used to upload performance audio and play back pinch-hitting performances using generative models, such as a smartphone or PC.
[1194] A "server" is a central management system for receiving, storing, and processing data via a network, and is a device that manages performance sound sources and feature data.
[1195] "Features" are characteristic values such as rhythm, tempo, and timbre extracted from audio data, and are important data used to train AI models.
[1196] A "dataset" is a collection of feature-rich data configured for use in training a generative AI model.
[1197] A "generative AI model" is a trained artificial intelligence model, a program that generates appropriate performances based on input data.
[1198] The "emotion engine" is a system that recognizes the emotional state of the user and audience and adjusts the performance style based on that data, and has the ability to analyze emotions from facial expressions and tone of voice.
[1199] "Real-time" refers to time processing that can instantly recognize the actions and emotions of users and spectators and generate instantaneous responses.
[1200] "Substitute performance" refers to performance data generated and played by AI in place of the original performing members, and is a means of solving problems when a member is suddenly absent or unavailable.
[1201] This invention relates to a system in which AI substitutes for a band member to solve the problem of sudden vacancies or the absence of a band member, making it difficult to perform. Furthermore, this invention combines an emotion engine that recognizes the emotions of the user and the audience to achieve a more human-like performance. This system consists of the following steps:
[1202] System Configuration
[1203] 1. Upload your performance audio
[1204] Users upload recordings of past performances using their smartphones, which are then saved in separate parts, such as guitar, bass, and drum parts.
[1205] 2. Saving the audio
[1206] The server receives the audio files uploaded by the user and stores them in a database, which allows data management for each part.
[1207] 3. Sound source preprocessing and feature extraction
[1208] The server performs preprocessing on the received audio source, such as noise reduction, normalization, and clipping, and then extracts features such as rhythm, tempo, and timbre. This process uses an audio processing library such as librosa.
[1209] 4. Training the AI model
[1210] The server uses the preprocessed dataset to train a generative AI model, which uses TensorFlow to learn the playing style and phrasing patterns of each part.
[1211] 5. Operation of the Emotion Engine
[1212] The device uses an emotion engine to recognize the emotions of users and spectators in real time. This engine analyzes the emotional state of users from their facial expressions and tone of voice, and uses a camera and microphone.
[1213] 6. Emotional Performance Adjustment
[1214] The device sends a request to the generative AI model on the server based on the emotional data recognized by the emotion engine. The generative AI model uses this emotional data to adjust the playing style.
[1215] 7. Generation and playback of pinch-hitting performances
[1216] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback, allowing the user to play in sync with other band members.
[1217] Specific examples
[1218] Example 1: Emotional responses during live performance
[1219] The user configures the emotion engine to detect facial and vocal reactions, such as surprise or cheers, from the audience during a live performance. For example, the user can configure the prompt as follows: "When facial and vocal reactions, such as surprise or cheers, are detected from the audience, please adjust your performance style to be more energetic to match that emotion." If this condition is met, a groove-like rhythm pattern is added, and an energetic performance is generated and played.
[1220] Example 2: Playing a ballad in a quiet scene
[1221] The user sets the goal to generate an intimate, moving ballad-style performance in a quiet scene. In this case, the prompt statement is "Please generate an intimate, moving ballad-style performance from a scene where the audience is quietly listening." If the emotion engine detects that the audience is sitting quietly, a slow, moving ballad-style performance is generated and played.
[1222] In this way, the present invention adjusts the performance style in real time based on the emotions of the audience or user, making it possible to smoothly carry out band activities or live music even when there is a sudden vacancy.
[1223] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1224] Step 1:
[1225] The device allows users to upload past performance recordings using a smartphone app. The input is a recorded music file, and the output is a music file sent to the server. This music file is divided into multiple parts (e.g., guitar part, bass part, drum part).
[1226] Step 2:
[1227] The server stores the performance sound files received from the terminal. The input is the music file sent from the terminal, and the output is the file stored in the database. This data is categorized and managed by part in the database.
[1228] Step 3:
[1229] The server performs preprocessing of received audio files by removing noise, normalizing, and clipping. The input is a saved music file, and the output is the preprocessed music data. Specifically, it uses the librosa library to remove noise and normalize the audio to improve the sound quality.
[1230] Step 4:
[1231] The server extracts features such as rhythm, tempo, and timbre from the preprocessed music data. The input is the preprocessed music data, and the output is a feature dataset. This extracts the necessary musical characteristics and constructs a dataset.
[1232] Step 5:
[1233] The server trains the generative AI model using the characteristic dataset. The input is the characteristic dataset, and the output is the trained generative AI model. Specifically, the AI model learns using TensorFlow.
[1234] Step 6:
[1235] The server deploys the trained generative AI model and configures it as an executable endpoint. The input is the trained generative AI model, and the output is the endpoint URL, which allows the AI model to be accessed from a terminal.
[1236] Step 7:
[1237] The device runs an emotion engine using a camera and microphone to recognize the emotions of users and spectators in real time. The input is real-time facial and voice data acquired from the camera and microphone, and the output is emotion data.
[1238] Step 8:
[1239] The device sends a request to the generative AI model based on the acquired emotion data. The input is emotion data, and the output is adjusted performance data. Specifically, the request including emotion data is sent to the endpoint, and the generative AI model adjusts the performance data.
[1240] Step 9:
[1241] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback. The input is the adjusted performance data, and the output is the music played in the physical store. This allows you to play in sync with other band members.
[1242] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1243] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1244] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1245] [Fourth embodiment]
[1246] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1247] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1248] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1249] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1250] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1251] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1252] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1253] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1254] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1255] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1256] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1257] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1258] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1259] The present invention relates to a system in which an AI plays as a substitute when a band suddenly becomes vacant. The following describes an embodiment of this system.
[1260] System Configuration
[1261] This system consists of the following main components:
[1262] 1. A device for uploading performance audio
[1263] 2. A server for storing the performance sound source received from the terminal.
[1264] 3. Processing means for extracting features from the performance sound source and creating a data set
[1265] 4. Processing means for training an AI model using said dataset.
[1266] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[1267] System Operation
[1268] The program of this system operates in the following procedure.
[1269] 1. Uploading the audio
[1270] Users use a smartphone app to upload recordings of their past performances, which provides the system with the audio for each part of the song they want to use.
[1271] 2. Receiving and saving audio files
[1272] The server receives the audio files uploaded by the user and stores them in a database. The stored audio files are then processed as required for each part.
[1273] 3. Sound source preprocessing and feature extraction
[1274] The server performs preprocessing such as noise removal and clipping on the received audio files, then extracts features such as rhythm, tempo, and timbre to create a data set for each part.
[1275] 4. Training the AI model
[1276] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[1277] 5. Deploying the AI model
[1278] The server deploys the trained AI model as an executable endpoint, making it accessible from a smartphone app.
[1279] 6. Generation and playback of pinch-hitting performances
[1280] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch performance in real time in response to the request and sends it back to the device. The device then plays the received performance data and outputs it as sound through an audio output device such as an amplifier.
[1281] Specific examples
[1282] Example 1: Guitarist sub-hiring scenario during band practice
[1283] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[1284] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[1285] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[1286] Example 2: Reproducing the style of a professional musician
[1287] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[1288] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[1289] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[1290] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs, and also maintains musical consistency of substitute members according to the performance style and direction of the band.
[1291] The processing flow will be explained below.
[1292] Step 1:
[1293] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[1294] Step 2:
[1295] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[1296] Step 3:
[1297] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[1298] Step 4:
[1299] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[1300] Step 5:
[1301] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[1302] Step 6:
[1303] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[1304] Step 7:
[1305] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[1306] Step 8:
[1307] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[1308] Step 9:
[1309] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[1310] Step 10:
[1311] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[1312] Step 11:
[1313] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[1314] Step 12:
[1315] Users can provide feedback on the playback of their performance, and if necessary, make fine adjustments within the app to achieve even more precise performance.
[1316] Through these steps, the system can use AI to quickly and effectively provide substitute performances even when a sudden vacancy occurs in a band activity.
[1317] Example 1
[1318] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1319] When a band member suddenly becomes unavailable, it is difficult to find a substitute musician to fill in, resulting in interruptions to the performance. Effective methods are needed to resolve this issue and maintain the consistency and quality of the performance. It can also be difficult to reproduce the performance style of a specific musician.
[1320] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1321] In this invention, the server includes a means for storing performance audio, a means for extracting features from the performance audio to create a dataset, a means for training an AI model using the dataset, and a means for generating a pinch performance using the trained AI model and transmitting it to a terminal. This allows the AI to generate a pinch performance in real time based on audio uploaded by a user and play it on the terminal. In particular, the server performs noise removal and extracts features, allowing the AI model to be trained using high-quality data, allowing the band to continue performing effectively even when a member is absent.
[1322] A "performance sound source" is a recording of music performed by a user, and is audio data stored in digital format.
[1323] A "terminal" is a device used by a user, including a smartphone, tablet, or PC, which is used to upload and play back performance audio.
[1324] A "server" is a computer system that processes information via a network and manages the storage of performance audio, feature extraction, training of AI models, and generation of pinch performances.
[1325] "Features" are quantitative attributes such as rhythm, tempo, and timbre extracted from audio data, and are used to train AI models.
[1326] A "dataset" is a collection of data that compiles features extracted from performance audio, and serves as basic information for training an AI model.
[1327] An "AI model" is an algorithmic system built using artificial intelligence technology, and is a program that learns specific performances and generates substitute performances.
[1328] "Training" is the process of building an AI model, where the AI learns specific patterns and styles using a dataset.
[1329] A "substitute performance" is a replacement performance for an absent or missing band member, generated by a trained AI model.
[1330] "Noise reduction" is a process that removes unnecessary background sounds and unpleasant acoustics from sound source data, and is a preprocessing step to improve data quality.
[1331] An "audio output device" is a device that is connected to a terminal and outputs the generated pinch-hit performance as sound, and includes a speaker, an amplifier, and the like.
[1332] The present invention relates to a system that uses AI to fill in for a band member when a sudden vacancy occurs during a performance. The configuration and operation of this system are described in detail below.
[1333] System configuration
[1334] This system consists of the following main components:
[1335] 1. A device for uploading performance audio
[1336] 2. Server for storing performance audio
[1337] 3. Processing method for extracting features from performance audio and creating a dataset
[1338] 4. A processing method for training an AI model using a dataset
[1339] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[1340] 6. Audio Output Devices
[1341] Hardware and software used
[1342] Devices: Smartphones, tablets, computers, etc.
[1343] Server: A computer system that processes information over a network (e.g., AWS EC2, Google Cloud)
[1344] Software: Python (audio source preprocessing and feature extraction), TensorFlow or PyTorch (AI model training), LibROSA (music signal processing), AWS S3 (audio source storage)
[1345] Audio output devices: speakers, amplifiers, etc.
[1346] System Operation
[1347] This system operates in the following procedure.
[1348] 1. Uploading the audio
[1349] Users use a smartphone app to upload recordings of past performances, such as recordings of band members performing or rehearsing, which provides the AI with data to learn from.
[1350] 2. Receiving and saving audio files
[1351] The server receives the audio files uploaded by the user and stores them using cloud storage such as AWS S3. Meta information (such as song title and part name) is added to the stored audio data.
[1352] 3. Sound source preprocessing and feature extraction
[1353] The server uses a music signal processing library such as LibROSA to perform preprocessing such as noise removal and clipping, then extracts features such as rhythm, tempo, and timbre to create a dataset corresponding to each part.
[1354] 4. Training the AI model
[1355] The server uses the preprocessed dataset to train an AI model in TensorFlow or PyTorch, which learns specific playing styles and phrasing patterns for each musical part.
[1356] 5. Deploying the AI model
[1357] The server deploys the trained AI model as an executable endpoint using AWS SageMaker or Google AI Platform, making it accessible from a smartphone app.
[1358] 6. Generation and playback of pinch-hitting performances
[1359] The device sends a request to the AI model endpoint based on the song and part specified by the user. The server generates a substitute performance in response to the request and returns it to the device in real time. The device then plays back the received performance data and outputs it as sound through an audio output device.
[1360] Examples of specific examples and prompts
[1361] Example 1: Guitarist sub-hiring scenario during band practice
[1362] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[1363] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[1364] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[1365] Example 2: Reproducing the style of a professional musician
[1366] If a user wants to reproduce the performance characteristics of a famous professional musician, they can upload the musician's performance recording to the app.
[1367] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[1368] The device uses the trained AI model to generate and play back pinch-hit performances in the style of professional musicians.
[1369] Prompt Sentence Examples
[1370] "Generate guitar parts based on this musician's playing style."
[1371] "Create a new performance based on the audio you uploaded."
[1372] This system allows the band to accommodate sudden absences of members and continue activities while maintaining consistency in performance.
[1373] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1374] Step 1: Upload your audio
[1375] Users use a smartphone app to upload audio recordings of past performances. Specifically, the user taps the "Upload Audio" button on the smartphone app and selects an audio file using a file browser. The selected file is also given meta information such as the song title and part name. The input is the audio file and meta information selected by the user, and the output is the audio file and meta information sent to the server.
[1376] Step 2: Receiving and saving the audio
[1377] The server receives audio files uploaded by users and stores them in a database. Specifically, the server receives an HTTP request and stores the audio files in cloud storage such as AWS S3. The input is the audio file and meta information sent by the user, and the output is the file path of the saved audio file and the meta information registered in the database.
[1378] Step 3: Preprocessing and feature extraction of the audio source
[1379] The server performs preprocessing such as noise removal and clipping on the received audio file. It then uses a music signal processing library such as LibROSA to extract features such as rhythm, tempo, and timbre and create a dataset. Specifically, the server loads the audio file and applies a noise removal algorithm (e.g., spectral subtraction). Next, it extracts features such as rhythm, tempo, and timbre using LibROSA. The input is the received audio file, and the output is the extracted features and a dataset.
[1380] Step 4: Training the AI model
[1381] The server trains the AI model using the preprocessed dataset. It uses TensorFlow or PyTorch to train a recurrent neural network (RNN) or a transformer model. Specifically, the server builds the RNN or transformer model, supplies it with the preprocessed dataset, and performs training. It monitors the training progress and accuracy. The input is the preprocessed dataset, and the output is the trained AI model.
[1382] Step 5: Deploying the AI model
[1383] The server deploys the trained AI model as an executable endpoint. The endpoint is built using AWS SageMaker or Google AI Platform and made accessible from a smartphone app. Specifically, the server saves the trained AI model and deploys it as an endpoint using AWS SageMaker. The input is the trained AI model, and the output is the endpoint URL.
[1384] Step 6: Generate and play back a pinch performance
[1385] The device sends a request to the server's AI model endpoint based on the song and part specified by the user. The server generates a pinch-hit performance in real time in response to the request and sends it back to the device. The device plays back the received performance data and outputs it as sound through an audio output device. Specifically, the user taps the "Generate Pinch-hit Performance" button in the smartphone app and selects the song and part. The device then sends a request to the AI model endpoint based on the selected information, the server sends the generated pinch-hit performance data to the device, and the device plays back the received performance data. The input is the song and part specified by the user, and the output is the generated pinch-hit performance data.
[1386] In this way, each processing step works in conjunction with each other to create a system that can also handle sudden vacancies in the band.
[1387] (Application example 1)
[1388] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] Currently, musical performances at music events, theme parks, and other venues rely heavily on human resources. This means that sudden staff shortages or unexpected problems can disrupt performances. Furthermore, reproducing the style of a professional performer requires advanced technology, which is difficult to replace immediately. Therefore, there is a need to solve these issues by introducing an automated performance system using robots.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1391] In this invention, the server includes a terminal for uploading performance sound sources, a server for storing the performance sound sources received from the terminal, processing means for extracting features from the performance sound sources and creating a data set, processing means for training an AI model using the data set, a terminal for generating and playing a pinch-hit performance using the trained AI model, and control means for giving instructions to an automatic musical performance robot used at music events and theme parks. This allows the robot to automatically fill in the gaps in performance even when a sudden vacancy occurs.
[1392] "Performance sound source" refers to audio data that records musical performances.
[1393] An "uploading device" is a device through which a user sends data over the Internet.
[1394] A "server" is a computer system that stores and processes received data.
[1395] "Feature extraction" means extracting important information such as rhythm, tempo, dynamics, and timbre from audio data.
[1396] "Creating a dataset" means organizing and aggregating the extracted features and constructing a set of data to be used for training an AI model.
[1397] "Training an AI model" means using machine learning algorithms to learn patterns from a dataset.
[1398] A "terminal for generating and playing pinch performances" is a device for playing music generated by AI.
[1399] An "automatic playing robot" is a robot designed to play music automatically.
[1400] "Control means" refers to a system or mechanism that gives specific instructions to the automatic musical robot.
[1401] "Playback" means outputting stored sound sources or generated music data as sound.
[1402] This invention is a system for robots to automatically perform music at music events and theme parks. The system consists of the following main components:
[1403] 1. A device for uploading performance audio
[1404] Users upload recordings of past performances to the system using smartphones or computers, which have the ability to transmit audio data to a server via the Internet.
[1405] 2. Server for storing performance audio
[1406] The server receives the audio files uploaded by the user and stores them in a database, which contains the audio data for each instrument part.
[1407] 3. Processing method for extracting features and creating a dataset
[1408] The server performs preprocessing such as noise reduction and clipping on the received audio files, then extracts features such as rhythm, tempo, dynamics, and timbre to create a data set for each part.
[1409] 4. Processing means for training AI models
[1410] The server uses the preprocessed dataset to train the AI model, which is done using machine learning libraries such as TensorFlow, so that the AI model learns the patterns and styles of musical performance.
[1411] 5. Device for generating and playing pinch-hitting performances
[1412] The trained AI model generates a pinch-hit performance based on the song and parts specified by the user. The generated performance data is sent to a terminal that provides instructions to the robot. The terminal then gives the robot specific instructions to play and plays the audio using speakers and amplifiers.
[1413] 6. Control means for giving instructions to automatic musical robots used in music events and theme parks
[1414] The control means connects the server and the robot, and is designed to allow the robot to receive accurate performance instructions. This control means sends the generated performance data to the robot in real time, ensuring smooth performance.
[1415] Specific examples
[1416] Music events at theme parks
[1417] Imagine an evening music event at a theme park where the drumming robot is suddenly absent. The user uploads past drum part audio using a dedicated smartphone app. The server receives this audio, extracts features, and creates a dataset. After the AI model is trained, a new drum part is generated and sent to the robot. The robot then automatically plays the drums based on the generated performance data.
[1418] Prompt Sentence Examples
[1419] The robot drummer for a nighttime music event at a theme park suddenly misses an opportunity, so we need to have an AI generate a drum part using past audio data. Please generate prompts to reproduce the drumming patterns and rhythms in real time.
[1420] Required files: Previous drum part sound files
[1421] Goal: Jazz-style drumming at a tempo of 120 BPM
[1422] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1423] Step 1:
[1424] Users upload audio files of past performances using smartphones or computers. The audio files uploaded by users are sent to a server via the Internet. The input is the audio files of past performances, and the output is the audio data stored on the server.
[1425] Step 2:
[1426] The server stores the audio files received from the user in a database. This stored audio data is used in subsequent processing. The input is the audio file sent from the terminal, and the output is the audio data stored in the database.
[1427] Step 3:
[1428] The server performs preprocessing such as noise removal and clipping on the received audio files. It then extracts features such as rhythm, tempo, dynamics, and timbre from the audio data. This creates a dataset. The input is the saved audio data, and the output is the preprocessed feature data.
[1429] Step 4:
[1430] The server uses the preprocessed feature data to train an AI model. In this step, it runs machine learning algorithms using libraries such as TensorFlow to learn musical patterns and styles. The input is the feature dataset, and the output is a trained AI model.
[1431] Step 5:
[1432] The server uses the trained AI model to generate a pinch-hit performance. The AI model generates a pinch-hit performance based on the song and parts specified by the user. This generated performance data is sent to the device for playback. The input is the song and parts specified by the user and the AI model, and the output is the generated pinch-hit performance data.
[1433] Step 6:
[1434] The terminal transmits the received pinch-hit performance data to the automatic musical robot. Specific performance instructions are given to the robot via the control means, and the robot plays music. The input is the pinch-hit performance data, and the output is the actual performance by the robot.
[1435] Step 7:
[1436] The control means sends the generated performance data to the robot in real time, supporting accurate performance. The input is the performance data, and the output is the robot's smooth playing movements.
[1437] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1438] The present invention relates to a system in which AI performs as a substitute to solve problems that arise when a band suddenly becomes empty or when it is difficult to recruit new members. In addition, by combining it with an emotion engine that recognizes the user's emotions, a more human-like performance can be achieved. The following describes in detail the mode for implementing this system.
[1439] System Configuration
[1440] This system consists of the following main components:
[1441] 1. A device for uploading performance audio
[1442] 2. A server for storing the performance sound source received from the terminal.
[1443] 3. Processing means for extracting features from the performance sound source and creating a data set
[1444] 4. Processing means for training an AI model using said dataset.
[1445] 5. A device for generating and playing pinch-hit performances using the trained AI model.
[1446] 6. Emotion engine that recognizes user emotions
[1447] System Operation
[1448] The program of this system operates in the following procedure.
[1449] 1. Uploading the audio
[1450] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[1451] 2. Receiving and saving audio files
[1452] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[1453] 3. Sound source preprocessing and feature extraction
[1454] The server performs preprocessing on the received audio files. This includes noise removal, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied. After that, feature quantities such as rhythm, tempo, and timbre are extracted, and a data set corresponding to each part is created.
[1455] 4. Training the AI model
[1456] The server uses the preprocessed dataset to train the AI model, which learns the playing style and phrase patterns of each part.
[1457] 5. Deploying the AI model
[1458] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[1459] 6. Operation of the Emotion Engine
[1460] The device uses an emotion engine to recognize the user's emotions in real time, which analyzes the user's facial expressions and tone of voice to detect their emotional state.
[1461] 7. Generating emotional performance
[1462] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the emotional data, and the AI model adjusts its playing style based on that data.
[1463] 8. Generating and Playing Pinch-Hit Performances
[1464] The device receives and plays the generated pinch-hitting performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play together with other band members.
[1465] Specific examples
[1466] Example 1: Guitarist sub-hiring scenario during band practice
[1467] If a guitarist is absent during band practice, the user can launch the smartphone app and upload a recording of the guitar part.
[1468] The server receives the uploaded guitar part audio, performs preprocessing and feature extraction, and trains the AI model.
[1469] The device sends a request to the server's AI model endpoint to generate a pinch-hit performance for the guitar part, plays the received performance data, and outputs the sound through an amplifier.
[1470] The device also detects the user's emotions and reflects them in the performance, creating a more realistic performance.
[1471] Example 2: Recreating the style of a professional musician in a stage performance
[1472] If a user wants to reproduce the performance characteristics of a famous professional musician on a live stage, they can upload the musician's performance recording to the app.
[1473] The server extracts features from the received audio source and trains an AI model to learn the playing style.
[1474] The device uses a trained AI model to generate and play back a pinch-hit performance in the style of a professional musician, and an emotion engine detects the reactions of the user and audience and reflects them in the performance.
[1475] With the above configuration, the present invention allows band activities to proceed smoothly even when a sudden vacancy occurs. In addition, by using an emotion engine, it is possible to provide a more human-like performance that responds to the user's emotions.
[1476] The processing flow will be explained below.
[1477] Step 1:
[1478] Users use a smartphone app to upload recordings of their past performances, which are divided into guitar parts, bass parts, drum parts, etc.
[1479] Step 2:
[1480] The server receives the audio files uploaded by the user and stores them in a database. The received files are managed as individual data for each part.
[1481] Step 3:
[1482] The server performs preprocessing on the received audio files, including noise reduction, normalization, clipping, etc. Since sound quality is particularly important, high-precision filtering is applied.
[1483] Step 4:
[1484] The server extracts features from the preprocessed audio data, analyzing features such as rhythm, tempo, dynamics, and timbre, and creates a dataset for each part.
[1485] Step 5:
[1486] The server uses the feature dataset to train an AI model, which uses techniques such as deep learning and recurrent neural networks (RNNs) to learn the playing style of each part.
[1487] Step 6:
[1488] The server deploys the trained AI model and configures it as an executable endpoint, allowing it to be accessed by a smartphone app.
[1489] Step 7:
[1490] The user opens the smartphone app and selects the song they want to play and the part they want to fill in. For example, if a guitarist is absent, they select the guitar part.
[1491] Step 8:
[1492] Based on the selected part, the device sends a request to the server's AI model endpoint, which includes information about the song to be played and the part specification.
[1493] Step 9:
[1494] In response to the received request, the server executes the AI model and generates a substitute performance for the specified part. The generated performance data is processed in real time.
[1495] Step 10:
[1496] The server returns the generated pinch hit performance data to the terminal. This data is audio data corresponding to the specified part of the song to be played.
[1497] Step 11:
[1498] The device then plays back the received performance data. By connecting the device to an amplifier via its audio output terminal, it becomes possible to play along with other band members.
[1499] Step 12:
[1500] The device uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's facial expressions and tone of voice to detect emotion data.
[1501] Step 13:
[1502] The device resends the request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes the user's emotional data, and the AI model adjusts its playing style based on that data.
[1503] Step 14:
[1504] The server generates new performance data that reflects the emotional data and returns it to the terminal. This data is audio data that includes a performance style that reflects the user's emotional state.
[1505] Step 15:
[1506] The terminal then plays back the received emotion-reflecting performance data, resulting in a more human-like performance that reflects the user's emotions being output through an amplifier.
[1507] Through these steps, this system can quickly and effectively provide substitute performances using AI even when a band member suddenly becomes unavailable. It also uses an emotion engine to generate performances that reflect the user's emotional state, providing a realistic and immersive band experience.
[1508] Example 2
[1509] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1510] There is a need to solve the problem of bands suddenly becoming empty or finding new members is difficult, and to provide more human-like performances that reflect the user's emotions. However, conventional technologies have difficulty in processing performance sound sources in real time and in realizing performances that reflect the user's emotions. Furthermore, there is a lack of a function to dynamically generate performances that respond to emotions while faithfully reproducing the performance style. It is necessary to solve these issues.
[1511] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1512] In this invention, the server includes a means for storing performance sound sources, a means for extracting features such as rhythm, tempo, and timbre from the performance sound sources to create a dataset, and a means for training an AI model using deep learning technology. This enables highly accurate pre-processing and feature extraction of performance sound sources, generation and playback of pinch hit performances in real time, and adjustment of performance style to reflect the user's emotions.
[1513] "Performance sound source" refers to performance data of instruments or vocals that have been recorded in the past.
[1514] "Terminal" refers to an electronic device such as a smartphone or computer used by a user.
[1515] "Server" refers to a networked hardware or software system for storing performance audio, processing, and training AI models.
[1516] "Features" refer to individual characteristic data such as rhythm, tempo, and timbre extracted from a performance sound source.
[1517] A "dataset" refers to a collection of data composed of feature quantities.
[1518] "Deep learning technology" is a method of training AI models using large datasets.
[1519] An "AI model" refers to an artificial intelligence system that has been trained to perform a specific task.
[1520] "Pinch performance" refers to performance data generated by an AI model to fill in for a missing player.
[1521] An "emotion engine" refers to software or hardware that recognizes a user's emotions in real time and generates that data.
[1522] "Processing means" refers to the technical mechanisms that receive, store, preprocess, extract features from performance audio, create datasets, train AI models, and analyze emotional data.
[1523] "Noise reduction" refers to the process of removing unwanted noise from a recorded performance sound source.
[1524] "Normalization" refers to an adjustment process for unifying the volume levels of the performance sound sources.
[1525] "Clipping" is a process of cutting off the peaks of an audio signal to adjust the volume.
[1526] "High-precision filtering processing" refers to a sound source processing procedure that uses advanced algorithms to maintain sound quality.
[1527] "Real time" means that processing occurs immediately without delay.
[1528] "Performance style" refers to a collection of specific rhythms, tempos, tones, phrases, etc. when performing.
[1529] "REST API endpoint" refers to an interface for accessing the functionality of an AI model over a network.
[1530] This system is designed to solve the problem of bands suddenly becoming vacant or finding new members is difficult, making it difficult to perform. In this system, AI substitutes and plays music, and by reflecting the user's emotions, it achieves a more human-like performance.
[1531] System configuration
[1532] This system consists of the following main components:
[1533] 1. A device for uploading performance audio
[1534] 2. A server for storing the performance sound source received from the terminal.
[1535] 3. A processing method for extracting features such as rhythm, tempo, and timbre from performance audio and creating a data set
[1536] 4. Processing means for training AI models using deep learning techniques
[1537] 5. A device for generating and playing back pinch-hitting performances using the trained AI model
[1538] 6. Emotion engine that recognizes user emotions in real time
[1539] 7. AI model integration to adjust playing style based on data from the emotion engine
[1540] System Operation Overview
[1541] The specific operation of this system is outlined below.
[1542] Uploading and analyzing audio files
[1543] Users use a smartphone app to upload recordings of their past performances. The uploaded audio files are divided into guitar parts, bass parts, drum parts, etc. Each audio file is sent to a server for analysis.
[1544] Saving and preprocessing audio sources
[1545] The server receives the uploaded audio files and stores them in a database. The received files are first preprocessed using noise removal, normalization, clipping, and other preprocessing techniques. This is done using a high-precision filtering algorithm (e.g., Butterworth filter). After preprocessing is complete, features such as rhythm, tempo, and timbre are extracted from the audio files, and a dataset is created.
[1546] Training an AI model
[1547] The server uses deep learning technologies (e.g., TensorFlow, PyTorch) to train an AI model using the dataset after preprocessing and feature extraction. This AI model learns the playing style and phrase patterns of each part, and once training is complete, it is deployed as a web service.
[1548] Emotion Engine Operation
[1549] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and uses emotion recognition algorithms to detect the user's emotional state, which is then sent to an AI model on the server to reflect a specific playing style.
[1550] Pinch-hitting performance generation and playback
[1551] The device sends a request from the server to the AI model based on the emotional data recognized by the emotion engine. The AI model takes the emotional data into account, adjusts the playing style, and generates appropriate performance data for the substitute. This performance data is played back on the device and connected to an amplifier via the audio output terminal, allowing the user to play along with other band members.
[1552] Specific examples
[1553] A specific example of the system is shown below.
[1554] Example 1: Guitarist sub-hiring scenario during band practice
[1555] 1. The user uploads a recording of their guitar part using a smartphone app.
[1556] 2. The server receives the audio files and stores and preprocesses them.
[1557] 3. The server trains the AI model and sets up the endpoint.
[1558] 4. The device detects the user's emotions using its emotion engine and sends the "joy" emotion data to the AI model.
[1559] 5. The terminal plays the received pinch-hit performance data and outputs the sound through the amplifier.
[1560] Example prompt sentence:
[1561] "Please upload a replacement guitar part for the absent guitarist. Please generate a performance that reflects the emotion of joy."
[1562] In this way, the system can quickly respond to sudden vacancies and provide human-like performances that reflect the user's emotions.
[1563] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1564] Step 1: Upload your audio
[1565] Users upload audio recordings of past performances using a smartphone app. Specifically, they open the app's "Audio Upload" function and select audio files from their PC or smartphone. Audio files are uploaded categorized into guitar parts, bass parts, drum parts, etc.
[1566] Input: Audio files on your PC or smartphone
[1567] Output: Uploaded audio file (sent to server)
[1568] Step 2: Receiving and saving the audio
[1569] The server receives audio files uploaded by users. The received files are stored in a database, and each part is managed as individual data. Specifically, the server verifies the file format (e.g., WAV, MP3), converts it to the appropriate data format, and saves it.
[1570] Input: Uploaded audio file
[1571] Output: Saved audio data (database)
[1572] Step 3: Preprocessing and feature extraction of the audio source
[1573] The server performs preprocessing on the received audio files, such as noise removal, normalization, and clipping. In particular, it applies a high-precision filtering algorithm (e.g., Butterworth filter) to remove noise. It then extracts features such as rhythm, tempo, and timbre, creating a dataset.
[1574] Input: Saved audio data
[1575] Output: A preprocessed dataset with features
[1576] Step 4: Training the AI model
[1577] The server uses the preprocessed dataset to train the AI model using deep learning techniques (e.g., TensorFlow, PyTorch). This involves learning the playing style and phrase patterns of each part. Once trained, the AI model is deployed as a web service.
[1578] Input: Preprocessed dataset
[1579] Output: A trained AI model
[1580] Step 5: Deploying the AI model
[1581] The server deploys the trained AI model as a REST API endpoint, which allows the AI model to be accessed from a smartphone app.
[1582] Input: A trained AI model
[1583] Output: REST API endpoint
[1584] Step 6: Emotion Engine in Action
[1585] The device uses the user's camera and microphone to analyze the user's facial expressions and tone of voice in real time, and detects their emotional state using an emotion recognition algorithm (e.g., Facial Action Coding System (FACS)). For example, if the user is smiling, emotion data of "joy" is generated.
[1586] Input: Real-time video and audio data of the user
[1587] Output: Emotion data
[1588] Step 7: Generating emotional performance
[1589] The device sends a request to the AI model on the server based on the emotional data recognized by the emotion engine. The request includes emotional data such as "joy," and the AI model adjusts its playing style based on this. The generated pinch-hit performance data is then sent to the device.
[1590] Input: Emotion data
[1591] Output: Pinch hit performance data
[1592] Step 8: Generate and play back a pinch performance
[1593] The device plays the performance data received from the server. Specifically, it outputs the sound using the built-in audio player and connects to an amplifier via the audio output terminal. This allows the user to play along with other band members.
[1594] Input: Pinch-hitting performance data sent from the AI model
[1595] Output: Playback audio
[1596] (Application example 2)
[1597] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1598] The problem that this invention aims to solve is the problem of bands being unable to perform due to sudden vacancies or the absence of a member. Another problem is the difficulty of realizing a more human-like performance that reflects the emotions of the audience and users. In particular, in physical venues that offer live music, it is difficult to quickly find a replacement member when a sudden vacancy occurs, making it difficult to maintain the quality of the performance.
[1599] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for uploading performance sound sources, a server for storing performance sound sources received from the terminal, means for extracting features from the performance sound sources and creating a dataset, means for training a generative AI model using the dataset, a terminal for generating and playing a pinch-hit performance using the trained generative AI model, an emotion engine that recognizes the emotions of the audience and users, and means for adjusting the generative AI model based on the emotion data recognized by the emotion engine. This makes it possible to smoothly carry out band activities and live music even when there is a sudden vacancy, and to provide a realistic performance that corresponds to the emotions of the audience and users.
[1600] A "performance sound source" is audio data generated by instruments, human voices, etc., and is a digital file of a recorded song or part.
[1601] A "terminal" is an electronic device used to upload performance audio and play back pinch-hitting performances using generative models, such as a smartphone or PC.
[1602] A "server" is a central management system for receiving, storing, and processing data via a network, and is a device that manages performance sound sources and feature data.
[1603] "Features" are characteristic values such as rhythm, tempo, and timbre extracted from audio data, and are important data used to train AI models.
[1604] A "dataset" is a collection of feature-rich data configured for use in training a generative AI model.
[1605] A "generative AI model" is a trained artificial intelligence model, a program that generates appropriate performances based on input data.
[1606] The "emotion engine" is a system that recognizes the emotional state of the user and audience and adjusts the performance style based on that data, and has the ability to analyze emotions from facial expressions and tone of voice.
[1607] "Real-time" refers to time processing that can instantly recognize the actions and emotions of users and spectators and generate instantaneous responses.
[1608] "Substitute performance" refers to performance data generated and played by AI in place of the original performing members, and is a means of solving problems when a member is suddenly absent or unavailable.
[1609] This invention relates to a system in which AI substitutes for a band member to solve the problem of sudden vacancies or the absence of a band member, making it difficult to perform. Furthermore, this invention combines an emotion engine that recognizes the emotions of the user and the audience to achieve a more human-like performance. This system consists of the following steps:
[1610] System Configuration
[1611] 1. Upload your performance audio
[1612] Users upload recordings of past performances using their smartphones, which are then saved in separate parts, such as guitar, bass, and drum parts.
[1613] 2. Saving the audio
[1614] The server receives the audio files uploaded by the user and stores them in a database, which allows data management for each part.
[1615] 3. Sound source preprocessing and feature extraction
[1616] The server performs preprocessing on the received audio source, such as noise reduction, normalization, and clipping, and then extracts features such as rhythm, tempo, and timbre. This process uses an audio processing library such as librosa.
[1617] 4. Training the AI model
[1618] The server uses the preprocessed dataset to train a generative AI model, which uses TensorFlow to learn the playing style and phrasing patterns of each part.
[1619] 5. Operation of the Emotion Engine
[1620] The device uses an emotion engine to recognize the emotions of users and spectators in real time. This engine analyzes the emotional state of users from their facial expressions and tone of voice, and uses a camera and microphone.
[1621] 6. Emotional Performance Adjustment
[1622] The device sends a request to the generative AI model on the server based on the emotional data recognized by the emotion engine. The generative AI model uses this emotional data to adjust the playing style.
[1623] 7. Generation and playback of pinch-hitting performances
[1624] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback, allowing the user to play in sync with other band members.
[1625] Specific examples
[1626] Example 1: Emotional responses during live performance
[1627] The user configures the emotion engine to detect facial and vocal reactions, such as surprise or cheers, from the audience during a live performance. For example, the user can configure the prompt as follows: "When facial and vocal reactions, such as surprise or cheers, are detected from the audience, please adjust your performance style to be more energetic to match that emotion." If this condition is met, a groove-like rhythm pattern is added, and an energetic performance is generated and played.
[1628] Example 2: Playing a ballad in a quiet scene
[1629] The user sets the goal to generate an intimate, moving ballad-style performance in a quiet scene. In this case, the prompt statement is "Please generate an intimate, moving ballad-style performance from a scene where the audience is quietly listening." If the emotion engine detects that the audience is sitting quietly, a slow, moving ballad-style performance is generated and played.
[1630] In this way, the present invention adjusts the performance style in real time based on the emotions of the audience or user, making it possible to smoothly carry out band activities or live music even when there is a sudden vacancy.
[1631] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1632] Step 1:
[1633] The device allows users to upload past performance recordings using a smartphone app. The input is a recorded music file, and the output is a music file sent to the server. This music file is divided into multiple parts (e.g., guitar part, bass part, drum part).
[1634] Step 2:
[1635] The server stores the performance sound files received from the terminal. The input is the music file sent from the terminal, and the output is the file stored in the database. This data is categorized and managed by part in the database.
[1636] Step 3:
[1637] The server performs preprocessing of received audio files by removing noise, normalizing, and clipping. The input is a saved music file, and the output is the preprocessed music data. Specifically, it uses the librosa library to remove noise and normalize the audio to improve the sound quality.
[1638] Step 4:
[1639] The server extracts features such as rhythm, tempo, and timbre from the preprocessed music data. The input is the preprocessed music data, and the output is a feature dataset. This extracts the necessary musical characteristics and constructs a dataset.
[1640] Step 5:
[1641] The server trains the generative AI model using the characteristic dataset. The input is the characteristic dataset, and the output is the trained generative AI model. Specifically, the AI model learns using TensorFlow.
[1642] Step 6:
[1643] The server deploys the trained generative AI model and configures it as an executable endpoint. The input is the trained generative AI model, and the output is the endpoint URL, which allows the AI model to be accessed from a terminal.
[1644] Step 7:
[1645] The device runs an emotion engine using a camera and microphone to recognize the emotions of users and spectators in real time. The input is real-time facial and voice data acquired from the camera and microphone, and the output is emotion data.
[1646] Step 8:
[1647] The device sends a request to the generative AI model based on the acquired emotion data. The input is emotion data, and the output is adjusted performance data. Specifically, the request including emotion data is sent to the endpoint, and the generative AI model adjusts the performance data.
[1648] Step 9:
[1649] The device receives the generated pinch-hitting performance data and connects it to an external speaker or amplifier for playback. The input is the adjusted performance data, and the output is the music played in the physical store. This allows you to play in sync with other band members.
[1650] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1651] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1652] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1653] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1654] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1655] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1656] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1657] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1658] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1659] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1660] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1661] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1662] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1663] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1664] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1665] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1666] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1667] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1668] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1669] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1670] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1671] The following is further disclosed regarding the above embodiment.
[1672] (Claim 1)
[1673] A device for uploading performance audio;
[1674] a server for storing the performance sound source received from the terminal;
[1675] a processing means for extracting features from the performance sound source and creating a data set;
[1676] processing means for training an AI model using said dataset;
[1677] a terminal for generating and playing a pinch-hit performance using the trained AI model;
[1678] A system including:
[1679] (Claim 2)
[1680] 2. The system according to claim 1, wherein the terminal comprises a processing means for generating a pinch-hit performance in real time based on the collected performance sound source.
[1681] (Claim 3)
[1682] 2. The system according to claim 1, wherein the server comprises a processing means for analyzing past performance sound sources and extracting features including rhythm, tempo, dynamics, and timbre.
[1683] "Example 1"
[1684] (Claim 1)
[1685] A device for uploading performance audio;
[1686] a server for storing the performance sound source received from the terminal;
[1687] a processing means for extracting features from the performance sound source and creating a data set;
[1688] processing means for training an AI model using said dataset;
[1689] A processing means for generating a pinch-hit performance using the trained AI model and transmitting the generated performance from the server to a terminal;
[1690] means for reproducing the pinch-hit performance on the terminal and outputting it as sound through an audio output device;
[1691] A system including:
[1692] (Claim 2)
[1693] a processing means for generating a pinch-hit performance in real time based on the collected performance sound source;
[1694] and an audio output device for playing the generated rendition.
[1695] (Claim 3)
[1696] 2. The system according to claim 1, wherein the server comprises processing means for analyzing past performance sound sources, removing noise, and extracting features including rhythm, tempo, and timbre.
[1697] "Application Example 1"
[1698] (Claim 1)
[1699] A device for uploading performance audio;
[1700] a server for storing the performance sound source received from the terminal;
[1701] a processing means for extracting features from the performance sound source and creating a data set;
[1702] processing means for training an AI model using said dataset;
[1703] a terminal for generating and playing a pinch-hit performance using the trained AI model;
[1704] A control means for giving instructions to an automatic playing robot used in music events and theme parks;
[1705] A system including:
[1706] (Claim 2)
[1707] 2. The system according to claim 1, wherein the terminal comprises a processing means for generating a pinch-hit performance in real time based on the collected performance sound source.
[1708] (Claim 3)
[1709] 2. The system according to claim 1, wherein the server comprises a processing means for analyzing past performance sound sources and extracting features including rhythm, tempo, dynamics, and timbre.
[1710] "Example 2: Combining Emotion Engines"
[1711] (Claim 1)
[1712] A device for uploading performance audio;
[1713] a server for storing the performance sound source received from the terminal;
[1714] a processing means for extracting features such as rhythm, tempo, and timbre from the performance sound source and creating a data set;
[1715] A processing means for training an AI model using deep learning techniques using the dataset;
[1716] a terminal for adjusting a playing style using the trained AI model to generate and play a pinch-hit performance;
[1717] An emotion engine that recognizes the user's emotions in real time;
[1718] a processing means for transmitting a request to an AI model based on the user's emotion data recognized by the emotion engine, and adjusting the playing style;
[1719] A system including:
[1720] (Claim 2)
[1721] 2. The system of claim 1, wherein the terminal comprises processing means for analyzing a user's real-time facial expressions and tone of voice to detect an emotional state.
[1722] (Claim 3)
[1723] The system according to claim 1, characterized in that the server is provided with a high-precision filtering processing means for performing pre-processing such as noise removal, normalization, and clipping on the uploaded performance sound source file.
[1724] "Application example 2 when combining emotion engines"
[1725] (Claim 1)
[1726] A device for uploading performance audio;
[1727] a server for storing the performance sound source received from the terminal;
[1728] a processing means for extracting features from the performance sound source and creating a data set;
[1729] processing means for training a generative AI model using the dataset;
[1730] a terminal for generating and playing back a pinch-hit performance using the trained generative AI model;
[1731] An emotion engine that recognizes the emotions of the audience and users;
[1732] A processing means for adjusting a generative AI model based on emotion data recognized by the emotion engine;
[1733] A system including:
[1734] (Claim 2)
[1735] 2. The system according to claim 1, wherein the terminal comprises a processing means for generating a pinch-hit performance based on collected performance sound sources and emotion data acquired in real time.
[1736] (Claim 3)
[1737] 2. The system according to claim 1, wherein the server comprises a processing means for analyzing past performance sound sources and extracting features including rhythm, tempo, dynamics, and timbre. [Explanation of symbols]
[1738] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A device for uploading performance audio; a server for storing the performance sound source received from the terminal; a processing means for extracting features from the performance sound source and creating a data set; processing means for training an AI model using said dataset; a terminal for generating and playing a pinch-hit performance using the trained AI model; A system including:
2. 2. The system according to claim 1, wherein the terminal comprises a processing means for generating a pinch-hit performance in real time based on the collected performance sound source.
3. 2. The system according to claim 1, wherein the server comprises processing means for analyzing a sound source of a past performance and extracting characteristics including rhythm, tempo, dynamics, and timbre.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A