System
The system addresses the lack of personalization in music generation by allowing users to engage in realistic musical sessions with virtual players that reflect their preferences and compensates performers, enhancing user experience and system improvement through feedback-based reward distribution.
Patent Information
- Application Number
- JP2024122872
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Current music generation technology lacks the ability to reflect personal elements and life experiences of individual performers, resulting in unrealistic performance experiences for users, and there is a need for a system that appropriately compensates performers for their original data.
A system that includes means for inputting user information, managing virtual player data, selecting a virtual player, generating performance data in real time, collecting feedback, and paying rewards based on usage, utilizing a machine learning model to analyze playing styles and provide personalized musical sessions.
Enables users to enjoy realistic musical sessions with virtual players that match their preferences and compensates performers for their data, improving the system through feedback and ensuring appropriate reward distribution.
Smart Images

Figure 2026021190000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current music generation technology is limited to mechanically generating basic scales and rhythms and is unable to reflect the personal elements and life experiences of individual performers. This makes it difficult to provide a realistic performance experience when users engage in a session. To solve this problem, there is a need to use virtual players that reflect the performer's personal elements, such as their religious views, political views, and favorite books and movies, allowing users to enjoy a session that is closer to a real experience. Another challenge is to secure new sources of revenue for performers by introducing a system that appropriately pays performers who provide original data. [Means for solving the problem]
[0005] The present invention solves the problem with a system that includes a means for inputting user information, a means for managing virtual player data including performance data and personal preference information, a means for a user to select a virtual player, a means for starting a session with the selected virtual player, a means for generating performance data in real time during the session, a means for collecting feedback after the session, and a means for paying a reward to the performer according to the degree of use. The system also includes a means for analyzing the performance data of the virtual player data and identifying a playing style using a machine learning model, and a means for analyzing the user's performance data during the session and generating performance data to which the virtual player responds in real time, thereby providing a more realistic playing experience.
[0006] "User Information" refers to personal information provided by a user when they start using an application, and includes information such as their name, email address, password, favorite instrument, music genre, religious views, political views, and favorite books and movies.
[0007] "Performance data" is data that records the musical performance provided by the performer who is the basis for the virtual player, and is composed of music files, MIDI data, etc.
[0008] "Personal preference information" is information about the performer's individual preferences, such as religious views, political views, favorite books and movies, etc.
[0009] "Virtual player data" refers to data for configuring a virtual player, including performance data and personal preference information.
[0010] A "virtual player" is a virtual performer that is generated based on performance data and personal preference information and with which a user can engage in a session.
[0011] A "machine learning model" is an algorithm or statistical model with data analysis and predictive capabilities that is used to identify a virtual player's playing style and generate performance data in real time.
[0012] A "session" is a musical activity in which a user and a virtual player perform together.
[0013] "Means for generating performance data in real time" refers to a technology that instantly generates performance data for a virtual player in response to a user's performance during a session, and responds in real time.
[0014] "Means for collecting feedback" refers to a mechanism for collecting evaluations and opinions from users about the virtual player's performance after the session ends.
[0015] The "means for paying remuneration" is a system for paying remuneration to the performer who provided the original data of the virtual player based on the degree of use. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention provides a system that allows users to engage in real-time music sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[0038] A natural language description of the program's operation
[0039] The overall system processing is as follows:
[0040] 1. Enter your user information
[0041] On the device: The user launches the application and creates an account. They enter their name, email address, and password, and then enter profile information such as their favorite instrument, music genre, religious views, political views, and favorite books and movies.
[0042] Server: Saves the entered user information in the database and generates a new user ID.
[0043] 2. Virtual Player Data Management
[0044] Server: Manages performance data provided by actual performers, as well as personal preference information such as their religious views, political views, and favorite books and movies, as virtual player data.
[0045] Server: Analyzes the collected performance data and uses a machine learning model to learn the playing style.
[0046] 3. Select a Virtual Player
[0047] On the device: The user views a list of virtual players within the app, each of which includes information about their playing style and personal preferences.
[0048] Server: Provides data on the virtual players viewed by users.
[0049] Device: The user selects the virtual player that best suits their preferences.
[0050] 4. Starting a session
[0051] Terminal: Enter basic settings such as tempo and key to start a session with the user's selected virtual player.
[0052] Server: Prepares the session based on the settings you enter.
[0053] 5. Generating performance data
[0054] Server: Receives user performance data in real time and generates performance data for the virtual player to respond to based on a machine learning model.
[0055] Server: Sends the generated performance data to the user's device in real time.
[0056] 6. Providing Feedback
[0057] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[0058] Server: Based on the collected feedback data, the machine learning model is updated to improve the quality of the virtual player's performance data.
[0059] 7. Reward Distribution
[0060] Server: Collects data such as the number of times each virtual player uses the service and the duration of their use, and calculates the reward for the provider of virtual player data.
[0061] Server: Pays the calculated reward to the performer.
[0062] Specific examples
[0063] Specific examples are shown below.
[0064] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player in the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the actual session, as the user improvises on the piano, the virtual player responds to their performance in real time and plays along.
[0065] After the session, the user provides feedback on the virtual player's responses, such as "I'd like the performance to be a little more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[0066] In addition, the provider of the virtual player data is paid a fee calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes and styles.
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] Entering user information
[0070] On the device: The user launches the application and begins creating an account. The user is prompted to enter their name, email address, and password.
[0071] Device: The user is taken to a screen where they can enter their preferred instruments, music genres, religious and political views, and favorite books and movies.
[0072] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0073] Step 2:
[0074] Collection and management of virtual player data
[0075] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[0076] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0077] Step 3:
[0078] Data analysis and machine learning model creation
[0079] Server: Analyzes the collected performance data using analysis tools to extract features such as rhythm, tempo, and pitch.
[0080] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0081] Step 4:
[0082] Virtual Player Selection
[0083] Device: The user views the virtual player list screen within the app.
[0084] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0085] Device: The user selects the virtual player that best suits their preferences.
[0086] Step 5:
[0087] Session Start Settings
[0088] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0089] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0090] Step 6:
[0091] Conducting real-time sessions
[0092] Terminal: A session begins when the user presses the "Start Session" button.
[0093] Server: Receives the user's performance data in real time and extracts performance characteristics using analysis tools.
[0094] Server: Based on the extracted features and machine learning model, it generates performance data in real time that the virtual player responds to.
[0095] Server: The generated response performance data is sent to the user's device in real time.
[0096] Terminal: The user plays in real time with the virtual player.
[0097] Step 7:
[0098] Gathering feedback
[0099] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[0100] Server: Receives feedback data from users and stores it in a database.
[0101] Server: Updates the machine learning model based on the collected feedback to improve the quality of future performance data.
[0102] Step 8:
[0103] Reward calculation and payment
[0104] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0105] Server: Based on the aggregated data, calculates the reward for the provider of virtual player data.
[0106] Server: Pays the calculated reward amount according to the payment method specified by the original data provider.
[0107] Example 1
[0108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0109] The present invention aims to provide a system that automatically allocates appropriate rewards to providers of virtual player data, while providing a performance experience tailored to the user's individual preferences, real-time responses during the session, and effective use of feedback after the session, which have been difficult to achieve with conventional systems when users enjoy real-time musical sessions with virtual players that suit their preferences.
[0110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0111] In this invention, the server includes a means for inputting user information, a means for managing virtual player data including performance data and personal preference information, and a means for the user to select a virtual player, thereby enabling the user to select a virtual player according to their individual preferences.
[0112] The server includes means for starting a session with a selected virtual player, means for generating performance data in real time during the session, and means for collecting feedback after the session, thereby enabling real-time generation of performance data during the session and improvement of the quality of the next session based on the feedback.
[0113] The server further includes a means for paying rewards to performers according to the degree of use, a means for providing data of a virtual player selected by a user, a means for receiving the user's performance data in real time and generating performance data to which the virtual player responds based on a machine learning model, and a means for transmitting the generated performance data to the user's device in real time, thereby enabling appropriate distribution of rewards to providers of virtual player data and real-time musical sessions between users and virtual players.
[0114] "User information" refers to personal data entered by a user when registering with the system, specifically including name, email address, password, and profile information such as favorite instruments and music genres, religious views, political views, and favorite books and movies.
[0115] "Virtual player data" refers to performance data provided by an actual performer and data including the performer's personal preference information, such as playing style and personal tastes and preferences.
[0116] A "machine learning model" refers to an algorithm or statistical model that analyzes collected performance data and learns the playing style of a virtual player.
[0117] A "prompt sentence" is a sentence containing specific instructions or questions to be input into the generated AI model, and describes a specific request from the user or system.
[0118] "Performance data" is music data generated by a user or a virtual player, and is data that electronically records a musical performance, such as an audio file or MIDI data.
[0119] "Real-time response" means that the virtual player responds instantly to the performance input by the user, generating music data in real time.
[0120] "Feedback" refers to opinions and evaluations provided by users after a session has ended, and is used to improve the system and adjust the virtual player's playing style.
[0121] "Reward allocation" refers to the process of calculating and paying rewards to providers of virtual player data, and determining the amount of reward based on the degree of usage and feedback.
[0122] The present invention provides a system that allows users to engage in real-time musical sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[0123] Hardware and software used
[0124] The system is implemented using the following hardware and software:
[0125] Server: Manages user information and virtual player data, generates performance data, and collects feedback.
[0126] Terminal: The device (smartphone, tablet, PC) on which the user operates the application.
[0127] Generative AI model: A model that uses machine learning algorithms to learn playing styles and generate responsive performance data in real time.
[0128] Database: Storage for user information, virtual player data, performance data, and feedback.
[0129] Specific operation of the system
[0130] 1. Entering user information: A user launches an application on their device and creates an account, entering their name, email address, password, favorite instrument and music genre, religious views, political views, and favorite books and movies.
[0131] 2. Virtual Player Data Management: The server collects performance data provided by real performers and their personal preferences. The collected data is analyzed using a machine learning model to learn their playing style.
[0132] 3. Virtual player selection: The user browses the list of virtual players on the device, checks their playing style and preferences, and selects the virtual player that best suits their preferences.
[0133] 4. Starting a session: To start a session, the user inputs settings such as tempo and key into the terminal. The server prepares the session based on the settings.
[0134] 5. Performance data generation: When a user starts playing on their device, the data is sent to the server in real time. The server uses a machine learning model to generate performance data that the virtual player responds to, and sends it to the device in real time.
[0135] 6. Providing feedback: After the session ends, the user enters their feedback into the device. The server collects the feedback and updates the machine learning model to improve the quality of the next session.
[0136] 7. Reward distribution: The server aggregates the usage status of each virtual player, calculates and pays rewards to the performers according to the degree of usage.
[0137] Specific examples
[0138] For example, if a user wants to play jazz piano, they can create an account by entering profile information such as "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the session, the user can improvise on the piano, and the virtual player will respond and play in real time. After the session ends, the user can provide feedback, such as "I'd like a more rhythmic performance." The server uses this feedback to update the AI model and provide a better playing experience in the next session.
[0139] This invention allows users to enjoy realistic music sessions with virtual players that suit their individual tastes and styles, and the provider of virtual player data is paid a fee according to the degree of use.
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Processing flow
[0142] Step 1: Enter your user information
[0143] 1.1 Device: The user downloads and installs the application.
[0144] Input: Application file
[0145] Output: Installed applications
[0146] Specific behavior: The user selects an application from the store and begins downloading it.
[0147] 1.2 Device: The user launches the app and creates an account.
[0148] Input: Name, Email Address, Password
[0149] Output: User account data
[0150] Specific operation: The user launches the application and enters the necessary information on the account creation screen.
[0151] 1.3 Device: User enters detailed profile information.
[0152] Input: Favorite instrument, music genre, religious views, political views, favorite books and movies
[0153] Output: Detailed user profile data
[0154] Specific operation: The user accesses the profile setting screen and enters each item.
[0155] 1.4 Server: Stores the entered user information in a database.
[0156] Input: User account data, detailed user profile data
[0157] Output: User information stored in the database
[0158] Specific operation: The server receives the data sent from the terminal and stores it in a database.
[0159] Step 2: Managing Virtual Player Data
[0160] 2.1 Server: Collects performance data and personal preference information from live performers.
[0161] Input: performance data, performer preference information
[0162] Output: Virtual player data
[0163] Specific operation: The performer uploads data through a dedicated interface.
[0164] 2.2 Server: Analyzes the collected performance data and learns using machine learning models.
[0165] Input: Performance data
[0166] Output: A trained machine learning model
[0167] Specific operation: The server processes the performance data, extracts characteristics of the performance style, and trains the model.
[0168] Step 3: Select a Virtual Player
[0169] 3.1 Terminal: The user browses the list of virtual players.
[0170] Input: None
[0171] Output: Virtual player list
[0172] Specific operation: A list of virtual players will be displayed on the device.
[0173] 3.2 Terminal: The user selects the preferred virtual player.
[0174] Input: User selection information
[0175] Output: Selected virtual player data
[0176] Specific operation: The user selects a virtual player on the screen and presses the select button.
[0177] 3.3 Server: Records the user's selections and stores them in a database.
[0178] Input: Selected Virtual Player Data
[0179] Output: Recorded selection data
[0180] Specific operation: The server receives the selection data sent from the terminal and stores it in a database.
[0181] Step 4: Starting a session
[0182] 4.1 Terminal: User enters session settings.
[0183] Input: Session settings data such as tempo, key, etc.
[0184] Output: Session initiation request
[0185] Specific operations: The user enters the required information on the session setting screen and presses the start button.
[0186] 4.2 Server: Prepares the session.
[0187] Input: Session configuration data
[0188] Output: Session ready notification
[0189] Specific operation: The server allocates resources based on the configuration and notifies the server that they are ready.
[0190] Step 5: Generate performance data
[0191] 5.1 Terminal: The user starts playing.
[0192] Input: User's performance data
[0193] Output: Real-time performance data
[0194] Specific operation: The user starts playing an instrument on the device, and data is recorded in real time.
[0195] 5.2 Server: Receives and analyzes user performance data.
[0196] Input: User's performance data
[0197] Output: Virtual player response data
[0198] Specific operation: Based on the performance data received by the server, a machine learning model is applied to generate response data.
[0199] 5.3 Server: Sends the generated performance data to the user.
[0200] Input: Virtual player response data
[0201] Output: Performance data to the user's device
[0202] Specific operation: The server streams data generated in real time to the device.
[0203] Step 6: Provide feedback
[0204] 6.1 Terminal: User enters feedback after the session ends.
[0205] Input: Feedback data
[0206] Output: Feedback sent
[0207] Specific operation: The user accesses the feedback input screen and enters their evaluation and opinion.
[0208] 6.2 Server: Collects and analyzes feedback data.
[0209] Input: Feedback data
[0210] Output: Updated machine learning model
[0211] Specific operation: The server updates the model based on the feedback and aims to improve the quality next time.
[0212] Step 7: Distributing rewards
[0213] 7.1 Server: Aggregates usage of Virtual Player data.
[0214] Input: Number of uses, duration of use, feedback data
[0215] Output: Aggregated data
[0216] Specific operation: The server aggregates the data and generates statistics.
[0217] 7.2 Server: Calculates the remuneration and pays the performers.
[0218] Input: Aggregate data
[0219] Output: Reward payment data
[0220] Specific operation: The server calculates the reward based on the aggregated results and pays the performer.
[0221] The above is the specific flow of processing of the program in this system.
[0222] (Application example 1)
[0223] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0224] Conventional music session systems with virtual players have had problems such as users finding a player that matches their preferences and a lack of real-time interaction, resulting in low satisfaction. Furthermore, there are insufficient means for efficiently collecting user feedback and using it to improve the system. Furthermore, there is no system in place for appropriately distributing rewards to providers of virtual player data. To solve these problems, it is necessary to provide a system that allows users to enjoy real-time sessions with virtual players that match their preferences and that continuously improves based on feedback.
[0225] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0226] In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for a user to select a virtual player, means for starting a session with the selected virtual player, means for generating performance data in real time during the session, means for collecting feedback after the session, means for paying a reward to the performer according to the degree of use, and means for conducting the session via an application installed on a smartphone or a head-mounted display. This enables users to enjoy real-time musical sessions with virtual players that suit their preferences, enables continuous improvement of the system based on feedback, and enables appropriate distribution of rewards to providers of virtual player data.
[0227] The "means for inputting user information" is a device or program that provides an interface for a user to input their profile information.
[0228] "Means for managing virtual player data" refers to a device or program that collects and stores the performance data and personal preference information of performers, and uses this information to individually manage virtual players.
[0229] The "means for a user to select a virtual player" is a device or program that provides an interface for a user to select a virtual player that suits his / her preferences from a list of provided virtual players.
[0230] A "means for initiating a session" is a device or program that prepares to initiate a musical session in real time with selected virtual players.
[0231] The "means for generating performance data in real time" is a device or program for generating performance data to which the virtual player responds in real time, based on the performance data of the user.
[0232] A "means for collecting feedback" is a device or program that collects opinions and impressions from users after a session ends, analyzes them, and uses them to improve the system.
[0233] The "means for paying a remuneration to a performer" is a device or program for calculating and paying a remuneration to the provider of virtual player data based on the degree of use of the virtual player data.
[0234] "Means for conducting a session via an application installed on a smartphone or head-mounted display" refers to a device or program that allows a user to conduct a musical session with a virtual player via an application installed on a smartphone or head-mounted display.
[0235] The present invention relates to a system for allowing a user to have a real-time music session with a virtual player that matches the user's preferences. The program for implementing this system is as follows.
[0236] Entering user information
[0237] The user launches a dedicated application installed on a smartphone or head-mounted display and enters information such as their name, email address, password, musical preferences (instruments, genres, religious views, political views, favorite books and movies), etc. This user information is then sent to the server and stored in a database. The server then generates a new user ID and associates it with the user information.
[0238] Managing Virtual Player Data
[0239] The server collects and centrally manages performance data provided by performers and their personal preferences. The collected performance data is analyzed and machine learning models are used to learn their playing style.
[0240] Virtual Player Selection
[0241] Through the application, the user can view a list of virtual players displayed on a panel. This list includes information about each virtual player's playing style and personal preferences. The user selects a virtual player that matches their preferences, and the server provides the data of the selected virtual player.
[0242] Starting a Session
[0243] To start a session with a selected virtual player, the user inputs the session's basic settings (tempo, key, etc.). The server prepares the session based on the input settings and processes requests for real-time performance.
[0244] Generating performance data
[0245] During a session, the user's performance data is sent in real time to the server, which uses a machine learning model to generate performance data for a virtual player that responds to the user's performance and immediately returns it to the user. This process establishes a real-time session between the user and the virtual player.
[0246] Providing Feedback
[0247] After the session, the user enters feedback on the virtual player's performance in the application, which is then sent to and collected by the server. The server uses this feedback to update the machine learning model and continuously improve the quality of the virtual player's performance data.
[0248] Reward Distribution
[0249] The server tally up the degree of use of the virtual player data (such as the number of uses and duration of use), calculates the reward for the provider, and pays the reward to the provider of the virtual player data based on the calculation results.
[0250] Specific examples
[0251] For example, consider a case where a user likes jazz piano and starts a session with a tempo of 120 BPM and the key of C major. The user selects a jazz piano virtual player and begins improvising. The virtual player then responds to the user's performance in real time to provide accompaniment. After the session ends, the user can provide feedback such as "I would like the performance to be a little more rhythmic," and this feedback will be used to improve the system. The following is an example of a prompt sentence.
[0252] "You are a virtual jazz pianist. If a user plays jazz piano in real time, you want to play along based on that performance. The user starts playing in the key of C major at a tempo of 120 BPM. How would you perform?"
[0253] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0254] Step 1:
[0255] A user launches an application and enters their user information. The user enters profile information such as their name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies, etc. The entered information is sent to the server and stored in a database. The server generates a new user ID and associates it with the user information.
[0256] Input: Name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies
[0257] Output: User ID, user information stored in the database
[0258] Step 2:
[0259] The server collects performance data and personal preference information provided by the performers, and centrally manages this data as virtual player data. The server analyzes this data and uses a machine learning model (e.g., PerformanceModel) to learn each performer's playing style.
[0260] Input: performance data, performer's personal preferences
[0261] Output: The playing style learned by the machine learning model
[0262] Step 3:
[0263] The user browses through the application a list of virtual players. During browsing, the server provides information about the virtual players' playing styles and personal preferences. The user selects a virtual player that best suits their preferences, and the selection is sent to the server.
[0264] Input: User's virtual player selection
[0265] Output: Data of the selected virtual player
[0266] Step 4:
[0267] To start a session with a selected virtual player, the user inputs basic settings such as tempo and key, and the server receives this information and prepares the session.
[0268] Input: Session settings such as tempo, key, etc.
[0269] Output: Session ready state
[0270] Step 5:
[0271] During a session, the user's performance data is sent in real time to the server, which processes the data and uses machine learning models to generate performance data that the virtual player responds to in real time. The generated performance data is then sent to the user's device.
[0272] Input: Real-time user performance data
[0273] Output: Virtual player response performance data
[0274] Step 6:
[0275] After the session ends, the user provides feedback on the virtual player's performance, which is then sent to the server, where it updates the machine learning model and improves the quality of the virtual player's performance data.
[0276] Input: User feedback
[0277] Output: An improved machine learning model
[0278] Step 7:
[0279] The server compiles the degree of use of each virtual player data (number of times used and duration of use), calculates a reward based on this data, and pays the reward to the provider of the virtual player data based on the calculation results.
[0280] Input: Virtual player data usage
[0281] Output: Calculated reward, reward payment completed
[0282] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0283] The present invention is a system that allows a user to engage in a musical session with a virtual player in real time, and further recognizes the user's emotions and adjusts the performance style accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[0284] A natural language description of the program's operation
[0285] The overall system processing is as follows:
[0286] 1. Enter your user information
[0287] On the device: The user launches the application and begins creating an account. The user enters their name, email address, and password.
[0288] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[0289] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0290] 2. Collection and Management of Virtual Player Data
[0291] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[0292] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0293] 3. Data analysis and machine learning model creation
[0294] Server: Analyzes performance data and extracts features such as rhythm, tempo, and pitch.
[0295] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0296] 4. Selecting a Virtual Player
[0297] Device: User views the list of virtual players within the app.
[0298] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0299] Device: The user selects the virtual player that best suits their preferences.
[0300] 5. Session Start Settings
[0301] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0302] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0303] 6. Conducting real-time sessions
[0304] Terminal: A session begins when the user presses the "Start Session" button.
[0305] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[0306] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0307] Server: Adjusts the virtual player's playing style in real time based on the emotion data from the emotion engine and the features of the performance data.
[0308] Server: The generated response performance data is sent to the user's device in real time.
[0309] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[0310] 7. Gathering Feedback
[0311] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[0312] Server: Collected feedback data is stored in a database and used for future model updates.
[0313] 8. Calculation and Payment of Rewards
[0314] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0315] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0316] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0317] Specific examples
[0318] For example, if a user wants to play jazz piano, they enter the following user information: "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. As the user plays the piano during the session, the emotion engine recognizes the emotion of joy from the user's facial expressions and voice, and adjusts the virtual player's playing style to be brighter and more energetic.
[0319] After the session, the user provides feedback on the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[0320] In addition, the provider of the virtual player data is paid a reward calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes, style, and even emotional state.
[0321] The processing flow will be explained below.
[0322] Step 1:
[0323] Entering user information
[0324] On the device: The user launches the application and proceeds to the account creation screen, where they enter their name, email address, and password.
[0325] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[0326] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0327] Step 2:
[0328] Collection and management of virtual player data
[0329] Server: Collects performance data (music files and MIDI data) from actual performers, as well as personal preference information such as religious views, political views, and favorite books and movies.
[0330] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0331] Step 3:
[0332] Data analysis and machine learning model creation
[0333] Server: Analyzes the performance data using analysis tools and extracts features such as rhythm, tempo, and pitch.
[0334] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0335] Step 4:
[0336] Virtual Player Selection
[0337] Device: The user views the virtual player list screen within the app.
[0338] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0339] Device: The user selects the virtual player that best suits their preferences.
[0340] Step 5:
[0341] Session Start Settings
[0342] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0343] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0344] Step 6:
[0345] Conducting real-time sessions
[0346] Terminal: A session begins when the user presses the "Start Session" button.
[0347] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[0348] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0349] Server: Adjusts the virtual player's playing style in real time based on the emotion data obtained from the emotion engine and the features of the performance data.
[0350] Server: The generated response performance data is sent to the user's device in real time.
[0351] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[0352] Step 7:
[0353] Gathering feedback
[0354] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[0355] Server: Receives feedback data from users and stores it in a database.
[0356] Server: Updates the machine learning model based on the collected feedback and uses it to improve the model in the future.
[0357] Step 8:
[0358] Reward calculation and payment
[0359] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0360] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0361] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0362] Example 2
[0363] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0364] In conventional music session systems, it has been difficult to provide a system that allows users to change their playing style in response to their emotions when engaging in a real-time music session. It has also been difficult to select and adjust virtual players to suit the user's personal preferences or specific music genres. This has led to problems such as a poor user experience and low satisfaction.
[0365] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0366] In this invention, the server includes means for recognizing a user's emotional state and dynamically adjusting a performance style based on the emotional state, means for identifying a performance style using a machine learning model, and means for generating responsive performance data in real time, thereby enabling a user to enjoy an interactive musical session with a virtual player in real time according to their emotions and preferences.
[0367] "User information" refers to data about individuals who use the system, such as their name, email address, password, favorite instrument, music genre, religious views, political views, favorite books and movies, etc.
[0368] "Virtual player data" is digital performance data with a variety of musical styles, including performance data collected from actual performers and personal preference information.
[0369] "Performance data" refers to data such as music files and MIDI data that digitally represents a specific musical performance.
[0370] "Personal preference information" is information about the performer's or user's individual preferences, such as religious views, political views, favorite books and movies, etc.
[0371] "Emotional state recognition" means determining a user's current emotions in real time based on data collected from the user's voice, facial expressions, biometric sensors, etc.
[0372] "Performance style" refers to the characteristics of performance based on musical features such as rhythm, tempo, pitch, and dynamics.
[0373] "Dynamic adjustment" means instantly modifying or changing the system's behavior or output in response to input data (e.g., user emotions or performance data) that changes in real time.
[0374] A "machine learning model" is an algorithm that learns patterns from large amounts of data and uses them to make predictions and classifications.
[0375] "Feedback" refers to data regarding ratings and opinions provided by users regarding the system and virtual players.
[0376] "Remuneration" means money or other benefits paid to a provider of virtual player data or in exchange for use of the system.
[0377] "Responding in real time" means instantly responding to user input and performance, and generating appropriate performance data and system operations.
[0378] The system of the present invention allows a user to engage in a musical session with a virtual player in real time, and further allows the user's emotions to be recognized and the performance style to be adjusted accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[0379] In terms of specific hardware and software, the devices used by users are computer devices such as smartphones, tablets, and PCs. The servers use cloud computing services to store databases and machine learning models. Programming languages and libraries such as Python and TensorFlow are used to train the machine learning models. In addition, the emotion engine uses voice recognition and facial recognition technologies, so cloud services such as Amazon Rekognition and Google Cloud Speech-to-Text are used.
[0380] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a C major key. As the user plays the piano during the session, the emotion engine recognizes joyful emotions from the user's facial expressions and voice, adjusting the virtual player's playing style to be more cheerful and lively. After the session ends, the user provides feedback about the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide more appropriate performances in future sessions. The provider of the virtual player data is also paid a reward calculated based on the degree of use of the session.
[0381] An example of a prompt is as follows:
[0382] "Jazz piano, I'm a Buddhist and a liberal, my favorite book is The Alchemist, and my favorite movie is Inception."
[0383] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0384] Step 1: Enter your user information
[0385] On the device: The user launches the application, enters their name, email address, and password in the form that appears on the screen, and then clicks the "Next" button.
[0386] Input: Name, Email Address, Password
[0387] Output: Sends the user's basic information to the server
[0388] Device: The next screen displays fields for the user to enter their favorite instrument, music genre, religious views, political views, and favorite books and movies. The user fills in each field and clicks the "Submit" button.
[0389] Input: musical instruments, music genres, religious views, political views, favorite books, movies
[0390] Output: Sends user details to server
[0391] Server: Receives the entered information, stores it in a database, generates a new user ID, and returns that information to the device.
[0392] Input: User information
[0393] Data processing: saving to database, generating new user ID
[0394] Output: New user ID
[0395] Step 2: Collecting and Managing Virtual Player Data
[0396] Server: Provides an interface with the functionality to collect performance data (music files and MIDI data) from real performers.
[0397] Input: Performance data
[0398] Output: Collected performance data
[0399] Server: Provides and collects a form for performers to enter their preferences (religious views, political views, favorite books and movies, etc.).
[0400] Input: Player preference information
[0401] Output: Collected preference information
[0402] Server: Collected performance data and performer preference information are stored in a database and associated with individual performer IDs.
[0403] Input: Collected performance data, preference information
[0404] Data processing: Save to database, associate with performer ID
[0405] Output: Data linked to performer ID
[0406] Step 3: Data analysis and machine learning model creation
[0407] Server: Executes data analysis algorithms to extract features such as rhythm, tempo, and pitch from performance data.
[0408] Input: Performance data
[0409] Data Calculation: Feature Extraction
[0410] Output: Extracted features
[0411] Server: Using the extracted features and personal preference information as input, the generative AI model is used to train a machine learning model to identify playing styles.
[0412] Input: extracted features, personal preference information
[0413] Data Computing: Training Machine Learning Models
[0414] Output: A trained machine learning model
[0415] Step 4: Select a Virtual Player
[0416] Device: The user accesses the "Virtual Player Selection" screen within the app.
[0417] Input: A view request from a user
[0418] Output: List of virtual players
[0419] Server: Provides virtual player data (playing style, preferences, sample performance) to the device.
[0420] Input: User ID
[0421] Data processing: filtering virtual players according to user preferences
[0422] Output: Virtual player data
[0423] Terminal: The user browses the provided data and selects the virtual player that best suits their preferences.
[0424] Input: View virtual player data
[0425] Output: Selected Virtual Player
[0426] Step 5: Session Start Settings
[0427] Device: The user goes to the "Session Start Settings" screen and enters basic settings such as tempo, key, and session time.
[0428] Input: Tempo, Key, Session Time
[0429] Output: Configuration information
[0430] Server: Receives this configuration information and begins preparing a session with the selected virtual player.
[0431] Input: Setting information, User ID, Virtual Player ID
[0432] Data processing: session preparation
[0433] Output: Session ready notification
[0434] Step 6: Conduct a real-time session
[0435] Terminal: A session begins when the user presses the "Start Session" button.
[0436] Input: Session start request from user
[0437] Output: Session start notification
[0438] Server: Receives the user's performance data in real time and extracts features such as rhythm, tempo, and pitch.
[0439] Input: User performance data
[0440] Data calculations: Real-time feature extraction
[0441] Output: Extracted features
[0442] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0443] Input: Voice data, facial expression data, biometric sensor data
[0444] Data Computing: Emotional State Recognition
[0445] Output: Emotion data
[0446] Server: Adjusts the virtual player's playing style in real time based on the features of emotional data and performance data.
[0447] Input: Emotion data, performance data features
[0448] Data calculation: Adjusting playing style
[0449] Output: Adjusted performance data
[0450] Server: The generated response performance data is sent to the user's device in real time.
[0451] Input: Adjusted performance data
[0452] Output: Response performance data
[0453] Terminal: Based on the response performance data generated by the user, the user continues to play in real time together with the virtual player.
[0454] Input: Response performance data
[0455] Output: Real-time collaboration
[0456] Step 7: Gather feedback
[0457] Terminal: After the session ends, a form will appear where the user can enter feedback about the virtual player's performance. The user enters their feedback and clicks the "Submit" button.
[0458] Input: Feedback information
[0459] Output: Feedback information sent to server
[0460] Server: Receives feedback data and stores it in a database. This feedback data is used to update the machine learning model in the future.
[0461] Input: Feedback data
[0462] Data processing: Saving to database
[0463] Output: Stored feedback data
[0464] Step 8: Calculating and paying compensation
[0465] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0466] Input: Number of uses, duration, feedback
[0467] Data calculation: Aggregation work
[0468] Output: Aggregated data
[0469] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0470] Input: Aggregate data
[0471] Data calculation: Reward calculation
[0472] Output: Calculated reward data
[0473] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0474] Input: Remuneration data, payment method
[0475] Output: Payment processing completion notification
[0476] (Application example 2)
[0477] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0478] Conventional music session systems have had problems in that they make it difficult for users to engage in emotionally-based interactive sessions with virtual players, and lack the functionality to dynamically adjust the virtual player's playing style in response to real-time changes in the user's emotions. Another issue is accurately recognizing the user's emotions and providing feedback accordingly.
[0479] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for the user to select a virtual player, means for generating performance data in real time during a session, means for collecting feedback after the session, means for paying a reward to the data provider according to the degree of use, and means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the playing style of the virtual player based on that. This allows the user to enjoy an interactive music session with a virtual player that responds to their emotions in real time.
[0480] The "means for inputting user information" is an interface that allows a user to input their basic information, musical preferences, etc. into the system.
[0481] The "means for managing virtual player data including performance data and personal preference information" is a system that collects performance data and preference information of virtual players and manages them in a database.
[0482] The "means for the user to select a virtual player" is an interface that allows the user to select one of the virtual players displayed in the system.
[0483] The "means for generating performance data in real time during a session" is a system that generates and analyzes performance data of the user and virtual players in real time during a session.
[0484] The "means for collecting feedback after a session" is an interface for collecting feedback from the user after the session has ended.
[0485] The "means for paying a reward to a data provider according to the degree of use" is a system for calculating and paying a reward to a data provider of a virtual player according to the degree of use.
[0486] "Means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the virtual player's playing style based on that" is a system that uses an emotion engine to recognize the user's emotions and dynamically adjusts the virtual player's playing style according to the user's emotional state.
[0487] This invention is a system that allows a user to engage in a real-time musical session with a virtual player, and further recognizes the user's emotions and adjusts the playing style accordingly. The system includes the following main components:
[0488] 1. Enter your user information
[0489] Hardware: Smartphones, tablets, computers
[0490] Software: Applications
[0491] What happens: A user creates an account and enters basic information such as name, email address, password, preferred instrument and music genre, etc. This information is stored in a cloud database and a new user ID is generated.
[0492] 2. Virtual Player Data Management
[0493] Hardware: Server
[0494] Software: Cloud Management System
[0495] Processing: Performance data (music files and MIDI data) collected from actual performers and personal preference information are stored in a database, and data about each performer is managed.
[0496] 3. Emotion Recognition and Data Analysis
[0497] Hardware: smartphone facial recognition cameras, microphones, and biometric sensors
[0498] Software: Emotion recognition engine (open source emotion recognition software, e.g., OpenFace)
[0499] Processing: Emotional data is collected from the user's facial expressions, voice, and biometric sensors, and analyzed in real time by an emotion recognition engine.
[0500] 4. Conducting real-time sessions
[0501] Hardware: Smartphone
[0502] Software: In-app real-time music session function
[0503] Processing: Receives and analyzes the user's performance data in real time to extract features. Dynamically adjusts the virtual player's performance style based on the data from the emotion engine. The generated response performance data is sent to the user's device in real time, allowing the user to engage in a session with the virtual player.
[0504] 5. Gathering Feedback
[0505] Hardware: Smartphone
[0506] Software: Feedback collection feature
[0507] Processing: After the session ends, the user provides feedback on the virtual player's performance. The collected feedback data is stored in a cloud database and used to update the machine learning model.
[0508] 6. Calculation and Payment of Rewards
[0509] Hardware: Server
[0510] Software: Remuneration calculation system
[0511] Processing: Compensation is calculated based on the number of uses, duration of use, and user feedback for each virtual player, and paid to the data provider.
[0512] Examples:
[0513] For example, if a user wants to play jazz guitar, they enter their user information as "jazz guitar, Catholic, liberal, favorite book is 'Norwegian Wood,' and favorite movie is 'La La Land.'" They then select a virtual jazz guitarist within the app and enter settings to start a session with a tempo of 140 BPM and a key of F major. During the session, the emotion engine recognizes the user's excitement level from their facial expressions and voice, adjusting the virtual player's performance to be more energetic. After the session ends, the user provides feedback such as "I'd like a more rhythmic performance," and the server uses this feedback to update the machine learning model.
[0514] Example prompt sentence:
[0515] The user begins a 140 BPM jazz session in F major with a virtual jazz band. The user sets their favorite music genre as jazz, is Catholic and liberal, and their favorite book is "Norwegian Wood," and their favorite movie is "La La Land." During the session, the system recognizes emotions from the user's smile and voice, adjusting the performance to be more energetic.
[0516] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0517] Step 1:
[0518] A user launches an application and creates an account. The user enters basic information such as their name, email address, password, favorite instrument, and music genre. The information entered is sent from the device to the server and stored in a cloud database. A new user ID is generated and returned to the device. The input for this step is the user information, and the output is the new user ID.
[0519] Step 2:
[0520] The server manages data about each performer by storing performance data (music files and MIDI data) and personal preference information collected from the actual performers in a database. The input is the performance data and preference information, and the output is the virtual player data stored in the database.
[0521] Step 3:
[0522] The user browses a list of virtual players in the app and selects one that suits their preferences. The server provides virtual player data (playing style, preferences, sample performance, etc.) to the device. The input is the virtual player data, and the output is the virtual player selected by the user.
[0523] Step 4:
[0524] The user proceeds to the session start screen and inputs basic settings such as tempo, key, etc. The device sends the input settings to the server and begins preparing for a session with the selected virtual player. The input is session setting information, and the output is a message indicating that the session is ready.
[0525] Step 5:
[0526] A session begins when the user presses the "Start Session" button. The device collects the user's performance data in real time and sends it to the server. The server analyzes the received performance data and extracts features such as rhythm and tempo. The input is the user's performance data, and the output is the analyzed features.
[0527] Step 6:
[0528] The emotion engine collects emotional data from the user's voice, facial expression, and biometric sensors to recognize the emotional state. The emotion engine analyzes the emotional state based on this data and identifies the type of emotion (happiness, sadness, excitement, etc.). The input is voice data, facial expression data, and biometric data, and the output is the recognized emotional state.
[0529] Step 7:
[0530] The server adjusts the virtual player's performance style in real time based on the emotion data from the emotion engine and the performance data features. This allows the performance style to reflect the user's emotions. The input is emotion data and performance data features, and the output is the adjusted performance data.
[0531] Step 8:
[0532] The server transmits the response performance data generated in real time to the user's terminal, and the user engages in a session in real time with the virtual player. The input is the adjusted performance data, and the output is the performance data transmitted to the user's terminal.
[0533] Step 9:
[0534] After the session ends, the user provides feedback on the virtual player's performance. The device sends the feedback data to the server and stores it in a database. The input is the feedback data, and the output is the feedback data stored in the database.
[0535] Step 10:
[0536] The server calculates rewards for each virtual player based on the number of uses, duration of use, and user feedback. The calculated rewards are paid to the data provider. The inputs are usage data and feedback data, and the output is the calculated rewards.
[0537] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0538] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0539] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0540] [Second embodiment]
[0541] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0542] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0543] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0544] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0545] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0546] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0547] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0548] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0549] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0550] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0551] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0552] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0553] The present invention provides a system that allows users to engage in real-time music sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[0554] A natural language description of the program's operation
[0555] The overall system processing is as follows:
[0556] 1. Enter your user information
[0557] On the device: The user launches the application and creates an account. They enter their name, email address, and password, and then enter profile information such as their favorite instrument, music genre, religious views, political views, and favorite books and movies.
[0558] Server: Saves the entered user information in the database and generates a new user ID.
[0559] 2. Virtual Player Data Management
[0560] Server: Along with performance data provided by actual performers, the server manages personal preference information such as the performers' religious views, political views, and favorite books and movies as virtual player data.
[0561] Server: Analyzes the collected performance data and uses a machine learning model to learn the playing style.
[0562] 3. Select a Virtual Player
[0563] On the device: The user views a list of virtual players within the app, each of which includes information about their playing style and personal preferences.
[0564] Server: Provides data on the virtual players viewed by users.
[0565] Device: The user selects the virtual player that best suits their preferences.
[0566] 4. Starting a session
[0567] Terminal: Enter basic settings such as tempo and key to start a session with the user's selected virtual player.
[0568] Server: Prepares the session based on the settings you enter.
[0569] 5. Generating performance data
[0570] Server: Receives user performance data in real time and generates performance data for the virtual player to respond to based on a machine learning model.
[0571] Server: Sends the generated performance data to the user's device in real time.
[0572] 6. Providing Feedback
[0573] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[0574] Server: Based on the collected feedback data, the machine learning model is updated to improve the quality of the virtual player's performance data.
[0575] 7. Reward Distribution
[0576] Server: Collects data such as the number of times each virtual player uses the service and the duration of their use, and calculates the reward for the provider of virtual player data.
[0577] Server: Pays the calculated reward to the performer.
[0578] Specific examples
[0579] Specific examples are shown below.
[0580] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player in the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the actual session, as the user improvises on the piano, the virtual player responds to their performance in real time and plays along.
[0581] After the session, the user provides feedback on the virtual player's responses, such as "I'd like the performance to be a little more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[0582] In addition, the provider of the virtual player data is paid a fee calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes and styles.
[0583] The processing flow will be explained below.
[0584] Step 1:
[0585] Entering user information
[0586] On the device: The user launches the application and begins creating an account. The user is prompted to enter their name, email address, and password.
[0587] Device: The user is taken to a screen where they can enter their preferred instruments, music genres, religious and political views, and favorite books and movies.
[0588] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0589] Step 2:
[0590] Collection and management of virtual player data
[0591] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[0592] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0593] Step 3:
[0594] Data analysis and machine learning model creation
[0595] Server: Analyzes the collected performance data using analysis tools to extract features such as rhythm, tempo, and pitch.
[0596] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0597] Step 4:
[0598] Virtual Player Selection
[0599] Device: The user views the virtual player list screen within the app.
[0600] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0601] Device: The user selects the virtual player that best suits their preferences.
[0602] Step 5:
[0603] Session Start Settings
[0604] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0605] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0606] Step 6:
[0607] Conducting real-time sessions
[0608] Terminal: A session begins when the user presses the "Start Session" button.
[0609] Server: Receives the user's performance data in real time and extracts performance characteristics using analysis tools.
[0610] Server: Based on the extracted features and machine learning model, it generates performance data in real time that the virtual player responds to.
[0611] Server: The generated response performance data is sent to the user's device in real time.
[0612] Terminal: The user plays in real time with the virtual player.
[0613] Step 7:
[0614] Gathering feedback
[0615] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[0616] Server: Receives feedback data from users and stores it in a database.
[0617] Server: Updates the machine learning model based on the collected feedback to improve the quality of future performance data.
[0618] Step 8:
[0619] Reward calculation and payment
[0620] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0621] Server: Based on the aggregated data, calculates the reward for the provider of virtual player data.
[0622] Server: Pays the calculated reward amount according to the payment method specified by the original data provider.
[0623] Example 1
[0624] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0625] The present invention aims to provide a system that automatically allocates appropriate rewards to providers of virtual player data, while providing a performance experience tailored to the user's individual preferences, real-time responses during the session, and effective use of feedback after the session, which have been difficult to achieve with conventional systems when users enjoy real-time musical sessions with virtual players that suit their preferences.
[0626] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0627] In this invention, the server includes a means for inputting user information, a means for managing virtual player data including performance data and personal preference information, and a means for the user to select a virtual player, thereby enabling the user to select a virtual player according to their individual preferences.
[0628] The server includes means for starting a session with a selected virtual player, means for generating performance data in real time during the session, and means for collecting feedback after the session, thereby enabling real-time generation of performance data during the session and improvement of the quality of the next session based on the feedback.
[0629] The server further includes a means for paying rewards to performers according to the degree of use, a means for providing data of a virtual player selected by a user, a means for receiving the user's performance data in real time and generating performance data to which the virtual player responds based on a machine learning model, and a means for transmitting the generated performance data to the user's device in real time, thereby enabling appropriate distribution of rewards to providers of virtual player data and real-time musical sessions between users and virtual players.
[0630] "User information" refers to personal data entered by a user when registering with the system, specifically including name, email address, password, and profile information such as favorite instruments and music genres, religious views, political views, and favorite books and movies.
[0631] "Virtual player data" refers to performance data provided by an actual performer and data including the performer's personal preference information, such as playing style and personal tastes and preferences.
[0632] A "machine learning model" refers to an algorithm or statistical model that analyzes collected performance data and learns the playing style of a virtual player.
[0633] A "prompt sentence" is a sentence containing specific instructions or questions to be input into the generated AI model, and describes a specific request from the user or system.
[0634] "Performance data" is music data generated by a user or a virtual player, and is data that electronically records a musical performance, such as an audio file or MIDI data.
[0635] "Real-time response" means that the virtual player responds instantly to the performance input by the user, generating music data in real time.
[0636] "Feedback" refers to opinions and evaluations provided by users after a session has ended, and is used to improve the system and adjust the virtual player's playing style.
[0637] "Reward allocation" refers to the process of calculating and paying rewards to providers of virtual player data, and determining the amount of reward based on the degree of usage and feedback.
[0638] The present invention provides a system that allows users to engage in real-time musical sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[0639] Hardware and software used
[0640] The system is implemented using the following hardware and software:
[0641] Server: Manages user information and virtual player data, generates performance data, and collects feedback.
[0642] Terminal: The device (smartphone, tablet, PC) on which the user operates the application.
[0643] Generative AI model: A model that uses machine learning algorithms to learn playing styles and generate responsive performance data in real time.
[0644] Database: Storage for user information, virtual player data, performance data, and feedback.
[0645] Specific operation of the system
[0646] 1. Entering user information: A user launches an application on their device and creates an account, entering their name, email address, password, favorite instrument and music genre, religious views, political views, and favorite books and movies.
[0647] 2. Virtual Player Data Management: The server collects performance data provided by real performers and their personal preferences. The collected data is analyzed using a machine learning model to learn their playing style.
[0648] 3. Virtual player selection: The user browses the list of virtual players on the device, checks their playing style and preferences, and selects the virtual player that best suits their preferences.
[0649] 4. Starting a session: To start a session, the user inputs settings such as tempo and key into the terminal. The server prepares the session based on the settings.
[0650] 5. Performance data generation: When a user starts playing on their device, the data is sent to the server in real time. The server uses a machine learning model to generate performance data that the virtual player responds to, and sends it to the device in real time.
[0651] 6. Providing feedback: After the session ends, the user enters their feedback into the device. The server collects the feedback and updates the machine learning model to improve the quality of the next session.
[0652] 7. Reward distribution: The server aggregates the usage status of each virtual player, calculates and pays rewards to the performers according to the degree of usage.
[0653] Specific examples
[0654] For example, if a user wants to play jazz piano, they can create an account by entering profile information such as "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the session, the user can improvise on the piano, and the virtual player will respond and play in real time. After the session ends, the user can provide feedback, such as "I'd like a more rhythmic performance." The server uses this feedback to update the AI model and provide a better playing experience in the next session.
[0655] This invention allows users to enjoy realistic music sessions with virtual players that suit their individual tastes and styles, and the provider of virtual player data is paid a fee according to the degree of use.
[0656] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0657] Processing flow
[0658] Step 1: Enter your user information
[0659] 1.1 Device: The user downloads and installs the application.
[0660] Input: Application file
[0661] Output: Installed applications
[0662] Specific behavior: The user selects an application from the store and begins downloading it.
[0663] 1.2 Device: The user launches the app and creates an account.
[0664] Input: Name, Email Address, Password
[0665] Output: User account data
[0666] Specific operation: The user launches the application and enters the necessary information on the account creation screen.
[0667] 1.3 Device: User enters detailed profile information.
[0668] Input: Favorite instrument, music genre, religious views, political views, favorite books and movies
[0669] Output: Detailed user profile data
[0670] Specific operation: The user accesses the profile setting screen and enters each item.
[0671] 1.4 Server: Stores the entered user information in a database.
[0672] Input: User account data, detailed user profile data
[0673] Output: User information stored in the database
[0674] Specific operation: The server receives the data sent from the terminal and stores it in a database.
[0675] Step 2: Managing Virtual Player Data
[0676] 2.1 Server: Collects performance data and personal preference information from live performers.
[0677] Input: performance data, performer preference information
[0678] Output: Virtual player data
[0679] Specific operation: The performer uploads data through a dedicated interface.
[0680] 2.2 Server: Analyzes the collected performance data and learns using machine learning models.
[0681] Input: Performance data
[0682] Output: A trained machine learning model
[0683] Specific operation: The server processes the performance data, extracts characteristics of the performance style, and trains the model.
[0684] Step 3: Select a Virtual Player
[0685] 3.1 Terminal: The user browses the list of virtual players.
[0686] Input: None
[0687] Output: Virtual player list
[0688] Specific operation: A list of virtual players will be displayed on the device.
[0689] 3.2 Terminal: The user selects the preferred virtual player.
[0690] Input: User selection information
[0691] Output: Selected virtual player data
[0692] Specific operation: The user selects a virtual player on the screen and presses the select button.
[0693] 3.3 Server: Records the user's selections and stores them in a database.
[0694] Input: Selected Virtual Player Data
[0695] Output: Recorded selection data
[0696] Specific operation: The server receives the selection data sent from the terminal and stores it in a database.
[0697] Step 4: Starting a session
[0698] 4.1 Terminal: User enters session settings.
[0699] Input: Session settings data such as tempo, key, etc.
[0700] Output: Session initiation request
[0701] Specific operations: The user enters the required information on the session setting screen and presses the start button.
[0702] 4.2 Server: Prepares the session.
[0703] Input: Session configuration data
[0704] Output: Session ready notification
[0705] Specific operation: The server allocates resources based on the configuration and notifies the server that they are ready.
[0706] Step 5: Generate performance data
[0707] 5.1 Terminal: The user starts playing.
[0708] Input: User's performance data
[0709] Output: Real-time performance data
[0710] Specific operation: The user starts playing an instrument on the device, and data is recorded in real time.
[0711] 5.2 Server: Receives and analyzes user performance data.
[0712] Input: User's performance data
[0713] Output: Virtual player response data
[0714] Specific operation: Based on the performance data received by the server, a machine learning model is applied to generate response data.
[0715] 5.3 Server: Sends the generated performance data to the user.
[0716] Input: Virtual player response data
[0717] Output: Performance data to the user's device
[0718] Specific operation: The server streams data generated in real time to the device.
[0719] Step 6: Provide feedback
[0720] 6.1 Terminal: User enters feedback after the session ends.
[0721] Input: Feedback data
[0722] Output: Feedback sent
[0723] Specific operation: The user accesses the feedback input screen and enters their evaluation and opinion.
[0724] 6.2 Server: Collects and analyzes feedback data.
[0725] Input: Feedback data
[0726] Output: Updated machine learning model
[0727] Specific operation: The server updates the model based on the feedback and aims to improve the quality next time.
[0728] Step 7: Distributing rewards
[0729] 7.1 Server: Aggregates usage of Virtual Player data.
[0730] Input: Number of uses, duration of use, feedback data
[0731] Output: Aggregated data
[0732] Specific operation: The server aggregates the data and generates statistics.
[0733] 7.2 Server: Calculates the remuneration and pays the performers.
[0734] Input: Aggregate data
[0735] Output: Reward payment data
[0736] Specific operation: The server calculates the reward based on the aggregated results and pays the performer.
[0737] The above is the specific flow of processing of the program in this system.
[0738] (Application example 1)
[0739] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0740] Conventional music session systems with virtual players have had problems such as users finding a player that matches their preferences and a lack of real-time interaction, resulting in low satisfaction. Furthermore, there are insufficient means for efficiently collecting user feedback and using it to improve the system. Furthermore, there is no system in place for appropriately distributing rewards to providers of virtual player data. To solve these problems, it is necessary to provide a system that allows users to enjoy real-time sessions with virtual players that match their preferences and that continuously improves based on feedback.
[0741] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0742] In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for a user to select a virtual player, means for starting a session with the selected virtual player, means for generating performance data in real time during the session, means for collecting feedback after the session, means for paying a reward to the performer according to the degree of use, and means for conducting the session via an application installed on a smartphone or a head-mounted display. This enables users to enjoy real-time musical sessions with virtual players that suit their preferences, enables continuous improvement of the system based on feedback, and enables appropriate distribution of rewards to providers of virtual player data.
[0743] The "means for inputting user information" is a device or program that provides an interface for a user to input their profile information.
[0744] "Means for managing virtual player data" refers to a device or program that collects and stores the performance data and personal preference information of performers, and uses this information to individually manage virtual players.
[0745] The "means for a user to select a virtual player" is a device or program that provides an interface for a user to select a virtual player that suits his / her preferences from a list of provided virtual players.
[0746] A "means for initiating a session" is a device or program that prepares to initiate a musical session in real time with selected virtual players.
[0747] The "means for generating performance data in real time" is a device or program for generating performance data to which the virtual player responds in real time, based on the performance data of the user.
[0748] A "means for collecting feedback" is a device or program that collects opinions and impressions from users after a session ends, analyzes them, and uses them to improve the system.
[0749] The "means for paying a remuneration to a performer" is a device or program for calculating and paying a remuneration to the provider of virtual player data based on the degree of use of the virtual player data.
[0750] "Means for conducting a session via an application installed on a smartphone or head-mounted display" refers to a device or program that allows a user to conduct a musical session with a virtual player via an application installed on a smartphone or head-mounted display.
[0751] The present invention relates to a system for allowing a user to have a real-time music session with a virtual player that matches the user's preferences. The program for implementing this system is as follows.
[0752] Entering user information
[0753] The user launches a dedicated application installed on a smartphone or head-mounted display and enters information such as their name, email address, password, musical preferences (instruments, genres, religious views, political views, favorite books and movies), etc. This user information is then sent to the server and stored in a database. The server then generates a new user ID and associates it with the user information.
[0754] Managing Virtual Player Data
[0755] The server collects and centrally manages performance data provided by performers and their personal preferences. The collected performance data is analyzed and machine learning models are used to learn their playing style.
[0756] Virtual Player Selection
[0757] Through the application, the user can view a list of virtual players displayed on a panel. This list includes information about each virtual player's playing style and personal preferences. The user selects a virtual player that matches their preferences, and the server provides the data of the selected virtual player.
[0758] Starting a Session
[0759] To start a session with a selected virtual player, the user inputs the session's basic settings (tempo, key, etc.). The server prepares the session based on the input settings and processes requests for real-time performance.
[0760] Generating performance data
[0761] During a session, the user's performance data is sent in real time to the server, which uses a machine learning model to generate performance data for a virtual player that responds to the user's performance and immediately returns it to the user. This process establishes a real-time session between the user and the virtual player.
[0762] Providing Feedback
[0763] After the session, the user enters feedback on the virtual player's performance in the application, which is then sent to and collected by the server. The server uses this feedback to update the machine learning model and continuously improve the quality of the virtual player's performance data.
[0764] Reward Distribution
[0765] The server tally up the degree of use of the virtual player data (such as the number of uses and duration of use), calculates the reward for the provider, and pays the reward to the provider of the virtual player data based on the calculation results.
[0766] Specific examples
[0767] For example, consider a case where a user likes jazz piano and starts a session with a tempo of 120 BPM and the key of C major. The user selects a jazz piano virtual player and begins improvising. The virtual player then responds to the user's performance in real time to provide accompaniment. After the session ends, the user can provide feedback such as "I would like the performance to be a little more rhythmic," and this feedback will be used to improve the system. The following is an example of a prompt sentence.
[0768] "You are a virtual jazz pianist. If a user plays jazz piano in real time, you want to play along based on that performance. The user starts playing in the key of C major at a tempo of 120 BPM. How would you perform?"
[0769] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0770] Step 1:
[0771] A user launches an application and enters their user information. The user enters profile information such as their name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies, etc. The entered information is sent to the server and stored in a database. The server generates a new user ID and associates it with the user information.
[0772] Input: Name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies
[0773] Output: User ID, user information stored in the database
[0774] Step 2:
[0775] The server collects performance data and personal preference information provided by the performers, and centrally manages this data as virtual player data. The server analyzes this data and uses a machine learning model (e.g., PerformanceModel) to learn each performer's playing style.
[0776] Input: performance data, performer's personal preferences
[0777] Output: The playing style learned by the machine learning model
[0778] Step 3:
[0779] The user browses through the application a list of virtual players. During browsing, the server provides information about the virtual players' playing styles and personal preferences. The user selects a virtual player that best suits their preferences, and the selection is sent to the server.
[0780] Input: User's virtual player selection
[0781] Output: Data of the selected virtual player
[0782] Step 4:
[0783] To start a session with a selected virtual player, the user inputs basic settings such as tempo and key, and the server receives this information and prepares the session.
[0784] Input: Session settings such as tempo, key, etc.
[0785] Output: Session ready state
[0786] Step 5:
[0787] During a session, the user's performance data is sent in real time to the server, which processes the data and uses machine learning models to generate performance data that the virtual player responds to in real time. The generated performance data is then sent to the user's device.
[0788] Input: Real-time user performance data
[0789] Output: Virtual player response performance data
[0790] Step 6:
[0791] After the session ends, the user provides feedback on the virtual player's performance, which is then sent to the server, where it updates the machine learning model and improves the quality of the virtual player's performance data.
[0792] Input: User feedback
[0793] Output: An improved machine learning model
[0794] Step 7:
[0795] The server compiles the degree of use of each virtual player data (number of times used and duration of use), calculates a reward based on this data, and pays the reward to the provider of the virtual player data based on the calculation results.
[0796] Input: Virtual player data usage
[0797] Output: Calculated reward, reward payment completed
[0798] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0799] The present invention is a system that allows a user to engage in a musical session with a virtual player in real time, and further recognizes the user's emotions and adjusts the performance style accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[0800] A natural language description of the program's operation
[0801] The overall system processing is as follows:
[0802] 1. Enter your user information
[0803] On the device: The user launches the application and begins creating an account. The user enters their name, email address, and password.
[0804] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[0805] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0806] 2. Collection and Management of Virtual Player Data
[0807] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[0808] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0809] 3. Data analysis and machine learning model creation
[0810] Server: Analyzes performance data and extracts features such as rhythm, tempo, and pitch.
[0811] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0812] 4. Selecting a Virtual Player
[0813] Device: User views the list of virtual players within the app.
[0814] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0815] Device: The user selects the virtual player that best suits their preferences.
[0816] 5. Session Start Settings
[0817] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0818] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0819] 6. Conducting real-time sessions
[0820] Terminal: A session begins when the user presses the "Start Session" button.
[0821] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[0822] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0823] Server: Adjusts the virtual player's playing style in real time based on the emotion data from the emotion engine and the features of the performance data.
[0824] Server: The generated response performance data is sent to the user's device in real time.
[0825] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[0826] 7. Gathering Feedback
[0827] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[0828] Server: Collected feedback data is stored in a database and used for future model updates.
[0829] 8. Calculation and Payment of Rewards
[0830] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0831] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0832] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0833] Specific examples
[0834] For example, if a user wants to play jazz piano, they enter the following user information: "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. As the user plays the piano during the session, the emotion engine recognizes the emotion of joy from the user's facial expressions and voice, and adjusts the virtual player's playing style to be brighter and more energetic.
[0835] After the session, the user provides feedback on the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[0836] In addition, the provider of the virtual player data is paid a reward calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes, style, and even emotional state.
[0837] The processing flow will be explained below.
[0838] Step 1:
[0839] Entering user information
[0840] On the device: The user launches the application and proceeds to the account creation screen, where they enter their name, email address, and password.
[0841] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[0842] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[0843] Step 2:
[0844] Collection and management of virtual player data
[0845] Server: Collects performance data (music files and MIDI data) from actual performers, as well as personal preference information such as religious views, political views, and favorite books and movies.
[0846] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[0847] Step 3:
[0848] Data analysis and machine learning model creation
[0849] Server: Analyzes the performance data using analysis tools and extracts features such as rhythm, tempo, and pitch.
[0850] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[0851] Step 4:
[0852] Virtual Player Selection
[0853] Device: The user views the virtual player list screen within the app.
[0854] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[0855] Device: The user selects the virtual player that best suits their preferences.
[0856] Step 5:
[0857] Session Start Settings
[0858] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[0859] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[0860] Step 6:
[0861] Conducting real-time sessions
[0862] Terminal: A session begins when the user presses the "Start Session" button.
[0863] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[0864] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0865] Server: Adjusts the virtual player's playing style in real time based on the emotion data obtained from the emotion engine and the features of the performance data.
[0866] Server: The generated response performance data is sent to the user's device in real time.
[0867] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[0868] Step 7:
[0869] Gathering feedback
[0870] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[0871] Server: Receives feedback data from users and stores it in a database.
[0872] Server: Updates the machine learning model based on the collected feedback and uses it to improve the model in the future.
[0873] Step 8:
[0874] Reward calculation and payment
[0875] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0876] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0877] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0878] Example 2
[0879] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0880] In conventional music session systems, it has been difficult to provide a system that allows users to change their playing style in response to their emotions when engaging in a real-time music session. It has also been difficult to select and adjust virtual players to suit the user's personal preferences or specific music genres. This has led to problems such as a poor user experience and low satisfaction.
[0881] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0882] In this invention, the server includes means for recognizing a user's emotional state and dynamically adjusting a performance style based on the emotional state, means for identifying a performance style using a machine learning model, and means for generating responsive performance data in real time, thereby enabling a user to enjoy an interactive musical session with a virtual player in real time according to their emotions and preferences.
[0883] "User information" refers to data about individuals who use the system, such as their name, email address, password, favorite instrument, music genre, religious views, political views, favorite books and movies, etc.
[0884] "Virtual player data" is digital performance data with a variety of musical styles, including performance data collected from actual performers and personal preference information.
[0885] "Performance data" refers to data such as music files and MIDI data that digitally represents a specific musical performance.
[0886] "Personal preference information" is information about the performer's or user's individual preferences, such as religious views, political views, favorite books and movies, etc.
[0887] "Emotional state recognition" means determining a user's current emotions in real time based on data collected from the user's voice, facial expressions, biometric sensors, etc.
[0888] "Performance style" refers to the characteristics of performance based on musical features such as rhythm, tempo, pitch, and dynamics.
[0889] "Dynamic adjustment" means instantly modifying or changing the system's behavior or output in response to input data (e.g., user emotions or performance data) that changes in real time.
[0890] A "machine learning model" is an algorithm that learns patterns from large amounts of data and uses them to make predictions and classifications.
[0891] "Feedback" refers to data regarding ratings and opinions provided by users regarding the system and virtual players.
[0892] "Remuneration" means money or other benefits paid to a provider of virtual player data or in exchange for use of the system.
[0893] "Responding in real time" means instantly responding to user input and performance, and generating appropriate performance data and system operations.
[0894] The system of the present invention allows a user to engage in a musical session with a virtual player in real time, and further allows the user's emotions to be recognized and the performance style to be adjusted accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[0895] In terms of specific hardware and software, the devices used by users are computer devices such as smartphones, tablets, and PCs. The servers use cloud computing services to store databases and machine learning models. Programming languages and libraries such as Python and TensorFlow are used to train the machine learning models. In addition, the emotion engine uses voice recognition and facial recognition technologies, so cloud services such as Amazon Rekognition and Google Cloud Speech-to-Text are used.
[0896] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a C major key. As the user plays the piano during the session, the emotion engine recognizes joyful emotions from the user's facial expressions and voice, adjusting the virtual player's playing style to be more cheerful and lively. After the session ends, the user provides feedback about the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide more appropriate performances in future sessions. The provider of the virtual player data is also paid a reward calculated based on the degree of use of the session.
[0897] An example of a prompt is as follows:
[0898] "Jazz piano, I'm a Buddhist and a liberal, my favorite book is The Alchemist, and my favorite movie is Inception."
[0899] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0900] Step 1: Enter your user information
[0901] On the device: The user launches the application, enters their name, email address, and password in the form that appears on the screen, and then clicks the "Next" button.
[0902] Input: Name, Email Address, Password
[0903] Output: Sends the user's basic information to the server
[0904] Device: The next screen displays fields for the user to enter their favorite instrument, music genre, religious views, political views, and favorite books and movies. The user fills in each field and clicks the "Submit" button.
[0905] Input: musical instruments, music genres, religious views, political views, favorite books, movies
[0906] Output: Sends user details to server
[0907] Server: Receives the entered information, stores it in a database, generates a new user ID, and returns that information to the device.
[0908] Input: User information
[0909] Data processing: saving to database, generating new user ID
[0910] Output: New user ID
[0911] Step 2: Collecting and Managing Virtual Player Data
[0912] Server: Provides an interface with the functionality to collect performance data (music files and MIDI data) from real performers.
[0913] Input: Performance data
[0914] Output: Collected performance data
[0915] Server: Provides and collects a form for performers to enter their preferences (religious views, political views, favorite books and movies, etc.).
[0916] Input: Player preference information
[0917] Output: Collected preference information
[0918] Server: Collected performance data and performer preference information are stored in a database and associated with individual performer IDs.
[0919] Input: Collected performance data, preference information
[0920] Data processing: Save to database, associate with performer ID
[0921] Output: Data linked to performer ID
[0922] Step 3: Data analysis and machine learning model creation
[0923] Server: Executes data analysis algorithms to extract features such as rhythm, tempo, and pitch from performance data.
[0924] Input: Performance data
[0925] Data Calculation: Feature Extraction
[0926] Output: Extracted features
[0927] Server: Using the extracted features and personal preference information as input, the generative AI model is used to train a machine learning model to identify playing styles.
[0928] Input: extracted features, personal preference information
[0929] Data Computing: Training Machine Learning Models
[0930] Output: A trained machine learning model
[0931] Step 4: Select a Virtual Player
[0932] Device: The user accesses the "Virtual Player Selection" screen within the app.
[0933] Input: A view request from a user
[0934] Output: List of virtual players
[0935] Server: Provides virtual player data (playing style, preferences, sample performance) to the device.
[0936] Input: User ID
[0937] Data processing: filtering virtual players according to user preferences
[0938] Output: Virtual player data
[0939] Terminal: The user browses the provided data and selects the virtual player that best suits their preferences.
[0940] Input: View virtual player data
[0941] Output: Selected Virtual Player
[0942] Step 5: Session Start Settings
[0943] Device: The user goes to the "Session Start Settings" screen and enters basic settings such as tempo, key, and session time.
[0944] Input: Tempo, Key, Session Time
[0945] Output: Configuration information
[0946] Server: Receives this configuration information and begins preparing a session with the selected virtual player.
[0947] Input: Setting information, User ID, Virtual Player ID
[0948] Data processing: session preparation
[0949] Output: Session ready notification
[0950] Step 6: Conduct a real-time session
[0951] Terminal: A session begins when the user presses the "Start Session" button.
[0952] Input: Session start request from user
[0953] Output: Session start notification
[0954] Server: Receives the user's performance data in real time and extracts features such as rhythm, tempo, and pitch.
[0955] Input: User performance data
[0956] Data calculations: Real-time feature extraction
[0957] Output: Extracted features
[0958] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[0959] Input: Voice data, facial expression data, biometric sensor data
[0960] Data Computing: Emotional State Recognition
[0961] Output: Emotion data
[0962] Server: Adjusts the virtual player's playing style in real time based on the features of emotional data and performance data.
[0963] Input: Emotion data, performance data features
[0964] Data calculation: Adjusting playing style
[0965] Output: Adjusted performance data
[0966] Server: The generated response performance data is sent to the user's device in real time.
[0967] Input: Adjusted performance data
[0968] Output: Response performance data
[0969] Terminal: Based on the response performance data generated by the user, the user continues to play in real time together with the virtual player.
[0970] Input: Response performance data
[0971] Output: Real-time collaboration
[0972] Step 7: Gather feedback
[0973] Terminal: After the session ends, a form will appear where the user can enter feedback about the virtual player's performance. The user enters their feedback and clicks the "Submit" button.
[0974] Input: Feedback information
[0975] Output: Feedback information sent to server
[0976] Server: Receives feedback data and stores it in a database. This feedback data is used to update the machine learning model in the future.
[0977] Input: Feedback data
[0978] Data processing: Saving to database
[0979] Output: Stored feedback data
[0980] Step 8: Calculating and paying compensation
[0981] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[0982] Input: Number of uses, duration, feedback
[0983] Data calculation: Aggregation work
[0984] Output: Aggregated data
[0985] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[0986] Input: Aggregate data
[0987] Data calculation: Reward calculation
[0988] Output: Calculated reward data
[0989] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[0990] Input: Remuneration data, payment method
[0991] Output: Payment processing completion notification
[0992] (Application example 2)
[0993] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0994] Conventional music session systems have had problems in that they make it difficult for users to engage in emotionally-based interactive sessions with virtual players, and lack the functionality to dynamically adjust the virtual player's playing style in response to real-time changes in the user's emotions. Another issue is accurately recognizing the user's emotions and providing feedback accordingly.
[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for the user to select a virtual player, means for generating performance data in real time during a session, means for collecting feedback after the session, means for paying a reward to the data provider according to the degree of use, and means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the playing style of the virtual player based on that. This allows the user to enjoy an interactive music session with a virtual player that responds to their emotions in real time.
[0996] The "means for inputting user information" is an interface that allows a user to input their basic information, musical preferences, etc. into the system.
[0997] The "means for managing virtual player data including performance data and personal preference information" is a system that collects performance data and preference information of virtual players and manages them in a database.
[0998] The "means for the user to select a virtual player" is an interface that allows the user to select one of the virtual players displayed in the system.
[0999] The "means for generating performance data in real time during a session" is a system that generates and analyzes performance data of the user and virtual players in real time during a session.
[1000] The "means for collecting feedback after a session" is an interface for collecting feedback from the user after the session has ended.
[1001] The "means for paying a reward to a data provider according to the degree of use" is a system for calculating and paying a reward to a data provider of a virtual player according to the degree of use.
[1002] "Means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the virtual player's playing style based on that" is a system that uses an emotion engine to recognize the user's emotions and dynamically adjusts the virtual player's playing style according to the user's emotional state.
[1003] This invention is a system that allows a user to engage in a real-time musical session with a virtual player, and further recognizes the user's emotions and adjusts the playing style accordingly. The system includes the following main components:
[1004] 1. Enter your user information
[1005] Hardware: Smartphones, tablets, computers
[1006] Software: Applications
[1007] What happens: A user creates an account and enters basic information such as name, email address, password, preferred instrument and music genre, etc. This information is stored in a cloud database and a new user ID is generated.
[1008] 2. Virtual Player Data Management
[1009] Hardware: Server
[1010] Software: Cloud Management System
[1011] Processing: Performance data (music files and MIDI data) collected from actual performers and personal preference information are stored in a database, and data about each performer is managed.
[1012] 3. Emotion Recognition and Data Analysis
[1013] Hardware: smartphone facial recognition cameras, microphones, and biometric sensors
[1014] Software: Emotion recognition engine (open source emotion recognition software, e.g., OpenFace)
[1015] Processing: Emotional data is collected from the user's facial expressions, voice, and biometric sensors, and analyzed in real time by an emotion recognition engine.
[1016] 4. Conducting real-time sessions
[1017] Hardware: Smartphone
[1018] Software: In-app real-time music session function
[1019] Processing: Receives and analyzes the user's performance data in real time to extract features. Dynamically adjusts the virtual player's performance style based on the data from the emotion engine. The generated response performance data is sent to the user's device in real time, allowing the user to engage in a session with the virtual player.
[1020] 5. Gathering Feedback
[1021] Hardware: Smartphone
[1022] Software: Feedback collection feature
[1023] Processing: After the session ends, the user provides feedback on the virtual player's performance. The collected feedback data is stored in a cloud database and used to update the machine learning model.
[1024] 6. Calculation and Payment of Rewards
[1025] Hardware: Server
[1026] Software: Remuneration calculation system
[1027] Processing: Compensation is calculated based on the number of uses, duration of use, and user feedback for each virtual player, and paid to the data provider.
[1028] Examples:
[1029] For example, if a user wants to play jazz guitar, they enter their user information as "jazz guitar, Catholic, liberal, favorite book is 'Norwegian Wood,' and favorite movie is 'La La Land.'" They then select a virtual jazz guitarist within the app and enter settings to start a session with a tempo of 140 BPM and a key of F major. During the session, the emotion engine recognizes the user's excitement level from their facial expressions and voice, adjusting the virtual player's performance to be more energetic. After the session ends, the user provides feedback such as "I'd like a more rhythmic performance," and the server uses this feedback to update the machine learning model.
[1030] Example prompt sentence:
[1031] The user begins a 140 BPM jazz session in F major with a virtual jazz band. The user sets their favorite music genre as jazz, is Catholic and liberal, and their favorite book is "Norwegian Wood," and their favorite movie is "La La Land." During the session, the system recognizes emotions from the user's smile and voice, adjusting the performance to be more energetic.
[1032] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1033] Step 1:
[1034] A user launches an application and creates an account. The user enters basic information such as their name, email address, password, favorite instrument, and music genre. The information entered is sent from the device to the server and stored in a cloud database. A new user ID is generated and returned to the device. The input for this step is the user information, and the output is the new user ID.
[1035] Step 2:
[1036] The server manages data about each performer by storing performance data (music files and MIDI data) and personal preference information collected from the actual performers in a database. The input is the performance data and preference information, and the output is the virtual player data stored in the database.
[1037] Step 3:
[1038] The user browses a list of virtual players in the app and selects one that suits their preferences. The server provides virtual player data (playing style, preferences, sample performance, etc.) to the device. The input is the virtual player data, and the output is the virtual player selected by the user.
[1039] Step 4:
[1040] The user proceeds to the session start screen and inputs basic settings such as tempo, key, etc. The device sends the input settings to the server and begins preparing for a session with the selected virtual player. The input is session setting information, and the output is a message indicating that the session is ready.
[1041] Step 5:
[1042] A session begins when the user presses the "Start Session" button. The device collects the user's performance data in real time and sends it to the server. The server analyzes the received performance data and extracts features such as rhythm and tempo. The input is the user's performance data, and the output is the analyzed features.
[1043] Step 6:
[1044] The emotion engine collects emotional data from the user's voice, facial expression, and biometric sensors to recognize the emotional state. The emotion engine analyzes the emotional state based on this data and identifies the type of emotion (happiness, sadness, excitement, etc.). The input is voice data, facial expression data, and biometric data, and the output is the recognized emotional state.
[1045] Step 7:
[1046] The server adjusts the virtual player's performance style in real time based on the emotion data from the emotion engine and the performance data features. This allows the performance style to reflect the user's emotions. The input is emotion data and performance data features, and the output is the adjusted performance data.
[1047] Step 8:
[1048] The server transmits the response performance data generated in real time to the user's terminal, and the user engages in a session in real time with the virtual player. The input is the adjusted performance data, and the output is the performance data transmitted to the user's terminal.
[1049] Step 9:
[1050] After the session ends, the user provides feedback on the virtual player's performance. The device sends the feedback data to the server and stores it in a database. The input is the feedback data, and the output is the feedback data stored in the database.
[1051] Step 10:
[1052] The server calculates rewards for each virtual player based on the number of uses, duration of use, and user feedback. The calculated rewards are paid to the data provider. The inputs are usage data and feedback data, and the output is the calculated rewards.
[1053] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1054] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1055] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1056] [Third embodiment]
[1057] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1058] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1059] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1060] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1061] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1062] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1063] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1064] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1065] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1066] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1067] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1068] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1069] The present invention provides a system that allows users to engage in real-time music sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[1070] A natural language description of the program's operation
[1071] The overall system processing is as follows:
[1072] 1. Enter your user information
[1073] On the device: The user launches the application and creates an account. They enter their name, email address, and password, and then enter profile information such as their favorite instrument, music genre, religious views, political views, and favorite books and movies.
[1074] Server: Saves the entered user information in the database and generates a new user ID.
[1075] 2. Virtual Player Data Management
[1076] Server: Along with performance data provided by actual performers, the server manages personal preference information such as the performers' religious views, political views, and favorite books and movies as virtual player data.
[1077] Server: Analyzes the collected performance data and uses a machine learning model to learn the playing style.
[1078] 3. Select a Virtual Player
[1079] On the device: The user views a list of virtual players within the app, each of which includes information about their playing style and personal preferences.
[1080] Server: Provides data on the virtual players viewed by users.
[1081] Device: The user selects the virtual player that best suits their preferences.
[1082] 4. Starting a session
[1083] Terminal: Enter basic settings such as tempo and key to start a session with the user's selected virtual player.
[1084] Server: Prepares the session based on the settings you enter.
[1085] 5. Generating performance data
[1086] Server: Receives user performance data in real time and generates performance data for the virtual player to respond to based on a machine learning model.
[1087] Server: Sends the generated performance data to the user's device in real time.
[1088] 6. Providing Feedback
[1089] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[1090] Server: Based on the collected feedback data, the machine learning model is updated to improve the quality of the virtual player's performance data.
[1091] 7. Reward Distribution
[1092] Server: Collects data such as the number of times each virtual player uses the service and the duration of their use, and calculates the reward for the provider of virtual player data.
[1093] Server: Pays the calculated reward to the performer.
[1094] Specific examples
[1095] Specific examples are shown below.
[1096] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player in the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the actual session, as the user improvises on the piano, the virtual player responds to their performance in real time and plays along.
[1097] After the session, the user provides feedback on the virtual player's responses, such as "I'd like the performance to be a little more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[1098] In addition, the provider of the virtual player data is paid a fee calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes and styles.
[1099] The processing flow will be explained below.
[1100] Step 1:
[1101] Entering user information
[1102] On the device: The user launches the application and begins creating an account. The user is prompted to enter their name, email address, and password.
[1103] Device: The user is taken to a screen where they can enter their preferred instruments, music genres, religious and political views, and favorite books and movies.
[1104] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1105] Step 2:
[1106] Collection and management of virtual player data
[1107] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[1108] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1109] Step 3:
[1110] Data analysis and machine learning model creation
[1111] Server: Analyzes the collected performance data using analysis tools to extract features such as rhythm, tempo, and pitch.
[1112] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1113] Step 4:
[1114] Virtual Player Selection
[1115] Device: The user views the virtual player list screen within the app.
[1116] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1117] Device: The user selects the virtual player that best suits their preferences.
[1118] Step 5:
[1119] Session Start Settings
[1120] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1121] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1122] Step 6:
[1123] Conducting real-time sessions
[1124] Terminal: A session begins when the user presses the "Start Session" button.
[1125] Server: Receives the user's performance data in real time and extracts performance characteristics using analysis tools.
[1126] Server: Based on the extracted features and machine learning model, it generates performance data in real time that the virtual player responds to.
[1127] Server: The generated response performance data is sent to the user's device in real time.
[1128] Terminal: The user plays in real time with the virtual player.
[1129] Step 7:
[1130] Gathering feedback
[1131] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[1132] Server: Receives feedback data from users and stores it in a database.
[1133] Server: Updates the machine learning model based on the collected feedback to improve the quality of future performance data.
[1134] Step 8:
[1135] Reward calculation and payment
[1136] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1137] Server: Based on the aggregated data, calculates the reward for the provider of virtual player data.
[1138] Server: Pays the calculated reward amount according to the payment method specified by the original data provider.
[1139] Example 1
[1140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1141] The present invention aims to provide a system that automatically allocates appropriate rewards to providers of virtual player data, while providing a performance experience tailored to the user's individual preferences, real-time responses during the session, and effective use of feedback after the session, which have been difficult to achieve with conventional systems when users enjoy real-time musical sessions with virtual players that suit their preferences.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1143] In this invention, the server includes a means for inputting user information, a means for managing virtual player data including performance data and personal preference information, and a means for the user to select a virtual player, thereby enabling the user to select a virtual player according to their individual preferences.
[1144] The server includes means for starting a session with a selected virtual player, means for generating performance data in real time during the session, and means for collecting feedback after the session, thereby enabling real-time generation of performance data during the session and improvement of the quality of the next session based on the feedback.
[1145] The server further includes a means for paying rewards to performers according to the degree of use, a means for providing data of a virtual player selected by a user, a means for receiving the user's performance data in real time and generating performance data to which the virtual player responds based on a machine learning model, and a means for transmitting the generated performance data to the user's device in real time, thereby enabling appropriate distribution of rewards to providers of virtual player data and real-time musical sessions between users and virtual players.
[1146] "User information" refers to personal data entered by a user when registering with the system, specifically including name, email address, password, and profile information such as favorite instruments and music genres, religious views, political views, and favorite books and movies.
[1147] "Virtual player data" refers to performance data provided by an actual performer and data including the performer's personal preference information, such as playing style and personal tastes and preferences.
[1148] A "machine learning model" refers to an algorithm or statistical model that analyzes collected performance data and learns the playing style of a virtual player.
[1149] A "prompt sentence" is a sentence containing specific instructions or questions to be input into the generated AI model, and describes a specific request from the user or system.
[1150] "Performance data" is music data generated by a user or a virtual player, and is data that electronically records a musical performance, such as an audio file or MIDI data.
[1151] "Real-time response" means that the virtual player responds instantly to the performance input by the user, generating music data in real time.
[1152] "Feedback" refers to opinions and evaluations provided by users after a session has ended, and is used to improve the system and adjust the virtual player's playing style.
[1153] "Reward allocation" refers to the process of calculating and paying rewards to providers of virtual player data, and determining the amount of reward based on the degree of usage and feedback.
[1154] The present invention provides a system that allows users to engage in real-time musical sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[1155] Hardware and software used
[1156] The system is implemented using the following hardware and software:
[1157] Server: Manages user information and virtual player data, generates performance data, and collects feedback.
[1158] Terminal: The device (smartphone, tablet, PC) on which the user operates the application.
[1159] Generative AI model: A model that uses machine learning algorithms to learn playing styles and generate responsive performance data in real time.
[1160] Database: Storage for user information, virtual player data, performance data, and feedback.
[1161] Specific operation of the system
[1162] 1. Entering user information: A user launches an application on their device and creates an account, entering their name, email address, password, favorite instrument and music genre, religious views, political views, and favorite books and movies.
[1163] 2. Virtual Player Data Management: The server collects performance data provided by real performers and their personal preferences. The collected data is analyzed using a machine learning model to learn their playing style.
[1164] 3. Virtual player selection: The user browses the list of virtual players on the device, checks their playing style and preferences, and selects the virtual player that best suits their preferences.
[1165] 4. Starting a session: To start a session, the user inputs settings such as tempo and key into the terminal. The server prepares the session based on the settings.
[1166] 5. Performance data generation: When a user starts playing on their device, the data is sent to the server in real time. The server uses a machine learning model to generate performance data that the virtual player responds to, and sends it to the device in real time.
[1167] 6. Providing feedback: After the session ends, the user enters their feedback into the device. The server collects the feedback and updates the machine learning model to improve the quality of the next session.
[1168] 7. Reward distribution: The server aggregates the usage status of each virtual player, calculates and pays rewards to the performers according to the degree of usage.
[1169] Specific examples
[1170] For example, if a user wants to play jazz piano, they can create an account by entering profile information such as "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the session, the user can improvise on the piano, and the virtual player will respond and play in real time. After the session ends, the user can provide feedback, such as "I'd like a more rhythmic performance." The server uses this feedback to update the AI model and provide a better playing experience in the next session.
[1171] This invention allows users to enjoy realistic music sessions with virtual players that suit their individual tastes and styles, and the provider of virtual player data is paid a fee according to the degree of use.
[1172] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1173] Processing flow
[1174] Step 1: Enter your user information
[1175] 1.1 Device: The user downloads and installs the application.
[1176] Input: Application file
[1177] Output: Installed applications
[1178] Specific behavior: The user selects an application from the store and begins downloading it.
[1179] 1.2 Device: The user launches the app and creates an account.
[1180] Input: Name, Email Address, Password
[1181] Output: User account data
[1182] Specific operation: The user launches the application and enters the necessary information on the account creation screen.
[1183] 1.3 Device: User enters detailed profile information.
[1184] Input: Favorite instrument, music genre, religious views, political views, favorite books and movies
[1185] Output: Detailed user profile data
[1186] Specific operation: The user accesses the profile setting screen and enters each item.
[1187] 1.4 Server: Stores the entered user information in a database.
[1188] Input: User account data, detailed user profile data
[1189] Output: User information stored in the database
[1190] Specific operation: The server receives the data sent from the terminal and stores it in a database.
[1191] Step 2: Managing Virtual Player Data
[1192] 2.1 Server: Collects performance data and personal preference information from live performers.
[1193] Input: performance data, performer preference information
[1194] Output: Virtual player data
[1195] Specific operation: The performer uploads data through a dedicated interface.
[1196] 2.2 Server: Analyzes the collected performance data and learns using machine learning models.
[1197] Input: Performance data
[1198] Output: A trained machine learning model
[1199] Specific operation: The server processes the performance data, extracts characteristics of the performance style, and trains the model.
[1200] Step 3: Select a Virtual Player
[1201] 3.1 Terminal: The user browses the list of virtual players.
[1202] Input: None
[1203] Output: Virtual player list
[1204] Specific operation: A list of virtual players will be displayed on the device.
[1205] 3.2 Terminal: The user selects the preferred virtual player.
[1206] Input: User selection information
[1207] Output: Selected virtual player data
[1208] Specific operation: The user selects a virtual player on the screen and presses the select button.
[1209] 3.3 Server: Records the user's selections and stores them in a database.
[1210] Input: Selected Virtual Player Data
[1211] Output: Recorded selection data
[1212] Specific operation: The server receives the selection data sent from the terminal and stores it in a database.
[1213] Step 4: Starting a session
[1214] 4.1 Terminal: User enters session settings.
[1215] Input: Session settings data such as tempo, key, etc.
[1216] Output: Session initiation request
[1217] Specific operations: The user enters the required information on the session setting screen and presses the start button.
[1218] 4.2 Server: Prepares the session.
[1219] Input: Session configuration data
[1220] Output: Session ready notification
[1221] Specific operation: The server allocates resources based on the configuration and notifies the server that they are ready.
[1222] Step 5: Generate performance data
[1223] 5.1 Terminal: The user starts playing.
[1224] Input: User's performance data
[1225] Output: Real-time performance data
[1226] Specific operation: The user starts playing an instrument on the device, and data is recorded in real time.
[1227] 5.2 Server: Receives and analyzes user performance data.
[1228] Input: User's performance data
[1229] Output: Virtual player response data
[1230] Specific operation: Based on the performance data received by the server, a machine learning model is applied to generate response data.
[1231] 5.3 Server: Sends the generated performance data to the user.
[1232] Input: Virtual player response data
[1233] Output: Performance data to the user's device
[1234] Specific operation: The server streams data generated in real time to the device.
[1235] Step 6: Provide feedback
[1236] 6.1 Terminal: User enters feedback after the session ends.
[1237] Input: Feedback data
[1238] Output: Feedback sent
[1239] Specific operation: The user accesses the feedback input screen and enters their evaluation and opinion.
[1240] 6.2 Server: Collects and analyzes feedback data.
[1241] Input: Feedback data
[1242] Output: Updated machine learning model
[1243] Specific operation: The server updates the model based on the feedback and aims to improve the quality next time.
[1244] Step 7: Distributing rewards
[1245] 7.1 Server: Aggregates usage of Virtual Player data.
[1246] Input: Number of uses, duration of use, feedback data
[1247] Output: Aggregated data
[1248] Specific operation: The server aggregates the data and generates statistics.
[1249] 7.2 Server: Calculates the remuneration and pays the performers.
[1250] Input: Aggregate data
[1251] Output: Reward payment data
[1252] Specific operation: The server calculates the reward based on the aggregated results and pays the performer.
[1253] The above is the specific flow of processing of the program in this system.
[1254] (Application example 1)
[1255] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1256] Conventional music session systems with virtual players have had problems such as users finding a player that matches their preferences and a lack of real-time interaction, resulting in low satisfaction. Furthermore, there are insufficient means for efficiently collecting user feedback and using it to improve the system. Furthermore, there is no system in place for appropriately distributing rewards to providers of virtual player data. To solve these problems, it is necessary to provide a system that allows users to enjoy real-time sessions with virtual players that match their preferences and that continuously improves based on feedback.
[1257] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1258] In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for a user to select a virtual player, means for starting a session with the selected virtual player, means for generating performance data in real time during the session, means for collecting feedback after the session, means for paying a reward to the performer according to the degree of use, and means for conducting the session via an application installed on a smartphone or a head-mounted display. This enables users to enjoy real-time musical sessions with virtual players that suit their preferences, enables continuous improvement of the system based on feedback, and enables appropriate distribution of rewards to providers of virtual player data.
[1259] The "means for inputting user information" is a device or program that provides an interface for a user to input their profile information.
[1260] "Means for managing virtual player data" refers to a device or program that collects and stores the performance data and personal preference information of performers, and uses this information to individually manage virtual players.
[1261] The "means for a user to select a virtual player" is a device or program that provides an interface for a user to select a virtual player that suits his / her preferences from a list of provided virtual players.
[1262] A "means for initiating a session" is a device or program that prepares to initiate a musical session in real time with selected virtual players.
[1263] The "means for generating performance data in real time" is a device or program for generating performance data to which the virtual player responds in real time, based on the performance data of the user.
[1264] A "means for collecting feedback" is a device or program that collects opinions and impressions from users after a session ends, analyzes them, and uses them to improve the system.
[1265] The "means for paying a remuneration to a performer" is a device or program for calculating and paying a remuneration to the provider of virtual player data based on the degree of use of the virtual player data.
[1266] "Means for conducting a session via an application installed on a smartphone or head-mounted display" refers to a device or program that allows a user to conduct a musical session with a virtual player via an application installed on a smartphone or head-mounted display.
[1267] The present invention relates to a system for allowing a user to have a real-time music session with a virtual player that matches the user's preferences. The program for implementing this system is as follows.
[1268] Entering user information
[1269] The user launches a dedicated application installed on a smartphone or head-mounted display and enters information such as their name, email address, password, musical preferences (instruments, genres, religious views, political views, favorite books and movies), etc. This user information is then sent to the server and stored in a database. The server then generates a new user ID and associates it with the user information.
[1270] Managing Virtual Player Data
[1271] The server collects and centrally manages performance data provided by performers and their personal preferences. The collected performance data is analyzed and machine learning models are used to learn their playing style.
[1272] Virtual Player Selection
[1273] Through the application, the user can view a list of virtual players displayed on a panel. This list includes information about each virtual player's playing style and personal preferences. The user selects a virtual player that matches their preferences, and the server provides the data of the selected virtual player.
[1274] Starting a Session
[1275] To start a session with a selected virtual player, the user inputs basic session settings (tempo, key, etc.) The server prepares the session based on the input settings and processes requests for real-time performance.
[1276] Generating performance data
[1277] During a session, the user's performance data is sent in real time to the server, which uses a machine learning model to generate performance data for a virtual player that responds to the user's performance and immediately returns it to the user. This process establishes a real-time session between the user and the virtual player.
[1278] Providing Feedback
[1279] After the session, the user enters feedback on the virtual player's performance in the application, which is then sent to and collected by the server. The server uses this feedback to update the machine learning model and continuously improve the quality of the virtual player's performance data.
[1280] Reward Distribution
[1281] The server tally up the degree of use of the virtual player data (such as the number of uses and duration of use), calculates the reward for the provider, and pays the reward to the provider of the virtual player data based on the calculation results.
[1282] Specific examples
[1283] For example, consider a case where a user likes jazz piano and starts a session with a tempo of 120 BPM and the key of C major. The user selects a jazz piano virtual player and begins improvising. The virtual player then responds to the user's performance in real time to provide accompaniment. After the session ends, the user can provide feedback such as "I would like the performance to be a little more rhythmic," and this feedback will be used to improve the system. The following is an example of a prompt sentence.
[1284] "You are a virtual jazz pianist. If a user plays jazz piano in real time, you want to play along based on their performance. The user starts playing in the key of C major at a tempo of 120 BPM. How would you perform?"
[1285] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1286] Step 1:
[1287] A user launches an application and enters their user information. The user enters profile information such as their name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies, etc. The entered information is sent to the server and stored in a database. The server generates a new user ID and associates it with the user information.
[1288] Input: Name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies
[1289] Output: User ID, user information stored in the database
[1290] Step 2:
[1291] The server collects the performance data and personal preference information provided by the performers, and centrally manages this data as virtual player data.The server analyzes this data and uses a machine learning model (e.g., PerformanceModel) to learn each performer's playing style.
[1292] Input: performance data, performer's personal preferences
[1293] Output: The playing style learned by the machine learning model
[1294] Step 3:
[1295] The user browses through the application a list of virtual players. During browsing, the server provides information about the virtual players' playing styles and personal preferences. The user selects a virtual player that best suits their preferences, and the selection is sent to the server.
[1296] Input: User's virtual player selection
[1297] Output: Data of the selected virtual player
[1298] Step 4:
[1299] To start a session with a selected virtual player, the user inputs basic settings such as tempo and key, and the server receives this information and prepares the session.
[1300] Input: Session settings such as tempo, key, etc.
[1301] Output: Session ready state
[1302] Step 5:
[1303] During a session, the user's performance data is sent in real time to the server, which processes the data and uses machine learning models to generate performance data that the virtual player responds to in real time. The generated performance data is then sent to the user's device.
[1304] Input: Real-time user performance data
[1305] Output: Virtual player response performance data
[1306] Step 6:
[1307] After the session ends, the user provides feedback on the virtual player's performance, which is then sent to the server, where it updates the machine learning model and improves the quality of the virtual player's performance data.
[1308] Input: User feedback
[1309] Output: An improved machine learning model
[1310] Step 7:
[1311] The server compiles the degree of use of each virtual player data (number of times used and duration of use), calculates a reward based on this data, and pays the reward to the provider of the virtual player data based on the calculation results.
[1312] Input: Virtual player data usage
[1313] Output: Calculated reward, reward payment completed
[1314] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1315] The present invention is a system that allows a user to engage in a musical session with a virtual player in real time, and further recognizes the user's emotions and adjusts the performance style accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[1316] A natural language description of the program's operation
[1317] The overall system processing is as follows:
[1318] 1. Enter your user information
[1319] On the device: The user launches the application and begins creating an account. The user enters their name, email address, and password.
[1320] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[1321] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1322] 2. Collection and Management of Virtual Player Data
[1323] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[1324] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1325] 3. Data analysis and machine learning model creation
[1326] Server: Analyzes performance data and extracts features such as rhythm, tempo, and pitch.
[1327] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1328] 4. Selecting a Virtual Player
[1329] Device: User views the list of virtual players within the app.
[1330] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1331] Device: The user selects the virtual player that best suits their preferences.
[1332] 5. Session Start Settings
[1333] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1334] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1335] 6. Conducting real-time sessions
[1336] Terminal: A session begins when the user presses the "Start Session" button.
[1337] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[1338] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1339] Server: Adjusts the virtual player's playing style in real time based on the emotion data from the emotion engine and the features of the performance data.
[1340] Server: The generated response performance data is sent to the user's device in real time.
[1341] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[1342] 7. Gathering Feedback
[1343] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[1344] Server: Collected feedback data is stored in a database and used for future model updates.
[1345] 8. Calculation and Payment of Rewards
[1346] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1347] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[1348] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[1349] Specific examples
[1350] For example, if a user wants to play jazz piano, they enter the following user information: "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. As the user plays the piano during the session, the emotion engine recognizes the emotion of joy from the user's facial expressions and voice, and adjusts the virtual player's playing style to be brighter and more energetic.
[1351] After the session, the user provides feedback on the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[1352] In addition, the provider of the virtual player data is paid a reward calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes, style, and even emotional state.
[1353] The processing flow will be explained below.
[1354] Step 1:
[1355] Entering user information
[1356] On the device: The user launches the application and proceeds to the account creation screen, where they enter their name, email address, and password.
[1357] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[1358] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1359] Step 2:
[1360] Collection and management of virtual player data
[1361] Server: Collects performance data (music files and MIDI data) from actual performers, as well as personal preference information such as religious views, political views, and favorite books and movies.
[1362] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1363] Step 3:
[1364] Data analysis and machine learning model creation
[1365] Server: Analyzes the performance data using analysis tools and extracts features such as rhythm, tempo, and pitch.
[1366] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1367] Step 4:
[1368] Virtual Player Selection
[1369] Device: The user views the virtual player list screen within the app.
[1370] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1371] Device: The user selects the virtual player that best suits their preferences.
[1372] Step 5:
[1373] Session Start Settings
[1374] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1375] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1376] Step 6:
[1377] Conducting real-time sessions
[1378] Terminal: A session begins when the user presses the "Start Session" button.
[1379] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[1380] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1381] Server: Adjusts the virtual player's playing style in real time based on the emotion data obtained from the emotion engine and the features of the performance data.
[1382] Server: The generated response performance data is sent to the user's device in real time.
[1383] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[1384] Step 7:
[1385] Collecting feedback
[1386] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[1387] Server: Receives feedback data from users and stores it in a database.
[1388] Server: Updates the machine learning model based on the collected feedback and uses it to improve the model in the future.
[1389] Step 8:
[1390] Reward calculation and payment
[1391] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1392] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[1393] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[1394] Example 2
[1395] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1396] In conventional music session systems, it has been difficult to provide a system that allows users to change their playing style in response to their emotions when engaging in a real-time music session. It has also been difficult to select and adjust virtual players to suit the user's personal preferences or specific music genres. This has led to problems such as a poor user experience and low satisfaction.
[1397] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1398] In this invention, the server includes means for recognizing a user's emotional state and dynamically adjusting a performance style based on the emotional state, means for identifying a performance style using a machine learning model, and means for generating responsive performance data in real time, thereby enabling a user to enjoy an interactive musical session with a virtual player in real time according to their emotions and preferences.
[1399] "User information" refers to data about individuals who use the system, such as their name, email address, password, favorite instrument, music genre, religious views, political views, favorite books and movies, etc.
[1400] "Virtual player data" is digital performance data with a variety of musical styles, including performance data collected from actual performers and personal preference information.
[1401] "Performance data" refers to data such as music files and MIDI data that digitally represents a specific musical performance.
[1402] "Personal preference information" is information about the performer's or user's individual preferences, such as religious views, political views, favorite books and movies, etc.
[1403] "Emotional state recognition" means determining a user's current emotions in real time based on data collected from the user's voice, facial expressions, biometric sensors, etc.
[1404] "Performance style" refers to the characteristics of performance based on musical features such as rhythm, tempo, pitch, and dynamics.
[1405] "Dynamic adjustment" means instantly modifying or changing the system's behavior or output in response to input data (e.g., user emotions or performance data) that changes in real time.
[1406] A "machine learning model" is an algorithm that learns patterns from large amounts of data and uses them to make predictions and classifications.
[1407] "Feedback" refers to data regarding ratings and opinions provided by users regarding the system and virtual players.
[1408] "Remuneration" means money or other benefits paid to a provider of virtual player data or in exchange for use of the system.
[1409] "Responding in real time" means instantly responding to user input and performance, and generating appropriate performance data and system operations.
[1410] The system of the present invention allows a user to engage in a musical session with a virtual player in real time, and further allows the user's emotions to be recognized and the performance style to be adjusted accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[1411] In terms of specific hardware and software, the devices used by users are computer devices such as smartphones, tablets, and PCs. The servers use cloud computing services to store databases and machine learning models. Programming languages and libraries such as Python and TensorFlow are used to train the machine learning models. In addition, the emotion engine uses voice recognition and facial recognition technologies, so cloud services such as Amazon Rekognition and Google Cloud Speech-to-Text are used.
[1412] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a C major key. As the user plays the piano during the session, the emotion engine recognizes joyful emotions from the user's facial expressions and voice, adjusting the virtual player's playing style to be more cheerful and lively. After the session ends, the user provides feedback about the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide more appropriate performances in future sessions. The provider of the virtual player data is also paid a reward calculated based on the degree of use of the session.
[1413] An example of a prompt is as follows:
[1414] "Jazz piano, I'm a Buddhist and a liberal, my favorite book is The Alchemist, and my favorite movie is Inception."
[1415] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1416] Step 1: Enter your user information
[1417] On the device: The user launches the application, enters their name, email address, and password in the form that appears on the screen, and then clicks the "Next" button.
[1418] Input: Name, Email Address, Password
[1419] Output: Sends the user's basic information to the server
[1420] Device: The next screen displays fields for the user to enter their favorite instrument, music genre, religious views, political views, and favorite books and movies. The user fills in each field and clicks the "Submit" button.
[1421] Input: musical instruments, music genres, religious views, political views, favorite books, movies
[1422] Output: Sends user details to server
[1423] Server: Receives the entered information, stores it in a database, generates a new user ID, and returns that information to the device.
[1424] Input: User information
[1425] Data processing: saving to database, generating new user ID
[1426] Output: New user ID
[1427] Step 2: Collecting and Managing Virtual Player Data
[1428] Server: Provides an interface with the functionality to collect performance data (music files and MIDI data) from real performers.
[1429] Input: Performance data
[1430] Output: Collected performance data
[1431] Server: Provides and collects a form for performers to enter their preferences (religious views, political views, favorite books and movies, etc.).
[1432] Input: Player preference information
[1433] Output: Collected preference information
[1434] Server: Collected performance data and performer preference information are stored in a database and associated with individual performer IDs.
[1435] Input: Collected performance data, preference information
[1436] Data processing: Save to database, associate with performer ID
[1437] Output: Data linked to performer ID
[1438] Step 3: Data analysis and machine learning model creation
[1439] Server: Executes data analysis algorithms to extract features such as rhythm, tempo, and pitch from performance data.
[1440] Input: Performance data
[1441] Data Calculation: Feature Extraction
[1442] Output: Extracted features
[1443] Server: Using the extracted features and personal preference information as input, the generative AI model is used to train a machine learning model to identify playing styles.
[1444] Input: extracted features, personal preference information
[1445] Data Computing: Training Machine Learning Models
[1446] Output: A trained machine learning model
[1447] Step 4: Select a Virtual Player
[1448] Device: The user accesses the "Virtual Player Selection" screen within the app.
[1449] Input: A view request from a user
[1450] Output: List of virtual players
[1451] Server: Provides virtual player data (playing style, preferences, sample performance) to the device.
[1452] Input: User ID
[1453] Data processing: filtering virtual players according to user preferences
[1454] Output: Virtual player data
[1455] Terminal: The user browses the provided data and selects the virtual player that best suits their preferences.
[1456] Input: View virtual player data
[1457] Output: Selected Virtual Player
[1458] Step 5: Session Start Settings
[1459] Device: The user goes to the "Session Start Settings" screen and enters basic settings such as tempo, key, and session time.
[1460] Input: Tempo, Key, Session Time
[1461] Output: Configuration information
[1462] Server: Receives this configuration information and begins preparing a session with the selected virtual player.
[1463] Input: Setting information, User ID, Virtual Player ID
[1464] Data processing: session preparation
[1465] Output: Session ready notification
[1466] Step 6: Conduct a real-time session
[1467] Terminal: A session begins when the user presses the "Start Session" button.
[1468] Input: Session start request from user
[1469] Output: Session start notification
[1470] Server: Receives the user's performance data in real time and extracts features such as rhythm, tempo, and pitch.
[1471] Input: User performance data
[1472] Data calculations: Real-time feature extraction
[1473] Output: Extracted features
[1474] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1475] Input: Voice data, facial expression data, biometric sensor data
[1476] Data Computing: Emotional State Recognition
[1477] Output: Emotion data
[1478] Server: Adjusts the virtual player's playing style in real time based on the features of emotional data and performance data.
[1479] Input: Emotion data, performance data features
[1480] Data calculation: Adjusting playing style
[1481] Output: Adjusted performance data
[1482] Server: The generated response performance data is sent to the user's device in real time.
[1483] Input: Adjusted performance data
[1484] Output: Response performance data
[1485] Terminal: Based on the response performance data generated by the user, the user continues to play in real time together with the virtual player.
[1486] Input: Response performance data
[1487] Output: Real-time collaboration
[1488] Step 7: Gather feedback
[1489] Terminal: After the session ends, a form will appear where the user can enter feedback about the virtual player's performance. The user enters their feedback and clicks the "Submit" button.
[1490] Input: Feedback information
[1491] Output: Feedback information sent to server
[1492] Server: Receives feedback data and stores it in a database. This feedback data is used to update the machine learning model in the future.
[1493] Input: Feedback data
[1494] Data processing: Saving to database
[1495] Output: Stored feedback data
[1496] Step 8: Calculating and paying compensation
[1497] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1498] Input: Number of uses, duration, feedback
[1499] Data calculation: Aggregation work
[1500] Output: Aggregated data
[1501] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[1502] Input: Aggregate data
[1503] Data calculation: Reward calculation
[1504] Output: Calculated reward data
[1505] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[1506] Input: Remuneration data, payment method
[1507] Output: Payment processing completion notification
[1508] (Application example 2)
[1509] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1510] Conventional music session systems have had problems in that they make it difficult for users to engage in emotionally-based interactive sessions with virtual players, and lack the functionality to dynamically adjust the virtual player's playing style in response to real-time changes in the user's emotions. Another issue is accurately recognizing the user's emotions and providing feedback accordingly.
[1511] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for the user to select a virtual player, means for generating performance data in real time during a session, means for collecting feedback after the session, means for paying a reward to the data provider according to the degree of use, and means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the playing style of the virtual player based on that. This allows the user to enjoy an interactive music session with a virtual player that responds to their emotions in real time.
[1512] The "means for inputting user information" is an interface that allows a user to input their basic information, musical preferences, etc. into the system.
[1513] The "means for managing virtual player data including performance data and personal preference information" is a system that collects performance data and preference information of virtual players and manages them in a database.
[1514] The "means for the user to select a virtual player" is an interface that allows the user to select one of the virtual players displayed in the system.
[1515] The "means for generating performance data in real time during a session" is a system that generates and analyzes performance data of the user and virtual players in real time during a session.
[1516] The "means for collecting feedback after a session" is an interface for collecting feedback from the user after the session has ended.
[1517] The "means for paying a reward to a data provider according to the degree of use" is a system for calculating and paying a reward to a data provider of a virtual player according to the degree of use.
[1518] "Means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the virtual player's playing style based on that" is a system that uses an emotion engine to recognize the user's emotions and dynamically adjusts the virtual player's playing style according to the user's emotional state.
[1519] This invention is a system that allows a user to engage in a real-time musical session with a virtual player, and further recognizes the user's emotions and adjusts the playing style accordingly. The system includes the following main components:
[1520] 1. Enter your user information
[1521] Hardware: Smartphones, tablets, computers
[1522] Software: Applications
[1523] What happens: A user creates an account and enters basic information such as name, email address, password, preferred instrument and music genre, etc. This information is stored in a cloud database and a new user ID is generated.
[1524] 2. Virtual Player Data Management
[1525] Hardware: Server
[1526] Software: Cloud Management System
[1527] Processing: Performance data (music files and MIDI data) collected from actual performers and personal preference information are stored in a database, and data about each performer is managed.
[1528] 3. Emotion Recognition and Data Analysis
[1529] Hardware: smartphone facial recognition cameras, microphones, and biometric sensors
[1530] Software: Emotion recognition engine (open source emotion recognition software, e.g., OpenFace)
[1531] Processing: Emotional data is collected from the user's facial expressions, voice, and biometric sensors, and analyzed in real time by an emotion recognition engine.
[1532] 4. Conducting real-time sessions
[1533] Hardware: Smartphone
[1534] Software: In-app real-time music session function
[1535] Processing: Receives and analyzes the user's performance data in real time to extract features. Dynamically adjusts the virtual player's performance style based on the data from the emotion engine. The generated response performance data is sent to the user's device in real time, allowing the user to engage in a session with the virtual player.
[1536] 5. Gathering Feedback
[1537] Hardware: Smartphone
[1538] Software: Feedback collection feature
[1539] Processing: After the session ends, the user provides feedback on the virtual player's performance. The collected feedback data is stored in a cloud database and used to update the machine learning model.
[1540] 6. Calculation and Payment of Rewards
[1541] Hardware: Server
[1542] Software: Remuneration calculation system
[1543] Processing: Compensation is calculated based on the number of uses, duration of use, and user feedback for each virtual player, and paid to the data provider.
[1544] Examples:
[1545] For example, if a user wants to play jazz guitar, they enter their user information as "jazz guitar, Catholic, liberal, favorite book is 'Norwegian Wood,' and favorite movie is 'La La Land.'" They then select a virtual jazz guitarist within the app and enter settings to start a session with a tempo of 140 BPM and a key of F major. During the session, the emotion engine recognizes the user's excitement level from their facial expressions and voice, adjusting the virtual player's performance to be more energetic. After the session ends, the user provides feedback such as "I'd like a more rhythmic performance," and the server uses this feedback to update the machine learning model.
[1546] Example prompt sentence:
[1547] The user begins a 140 BPM jazz session in F major with a virtual jazz band. The user sets their favorite music genre as jazz, is Catholic and liberal, and their favorite book is "Norwegian Wood," and their favorite movie is "La La Land." During the session, the system recognizes emotions from the user's smile and voice, adjusting the performance to be more energetic.
[1548] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1549] Step 1:
[1550] A user launches an application and creates an account. The user enters basic information such as their name, email address, password, favorite instrument, and music genre. The information entered is sent from the device to the server and stored in a cloud database. A new user ID is generated and returned to the device. The input for this step is the user information, and the output is the new user ID.
[1551] Step 2:
[1552] The server manages data about each performer by storing performance data (music files and MIDI data) and personal preference information collected from the actual performers in a database. The input is the performance data and preference information, and the output is the virtual player data stored in the database.
[1553] Step 3:
[1554] The user browses a list of virtual players in the app and selects one that suits their preferences. The server provides virtual player data (playing style, preferences, sample performance, etc.) to the device. The input is the virtual player data, and the output is the virtual player selected by the user.
[1555] Step 4:
[1556] The user proceeds to the session start screen and inputs basic settings such as tempo, key, etc. The device sends the input settings to the server and begins preparing for a session with the selected virtual player. The input is session setting information, and the output is a message indicating that the session is ready.
[1557] Step 5:
[1558] A session begins when the user presses the "Start Session" button. The device collects the user's performance data in real time and sends it to the server. The server analyzes the received performance data and extracts features such as rhythm and tempo. The input is the user's performance data, and the output is the analyzed features.
[1559] Step 6:
[1560] The emotion engine collects emotional data from the user's voice, facial expression, and biometric sensors to recognize the emotional state. The emotion engine analyzes the emotional state based on this data and identifies the type of emotion (happiness, sadness, excitement, etc.). The input is voice data, facial expression data, and biometric data, and the output is the recognized emotional state.
[1561] Step 7:
[1562] The server adjusts the virtual player's performance style in real time based on the emotion data from the emotion engine and the performance data features. This allows the performance style to reflect the user's emotions. The input is emotion data and performance data features, and the output is the adjusted performance data.
[1563] Step 8:
[1564] The server transmits the response performance data generated in real time to the user's terminal, and the user engages in a session in real time with the virtual player. The input is the adjusted performance data, and the output is the performance data transmitted to the user's terminal.
[1565] Step 9:
[1566] After the session ends, the user provides feedback on the virtual player's performance. The device sends the feedback data to the server and stores it in a database. The input is the feedback data, and the output is the feedback data stored in the database.
[1567] Step 10:
[1568] The server calculates rewards for each virtual player based on the number of uses, duration of use, and user feedback. The calculated rewards are paid to the data provider. The inputs are usage data and feedback data, and the output is the calculated rewards.
[1569] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1570] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1571] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1572] [Fourth embodiment]
[1573] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1574] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1575] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1576] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1577] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1578] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1579] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1580] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1581] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1582] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1583] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1584] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1585] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1586] The present invention provides a system that allows users to engage in real-time music sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[1587] A natural language description of the program's operation
[1588] The overall system processing is as follows:
[1589] 1. Enter your user information
[1590] On the device: The user launches the application and creates an account. They enter their name, email address, and password, and then enter profile information such as their favorite instrument, music genre, religious views, political views, and favorite books and movies.
[1591] Server: Saves the entered user information in the database and generates a new user ID.
[1592] 2. Virtual Player Data Management
[1593] Server: Along with performance data provided by actual performers, the server manages personal preference information such as the performers' religious views, political views, and favorite books and movies as virtual player data.
[1594] Server: Analyzes the collected performance data and uses a machine learning model to learn the playing style.
[1595] 3. Select a Virtual Player
[1596] On the device: The user views a list of virtual players within the app, each of which includes information about their playing style and personal preferences.
[1597] Server: Provides data on the virtual players viewed by users.
[1598] Device: The user selects the virtual player that best suits their preferences.
[1599] 4. Starting a session
[1600] Terminal: Enter basic settings such as tempo and key to start a session with the user's selected virtual player.
[1601] Server: Prepares the session based on the settings you enter.
[1602] 5. Generating performance data
[1603] Server: Receives user performance data in real time and generates performance data for the virtual player to respond to based on a machine learning model.
[1604] Server: Sends the generated performance data to the user's device in real time.
[1605] 6. Providing Feedback
[1606] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[1607] Server: Based on the collected feedback data, the machine learning model is updated to improve the quality of the virtual player's performance data.
[1608] 7. Reward Distribution
[1609] Server: Collects data such as the number of times each virtual player uses the service and the duration of their use, and calculates the reward for the provider of virtual player data.
[1610] Server: Pays the calculated reward to the performer.
[1611] Specific examples
[1612] Specific examples are shown below.
[1613] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player in the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the actual session, as the user improvises on the piano, the virtual player responds to their performance in real time and plays along.
[1614] After the session, the user provides feedback on the virtual player's responses, such as "I'd like the performance to be a little more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[1615] In addition, the provider of the virtual player data is paid a fee calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes and styles.
[1616] The processing flow will be explained below.
[1617] Step 1:
[1618] Entering user information
[1619] On the device: The user launches the application and begins creating an account. The user is prompted to enter their name, email address, and password.
[1620] Device: The user is taken to a screen where they can enter their preferred instruments, music genres, religious and political views, and favorite books and movies.
[1621] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1622] Step 2:
[1623] Collection and management of virtual player data
[1624] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[1625] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1626] Step 3:
[1627] Data analysis and machine learning model creation
[1628] Server: Analyzes the collected performance data using analysis tools to extract features such as rhythm, tempo, and pitch.
[1629] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1630] Step 4:
[1631] Virtual Player Selection
[1632] Device: The user views the virtual player list screen within the app.
[1633] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1634] Device: The user selects the virtual player that best suits their preferences.
[1635] Step 5:
[1636] Session Start Settings
[1637] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1638] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1639] Step 6:
[1640] Conducting real-time sessions
[1641] Terminal: A session begins when the user presses the "Start Session" button.
[1642] Server: Receives the user's performance data in real time and extracts performance characteristics using analysis tools.
[1643] Server: Based on the extracted features and machine learning model, it generates performance data in real time that the virtual player responds to.
[1644] Server: The generated response performance data is sent to the user's device in real time.
[1645] Terminal: The user plays in real time with the virtual player.
[1646] Step 7:
[1647] Gathering feedback
[1648] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[1649] Server: Receives feedback data from users and stores it in a database.
[1650] Server: Updates the machine learning model based on the collected feedback to improve the quality of future performance data.
[1651] Step 8:
[1652] Reward calculation and payment
[1653] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1654] Server: Based on the aggregated data, calculates the reward for the provider of virtual player data.
[1655] Server: Pays the calculated reward amount according to the payment method specified by the original data provider.
[1656] Example 1
[1657] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1658] The present invention aims to provide a system that automatically allocates appropriate rewards to providers of virtual player data, while providing a performance experience tailored to the user's individual preferences, real-time responses during the session, and effective use of feedback after the session, which have been difficult to achieve with conventional systems when users enjoy real-time musical sessions with virtual players that suit their preferences.
[1659] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1660] In this invention, the server includes a means for inputting user information, a means for managing virtual player data including performance data and personal preference information, and a means for the user to select a virtual player, thereby enabling the user to select a virtual player according to their individual preferences.
[1661] The server includes means for starting a session with a selected virtual player, means for generating performance data in real time during the session, and means for collecting feedback after the session, thereby enabling real-time generation of performance data during the session and improvement of the quality of the next session based on the feedback.
[1662] The server further includes a means for paying rewards to performers according to the degree of use, a means for providing data of a virtual player selected by a user, a means for receiving the user's performance data in real time and generating performance data to which the virtual player responds based on a machine learning model, and a means for transmitting the generated performance data to the user's device in real time, thereby enabling appropriate distribution of rewards to providers of virtual player data and real-time musical sessions between users and virtual players.
[1663] "User information" refers to personal data entered by a user when registering with the system, specifically including name, email address, password, and profile information such as favorite instruments and music genres, religious views, political views, and favorite books and movies.
[1664] "Virtual player data" refers to performance data provided by an actual performer and data including the performer's personal preference information, such as playing style and personal tastes and preferences.
[1665] A "machine learning model" refers to an algorithm or statistical model that analyzes collected performance data and learns the playing style of a virtual player.
[1666] A "prompt sentence" is a sentence containing specific instructions or questions to be input into the generated AI model, and describes a specific request from the user or system.
[1667] "Performance data" is music data generated by a user or a virtual player, and is data that electronically records a musical performance, such as an audio file or MIDI data.
[1668] "Real-time response" means that the virtual player responds instantly to the performance input by the user, generating music data in real time.
[1669] "Feedback" refers to opinions and evaluations provided by users after a session has ended, and is used to improve the system and adjust the virtual player's playing style.
[1670] "Reward allocation" refers to the process of calculating and paying rewards to providers of virtual player data, and determining the amount of reward based on the degree of usage and feedback.
[1671] The present invention provides a system that allows users to engage in real-time musical sessions with virtual players that match their preferences. The system includes a means for inputting user information, a means for managing virtual player data, a means for the user to select a virtual player, a means for starting a session, a means for generating performance data in real time, a means for collecting feedback after the session, and a means for paying rewards according to the degree of use.
[1672] Hardware and software used
[1673] The system is implemented using the following hardware and software:
[1674] Server: Manages user information and virtual player data, generates performance data, and collects feedback.
[1675] Terminal: The device (smartphone, tablet, PC) on which the user operates the application.
[1676] Generative AI model: A model that uses machine learning algorithms to learn playing styles and generate responsive performance data in real time.
[1677] Database: Storage for user information, virtual player data, performance data, and feedback.
[1678] Specific operation of the system
[1679] 1. Entering user information: A user launches an application on their device and creates an account, entering their name, email address, password, favorite instrument and music genre, religious views, political views, and favorite books and movies.
[1680] 2. Virtual Player Data Management: The server collects performance data provided by real performers and their personal preferences. The collected data is analyzed using a machine learning model to learn their playing style.
[1681] 3. Virtual player selection: The user browses the list of virtual players on the device, checks their playing style and preferences, and selects the virtual player that best suits their preferences.
[1682] 4. Starting a session: To start a session, the user inputs settings such as tempo and key into the terminal. The server prepares the session based on the settings.
[1683] 5. Performance data generation: When a user starts playing on their device, the data is sent to the server in real time. The server uses a machine learning model to generate performance data that the virtual player responds to, and sends it to the device in real time.
[1684] 6. Providing feedback: After the session ends, the user enters their feedback into the device. The server collects the feedback and updates the machine learning model to improve the quality of the next session.
[1685] 7. Reward distribution: The server aggregates the usage status of each virtual player, calculates and pays rewards to the performers according to the degree of usage.
[1686] Specific examples
[1687] For example, if a user wants to play jazz piano, they can create an account by entering profile information such as "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. During the session, the user can improvise on the piano, and the virtual player will respond and play in real time. After the session ends, the user can provide feedback, such as "I'd like a more rhythmic performance." The server uses this feedback to update the AI model and provide a better playing experience in the next session.
[1688] This invention allows users to enjoy realistic music sessions with virtual players that suit their individual tastes and styles, and the provider of virtual player data is paid a fee according to the degree of use.
[1689] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1690] Processing flow
[1691] Step 1: Enter your user information
[1692] 1.1 Device: The user downloads and installs the application.
[1693] Input: Application file
[1694] Output: Installed applications
[1695] Specific behavior: The user selects an application from the store and begins downloading it.
[1696] 1.2 Device: The user launches the app and creates an account.
[1697] Input: Name, Email Address, Password
[1698] Output: User account data
[1699] Specific operation: The user launches the application and enters the necessary information on the account creation screen.
[1700] 1.3 Device: User enters detailed profile information.
[1701] Input: Favorite instrument, music genre, religious views, political views, favorite books and movies
[1702] Output: Detailed user profile data
[1703] Specific operation: The user accesses the profile setting screen and enters each item.
[1704] 1.4 Server: Stores the entered user information in a database.
[1705] Input: User account data, detailed user profile data
[1706] Output: User information stored in the database
[1707] Specific operation: The server receives the data sent from the terminal and stores it in a database.
[1708] Step 2: Managing Virtual Player Data
[1709] 2.1 Server: Collects performance data and personal preference information from live performers.
[1710] Input: performance data, performer preference information
[1711] Output: Virtual player data
[1712] Specific operation: The performer uploads data through a dedicated interface.
[1713] 2.2 Server: Analyzes the collected performance data and learns using machine learning models.
[1714] Input: Performance data
[1715] Output: A trained machine learning model
[1716] Specific operation: The server processes the performance data, extracts characteristics of the performance style, and trains the model.
[1717] Step 3: Select a Virtual Player
[1718] 3.1 Terminal: The user browses the list of virtual players.
[1719] Input: None
[1720] Output: Virtual player list
[1721] Specific operation: A list of virtual players will be displayed on the device.
[1722] 3.2 Terminal: The user selects the preferred virtual player.
[1723] Input: User selection information
[1724] Output: Selected virtual player data
[1725] Specific operation: The user selects a virtual player on the screen and presses the select button.
[1726] 3.3 Server: Records the user's selections and stores them in a database.
[1727] Input: Selected Virtual Player Data
[1728] Output: Recorded selection data
[1729] Specific operation: The server receives the selection data sent from the terminal and stores it in a database.
[1730] Step 4: Starting a session
[1731] 4.1 Terminal: User enters session settings.
[1732] Input: Session settings data such as tempo, key, etc.
[1733] Output: Session initiation request
[1734] Specific operations: The user enters the required information on the session setting screen and presses the start button.
[1735] 4.2 Server: Prepares the session.
[1736] Input: Session configuration data
[1737] Output: Session ready notification
[1738] Specific operation: The server allocates resources based on the configuration and notifies the server that they are ready.
[1739] Step 5: Generate performance data
[1740] 5.1 Terminal: The user starts playing.
[1741] Input: User's performance data
[1742] Output: Real-time performance data
[1743] Specific operation: The user starts playing an instrument on the device, and data is recorded in real time.
[1744] 5.2 Server: Receives and analyzes user performance data.
[1745] Input: User's performance data
[1746] Output: Virtual player response data
[1747] Specific operation: Based on the performance data received by the server, a machine learning model is applied to generate response data.
[1748] 5.3 Server: Sends the generated performance data to the user.
[1749] Input: Virtual player response data
[1750] Output: Performance data to the user's device
[1751] Specific operation: The server streams data generated in real time to the device.
[1752] Step 6: Provide feedback
[1753] 6.1 Terminal: User enters feedback after the session ends.
[1754] Input: Feedback data
[1755] Output: Feedback sent
[1756] Specific operation: The user accesses the feedback input screen and enters their evaluation and opinion.
[1757] 6.2 Server: Collects and analyzes feedback data.
[1758] Input: Feedback data
[1759] Output: Updated machine learning model
[1760] Specific operation: The server updates the model based on the feedback and aims to improve the quality next time.
[1761] Step 7: Distributing rewards
[1762] 7.1 Server: Aggregates usage of Virtual Player data.
[1763] Input: Number of uses, duration of use, feedback data
[1764] Output: Aggregated data
[1765] Specific operation: The server aggregates the data and generates statistics.
[1766] 7.2 Server: Calculates the remuneration and pays the performers.
[1767] Input: Aggregate data
[1768] Output: Reward payment data
[1769] Specific operation: The server calculates the reward based on the aggregated results and pays the performer.
[1770] The above is the specific flow of processing of the program in this system.
[1771] (Application example 1)
[1772] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1773] Conventional music session systems with virtual players have had problems such as users finding a player that matches their preferences and a lack of real-time interaction, resulting in low satisfaction. Furthermore, there are insufficient means for efficiently collecting user feedback and using it to improve the system. Furthermore, there is no system in place for appropriately distributing rewards to providers of virtual player data. To solve these problems, it is necessary to provide a system that allows users to enjoy real-time sessions with virtual players that match their preferences and that continuously improves based on feedback.
[1774] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1775] In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for a user to select a virtual player, means for starting a session with the selected virtual player, means for generating performance data in real time during the session, means for collecting feedback after the session, means for paying a reward to the performer according to the degree of use, and means for conducting the session via an application installed on a smartphone or a head-mounted display. This enables users to enjoy real-time musical sessions with virtual players that suit their preferences, enables continuous improvement of the system based on feedback, and enables appropriate distribution of rewards to providers of virtual player data.
[1776] The "means for inputting user information" is a device or program that provides an interface for a user to input their profile information.
[1777] "Means for managing virtual player data" refers to a device or program that collects and stores the performance data and personal preference information of performers, and uses this information to individually manage virtual players.
[1778] The "means for a user to select a virtual player" is a device or program that provides an interface for a user to select a virtual player that suits his / her preferences from a list of provided virtual players.
[1779] A "means for initiating a session" is a device or program that prepares to initiate a musical session in real time with selected virtual players.
[1780] The "means for generating performance data in real time" is a device or program for generating performance data to which the virtual player responds in real time, based on the performance data of the user.
[1781] A "means for collecting feedback" is a device or program that collects opinions and impressions from users after a session ends, analyzes them, and uses them to improve the system.
[1782] The "means for paying a remuneration to a performer" is a device or program for calculating and paying a remuneration to the provider of virtual player data based on the degree of use of the virtual player data.
[1783] "Means for conducting a session via an application installed on a smartphone or head-mounted display" refers to a device or program that allows a user to conduct a musical session with a virtual player via an application installed on a smartphone or head-mounted display.
[1784] The present invention relates to a system for allowing a user to have a real-time music session with a virtual player that matches the user's preferences. The program for implementing this system is as follows.
[1785] Entering user information
[1786] The user launches a dedicated application installed on a smartphone or head-mounted display and enters information such as their name, email address, password, musical preferences (instruments, genres, religious views, political views, favorite books and movies), etc. This user information is then sent to the server and stored in a database. The server then generates a new user ID and associates it with the user information.
[1787] Managing Virtual Player Data
[1788] The server collects and centrally manages performance data provided by performers and their personal preferences. The collected performance data is analyzed and machine learning models are used to learn their playing style.
[1789] Virtual Player Selection
[1790] Through the application, the user can view a list of virtual players displayed on a panel. This list includes information about each virtual player's playing style and personal preferences. The user selects a virtual player that matches their preferences, and the server provides the data of the selected virtual player.
[1791] Starting a Session
[1792] To start a session with a selected virtual player, the user inputs basic session settings (tempo, key, etc.) The server prepares the session based on the input settings and processes requests for real-time performance.
[1793] Generating performance data
[1794] During a session, the user's performance data is sent in real time to the server, which uses a machine learning model to generate performance data for a virtual player that responds to the user's performance and immediately returns it to the user. This process establishes a real-time session between the user and the virtual player.
[1795] Providing Feedback
[1796] After the session, the user enters feedback on the virtual player's performance in the application, which is then sent to and collected by the server. The server uses this feedback to update the machine learning model and continuously improve the quality of the virtual player's performance data.
[1797] Reward Distribution
[1798] The server tally up the degree of use of the virtual player data (such as the number of uses and duration of use), calculates the reward for the provider, and pays the reward to the provider of the virtual player data based on the calculation results.
[1799] Specific examples
[1800] For example, consider a case where a user likes jazz piano and starts a session with a tempo of 120 BPM and the key of C major. The user selects a jazz piano virtual player and begins improvising. The virtual player then responds to the user's performance in real time to provide accompaniment. After the session ends, the user can provide feedback such as "I would like the performance to be a little more rhythmic," and this feedback will be used to improve the system. The following is an example of a prompt sentence.
[1801] "You are a virtual jazz pianist. If a user plays jazz piano in real time, you want to play along based on that performance. The user starts playing in the key of C major at a tempo of 120 BPM. How would you perform?"
[1802] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1803] Step 1:
[1804] A user launches an application and enters their user information. The user enters profile information such as their name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies, etc. The entered information is sent to the server and stored in a database. The server generates a new user ID and associates it with the user information.
[1805] Input: Name, email address, password, musical instrument, music genre, religious views, political views, favorite books and movies
[1806] Output: User ID, user information stored in the database
[1807] Step 2:
[1808] The server collects performance data and personal preference information provided by the performers, and centrally manages this data as virtual player data. The server analyzes this data and uses a machine learning model (e.g., PerformanceModel) to learn each performer's playing style.
[1809] Input: performance data, performer's personal preferences
[1810] Output: The playing style learned by the machine learning model
[1811] Step 3:
[1812] The user browses through the application a list of virtual players. During browsing, the server provides information about the virtual players' playing styles and personal preferences. The user selects a virtual player that best suits their preferences, and the selection is sent to the server.
[1813] Input: User's virtual player selection
[1814] Output: Data of the selected virtual player
[1815] Step 4:
[1816] To start a session with a selected virtual player, the user inputs basic settings such as tempo and key, and the server receives this information and prepares the session.
[1817] Input: Session settings such as tempo, key, etc.
[1818] Output: Session ready state
[1819] Step 5:
[1820] During a session, the user's performance data is sent in real time to the server, which processes the data and uses machine learning models to generate performance data that the virtual player responds to in real time. The generated performance data is then sent to the user's device.
[1821] Input: Real-time user performance data
[1822] Output: Virtual player response performance data
[1823] Step 6:
[1824] After the session ends, the user provides feedback on the virtual player's performance, which is then sent to the server, where it updates the machine learning model and improves the quality of the virtual player's performance data.
[1825] Input: User feedback
[1826] Output: An improved machine learning model
[1827] Step 7:
[1828] The server compiles the degree of use of each virtual player data (number of times used and duration of use), calculates a reward based on this data, and pays the reward to the provider of the virtual player data based on the calculation results.
[1829] Input: Virtual player data usage
[1830] Output: Calculated reward, reward payment completed
[1831] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1832] The present invention is a system that allows a user to engage in a musical session with a virtual player in real time, and further recognizes the user's emotions and adjusts the performance style accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[1833] A natural language description of the program's operation
[1834] The overall system processing is as follows:
[1835] 1. Enter your user information
[1836] On the device: The user launches the application and begins creating an account. The user enters their name, email address, and password.
[1837] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[1838] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1839] 2. Collection and Management of Virtual Player Data
[1840] Server: Collects performance data (music files and MIDI data) and personal preference information (religious views, political views, favorite books and movies) from actual performers.
[1841] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1842] 3. Data analysis and machine learning model creation
[1843] Server: Analyzes performance data and extracts features such as rhythm, tempo, and pitch.
[1844] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1845] 4. Selecting a Virtual Player
[1846] Device: User views the list of virtual players within the app.
[1847] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1848] Device: The user selects the virtual player that best suits their preferences.
[1849] 5. Session Start Settings
[1850] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1851] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1852] 6. Conducting real-time sessions
[1853] Terminal: A session begins when the user presses the "Start Session" button.
[1854] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[1855] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1856] Server: Adjusts the virtual player's playing style in real time based on the emotion data from the emotion engine and the features of the performance data.
[1857] Server: The generated response performance data is sent to the user's device in real time.
[1858] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[1859] 7. Gathering Feedback
[1860] Terminal: After the session ends, the user enters feedback about the virtual player's performance.
[1861] Server: Collected feedback data is stored in a database and used for future model updates.
[1862] 8. Calculation and Payment of Rewards
[1863] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1864] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[1865] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[1866] Specific examples
[1867] For example, if a user wants to play jazz piano, they enter the following user information: "Jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They then select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a key of C major. As the user plays the piano during the session, the emotion engine recognizes the emotion of joy from the user's facial expressions and voice, and adjusts the virtual player's playing style to be brighter and more energetic.
[1868] After the session, the user provides feedback on the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide a more appropriate performance in future sessions.
[1869] In addition, the provider of the virtual player data is paid a reward calculated according to the degree of use of the session. In this way, users can enjoy realistic music sessions with virtual players that suit their individual tastes, style, and even emotional state.
[1870] The processing flow will be explained below.
[1871] Step 1:
[1872] Entering user information
[1873] On the device: The user launches the application and proceeds to the account creation screen, where they enter their name, email address, and password.
[1874] Device: The user enters their favorite instrument, music genre, religious views, political views, favorite books and movies.
[1875] Server: Saves the entered user information in a database, generates a new user ID, and returns it to the terminal.
[1876] Step 2:
[1877] Collection and management of virtual player data
[1878] Server: Collects performance data (music files and MIDI data) from actual performers, as well as personal preference information such as religious views, political views, and favorite books and movies.
[1879] Server: Collected performance data is stored in a database and associated with individual performer IDs.
[1880] Step 3:
[1881] Data analysis and machine learning model creation
[1882] Server: Analyzes the performance data using analysis tools and extracts features such as rhythm, tempo, and pitch.
[1883] Server: Trains a machine learning model based on the extracted features and personal preference information to identify playing styles.
[1884] Step 4:
[1885] Virtual Player Selection
[1886] Device: The user views the virtual player list screen within the app.
[1887] Server: Provides virtual player data (playing style, preferences, sample performances, etc.) to the user's device.
[1888] Device: The user selects the virtual player that best suits their preferences.
[1889] Step 5:
[1890] Session Start Settings
[1891] Terminal: Proceed to a session start screen where the user enters basic settings such as tempo, key, etc.
[1892] Server: Receives the entered settings and starts preparing a session with the selected virtual player.
[1893] Step 6:
[1894] Conducting real-time sessions
[1895] Terminal: A session begins when the user presses the "Start Session" button.
[1896] Server: Receives the user's performance data in real time, analyzes it, and extracts features.
[1897] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1898] Server: Adjusts the virtual player's playing style in real time based on the emotion data obtained from the emotion engine and the features of the performance data.
[1899] Server: The generated response performance data is sent to the user's device in real time.
[1900] Terminal: The user performs in real time with a virtual player that dynamically adjusts according to emotions.
[1901] Step 7:
[1902] Gathering feedback
[1903] Terminal: After the session ends, the user is taken to a screen where they can enter feedback about the virtual player's performance.
[1904] Server: Receives feedback data from users and stores it in a database.
[1905] Server: Updates the machine learning model based on the collected feedback and uses it to improve the model in the future.
[1906] Step 8:
[1907] Reward calculation and payment
[1908] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[1909] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[1910] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[1911] Example 2
[1912] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1913] In conventional music session systems, it has been difficult to provide a system that allows users to change their playing style in response to their emotions when engaging in a real-time music session. It has also been difficult to select and adjust virtual players to suit the user's personal preferences or specific music genres. This has led to problems such as a poor user experience and low satisfaction.
[1914] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1915] In this invention, the server includes means for recognizing a user's emotional state and dynamically adjusting a performance style based on the emotional state, means for identifying a performance style using a machine learning model, and means for generating responsive performance data in real time, thereby enabling a user to enjoy an interactive musical session with a virtual player in real time according to their emotions and preferences.
[1916] "User information" refers to data about individuals who use the system, such as their name, email address, password, favorite instrument, music genre, religious views, political views, favorite books and movies, etc.
[1917] "Virtual player data" is digital performance data with a variety of musical styles, including performance data collected from actual performers and personal preference information.
[1918] "Performance data" refers to data such as music files and MIDI data that digitally represents a specific musical performance.
[1919] "Personal preference information" is information about the performer's or user's individual preferences, such as religious views, political views, favorite books and movies, etc.
[1920] "Emotional state recognition" means determining a user's current emotions in real time based on data collected from the user's voice, facial expressions, biometric sensors, etc.
[1921] "Performance style" refers to the characteristics of performance based on musical features such as rhythm, tempo, pitch, and dynamics.
[1922] "Dynamic adjustment" means instantly modifying or changing the system's behavior or output in response to input data (e.g., user emotions or performance data) that changes in real time.
[1923] A "machine learning model" is an algorithm that learns patterns from large amounts of data and uses them to make predictions and classifications.
[1924] "Feedback" refers to data regarding ratings and opinions provided by users regarding the system and virtual players.
[1925] "Remuneration" means money or other benefits paid to a provider of virtual player data or in exchange for use of the system.
[1926] "Responding in real time" means instantly responding to user input and performance, and generating appropriate performance data and system operations.
[1927] The system of the present invention allows a user to engage in a musical session with a virtual player in real time, and further allows the user's emotions to be recognized and the performance style to be adjusted accordingly. This system includes means for inputting user information, means for managing virtual player data, means for the user to select a virtual player, means for starting a session, means for generating performance data in real time, means for collecting feedback, and means for paying rewards, as well as means for recognizing the user's emotions using an emotion engine and dynamically adjusting the performance style based on the emotions.
[1928] In terms of specific hardware and software, the devices used by users are computer devices such as smartphones, tablets, and PCs. The servers use cloud computing services to store databases and machine learning models. Programming languages and libraries such as Python and TensorFlow are used to train the machine learning models. In addition, the emotion engine uses voice recognition and facial recognition technologies, so cloud services such as Amazon Rekognition and Google Cloud Speech-to-Text are used.
[1929] For example, if a user wants to play jazz piano, they enter their user information as "jazz piano, I'm a Buddhist and liberal, my favorite book is The Alchemist, and my favorite movie is Inception." They select a jazz piano virtual player within the app and enter settings to start a session with a tempo of 120 BPM and a C major key. As the user plays the piano during the session, the emotion engine recognizes joyful emotions from the user's facial expressions and voice, adjusting the virtual player's playing style to be more cheerful and lively. After the session ends, the user provides feedback about the virtual player's performance, such as "I'd like it to be a bit more rhythmic." The server uses this feedback to update the machine learning model and provide more appropriate performances in future sessions. The provider of the virtual player data is also paid a reward calculated based on the degree of use of the session.
[1930] An example of a prompt is as follows:
[1931] "Jazz piano, I'm a Buddhist and a liberal, my favorite book is The Alchemist, and my favorite movie is Inception."
[1932] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1933] Step 1: Enter your user information
[1934] On the device: The user launches the application, enters their name, email address, and password in the form that appears on the screen, and then clicks the "Next" button.
[1935] Input: Name, Email Address, Password
[1936] Output: Sends the user's basic information to the server
[1937] Device: The next screen displays fields for the user to enter their favorite instrument, music genre, religious views, political views, and favorite books and movies. The user fills in each field and clicks the "Submit" button.
[1938] Input: musical instruments, music genres, religious views, political views, favorite books, movies
[1939] Output: Sends user details to server
[1940] Server: Receives the entered information, stores it in a database, generates a new user ID, and returns that information to the device.
[1941] Input: User information
[1942] Data processing: saving to database, generating new user ID
[1943] Output: New user ID
[1944] Step 2: Collecting and Managing Virtual Player Data
[1945] Server: Provides an interface with the functionality to collect performance data (music files and MIDI data) from real performers.
[1946] Input: Performance data
[1947] Output: Collected performance data
[1948] Server: Provides and collects a form for performers to enter their preferences (religious views, political views, favorite books and movies, etc.).
[1949] Input: Player preference information
[1950] Output: Collected preference information
[1951] Server: Collected performance data and performer preference information are stored in a database and associated with individual performer IDs.
[1952] Input: Collected performance data, preference information
[1953] Data processing: Save to database, associate with performer ID
[1954] Output: Data linked to performer ID
[1955] Step 3: Data analysis and machine learning model creation
[1956] Server: Executes data analysis algorithms to extract features such as rhythm, tempo, and pitch from performance data.
[1957] Input: Performance data
[1958] Data Calculation: Feature Extraction
[1959] Output: Extracted features
[1960] Server: Using the extracted features and personal preference information as input, the generative AI model is used to train a machine learning model to identify playing styles.
[1961] Input: extracted features, personal preference information
[1962] Data Computing: Training Machine Learning Models
[1963] Output: A trained machine learning model
[1964] Step 4: Select a Virtual Player
[1965] Device: The user accesses the "Virtual Player Selection" screen within the app.
[1966] Input: A view request from a user
[1967] Output: List of virtual players
[1968] Server: Provides virtual player data (playing style, preferences, sample performance) to the device.
[1969] Input: User ID
[1970] Data processing: filtering virtual players according to user preferences
[1971] Output: Virtual player data
[1972] Terminal: The user browses the provided data and selects the virtual player that best suits their preferences.
[1973] Input: View virtual player data
[1974] Output: Selected Virtual Player
[1975] Step 5: Session Start Settings
[1976] Device: The user goes to the "Session Start Settings" screen and enters basic settings such as tempo, key, and session time.
[1977] Input: Tempo, Key, Session Time
[1978] Output: Configuration information
[1979] Server: Receives this configuration information and begins preparing a session with the selected virtual player.
[1980] Input: Setting information, User ID, Virtual Player ID
[1981] Data processing: session preparation
[1982] Output: Session ready notification
[1983] Step 6: Conduct a real-time session
[1984] Terminal: A session begins when the user presses the "Start Session" button.
[1985] Input: Session start request from user
[1986] Output: Session start notification
[1987] Server: Receives the user's performance data in real time and extracts features such as rhythm, tempo, and pitch.
[1988] Input: User performance data
[1989] Data calculations: Real-time feature extraction
[1990] Output: Extracted features
[1991] Emotion Engine: Collects emotional data from the user's voice, facial expressions, and biometric sensors to recognize their emotional state.
[1992] Input: Voice data, facial expression data, biometric sensor data
[1993] Data Computing: Emotional State Recognition
[1994] Output: Emotion data
[1995] Server: Adjusts the virtual player's playing style in real time based on the features of emotional data and performance data.
[1996] Input: Emotion data, performance data features
[1997] Data calculation: Adjusting playing style
[1998] Output: Adjusted performance data
[1999] Server: The generated response performance data is sent to the user's device in real time.
[2000] Input: Adjusted performance data
[2001] Output: Response performance data
[2002] Terminal: Based on the response performance data generated by the user, the user continues to play in real time together with the virtual player.
[2003] Input: Response performance data
[2004] Output: Real-time collaboration
[2005] Step 7: Gather feedback
[2006] Terminal: After the session ends, a form will appear where the user can enter feedback about the virtual player's performance. The user enters their feedback and clicks the "Submit" button.
[2007] Input: Feedback information
[2008] Output: Feedback information sent to server
[2009] Server: Receives feedback data and stores it in a database. This feedback data is used to update the machine learning model in the future.
[2010] Input: Feedback data
[2011] Data processing: Saving to database
[2012] Output: Stored feedback data
[2013] Step 8: Calculating and paying compensation
[2014] Server: Collects data such as the number of times each virtual player has been used, the duration of use, and user feedback.
[2015] Input: Number of uses, duration, feedback
[2016] Data calculation: Aggregation work
[2017] Output: Aggregated data
[2018] Server: Based on the aggregated data, calculates the rewards for the providers of virtual player data.
[2019] Input: Aggregate data
[2020] Data calculation: Reward calculation
[2021] Output: Calculated reward data
[2022] Server: Pays the calculated reward according to the payment method specified by the original data provider.
[2023] Input: Remuneration data, payment method
[2024] Output: Payment processing completion notification
[2025] (Application example 2)
[2026] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2027] Conventional music session systems have had problems in that they make it difficult for users to engage in emotionally-based interactive sessions with virtual players, and lack the functionality to dynamically adjust the virtual player's playing style in response to real-time changes in the user's emotions. Another issue is accurately recognizing the user's emotions and providing feedback accordingly.
[2028] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting user information, means for managing virtual player data including performance data and personal preference information, means for the user to select a virtual player, means for generating performance data in real time during a session, means for collecting feedback after the session, means for paying a reward to the data provider according to the degree of use, and means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the playing style of the virtual player based on that. This allows the user to enjoy an interactive music session with a virtual player that responds to their emotions in real time.
[2029] The "means for inputting user information" is an interface that allows a user to input their basic information, musical preferences, etc. into the system.
[2030] The "means for managing virtual player data including performance data and personal preference information" is a system that collects performance data and preference information of virtual players and manages them in a database.
[2031] The "means for the user to select a virtual player" is an interface that allows the user to select one of the virtual players displayed in the system.
[2032] The "means for generating performance data in real time during a session" is a system that generates and analyzes performance data of the user and virtual players in real time during a session.
[2033] The "means for collecting feedback after a session" is an interface for collecting feedback from the user after the session has ended.
[2034] The "means for paying a reward to a data provider according to the degree of use" is a system for calculating and paying a reward to a data provider of a virtual player according to the degree of use.
[2035] "Means for recognizing the user's emotional state using an emotion engine and dynamically adjusting the virtual player's playing style based on that" is a system that uses an emotion engine to recognize the user's emotions and dynamically adjusts the virtual player's playing style according to the user's emotional state.
[2036] This invention is a system that allows a user to engage in a real-time musical session with a virtual player, and further recognizes the user's emotions and adjusts the playing style accordingly. The system includes the following main components:
[2037] 1. Enter your user information
[2038] Hardware: Smartphones, tablets, computers
[2039] Software: Applications
[2040] What happens: A user creates an account and enters basic information such as name, email address, password, preferred instrument and music genre, etc. This information is stored in a cloud database and a new user ID is generated.
[2041] 2. Virtual Player Data Management
[2042] Hardware: Server
[2043] Software: Cloud Management System
[2044] Processing: Performance data (music files and MIDI data) collected from actual performers and personal preference information are stored in a database, and data about each performer is managed.
[2045] 3. Emotion Recognition and Data Analysis
[2046] Hardware: smartphone facial recognition cameras, microphones, and biometric sensors
[2047] Software: Emotion recognition engine (open source emotion recognition software, e.g., OpenFace)
[2048] Processing: Emotional data is collected from the user's facial expressions, voice, and biometric sensors, and analyzed in real time by an emotion recognition engine.
[2049] 4. Conducting real-time sessions
[2050] Hardware: Smartphone
[2051] Software: In-app real-time music session function
[2052] Processing: Receives and analyzes the user's performance data in real time to extract features. Dynamically adjusts the virtual player's performance style based on the data from the emotion engine. The generated response performance data is sent to the user's device in real time, allowing the user to engage in a session with the virtual player.
[2053] 5. Gathering Feedback
[2054] Hardware: Smartphone
[2055] Software: Feedback collection feature
[2056] Processing: After the session ends, the user provides feedback on the virtual player's performance. The collected feedback data is stored in a cloud database and used to update the machine learning model.
[2057] 6. Calculation and Payment of Rewards
[2058] Hardware: Server
[2059] Software: Remuneration calculation system
[2060] Processing: Compensation is calculated based on the number of uses, duration of use, and user feedback for each virtual player, and paid to the data provider.
[2061] Examples:
[2062] For example, if a user wants to play jazz guitar, they enter their user information as "jazz guitar, Catholic, liberal, favorite book is 'Norwegian Wood,' and favorite movie is 'La La Land.'" They then select a virtual jazz guitarist within the app and enter settings to start a session with a tempo of 140 BPM and a key of F major. During the session, the emotion engine recognizes the user's excitement level from their facial expressions and voice, adjusting the virtual player's performance to be more energetic. After the session ends, the user provides feedback such as "I'd like a more rhythmic performance," and the server uses this feedback to update the machine learning model.
[2063] Example prompt sentence:
[2064] The user begins a 140 BPM jazz session in F major with a virtual jazz band. The user sets their favorite music genre as jazz, is Catholic and liberal, and their favorite book is "Norwegian Wood," and their favorite movie is "La La Land." During the session, the system recognizes emotions from the user's smile and voice, adjusting the performance to be more energetic.
[2065] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2066] Step 1:
[2067] A user launches an application and creates an account. The user enters basic information such as their name, email address, password, favorite instrument, and music genre. The information entered is sent from the device to the server and stored in a cloud database. A new user ID is generated and returned to the device. The input for this step is the user information, and the output is the new user ID.
[2068] Step 2:
[2069] The server manages data about each performer by storing performance data (music files and MIDI data) and personal preference information collected from the actual performers in a database. The input is the performance data and preference information, and the output is the virtual player data stored in the database.
[2070] Step 3:
[2071] The user browses a list of virtual players in the app and selects one that suits their preferences. The server provides virtual player data (playing style, preferences, sample performance, etc.) to the device. The input is the virtual player data, and the output is the virtual player selected by the user.
[2072] Step 4:
[2073] The user proceeds to the session start screen and inputs basic settings such as tempo, key, etc. The device sends the input settings to the server and begins preparing for a session with the selected virtual player. The input is session setting information, and the output is a message indicating that the session is ready.
[2074] Step 5:
[2075] A session begins when the user presses the "Start Session" button. The device collects the user's performance data in real time and sends it to the server. The server analyzes the received performance data and extracts features such as rhythm and tempo. The input is the user's performance data, and the output is the analyzed features.
[2076] Step 6:
[2077] The emotion engine collects emotional data from the user's voice, facial expression, and biometric sensors to recognize the emotional state. The emotion engine analyzes the emotional state based on this data and identifies the type of emotion (happiness, sadness, excitement, etc.). The input is voice data, facial expression data, and biometric data, and the output is the recognized emotional state.
[2078] Step 7:
[2079] The server adjusts the virtual player's performance style in real time based on the emotion data from the emotion engine and the performance data features. This allows the performance style to reflect the user's emotions. The input is emotion data and performance data features, and the output is the adjusted performance data.
[2080] Step 8:
[2081] The server transmits the response performance data generated in real time to the user's terminal, and the user engages in a session in real time with the virtual player. The input is the adjusted performance data, and the output is the performance data transmitted to the user's terminal.
[2082] Step 9:
[2083] After the session ends, the user provides feedback on the virtual player's performance. The device sends the feedback data to the server and stores it in a database. The input is the feedback data, and the output is the feedback data stored in the database.
[2084] Step 10:
[2085] The server calculates rewards for each virtual player based on the number of uses, duration of use, and user feedback. The calculated rewards are paid to the data provider. The inputs are usage data and feedback data, and the output is the calculated rewards.
[2086] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2087] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2088] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2089] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2090] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2091] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2092] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2093] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2094] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2095] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2096] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2097] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2098] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2099] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2100] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2101] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2102] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2103] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2104] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2105] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2106] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2107] The following is further disclosed regarding the above embodiment.
[2108] (Claim 1)
[2109] a means for inputting user infor...
Claims
1. a means for inputting user information; means for managing virtual player data including performance data and personal preference information; a means for a user to select a virtual player; means for initiating a session with a selected virtual player; means for generating performance data in real time during a session; a means of collecting feedback after the session; a means for paying remuneration to performers according to the degree of use; A system including:
2. The system according to claim 1 , further comprising means for analyzing performance data of the virtual player data and identifying a performance style using a machine learning model.
3. 2. The system according to claim 1, further comprising means for analyzing the user's performance data during a session and generating performance data to which the virtual player responds in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A